# Language implementation design decisions and trade offs

**URL:** <https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989>\
**Category:** Questions & Answers\
**Tags:** question\
**Created:** [May 12, 2022, 11:52am UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989 "2022-05-12T11:52:05Z")\
**Posts on this page:** 12\
**Page:** 1

<div class="post-metadata">

**Author:** ![spdegabrielle](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/spdegabrielle/32/95_2.png) [@spdegabrielle](https://racket.discourse.group/u/spdegabrielle)\
**Post date:** [May 12, 2022, 11:52am UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/1 "2022-05-12T11:52:05Z")

</div>

Hi

@lexi\_lambda wrote an interesting comment on language implementation trade-off’s:

> <https://twitter.com/lexi_lambda/status/1524374359871209472?s=21&t=-DKkAb1Knnq3atuD9ziCmw>

I recently learnt in chat(discord); roughly the racket compiler produces safer executables at the cost of a little (runtime)speed. (Thank you @samth )

I’d be interested to know more about these sort of design decisions?

I’m sure there are thousands - I'm interested in knowing what the top 5 most important or interesting ones are for professional language designers.

Pointers to texts, courses(course notes), papers and dissertations also appreciated.

Bw

Stephen

---

<div class="post-metadata">

**Author:** ![gus-massa](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/gus-massa/32/507_2.png) [@gus-massa](https://racket.discourse.group/u/gus-massa)\
**Post date:** [May 12, 2022, 10:06pm UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/2 "2022-05-12T22:06:12Z")

</div>

I have no formal notes, but my experience from trying to write solutions for [The Computer Language Benchmarks Game](https://benchmarksgame-team.pages.debian.net/benchmarksgame/index.html):

- Racket has thread safe `hash`es. This is a problem for toy programs becasue in a toy program you can put each `hash` in a different `place`, so it is "better" to implement your own thread-unsafe `hash`, but the "game" ask for idiomatic solutions.

- Strings are very slow in Racket BC. (I didn't try too many microbenchmarks in Racket CS.) As a rule of thumb, Racket is 5x faster than Python for number crushing, but x1 (i.e. equal) than Python for programs that use strings. I'm not sure about the detials here, but my guess is that Racket has Unicode support by default but other languages in the "game" use the fact that the data is just ASCII. Only ASCII is a good guess for technical files, but once you are nearby real humans, they love to add weird things arround letters. [¡Hi from Argentina! Don't ask why "ñ" can't be safetly replaced by "n".]

- In `vector`s and other things, Racket can mix `fixnum`, `flonum`s, `struct`s, `box`es. So a `fixnum`s `N` is represented in BC as `2*N+1` and by `8N` in CS. [I hope I remember the detaisl correctly.] So `A+B` is implemented as `A+B-1` or `(A-1)+(B-1)+1` or something like that, and each `fixnum` operations has a small overhead. It's worse for `flonum`s because each flonum has it's own secret hidden flonum-only-box that must be alloccated and garbage-collected. Luckly in both BC and CS the flonum’s can be unboxed by the compiler (@mflatt wrote that part). If you chain `flsomething` operations, the compiler may skip the intermediate boxes and reduce the overhead.

---

<div class="post-metadata">

**Author:** ![benknoble](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/benknoble/32/16_2.png) [@benknoble](https://racket.discourse.group/u/benknoble)\
**Post date:** [May 13, 2022, 1:56am UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/3 "2022-05-13T01:56:55Z")

</div>

RE: vectors and nums, I believe `flvector` and `fxvector` are relatively idiomatic for perf-sensitive code. They have (IIUC) more compact representations and operate faster in general?

---

<div class="post-metadata">

**Author:** ![sschwarzer](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/sschwarzer/32/1940_2.png) [@sschwarzer](https://racket.discourse.group/u/sschwarzer)\
**Post date:** [May 13, 2022, 10:09am UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/4 "2022-05-13T10:09:02Z")

</div>

> [@gus-massa](#):
>
> As a rule of thumb, Racket is 5x faster than Python for number crushing, but x1 (i.e. equal) than Python for programs that use strings.

By number crunching, do you mean programs using NumPy/Pandas for the numeric operations or programs where loops, conditions etc. are in Python? If the latter, I'm surprised if the factor to Racket isn't larger. 🙂

On the other hand, the builtin string operations in Python are all in C code "below" the Python API, so I'm _not_ surprised they're equally fast as in Racket.

> [@gus-massa](#):
>
> I'm not sure about the detials here, but my guess is that Racket has Unicode support by default but other languages in the "game" use the fact that the data is just ASCII.

Python uses unicode strings by default, like Racket.

> [@gus-massa](#):
>
> [...] but other languages in the "game" use the fact that the data is just ASCII. Only ASCII is a good guess for technical files, but once you are nearby real humans, they love to add weird things arround letters.

It really depends on what you're doing. It wouldn't make sense to decode and re-encode a, say, UTF-8 byte stream, if you just copy a file.

Also, if you know that the input bytes are UTF-8 (not just ASCII), you can still safely run certain operations on them without decoding them. For example, you can split paths at `/` separators or split off drive letters at `:`. This works because UTF-8 multibyte sequences [don't contain bytes that by themselves could be 7-bit ASCII characters](https://en.wikipedia.org/wiki/Utf-8#Encoding).

> [@benknoble](#):
>
> They have (IIUC) more compact representations and operate faster in general?

I think `flvector` and `fxvector` contain contiguous sequences of the low-level datatypes, so they would be equivalent to a C arrays of `double` or `int` (for example, depending on the size of a fixnum). Especially if you combine `flvector` and `fxvector` with their "corresponding" [unsafe operations](https://docs.racket-lang.org/reference/unsafe.html), you can get a nice speed-up. 🙂

---

<div class="post-metadata">

**Author:** ![maketo](https://avatars.discourse-cdn.com/v4/letter/m/a698b9/32.png) [@maketo](https://racket.discourse.group/u/maketo)\
**Post date:** [May 13, 2022, 11:26am UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/5 "2022-05-13T11:26:50Z")

</div>

> [@sschwarzer](#):
>
> By number crunching, do you mean programs using NumPy/Pandas for the numeric operations or programs where loops, conditions etc. are in Python? If the latter, I'm surprised if the factor to Racket isn't larger. 🙂

If memory serves me - most of Numpy is C/C++ and has been heavily optimized over the years by virtue of being exposed to production stress by many organizations. I cannot speak to Pandas. Vanilla Python - I would not be surprised if it is slow but nobody uses vanilla Python to do number crunching in any serious production code.

---

<div class="post-metadata">

**Author:** ![sschwarzer](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/sschwarzer/32/1940_2.png) [@sschwarzer](https://racket.discourse.group/u/sschwarzer)\
**Post date:** [May 13, 2022, 11:57am UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/6 "2022-05-13T11:57:13Z")

</div>

> [@maketo](#):
>
> nobody uses vanilla Python to do number crunching in any serious production code.

Depends on the extent of number crunching. 😉 If it's an algorithm you can't easily implement with NumPy and the "number crunching" part isn't a big part of the program or can be done "offline" (during the night etc.), you may get away with it. 😉

I agree that _if_ you have relatively straightforward numerical code you can implement with NumPy or SciPy, plain Python would be _much, much_ slower.

---

<div class="post-metadata">

**Author:** ![benknoble](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/benknoble/32/16_2.png) [@benknoble](https://racket.discourse.group/u/benknoble)\
**Post date:** [May 13, 2022, 2:15pm UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/7 "2022-05-13T14:15:50Z")

</div>

Pandas is considered slow and unscalable, I believe [citation needed].

---

<div class="post-metadata">

**Author:** ![gus-massa](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/gus-massa/32/507_2.png) [@gus-massa](https://racket.discourse.group/u/gus-massa)\
**Post date:** [May 13, 2022, 4:12pm UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/8 "2022-05-13T16:12:59Z")

</div>

> [@sschwarzer](#):
>
> By number crunching, do you mean programs using NumPy/Pandas for the numeric operations or programs where loops, conditions etc. are in Python? If the latter, I'm surprised if the factor to Racket isn't larger. 🙂

The rules of the "game" ask to use plain Python. Looking again at [https://benchmarksgame-team.pages.debian.net/benchmarksgame/fastest/racket-python3.html](https://benchmarksgame-team.pages.debian.net/benchmarksgame/fastest/racket-python3.html) there are a few x10, an even x20 or x40. Sometimes the difference is in the paralelization.

I expect less difference if the programs uses NumPy instead, but it's a lot of work to rewrite them. I only modified a few programs with NumPy for unrelated work, but I just followed the style of the previous code, searched in Stack Overflow and hopped that my coded didn't break the performance too much.

---

<div class="post-metadata">

**Author:** ![samth](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/samth/32/3_2.png) [@samth](https://racket.discourse.group/u/samth)\
**Post date:** [May 13, 2022, 6:37pm UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/9 "2022-05-13T18:37:22Z")

</div>

On strings specifically, CPython has multiple representations of strings, including more compact representations for ASCII-only strings. It is likely to be more efficient than Racket for such programs, since Racket uses UTF-32 for all strings in memory.

---

<div class="post-metadata">

**Author:** ![soegaard](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/soegaard/32/19_2.png) [@soegaard](https://racket.discourse.group/u/soegaard)\
**Post date:** [May 14, 2022, 7:51am UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/10 "2022-05-14T07:51:26Z")

</div>

Is it cheating to use byte strings in the benchmarks?

---

<div class="post-metadata">

**Author:** ![spdegabrielle](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/spdegabrielle/32/95_2.png) [@spdegabrielle](https://racket.discourse.group/u/spdegabrielle)\
**Post date:** [May 14, 2022, 8:54am UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/11 "2022-05-14T08:54:05Z")

</div>

I don’t think so. How can you cheat in an exercise where the goal is to cheat as much as possible?

---

<div class="post-metadata">

**Author:** ![gus-massa](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/gus-massa/32/507_2.png) [@gus-massa](https://racket.discourse.group/u/gus-massa)\
**Post date:** [May 14, 2022, 2:02pm UTC](https://racket.discourse.group/t/language-implementation-design-decisions-and-trade-offs/989/12 "2022-05-14T14:02:41Z")

</div>

It's more complicated, because the inputs of the programs are fixed and public, so you could write

```
#lang racket
(display 42)

```

and call it a day.

So the rules ask to use the "same" algorithm, whatever that means when you are using very different languages. [My guess is that this is a huge source of nasty emails send to the organizer, becuase people claim their creative solution is not cheating or claim that other languages are cheating.]

I think that using byte string is fine, but IIRC it didn't change the time too much. It was weird, and perhaps I was doing something wrong or perhaps I'm misremembering. I don't have too much spare time now, so if someone likes these challenges, it's possible to try agian.
