# Yet another Racket benchmark comparison

**URL:** <https://racket.discourse.group/t/yet-another-racket-benchmark-comparison/4335>\
**Category:** Show & Tell\
**Created:** [July 30, 2026, 2:20pm UTC](https://racket.discourse.group/t/yet-another-racket-benchmark-comparison/4335 "2026-07-30T14:20:03Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![drasken](https://avatars.discourse-cdn.com/v4/letter/d/f07891/32.png) [@drasken](https://racket.discourse.group/u/drasken)\
**Post date:** [July 30, 2026, 2:20pm UTC](https://racket.discourse.group/t/yet-another-racket-benchmark-comparison/4335/1 "2026-07-30T14:20:03Z")

</div>

After reading [this](https://www.stylewarning.com/posts/nbody/) blog post about a comparison using C and CommonLisp (SBCL) I quickly wrote a Racket version just out of curiosity to make yet another unnecessary comparison 🙂

I just cloned the directory linked in the post plus adding a .rkt file and the lines in the Makefile to run the benchmark, [here](https://github.com/drasken/lisp-stuff/tree/main/rnbody) is my version. At the end, the resulting Racket code run one order of magnitude slower than C and CL on with 5M parameter, two with 500k as parameter.

My code is surely far from optimal and have space for optimization so I just wanted to share to get some advice on how to possibly improve the code.

---

<div class="post-metadata">

**Author:** ![LiberalArtist](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/liberalartist/32/151_2.png) [@LiberalArtist](https://racket.discourse.group/u/LiberalArtist)\
**Post date:** [July 30, 2026, 3:21pm UTC](https://racket.discourse.group/t/yet-another-racket-benchmark-comparison/4335/2 "2026-07-30T15:21:02Z")

</div>

I've only looked at this quickly, but one thing to bear in mind is that `set!` is generally a red flag in performance-sensitive code. Since `set!` is rare in idiomatic Racket, the compiler doesn't try to do sophisticated analysis about assignments: in fact, any variable that appears on the left-hand side of a `set!` expression is effectively turned into a `box`, which makes the implementation of closures and continuations simple and efficient. It looks like many of your uses of `set!` could be replaced by changing your `for` loops to `for/fold`, perhaps with a `#:result` clause.

Beyond that, when doing performance-sensitive floating-point math, you generally want to replace generic arithmetic operations like `*` with specialized variants like `fl*`. Beyond avoiding unnecessary generic dispatch, the floating-point–specific operators can “unbox” their arguments and intermediate results, keeping them in machine floating-point registers rather than storing them on the heap. One way to get this benefit mostly automatically is to use Typed Racket with floating-point types: Typed Racket can compile `*` to `unsafe-fl*`, and the Optimization Coach can warn you about places where you may be missing the optimization.

A general note for benchmarks is that you can consider [`(#%declare #:unsafe)`](https://docs.racket-lang.org/reference/module.html#%28idx._%28gentag._115._%28lib._scribblings%2Freference%2Freference..scrbl%29%29%29)—but you can decide whether that makes a fair comparison or not, and, anyway, it’s better to start by optimizing your safe code.

---

<div class="post-metadata">

**Author:** ![6cdh](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/6cdh/32/2311_2.png) [@6cdh](https://racket.discourse.group/u/6cdh)\
**Post date:** [July 30, 2026, 5:05pm UTC](https://racket.discourse.group/t/yet-another-racket-benchmark-comparison/4335/3 "2026-07-30T17:05:00Z")

</div>

You can let it 6x faster by add more code, comment some old code, without touching the hot loop.

I also added a Julia version (same as C code) for compare, with 500k as parameter.

First version on my computer:

```plaintext
=== C (GCC) ===
-0.169237033
timing: 24 ms

=== Common Lisp (SBCL) ===
-0.169237033
timing: 32 ms

=== Racket ===
-0.169237033
timing: 875 ms

=== Julia ===
-0.169237033
timing: 29 ms

```

First optimization is, define `body` as flvector instead of struct, then comment the struct definition:

```scheme
; (struct body (x y z vx vy vz mass) #:prefab #:mutable)

(define (body x y z vx vy vz mass)
  (flvector (real->double-flonum x)
            (real->double-flonum y)
            (real->double-flonum z)
            (real->double-flonum vx)
            (real->double-flonum vy)
            (real->double-flonum vz)
            (real->double-flonum mass)))

(define (body-x b) (unsafe-flvector-ref b 0))
(define (body-y b) (unsafe-flvector-ref b 1))
(define (body-z b) (unsafe-flvector-ref b 2))
(define (body-vx b) (unsafe-flvector-ref b 3))
(define (body-vy b) (unsafe-flvector-ref b 4))
(define (body-vz b) (unsafe-flvector-ref b 5))
(define (body-mass b) (unsafe-flvector-ref b 6))

(define (set-body-x! b v) (unsafe-flvector-set! b 0 v))
(define (set-body-y! b v) (unsafe-flvector-set! b 1 v))
(define (set-body-z! b v) (unsafe-flvector-set! b 2 v))
(define (set-body-vx! b v) (unsafe-flvector-set! b 3 v))
(define (set-body-vy! b v) (unsafe-flvector-set! b 4 v))
(define (set-body-vz! b v) (unsafe-flvector-set! b 5 v))
(define (set-body-mass! b v) (unsafe-flvector-set! b 6 v))

```

The result is 666ms, 1.3x faster.

Second optimization is, use flonum operations to replace old arithmetic operations in hot loop. Require `racket/flonum`, then add these code to `advance` function, before any other code:

```scheme
(define + fl+)
(define - fl-)
(define * fl*)
(define / fl/)
(define sqrt flsqrt)

```

The result is 231ms, 3.78x faster than the first version.

In the first for loop of `advance`, you used `set!` to modify `dx`, `dy`, `dz`, `dsq`, and `mag`. But you can actually define them in the loop, and get rid of `set!`. The third optimization is, add an evil macro to `advance` function:

```scheme
(define-syntax-rule (set! x val) (define x val))

```

It replaces `set!` with `define` (never do this in real code). The result is 136ms, 6.4x faster than the first version! It still cannot compare to C, Common Lisp or Julia, but not that slow.

My extra experiment is extract many duplicate `(vector-ref bod-vec i)` to a variable, then use this variable:

```scheme
(for* ([i (in-range NBODIES)]
       [j (in-range (add1 i) NBODIES)])
  (define bi (vector-ref bod-vec i))
  (define bj (vector-ref bod-vec j))

  (define dx (- (body-x bi) (body-x bj)))
  (define dy (- (body-y bi) (body-y bj)))
  (define dz (- (body-z bi) (body-z bj)))
  (define dsq (+ (* dx dx) (* dy dy) (* dz dz)))
  (define mag (/ dt (* dsq (sqrt dsq))))

  (set-body-vx! bi (- (body-vx bi) (* dx mag (body-mass bj))))
  (set-body-vy! bi (- (body-vy bi) (* dy mag (body-mass bj))))
  (set-body-vz! bi (- (body-vz bi) (* dz mag (body-mass bj))))
  (set-body-vx! bj (+ (body-vx bj) (* dx mag (body-mass bi))))
  (set-body-vy! bj (+ (body-vy bj) (* dy mag (body-mass bi))))
  (set-body-vz! bj (+ (body-vz bj) (* dz mag (body-mass bi))))))

```

This result is 49ms. Only 2x slower than C.

---

<div class="post-metadata">

**Author:** ![drasken](https://avatars.discourse-cdn.com/v4/letter/d/f07891/32.png) [@drasken](https://racket.discourse.group/u/drasken)\
**Post date:** [July 30, 2026, 5:24pm UTC](https://racket.discourse.group/t/yet-another-racket-benchmark-comparison/4335/4 "2026-07-30T17:24:58Z")

</div>

> one thing to bear in mind is that `set!` is generally a red flag ... many of your uses of `set!` could be replaced by changing your `for` loops to `for/fold`, perhaps with a `#:result` clause

Yeah, this is one of the things I was thinking of implementing since I've read about the `set!` red flag but Couldn't figure out a immediate conversion from C...

> when doing performance-sensitive floating-point math, you generally want to replace generic arithmetic operations like `*` with specialized variants like `fl*`.

TIL: I didn't know about the fl\* operator, always used \*, thanks for letting me discoveer it

---

<div class="post-metadata">

**Author:** ![drasken](https://avatars.discourse-cdn.com/v4/letter/d/f07891/32.png) [@drasken](https://racket.discourse.group/u/drasken)\
**Post date:** [July 30, 2026, 5:30pm UTC](https://racket.discourse.group/t/yet-another-racket-benchmark-comparison/4335/5 "2026-07-30T17:30:58Z")

</div>

> First optimization is, define `body` as flvector instead of struct, then comment the struct definition:

Thanks to you too, didn't think of it at all, as I said I was too distracted by the C code and ended up as a code mimicking it, surprised that just that made a significant improvement in speed execution

> Second optimization is, use flonum operations to replace old arithmetic operations in hot loop. Require `racket/flonum`, then add these code to `advance` function, before any other code:

And thanks for this too, I didn't know about this, always used Racket for personal projects and stuck to the simple usual operators, great discovery

> This result is 49ms. Only 2x slower than C.

Cool, wasn't expecting such speed, great result!

---

<div class="post-metadata">

**Author:** ![LiberalArtist](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/liberalartist/32/151_2.png) [@LiberalArtist](https://racket.discourse.group/u/LiberalArtist)\
**Post date:** [July 30, 2026, 6:03pm UTC](https://racket.discourse.group/t/yet-another-racket-benchmark-comparison/4335/6 "2026-07-30T18:03:56Z")

</div>

> [@6cdh](#):
>
> My extra experiment is extract many duplicate `(vector-ref bod-vec i)` to a variable, then use this variable:

Good catch! Even better, you can replace `in-range` with `in-vector` like this, which will minimize bounds checks:

```scheme
(for* ([bi (in-vector bod-vec)]
       [bj (in-vector bod-vec 1)])

  (define dx (- (body-x bi) (body-x bj)))
  (define dy (- (body-y bi) (body-y bj)))
  (define dz (- (body-z bi) (body-z bj)))
  (define dsq (+ (* dx dx) (* dy dy) (* dz dz)))
  (define mag (/ dt (* dsq (sqrt dsq))))

  (set-body-vx! bi (- (body-vx bi) (* dx mag (body-mass bj))))
  (set-body-vy! bi (- (body-vy bi) (* dy mag (body-mass bj))))
  (set-body-vz! bi (- (body-vz bi) (* dz mag (body-mass bj))))
  (set-body-vx! bj (+ (body-vx bj) (* dx mag (body-mass bi))))
  (set-body-vy! bj (+ (body-vy bj) (* dy mag (body-mass bi))))
  (set-body-vz! bj (+ (body-vz bj) (* dz mag (body-mass bi))))))

```

---

<div class="post-metadata">

**Author:** ![gus-massa](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/gus-massa/32/507_2.png) [@gus-massa](https://racket.discourse.group/u/gus-massa)\
**Post date:** [August 3, 2026, 4:25pm UTC](https://racket.discourse.group/t/yet-another-racket-benchmark-comparison/4335/7 "2026-08-03T16:25:10Z")

</div>

I agree. Some remarks.

I've been doing a bunch of experiment with a throwaway project early this year. I should write about it, but mean while...

- Changing `*` to `fl*` is a big win. Many times the compiler can't figure that the value will be a flonum and uses the slow version, and also there are a few annoying corner cases when the arguments are a mix of flonum and fixnum. (I used `(define * fl*)` as @6cdh said in another comment.)

- Most of the times changing `fl*` to `unsafe-fl*` is only a minor improvement. In my case, many times the compiler can prove the argument is a flonum and fix it. The big problem is that every time I make a mistake it crash the program. I'd ignore it until the final version.

- `(expt x 2)` that is used a few times in this program is very slow. I tried a few variants. The fastets one (like a x10) is to create an auxiliary function `(define-inline (flsqr x) (fl* x x))`. (Using `define` is fine too.)

- ~~`unsafe-flsqrt` is like a 20% faster than `flsqrt` even when it's clear that the argument is a flonum. I'll take a look (in case someone else does not fix it before).~~

---

<div class="post-metadata">

**Author:** ![gus-massa](https://yyz2.discourse-cdn.com/free1/user_avatar/racket.discourse.group/gus-massa/32/507_2.png) [@gus-massa](https://racket.discourse.group/u/gus-massa)\
**Post date:** [August 4, 2026, 4:54pm UTC](https://racket.discourse.group/t/yet-another-racket-benchmark-comparison/4335/8 "2026-08-04T16:54:15Z")

</div>

I retract my comment about the difference of speed of `flsqrt`. I was using something like

```
(define n 3.0)
(set! n 3.0)
(when (flonum? n)
  (flsqrt n))

```

but `n` can potentially be mutated from another module, so the compiler does not apply the optimization. The difference disapears with the correct code

```
(define n (black-box 3.0))
(when (flonum? n)
  (flsqrt n))

```
