Source

Thread

1/ @NOTimothyLottes

For the love of octahedrons: https://fileadmin.cs.lth.se/graphics/research/papers/2008/simdmapping/clarberg_simdmapping08_preprint.pdf

Mapping this well to the GPU follows …

2/ @NOTimothyLottes

At some point the standard Octahedron mapping leaves a lot to be desired due to highly variable texel sizing, so have been using that paper’s equal area mapping instead. Spent some time optimizing today, got to this which is in theory just 31ops for the 2D to 3D transform

Branches: 2025-01-06-JBrooksBSI-some-optimization-ideas-1-for-rdna2-simd-insn

3/ @NOTimothyLottes

AMD’s latest driver actually does mostly a good job compiling that (32 op clocks compiled, just 1 more instruction somewhere). I’m surprised actually it’s now picking up ‘cos(x*2pi)’ and pattern matching that to just V_COS_F32!

4/ @NOTimothyLottes

And the inverse which should also be around 32 op clk (VALU). This uses the Horner form for the papers atan approximation. Could perhaps make that less or more accurate if desired.

5/ @NOTimothyLottes

AMD’s compiler also does a good job there.

6/ @NOTimothyLottes

If AMD evolved to a fused floating point compare and select instead of a separate V_CMP* and V_CNDMASK_B32, it would shave a few cycles. Also if they could push 2 results into a result cache in 1 clk, doing the MIN and MAX in one op would shave a cycle.

7/ @NOTimothyLottes

Related> Great ref for fast atan2: https://mazzo.li/posts/vectorized-atan2.html
I’m using the simplest form from the 1955 paper in horner form scaled by /pi for cylindrical view projection code

Branches: 2025-01-04-marc_b_reynolds-its-pretty-amazing-how-good-the-hastings, 2025-01-06-pixelmager-i-was-looking-at-a-very-similar-problem-for-hemi