Source

Branch

1/ @auburnsounds

@Jonathan_Blow @rflaherty71

FWIW I implement SIMD for Dlang. D has inline assembly, but minimal use. inline asm end up almost always slower than intrinsics. The exception to that is LLVM and GCC __asm that can inline in caller, and works in debug mode since no optimize needed. Else intrinsics win.