Glad they set a time rather than leaving it open

  • melfie@lemmy.zip
    link
    fedilink
    English
    arrow-up
    2
    ·
    9 days ago

    I got a 7900 XTX a while back to run 3.6 27B assuming it might be a while before there’s anything better, but the 3.8 announcement is a pleasant surprise. 3.6 27B is more or less on par with Claude Sonnet 4.6 for coding according to benchmarks and my own experience and I’m hoping 3.8 will be more like Sonnet 5.

    Incidentally, Newegg isn’t the greatest, but I did get a new 7900 XTX for like $730 after trading in my old RTX 3070. With used 3090s going for $1200 / $2k new or 4090s going for $3500 new, I think the 7900 XTX is really the only sanely priced 24GB GPU left. ROCm is slower than CUDA due to less maturity and you might be on your own getting certain things working that only support CUDA, but man does llama.cpp work well. With 3.6 27B, I’m getting 200k context, 700 t/s prefill on average (1000+ with small context) and 20 t/s decode without MTP or much optimization (hoping to get like 40-50 decode after spending more time).