Summary:

  • GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness.
  • GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels.
  • A key behavior observed in GPT-6 Astra was its ability to turn unfamiliar environments into compact symbolic world models. It represented game mechanics as logical rules and developed its own domain-specific language shorthand to track state and plan actions.

For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.
Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted.1

Astra’s results are also a major milestone worth celebrating. From our perspective, Astra represents a noticeable step-function change in frontier model capabilities.

When we launched ARC-AGI-3, we made it clear that saturating the benchmark would not represent “proof of achieving AGI.” Therefore, while we believe Astra represents meaningful progress towards generalization, we are not claiming that it is AGI.

    • brianpeiris@lemmy.caOP
      link
      fedilink
      English
      arrow-up
      2
      ·
      18 hours ago

      I’m okay with the downvotes. There’s a stronger anti-AI sentiment now than even just a month ago, and on the whole I think that’s going to be a good thing.

  • brianpeiris@lemmy.caOP
    link
    fedilink
    English
    arrow-up
    19
    arrow-down
    5
    ·
    3 days ago

    I guess this is being down-voted just because it feels pro-AI. Believe me I get it, you can look at my anti-AI post history. But I think it’s important to also keep an eye out for real signals towards general intelligence, which is why I wanted to share this. I’m not celebrating it, just bringing awareness.

    • gankouskhan@bookwyr.me
      link
      fedilink
      English
      arrow-up
      9
      arrow-down
      3
      ·
      3 days ago

      For as long as it is an LLM I will not call it intelligence for it does not think. It’s an approximation or an estimate at best. While capable this is a major hurdle these companies will face to replace skilled workers. We are a generation at least and a ways off from general intelligence, and not algorithmic output for a given input (idempotency).

      I do however agree that we would remain informed about the ever changing world around us; this however, is until reviewed by an unbiased third party in masse a marketing document. I will be taking it with a grain of salt, but will take it as an improvement over the prior versions.

      Now with that out of the way… interesting. I will be watching this progress, but am far less optimistic in this being as capable as they are suggesting. It’s more than likely another Mythos hype attempt that is not revolutionary, but a nice addition to existing systems or augments to staff. Now if we assume its 100% as capable as this suggests then we are in for a ride.

      • brianpeiris@lemmy.caOP
        link
        fedilink
        English
        arrow-up
        5
        arrow-down
        2
        ·
        3 days ago

        I agree with most of this. I’ve also said in the past that LLMs cannot think, and I think that’s still true for most models. The reason ARC-AGI-3 is interesting is that it was specifically designed to test reasoning, adaptability, novel problem solving, planning, memory, etc. So it was a surprise to me that Astra was able to defeat it so effectively, and that Astra invents algebras for each novel task.

        But I agree we can’t trust OpenAI if these results are self-reported, and we may not be able to trust the ARC Prize Foundation fully either. Extraordinary claims require extraordinary evidence, so we need replication, transparency, and proper open science to confirm things.

        I also agree with ARC Prize’s conclusion, that there are still capabilities any AI system would need to demonstrate before we can claim a full general intelligence.

  • hirihit640@sh.itjust.works
    link
    fedilink
    English
    arrow-up
    7
    arrow-down
    2
    ·
    3 days ago

    Wow they benchmaxxed this pretty quickly. Wasn’t this benchmark released only a few months ago with like 0% success rate from all major AI models?

    • brianpeiris@lemmy.caOP
      link
      fedilink
      English
      arrow-up
      4
      arrow-down
      1
      ·
      3 days ago

      Yes, it was released in March. One important point is that these results are based on the semi-private test set. Perhaps it’s best to wait until it is measured against the fully private test set, but that might not happen for months.