• A 150M model reached 29.5% while costing just $0.0007 per task
  • ChatGPT scored higher, yet its comparable reasoning runs cost substantially more
  • BDH-CQ performs reasoning internally instead of generating lengthy intermediate text

Pathway, an AI lab focused on building Post-Transformer architectures, has released new benchmark results for its BDH-CQ reasoning model.

According to the researchers, their 150M-parameter model scored 29.5% pass@2 on the public ARC-AGI-1 evaluation set.



Source link

Podcast also available on PocketCasts, SoundCloud, Spotify, Google Podcasts, Apple Podcasts, and RSS.