What Is AGI?

AI & HCI
"AGI 시대가 왔다"는 말은 결론이 아니라, 어떤 정의를 기준으로 삼느냐에 따라 달라지는 주장이다.
Posted on Sept. 14, 2026, 4:22 a.m. by SANGJIN
random_image

OpenAI's September 2026 launch of GPT-6 Astra came with a bold claim from President Greg Brockman: "we've entered the AGI era." Nvidia CEO Jensen Huang echoed it days later. But the term AGI has no single agreed-upon definition — so the same announcement means different things depending on whose yardstick is used.

Five ways AGI gets defined

Perspective Criterion Note
OpenAI (charter) Outperforms humans at most economically valuable work Brockman now calls it more of a "mission concept" than a technical bar
Google DeepMind 5-stage spectrum: Emerging → Competent → Expert → Virtuoso → Superhuman Treats AGI as a matter of degree, not a yes/no switch
Academic (CHC-based) Quantified scoring across 10 cognitive domains, benchmarked against human psychometrics Current models show "jagged" profiles — strong on knowledge, weak on long-term memory
Economic AI performs 80% of economically valuable jobs Skips the philosophy, measures labor-market impact directly
Skeptical / philosophical Turing test, consciousness, embodiment Questions whether benchmark scores equal genuine general intelligence

Why Astra is cited as evidence

  • Benchmark jumps: OpenAI reported 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, and 100% on ExploitBench (up from 78.5% for the prior model), including discovering two previously unknown zero-day vulnerabilities during evaluation.
  • Cross-domain computer use: Demos spanned PCB design in KiCad, 3D city scenes in Unity, an animated car transmission in FreeCAD/Blender, and drafting a tax return from a W-2 — domains chosen specifically to test breadth, not depth in one field. Simulated knowledge-work tasks reportedly took ~47% less time than with the prior model.
  • Industry framing: Jensen Huang tied the "AGI has arrived" claim to Astra's training scale (100,000+ Nvidia Grace Blackwell NVL72 systems), a framing some read as being as much about GPU demand as about the model itself.

The pushback

No standardized AGI test exists, so passing a handful of benchmarks doesn't settle the question. Anthropic didn't describe its own recent models as reaching AGI, and some observers instead point to an earlier Anthropic model as the real starting point — a sign that even the "which model first" question has no consensus. Astra also became the first OpenAI model to receive the company's highest cybersecurity risk rating, since it can find and exploit vulnerabilities without guidance — a capability jump that raises control concerns as much as it demonstrates generality.

====

2026년 9월 오픈AI가 GPT-6 아스트라를 공개하며 그렉 브록먼 사장은 "AGI 시대에 진입했다"고 선언했다. 며칠 뒤 젠슨 황 엔비디아 CEO도 같은 주장을 반복했다. 그런데 AGI라는 용어 자체에 합의된 정의가 없다 보니, 같은 발표도 어떤 기준을 대느냐에 따라 전혀 다른 의미가 된다.

AGI를 정의하는 다섯 가지 관점

관점 판정 기준 비고
오픈AI (헌장) 경제적으로 가치 있는 대부분 업무에서 인간 능가 브록먼은 최근 이를 기술 기준보다 '미션 개념'에 가깝다고 설명
구글 딥마인드 5단계 스펙트럼(초기→유능→전문가→거장→초인) AGI를 이분법이 아닌 정도의 문제로 취급
학계 (CHC 기반) 10개 인지 영역을 인간 심리측정 도구로 정량 채점 현재 모델은 지식 영역은 강하지만 장기 기억은 취약한 '들쭉날쭉'한 프로필
경제학적 관점 가치 있는 일자리의 80%를 AI가 수행 철학적 논쟁 대신 노동시장 충격을 직접 측정
회의적/철학적 관점 튜링 테스트, 의식, 체화(embodiment) 벤치마크 점수가 진짜 일반 지능을 뜻하는지에 의문 제기

아스트라가 근거로 지목되는 이유

  • 벤치마크 급상승: FrontierMath Tier 4 98%, ARC-AGI-3 99.9%, ExploitBench 100%(직전 모델은 78.5%)를 기록했고, 평가 과정에서 미공개 제로데이 취약점 2건을 직접 발견했다고 보고됐다.
  • 도메인을 넘나드는 컴퓨터 사용: KiCad PCB 설계, 유니티 3D 도시 장면, FreeCAD·블렌더 자동차 변속기 애니메이션, W-2 기반 세금 신고서 초안 작성까지 — 한 분야의 깊이가 아니라 여러 분야를 넘나드는 폭을 보여주려 설계된 시연이다. 실제 지식노동 과제 시뮬레이션에서는 이전 모델 대비 작업 시간이 약 47% 단축됐다.
  • 업계 리더의 프레이밍: 젠슨 황은 "AGI 도래" 주장을 아스트라의 학습 규모(엔비디아 그레이스 블랙웰 NVL72 10만 대 이상)와 연결지었는데, 이는 모델 자체보다 GPU 수요 관점의 발언이라는 해석도 있다.

반론

표준화된 AGI 판정 기준이 없는 상태에서 몇 개 벤치마크 통과만으로 결론을 내리기는 이르다. 앤트로픽은 자사 최신 모델을 AGI 도달로 표현하지 않았고, 오히려 이전 모델을 진짜 시작점으로 보는 시각도 있어 "어느 모델이 먼저였는가"조차 합의가 없다. 또한 아스트라는 가이드라인 없이도 취약점을 찾아 공격 도구를 만들 수 있어 오픈AI 모델 중 처음으로 최고 사이버보안 위험 등급을 받았는데, 이는 범용성의 증거이자 동시에 통제 가능성에 대한 우려이기도 하다.

Leave a Comment: