---
title: "When 1 is Bigger than 4 for AI"
description: "Learn how AI systems like ChatGPT can produce inconsistent answers and why this matters for B2B applications, especially in areas requiring consistent outputs."
categories: ["AI"]
keywords: ["venture capital","AI inconsistency","stochastic responses","ChatGPT","B2B applications","Theory Ventures","Tomasz Tunguz","genAI","revenue by region","password reset"]
ai_summary: "Learn how AI systems like ChatGPT can produce inconsistent answers and why this matters for B2B applications, especially in areas requiring consistent outputs."
date: 2023-05-03
lastmod: 2026-07-31
canonical_url: https://www.tomtunguz.com/yes-or-no-chatgpt/
author: "Tomasz Tunguz"
---

I asked ChatGPT about the numbers 1 & 4. Which one is bigger? 

Sometimes, 1 was bigger. Othertimes, 4 was bigger. [Sharon Zhou ran this experiment at scale](https://twitter.com/realsharonzhou/status/1614500046824955907?s=46&t=AMQH7tfS4mRCdodm0HHrJQ) to showing the order of yes & no matters in the response.

![image](https://res.cloudinary.com/dzawgnnlr/image/upload/wjyi6p7aonmmlga2orw9.png)
This is called a non-deterministic or stochastic answer. Similar inputs do not consistently produce identical outputs. The answers have inconsistent logic. 


We live with stochastic systems daily : weather reports, ETAs on Google maps, stock portfolio construction. We are stochastic - humans can be moody, err in our calculations, or change our minds with new information. 

In these conversations, the robot is sometimes wrong, but never in doubt. When a system produces an answer, we should verify the answer is correct. It's not just logical errors that occur: hallucinations, when the system invents answers that don't exist, [plagued about half of Bing chat results in this Stanford study.](https://arxiv.org/pdf/2304.09848v1.pdf)

We haven't calibrated ourselves to the level of doubt to express, yet. Like working with a new colleague, we need to understand their strengths & weaknesses. 

For consumers, the universe of acceptable outcomes can be quite broad. A [rabbit on top of a fire truck](https://tomtunguz.com/gpt-stock-images/) has many acceptable answers.

But in the B2B world, consistency matters. Businesses using genAI will demand consistent answers to prompts like these : what is the company's revenue by region? Or how do I reset my password? Or how much would I pay if I used a 1000 units of a product? 

GenAI will need to write, create, & calculate with a significantly better error rate than humans.

I'm working with [ProductBoard to understand how different B2B startups are planning to leverage AI with a survey](https://productexcellence.typeform.com/to/uQHfdW4k). If you're integrating GenAI into your product & interested to hear others' plans, please fill it out, & we'll send you the anonymized raw data. Look for the results to be published in a few weeks.

