From the Lab:The Joule Hypothesis is our first Tech Lab Report.
Tech Lab Report001

Public research / Joule Capacity

The Joule Hypothesis

How much useful AI work can a person, an AI system, or a human-AI team produce from a constrained and auditable amount of capacity?

Status
Active
Version
1.1
Approved
Aug 09, 2026
Reading time
21 min

Tech Lab Report 001: The Joule Hypothesis

This report establishes the first approved hypothesis, accounting baseline, and access conditions for Joule Capacity. Future reports will test the hypothesis against observed experiments; revisions to the accounting or access policy must preserve the version used by historical results.

Abstract

Artificial intelligence is usually presented as a capability: a model can write, reason, create, analyze, or play. The resource required to produce that work is often hidden behind subscriptions, token limits, and provider bills. That makes it difficult for a person to understand how much AI capacity they have, what an activity consumes, or whether one result was worth more than another.

Raw token counts cannot solve that problem because tokens are not standardized across models or workload types. Currency reveals the financial constraint but also imports changing prices, contracts, purchasing power, and market strategy into the experiment. The lab therefore preserves tokens and cost as evidence while using a versioned Joule policy to allocate experimental capacity.

BK Tech Labs begins with a different premise: AI capacity should be finite, visible, and measurable. We call the unit of that capacity a Joule.

The Tech Lab hypothesis is that a shared Joule budget can make AI use more understandable and AI evaluation more meaningful. Instead of asking only what an AI system produced, the lab can also ask what capacity it used, under what conditions, and whether the result justified the expenditure.

This first report defines that hypothesis, identifies the initial experimental system, and establishes three access conditions: guests, free members, and paid members.

Research Question

How much useful AI work can a person, an AI system, or a human-AI team produce from a constrained and auditable amount of capacity?

This question is deliberately broader than cost and narrower than intelligence. Joules do not claim to measure intelligence itself. They measure the capacity made available for AI work and the amount consumed while producing an observable result.

Once the resource constraint is visible, new comparisons become possible:

  • Which approach reaches a useful result with less capacity?
  • When does additional capacity materially improve the result?
  • How repeatable is the result under the same conditions?
  • How should people allocate limited AI capacity across play, conversation, research, and creation?
  • Does making capacity visible change how people use AI?

The Joule

A Joule is the common unit of AI capacity across the Tech Lab ecosystem. Kilojoules and megajoules make larger balances easier to read:

  • 1 kJ = 1,000 J
  • 1 MJ = 1,000,000 J

In the initial system, Joules are an accounting measure derived from AI usage and provider cost. They are not a claim that the lab has directly measured the electrical energy consumed by a model, data center, or individual piece of hardware. Keeping that distinction explicit is part of the experiment.

Why raw token usage is not enough

Token counts are necessary evidence, but they are not a common unit of AI capacity.

A token is defined by a model's tokenizer rather than by a universal physical or computational standard. The same prompt can be divided into a different number of tokens by different models. Even when two requests report the same token count, they may use different models, context windows, reasoning processes, caching behavior, modalities, tools, and input-to-output ratios. Their economic cost and experimental result can therefore be materially different.

Raw token totals also make unlike activity look alike. One thousand cached input tokens, one thousand newly processed input tokens, and one thousand generated output tokens may create different provider charges. Images, audio, tool calls, routing, and other workload components may not fit cleanly into a text-token total at all. A count that ignores those distinctions is easy to reproduce but incomplete as a resource measure.

This is not only an accounting concern. Microsoft's 2026 study of AI inference estimates that optimized frontier-scale inference has a median energy use of 0.31 Wh per query, with an interquartile range of 0.16-0.60 Wh, under its production-oriented assumptions. The study also finds that long reasoning and agentic queries can require more than an order of magnitude more energy because they generate more tokens while reducing serving concurrency. Token volume matters, but model design, hardware, system utilization, serving method, and query shape determine how that volume becomes energy use. A token count alone does not preserve those conditions.

There is also a behavioral problem. A longer answer consumes more tokens, but length is not the same as usefulness, reasoning quality, or difficulty. An experiment that treats raw token count as capacity may reward brevity when detail is necessary, penalize explanation, or make a verbose but inexpensive model appear more resource-intensive than a concise but expensive one.

Tokens should therefore remain in the experimental receipt. They help explain what happened and allow later reanalysis. They should not, by themselves, determine how much cross-model capacity a person has available.

Why financial cost must be included

Every experiment that calls a paid AI service consumes a real economic resource. Ignoring that cost would make the experimental budget fictional. Two models might consume the same number of tokens while imposing very different costs on the lab; treating them as equivalent would hide whether an experiment can be repeated or offered sustainably.

Provider-reported cost supplies a practical bridge across model-specific prices and workload types. It converts the provider's distinct treatment of input, output, cached context, media, tools, and routing into an amount the lab must actually settle. This makes financial cost an essential accounting input, even though cost is not the final experimental unit.

Why currency is not the experimental unit

Using dollars directly would introduce a different set of distortions. Currency is designed for exchange, not for controlled comparison. A displayed dollar amount can anchor participants on affordability, purchasing power, or perceived value before they evaluate the result. It can encourage the false conclusion that a more expensive model is more intelligent, or that a less expensive result is necessarily more efficient.

Provider prices also reflect commercial decisions as well as technical resources. They can change with time, region, contract, volume discount, promotion, subsidy, exchange rate, or competitive strategy. Taxes and payment fees may affect what one participant pays without changing the AI work itself. An experiment expressed only in currency can therefore become difficult to compare across providers, participants, or dates.

The lab cannot remove those biases merely by renaming dollars as Joules. The initial Joule conversion is linear, so provider-price assumptions remain part of the measurement. Joules instead provide a controlled separation between three records:

RecordPurpose
Tokens and workload detailsPreserve technical usage evidence
Provider cost and price componentsPreserve the economic settlement
Joules and policy versionAllocate and compare experimental capacity

Every result should retain all three. The Joule balance gives participants a stable experimental budget without making currency the foreground of every decision. The receipt keeps the underlying tokens, provider prices, actual cost, and conversion policy available for audit. Versioning the conversion allows later reports to identify price drift, challenge the baseline, or recalculate a comparison without rewriting the historical record.

Literature and Calibration

The Joule Hypothesis begins as a policy-backed accounting experiment, but it should become increasingly grounded in observed energy use. Three external references establish the first calibration path:

ReferenceContribution to the hypothesisBoundary
Tesla's 2006 Master PlanDemonstrates how a concise public plan can connect a long-term mission to measurable efficiency, and provides an automotive scale for interpreting megajoulesIts Roadster figures are historical, company-reported automotive estimates, not measurements of AI
Microsoft's 2026 AI inference studyProvides a production-oriented estimate of inference energy per query and shows how workload and serving conditions change that energyIts estimates describe specified model, hardware, utilization, and power-usage-effectiveness assumptions; they are not a universal constant per query or token
EIA's July 2026 Electric Power Monthly, Table 5.3Provides a public reference price for electricity in the United StatesThe national retail average is not the contracted price paid by a particular AI provider or data center

The supporting unit conversions, electricity-cost estimate, automotive comparison, and methodological boundaries are documented separately in Appendix A: Energy and Cost Calibration.

What the references change

The literature strengthens the case for Joules while narrowing what the lab may claim. Energy is a real physical quantity that can connect otherwise different activities, but inference energy varies with the task and system, and its financial value varies with place, time, and contract. The lab should therefore prefer measured energy when a provider or experiment can supply it, use documented estimates when direct measurement is unavailable, and retain the present provider-cost proxy until those measurements are reliable enough to govern member balances.

Initial conversion baseline

The lab begins with a simple reference conversion:

$1 of provider-reported AI usage cost
= 1 kWh reference value
= 3.6 MJ of Joule Capacity

The energy conversion is exact: 1 kWh = 3.6 MJ. The connection to provider spend is a policy baseline chosen by the lab, not a finding that one dollar of AI usage physically consumes one kilowatt-hour of electricity.

Appendix A compares this policy baseline with the EIA retail benchmark and explains why the two should not be treated as the same price.

The initial debit formula is therefore:

Joules consumed = provider cost in dollars x 3,600,000

This baseline makes the accounting inspectable. If provider-reported usage costs $0.25, the activity consumes 900 kJ. If it costs $1, it consumes 3.6 MJ. The conversion can be revised as the lab gathers better evidence, but historical results must retain the baseline used when they were produced.

The unit gives otherwise different AI activities a shared constraint. A game move, a generated conversation, or a future experiment can each draw from the same visible capacity balance while retaining a receipt of what produced the debit.

The balance is therefore not merely a quota. It is part of the observation.

Hypothesis

Under controlled and documented conditions, recording both an AI result and the Joules consumed to produce it will yield a more useful evaluation than recording the result alone.

The hypothesis has four parts:

  1. Visibility improves judgment. People make more deliberate choices when they can see the capacity available and the cost of an action.
  2. Constraints improve comparison. Results become more informative when the resource budget is held constant or reported alongside them.
  3. One unit connects different activities. A shared Joule balance makes AI use across the lab understandable as one system rather than a collection of unrelated product limits.
  4. Capacity can support participation. A meaningful free-member allowance lets anyone with an account run a real experiment, while a larger paid-member allowance supports longer, repeated, and more ambitious work.

Experimental System

The Joule Hypothesis is designed to be tested across a connected experimental environment. Each element has a different role. These roles describe the experimental design; they do not claim that every capability is already available in every product.

Tech Lab OS: the control surface

Tech Lab OS is intended to be the common entry point. It will make available capacity, recharge, and usage visible and provide a path into the lab's applications. Its role is to help a person understand what they can do, where their Joules went, and what they learned or produced.

BarKade: the strategy chamber

BarKade uses classic games as constrained environments for observing human and artificial reasoning. The rules, positions, moves, outcomes, and capacity used can all be recorded. This makes BarKade the first practical test of whether strategic performance can be evaluated alongside the resources required to produce it.

ReChat: the conversation chamber

ReChat is designed to preserve and share thoughtful conversations with AI. It extends the experiment beyond games: can a useful line of inquiry be reproduced, examined, and shared, and how much capacity did it require? The conversation becomes an artifact instead of disappearing into a private history.

The Tech Lab Brief: the observation loop

The Tech Lab Brief connects current events, experimental results, and larger questions about AI. It explains what the lab is testing, reports what happened, and invites scrutiny. The Brief turns isolated product activity into an ongoing public investigation.

Together, these elements create a loop:

Question
-> Allocate Joules
-> Run an experiment
-> Preserve the result
-> Interpret the evidence
-> Report what was learned
-> Form the next question

Approved Initial Access Conditions

The first comparison uses three approved capacity conditions. They define the experiment; they are not a claim that every condition is already available in every BK Tech Labs product.

A guest condition is anonymous and receives 50 kJ per week without an account during the experiment. This is a bounded demonstration allocation, equal to about $0.014 of provider spend at the initial baseline. Unused guest capacity does not roll into a later week.

The persistent experimental conditions begin with free membership:

ConditionMaximum capacityInitial capacityDaily recharge
Free member5 MJ3 MJ200 kJ per day
Paid member50 MJ30 MJ2 MJ per day

The paid-member condition is a consistent tenfold increase across all three variables. Both conditions begin at 60% of maximum capacity and recharge at 4% of maximum capacity per day.

With no usage:

  • a free member moves from the initial 3 MJ to the 5 MJ maximum in 10 days;
  • a paid member moves from the initial 30 MJ to the 50 MJ maximum in 10 days; and
  • either condition takes 25 days to recharge from empty to maximum.

This symmetry is intentional. Free and paid members encounter the same capacity model at different experimental scales. A free member receives enough capacity to do real work and understand the system. A paid member can run more trials, sustain longer investigations, and compare more alternatives before waiting for capacity to recharge.

Capacity is the first measurable difference between free and paid membership. It is not, by itself, the complete meaning of paid Tech Lab membership. A durable paid membership should ultimately be judged by the quality of participation and learning it enables, not simply by how much AI it allows someone to consume.

What The Baseline Implies

The conversion makes the capacity policy and its economics transparent:

ConditionAllocationMaximum provider spend at the baseline
Guest50 kJ per weekAbout $0.014 per week
Free member3 MJ initialAbout $0.83
Free member5 MJ maximum balanceAbout $1.39 at one time
Free member200 kJ daily rechargeAbout $1.67 per 30 days if continuously consumed
Paid member30 MJ initialAbout $8.33
Paid member50 MJ maximum balanceAbout $13.89 at one time
Paid member2 MJ daily rechargeAbout $16.67 per 30 days if continuously consumed

As of August 10, 2026, paid Tech Lab membership is offered at $20 per month or $200 per year. At full, continuous use, the recurring paid-member recharge represents about $16.67 of provider spend every 30 days. That leaves only $3.33 before payment processing, infrastructure, support, and the rest of the ecosystem. Economic headroom comes mainly from recharge that reaches the capacity ceiling without being consumed, while unusually active experimenters may be subsidized by lighter users.

The one-time initial allocation changes the first-period ceiling. A paid member who continuously consumes every available Joule could use 90 MJ, or $25 of provider spend, during the first 30 days. Across a 365-day first year, the 30 MJ initial allocation plus 730 MJ of continuously consumed recharge would represent about $211.11 of provider spend, above the $200 annual membership price.

Those figures are maximum-use scenarios, not expected bills. Recharge stops at the capacity ceiling when it is not consumed, and only actual AI work creates provider spend. The design therefore depends on observed utilization. Whether that balance is sustainable and whether the resulting work justifies it are part of the Tech Lab hypothesis, not assumptions to conceal.

The lab should report at least four economic observations over time:

  • actual provider spend per active free and paid member;
  • the proportion of recharged capacity that is consumed;
  • the proportion of accounts that reach the capacity ceiling; and
  • the cost of moving a guest to free or paid membership.

Initial Method

Each experiment should preserve enough information to answer five questions:

  1. What was attempted? Record the question, task, rules, or desired result.
  2. Under what conditions? Record the relevant application, model, settings, and starting state.
  3. What happened? Preserve the moves, conversation, output, or other observable result.
  4. What did it consume? Record the token and workload details, provider prices and actual cost, Joule conversion version, and Joules charged.
  5. What was learned? Separate the observed result from the interpretation and identify the next test.

Useful comparisons should hold the task and conditions steady wherever possible. When they cannot be held steady, the differences should remain visible rather than being compressed into a single score.

The first measurements should include:

  • token usage and other workload components;
  • provider prices and actual provider cost;
  • Joules consumed and the conversion policy version;
  • whether the task completed;
  • the quality or strategic value of the result under a stated method;
  • variation across repeated trials;
  • capacity remaining; and
  • time required to recharge enough capacity for another attempt.

No single number should be allowed to stand in for the whole result. Efficiency without quality is not success, and quality without resource accounting is not a complete observation.

Predictions

If the hypothesis is useful, the lab expects to observe that:

  • people make different choices when capacity and consumption are visible;
  • comparisons made under a shared Joule budget reveal differences hidden by outcome-only rankings;
  • BarKade produces reproducible evidence about strategy and efficiency;
  • ReChat produces inspectable evidence about the development of an idea;
  • Tech Lab OS makes activity across the ecosystem feel like one coherent body of work; and
  • the Tech Lab Brief improves the experiments by exposing methods, results, and interpretations to public scrutiny.

The guest condition should demonstrate the premise. Free membership should be sufficient to produce real value. Paid membership should enable meaningfully deeper experimentation, not merely more clicks.

Failure Conditions

This is a hypothesis, not a slogan. It should be revised or rejected if the evidence shows that:

  • Joules do not help people understand or manage AI use;
  • Joule receipts obscure rather than preserve raw usage, prices, or actual provider cost;
  • the accounting method cannot compare activities without creating false equivalence;
  • changes in commercial pricing dominate a comparison without being identified and controlled;
  • the same reported result produces materially inconsistent charges without a defensible explanation;
  • larger capacity produces only more consumption rather than better experiments or learning;
  • the guest allocation is too small to demonstrate the premise;
  • the free-member allowance is too small to demonstrate value; or
  • the paid-member allowance is larger but does not enable a meaningfully different class of work or sustainable membership economics.

The lab should also resist a subtler failure: treating Joules as a measure of human worth, intelligence, or effort. They are none of those things. Joules describe the AI capacity made available and consumed within this system.

Initial Experimental Plan

The first plan is simple:

  1. Give every guest enough Joules to see the premise work without an account.
  2. Give every free member enough Joules to run a real AI experiment.
  3. Make the capacity used and the result produced visible together.
  4. Give paid members ten times the free-member capacity to repeat, extend, and compare their work.
  5. Use BarKade and ReChat to generate observable artifacts.
  6. Use Tech Lab OS to connect access, accounting, and results.
  7. Use the Tech Lab Brief to explain the method, report findings, and define the next experiment.

The purpose of the Tech Lab is not to make unlimited AI feel free. It is to learn what people and machines can accomplish when AI capacity becomes visible, comparable, and open to examination.

References

  1. Musk, Elon. The Secret Tesla Motors Master Plan (just between you and me). Tesla Motors, August 2, 2006.
  2. Oviedo, Felipe, et al. Energy use of AI inference, efficiency pathways, and test-time scaling. Joule, April 2026.
  3. U.S. Energy Information Administration. Electric Power Monthly, July 2026, Table 5.3, "Average Price of Electricity to Ultimate Customers: Total by End-Use Sector, 2016-May 2026."

Report 001 / Appendix A

Tech Lab Report 001, Appendix A: Energy and Cost Calibration

This appendix supports Tech Lab Report 001: The Joule Hypothesis. It preserves the calculations used to relate an AI inference-energy estimate, a U.S. retail electricity-price benchmark, and a historical automotive-efficiency comparison.

These calculations provide scale and a path for future calibration. They do not replace the report's approved provider-cost accounting policy.

From Inference Energy to Joules

Microsoft's 2026 study reports a median of 0.31 Wh per optimized frontier-scale query under its production-oriented assumptions. Expressed in joules, that is approximately 1.116 kJ of estimated electrical energy:

0.31 Wh x 3.6 kJ/Wh = 1.116 kJ

The study reports an interquartile range of 0.16-0.60 Wh and finds that long reasoning and agentic queries can require more than an order of magnitude more energy. The median is therefore a reference point, not a universal energy cost for a query or token.

From Electrical Energy to a Cost Reference

Table 5.3 of the EIA's July 2026 Electric Power Monthly reports a May 2026 U.S. average retail electricity price across all sectors of 13.83 cents/kWh. The 2026 figure is preliminary and covers the 50 states and the District of Columbia.

Combining that benchmark with Microsoft's median produces an illustrative electricity-cost estimate:

0.00031 kWh/query x $0.1383/kWh
= $0.0000429/query
= about 0.0043 cents/query

This is not an estimate of the provider's price to the lab. It excludes hardware, networking, engineering, redundancy, idle capacity, financing, margins, and other costs embedded in an AI service. The intended measurement chain is:

workload and system conditions
-> estimated or measured electrical energy
-> reference electricity price
-> provider cost
-> versioned Joule Capacity debit

Comparison With the Initial Policy Baseline

The report's initial policy establishes:

$1 of provider-reported AI usage cost
= 1 kWh reference value
= 3.6 MJ of Joule Capacity

At the EIA's May 2026 all-sector average of $0.1383/kWh, one dollar would purchase about 7.23 kWh, or 26.0 MJ, of retail electricity. The lab's $1 = 1 kWh policy therefore does not reproduce the retail electricity market. It allocates capacity against a provider bill that contains far more than electricity while retaining a simple physical-unit reference that can be tested and revised.

Automotive Scale Comparison

Tesla's 2006 Master Plan reports that the Roadster required 0.4 MJ/km of electricity at the outlet, or traveled 2.53 km/MJ, before its well-to-wheel adjustment. On that separate scale, Microsoft's median inference estimate of 1.116 kJ is approximately the electricity required by the reported Roadster to travel 2.8 meters:

0.001116 MJ/query x 2.53 km/MJ
= 0.0028 km/query
= about 2.8 meters/query

Using Tesla's historical well-to-wheel figure of 1.14 km/MJ reduces the comparison to about 1.3 meters.

Neither number says that an AI query and a meter of driving have equal usefulness, emissions, or lifecycle impact. The comparison only gives readers an intuitive second scale for the same energy unit. Tesla's Roadster figures are historical, company-reported estimates rather than independently measured AI data.

Methodological Boundaries

  • Microsoft's estimates depend on specified model, hardware, utilization, serving, workload, and power-usage-effectiveness assumptions.
  • The EIA rate is a national retail benchmark, not the electricity contract of Microsoft, another AI provider, or a particular data center.
  • Provider charges include costs beyond electricity and may change without a proportional change in electrical energy.
  • The Tesla comparison translates units for intuition; it is not an emissions, lifecycle, utility, or value equivalence.
  • Historical results must retain the measurement source, assumptions, and Joule policy version used when they were produced.

References

  1. Musk, Elon. The Secret Tesla Motors Master Plan (just between you and me). Tesla Motors, August 2, 2006.
  2. Oviedo, Felipe, et al. Energy use of AI inference, efficiency pathways, and test-time scaling. Joule, April 2026.
  3. U.S. Energy Information Administration. Electric Power Monthly, July 2026, Table 5.3, "Average Price of Electricity to Ultimate Customers: Total by End-Use Sector, 2016-May 2026."