Tech Lab Report 001: The Joule Hypothesis
This report establishes the first approved hypothesis, accounting baseline, and access conditions for Joule Capacity. Future reports will test the hypothesis against observed experiments; revisions to the accounting or access policy must preserve the version used by historical results.
Abstract
Artificial intelligence is usually presented as a capability: a model can write, reason, create, analyze, or play. The resource required to produce that work is often hidden behind subscriptions, token limits, and provider bills. That makes it difficult for a person to understand how much AI capacity they have, what an activity consumes, or whether one result was worth more than another.
Raw token counts cannot solve that problem because tokens are not standardized across models or workload types. Currency reveals the financial constraint but also imports changing prices, contracts, purchasing power, and market strategy into the experiment. The lab therefore preserves tokens and cost as evidence while using a versioned Joule policy to allocate experimental capacity.
BK Tech Labs begins with a different premise: AI capacity should be finite, visible, and measurable. We call the unit of that capacity a Joule.
The Tech Lab hypothesis is that a shared Joule budget can make AI use more understandable and AI evaluation more meaningful. Instead of asking only what an AI system produced, the lab can also ask what capacity it used, under what conditions, and whether the result justified the expenditure.
This first report defines that hypothesis, identifies the initial experimental system, and establishes three access conditions: guests, free members, and paid members.
Research Question
How much useful AI work can a person, an AI system, or a human-AI team produce from a constrained and auditable amount of capacity?
This question is deliberately broader than cost and narrower than intelligence. Joules do not claim to measure intelligence itself. They measure the capacity made available for AI work and the amount consumed while producing an observable result.
Once the resource constraint is visible, new comparisons become possible:
- Which approach reaches a useful result with less capacity?
- When does additional capacity materially improve the result?
- How repeatable is the result under the same conditions?
- How should people allocate limited AI capacity across play, conversation, research, and creation?
- Does making capacity visible change how people use AI?
The Joule
A Joule is the common unit of AI capacity across the Tech Lab ecosystem. Kilojoules and megajoules make larger balances easier to read:
1 kJ = 1,000 J1 MJ = 1,000,000 J
In the initial system, Joules are an accounting measure derived from AI usage and provider cost. They are not a claim that the lab has directly measured the electrical energy consumed by a model, data center, or individual piece of hardware. Keeping that distinction explicit is part of the experiment.
Why raw token usage is not enough
Token counts are necessary evidence, but they are not a common unit of AI capacity.
A token is defined by a model's tokenizer rather than by a universal physical or computational standard. The same prompt can be divided into a different number of tokens by different models. Even when two requests report the same token count, they may use different models, context windows, reasoning processes, caching behavior, modalities, tools, and input-to-output ratios. Their economic cost and experimental result can therefore be materially different.
Raw token totals also make unlike activity look alike. One thousand cached input tokens, one thousand newly processed input tokens, and one thousand generated output tokens may create different provider charges. Images, audio, tool calls, routing, and other workload components may not fit cleanly into a text-token total at all. A count that ignores those distinctions is easy to reproduce but incomplete as a resource measure.
This is not only an accounting concern. Microsoft's 2026 study of AI inference
estimates that optimized frontier-scale inference has a median energy use of
0.31 Wh per query, with an interquartile range of 0.16-0.60 Wh, under its
production-oriented assumptions. The study also finds that long reasoning and
agentic queries can require more than an order of magnitude more energy because
they generate more tokens while reducing serving concurrency. Token volume
matters, but model design, hardware, system utilization, serving method, and
query shape determine how that volume becomes energy use. A token count alone
does not preserve those conditions.
There is also a behavioral problem. A longer answer consumes more tokens, but length is not the same as usefulness, reasoning quality, or difficulty. An experiment that treats raw token count as capacity may reward brevity when detail is necessary, penalize explanation, or make a verbose but inexpensive model appear more resource-intensive than a concise but expensive one.
Tokens should therefore remain in the experimental receipt. They help explain what happened and allow later reanalysis. They should not, by themselves, determine how much cross-model capacity a person has available.
Why financial cost must be included
Every experiment that calls a paid AI service consumes a real economic resource. Ignoring that cost would make the experimental budget fictional. Two models might consume the same number of tokens while imposing very different costs on the lab; treating them as equivalent would hide whether an experiment can be repeated or offered sustainably.
Provider-reported cost supplies a practical bridge across model-specific prices and workload types. It converts the provider's distinct treatment of input, output, cached context, media, tools, and routing into an amount the lab must actually settle. This makes financial cost an essential accounting input, even though cost is not the final experimental unit.
Why currency is not the experimental unit
Using dollars directly would introduce a different set of distortions. Currency is designed for exchange, not for controlled comparison. A displayed dollar amount can anchor participants on affordability, purchasing power, or perceived value before they evaluate the result. It can encourage the false conclusion that a more expensive model is more intelligent, or that a less expensive result is necessarily more efficient.
Provider prices also reflect commercial decisions as well as technical resources. They can change with time, region, contract, volume discount, promotion, subsidy, exchange rate, or competitive strategy. Taxes and payment fees may affect what one participant pays without changing the AI work itself. An experiment expressed only in currency can therefore become difficult to compare across providers, participants, or dates.
The lab cannot remove those biases merely by renaming dollars as Joules. The initial Joule conversion is linear, so provider-price assumptions remain part of the measurement. Joules instead provide a controlled separation between three records:
| Record | Purpose |
|---|---|
| Tokens and workload details | Preserve technical usage evidence |
| Provider cost and price components | Preserve the economic settlement |
| Joules and policy version | Allocate and compare experimental capacity |
Every result should retain all three. The Joule balance gives participants a stable experimental budget without making currency the foreground of every decision. The receipt keeps the underlying tokens, provider prices, actual cost, and conversion policy available for audit. Versioning the conversion allows later reports to identify price drift, challenge the baseline, or recalculate a comparison without rewriting the historical record.
Literature and Calibration
The Joule Hypothesis begins as a policy-backed accounting experiment, but it should become increasingly grounded in observed energy use. Three external references establish the first calibration path:
| Reference | Contribution to the hypothesis | Boundary |
|---|---|---|
| Tesla's 2006 Master Plan | Demonstrates how a concise public plan can connect a long-term mission to measurable efficiency, and provides an automotive scale for interpreting megajoules | Its Roadster figures are historical, company-reported automotive estimates, not measurements of AI |
| Microsoft's 2026 AI inference study | Provides a production-oriented estimate of inference energy per query and shows how workload and serving conditions change that energy | Its estimates describe specified model, hardware, utilization, and power-usage-effectiveness assumptions; they are not a universal constant per query or token |
| EIA's July 2026 Electric Power Monthly, Table 5.3 | Provides a public reference price for electricity in the United States | The national retail average is not the contracted price paid by a particular AI provider or data center |
The supporting unit conversions, electricity-cost estimate, automotive comparison, and methodological boundaries are documented separately in Appendix A: Energy and Cost Calibration.
What the references change
The literature strengthens the case for Joules while narrowing what the lab may claim. Energy is a real physical quantity that can connect otherwise different activities, but inference energy varies with the task and system, and its financial value varies with place, time, and contract. The lab should therefore prefer measured energy when a provider or experiment can supply it, use documented estimates when direct measurement is unavailable, and retain the present provider-cost proxy until those measurements are reliable enough to govern member balances.
Initial conversion baseline
The lab begins with a simple reference conversion:
$1 of provider-reported AI usage cost
= 1 kWh reference value
= 3.6 MJ of Joule Capacity
The energy conversion is exact: 1 kWh = 3.6 MJ. The connection to provider
spend is a policy baseline chosen by the lab, not a finding that one dollar of
AI usage physically consumes one kilowatt-hour of electricity.
Appendix A compares this policy baseline with the EIA retail benchmark and explains why the two should not be treated as the same price.
The initial debit formula is therefore:
Joules consumed = provider cost in dollars x 3,600,000
This baseline makes the accounting inspectable. If provider-reported usage
costs $0.25, the activity consumes 900 kJ. If it costs $1, it consumes
3.6 MJ. The conversion can be revised as the lab gathers better evidence,
but historical results must retain the baseline used when they were produced.
The unit gives otherwise different AI activities a shared constraint. A game move, a generated conversation, or a future experiment can each draw from the same visible capacity balance while retaining a receipt of what produced the debit.
The balance is therefore not merely a quota. It is part of the observation.
Hypothesis
Under controlled and documented conditions, recording both an AI result and the Joules consumed to produce it will yield a more useful evaluation than recording the result alone.
The hypothesis has four parts:
- Visibility improves judgment. People make more deliberate choices when they can see the capacity available and the cost of an action.
- Constraints improve comparison. Results become more informative when the resource budget is held constant or reported alongside them.
- One unit connects different activities. A shared Joule balance makes AI use across the lab understandable as one system rather than a collection of unrelated product limits.
- Capacity can support participation. A meaningful free-member allowance lets anyone with an account run a real experiment, while a larger paid-member allowance supports longer, repeated, and more ambitious work.
Experimental System
The Joule Hypothesis is designed to be tested across a connected experimental environment. Each element has a different role. These roles describe the experimental design; they do not claim that every capability is already available in every product.
Tech Lab OS: the control surface
Tech Lab OS is intended to be the common entry point. It will make available capacity, recharge, and usage visible and provide a path into the lab's applications. Its role is to help a person understand what they can do, where their Joules went, and what they learned or produced.
BarKade: the strategy chamber
BarKade uses classic games as constrained environments for observing human and artificial reasoning. The rules, positions, moves, outcomes, and capacity used can all be recorded. This makes BarKade the first practical test of whether strategic performance can be evaluated alongside the resources required to produce it.
ReChat: the conversation chamber
ReChat is designed to preserve and share thoughtful conversations with AI. It extends the experiment beyond games: can a useful line of inquiry be reproduced, examined, and shared, and how much capacity did it require? The conversation becomes an artifact instead of disappearing into a private history.
The Tech Lab Brief: the observation loop
The Tech Lab Brief connects current events, experimental results, and larger questions about AI. It explains what the lab is testing, reports what happened, and invites scrutiny. The Brief turns isolated product activity into an ongoing public investigation.
Together, these elements create a loop:
Question
-> Allocate Joules
-> Run an experiment
-> Preserve the result
-> Interpret the evidence
-> Report what was learned
-> Form the next question
Approved Initial Access Conditions
The first comparison uses three approved capacity conditions. They define the experiment; they are not a claim that every condition is already available in every BK Tech Labs product.
A guest condition is anonymous and receives 50 kJ per week without an
account during the experiment. This is a bounded demonstration allocation,
equal to about $0.014 of provider spend at the initial baseline. Unused guest
capacity does not roll into a later week.
The persistent experimental conditions begin with free membership:
| Condition | Maximum capacity | Initial capacity | Daily recharge |
|---|---|---|---|
| Free member | 5 MJ | 3 MJ | 200 kJ per day |
| Paid member | 50 MJ | 30 MJ | 2 MJ per day |
The paid-member condition is a consistent tenfold increase across all three variables. Both conditions begin at 60% of maximum capacity and recharge at 4% of maximum capacity per day.
With no usage:
- a free member moves from the initial 3 MJ to the 5 MJ maximum in 10 days;
- a paid member moves from the initial 30 MJ to the 50 MJ maximum in 10 days; and
- either condition takes 25 days to recharge from empty to maximum.
This symmetry is intentional. Free and paid members encounter the same capacity model at different experimental scales. A free member receives enough capacity to do real work and understand the system. A paid member can run more trials, sustain longer investigations, and compare more alternatives before waiting for capacity to recharge.
Capacity is the first measurable difference between free and paid membership. It is not, by itself, the complete meaning of paid Tech Lab membership. A durable paid membership should ultimately be judged by the quality of participation and learning it enables, not simply by how much AI it allows someone to consume.
What The Baseline Implies
The conversion makes the capacity policy and its economics transparent:
| Condition | Allocation | Maximum provider spend at the baseline |
|---|---|---|
| Guest | 50 kJ per week | About $0.014 per week |
| Free member | 3 MJ initial | About $0.83 |
| Free member | 5 MJ maximum balance | About $1.39 at one time |
| Free member | 200 kJ daily recharge | About $1.67 per 30 days if continuously consumed |
| Paid member | 30 MJ initial | About $8.33 |
| Paid member | 50 MJ maximum balance | About $13.89 at one time |
| Paid member | 2 MJ daily recharge | About $16.67 per 30 days if continuously consumed |
As of August 10, 2026, paid Tech Lab membership is offered at
$20 per month or $200 per year. At full,
continuous use, the recurring paid-member recharge represents about $16.67
of provider spend every 30 days. That leaves only $3.33 before payment
processing, infrastructure, support, and the rest of the ecosystem. Economic
headroom comes mainly from recharge that reaches the capacity ceiling without
being consumed, while unusually active experimenters may be subsidized by
lighter users.
The one-time initial allocation changes the first-period ceiling. A paid member
who continuously consumes every available Joule could use 90 MJ, or $25 of
provider spend, during the first 30 days. Across a 365-day first year, the
30 MJ initial allocation plus 730 MJ of continuously consumed recharge
would represent about $211.11 of provider spend, above the $200 annual
membership price.
Those figures are maximum-use scenarios, not expected bills. Recharge stops at the capacity ceiling when it is not consumed, and only actual AI work creates provider spend. The design therefore depends on observed utilization. Whether that balance is sustainable and whether the resulting work justifies it are part of the Tech Lab hypothesis, not assumptions to conceal.
The lab should report at least four economic observations over time:
- actual provider spend per active free and paid member;
- the proportion of recharged capacity that is consumed;
- the proportion of accounts that reach the capacity ceiling; and
- the cost of moving a guest to free or paid membership.
Initial Method
Each experiment should preserve enough information to answer five questions:
- What was attempted? Record the question, task, rules, or desired result.
- Under what conditions? Record the relevant application, model, settings, and starting state.
- What happened? Preserve the moves, conversation, output, or other observable result.
- What did it consume? Record the token and workload details, provider prices and actual cost, Joule conversion version, and Joules charged.
- What was learned? Separate the observed result from the interpretation and identify the next test.
Useful comparisons should hold the task and conditions steady wherever possible. When they cannot be held steady, the differences should remain visible rather than being compressed into a single score.
The first measurements should include:
- token usage and other workload components;
- provider prices and actual provider cost;
- Joules consumed and the conversion policy version;
- whether the task completed;
- the quality or strategic value of the result under a stated method;
- variation across repeated trials;
- capacity remaining; and
- time required to recharge enough capacity for another attempt.
No single number should be allowed to stand in for the whole result. Efficiency without quality is not success, and quality without resource accounting is not a complete observation.
Predictions
If the hypothesis is useful, the lab expects to observe that:
- people make different choices when capacity and consumption are visible;
- comparisons made under a shared Joule budget reveal differences hidden by outcome-only rankings;
- BarKade produces reproducible evidence about strategy and efficiency;
- ReChat produces inspectable evidence about the development of an idea;
- Tech Lab OS makes activity across the ecosystem feel like one coherent body of work; and
- the Tech Lab Brief improves the experiments by exposing methods, results, and interpretations to public scrutiny.
The guest condition should demonstrate the premise. Free membership should be sufficient to produce real value. Paid membership should enable meaningfully deeper experimentation, not merely more clicks.
Failure Conditions
This is a hypothesis, not a slogan. It should be revised or rejected if the evidence shows that:
- Joules do not help people understand or manage AI use;
- Joule receipts obscure rather than preserve raw usage, prices, or actual provider cost;
- the accounting method cannot compare activities without creating false equivalence;
- changes in commercial pricing dominate a comparison without being identified and controlled;
- the same reported result produces materially inconsistent charges without a defensible explanation;
- larger capacity produces only more consumption rather than better experiments or learning;
- the guest allocation is too small to demonstrate the premise;
- the free-member allowance is too small to demonstrate value; or
- the paid-member allowance is larger but does not enable a meaningfully different class of work or sustainable membership economics.
The lab should also resist a subtler failure: treating Joules as a measure of human worth, intelligence, or effort. They are none of those things. Joules describe the AI capacity made available and consumed within this system.
Initial Experimental Plan
The first plan is simple:
- Give every guest enough Joules to see the premise work without an account.
- Give every free member enough Joules to run a real AI experiment.
- Make the capacity used and the result produced visible together.
- Give paid members ten times the free-member capacity to repeat, extend, and compare their work.
- Use BarKade and ReChat to generate observable artifacts.
- Use Tech Lab OS to connect access, accounting, and results.
- Use the Tech Lab Brief to explain the method, report findings, and define the next experiment.
The purpose of the Tech Lab is not to make unlimited AI feel free. It is to learn what people and machines can accomplish when AI capacity becomes visible, comparable, and open to examination.
References
- Musk, Elon. The Secret Tesla Motors Master Plan (just between you and me). Tesla Motors, August 2, 2006.
- Oviedo, Felipe, et al. Energy use of AI inference, efficiency pathways, and test-time scaling. Joule, April 2026.
- U.S. Energy Information Administration. Electric Power Monthly, July 2026, Table 5.3, "Average Price of Electricity to Ultimate Customers: Total by End-Use Sector, 2016-May 2026."
