Coffee Break: Armed Madhouse – The Lonely Voice of Ted Postol

Coffee Break: Armed Madhouse – The Lonely Voice of Ted Postol

For more than three decades, Theodore Postol, a physicist and professor emeritus at MIT, has been an unusually persistent critic of U.S. missile-defense programs. A former scientific adviser to the Chief of Naval Operations and analyst at the Congressional Office of Technology Assessment, Postol possesses precisely the technical credentials one might expect the defense establishment to value. Yet his repeated challenges to official claims about missile-defense performance have often placed him outside the institutional consensus rather than at its center.

Postol’s long campaign raises a question extending far beyond missile defense. The United States has an enormous defense technical establishment populated by engineers, scientists, military officers, acquisition professionals, contractors, and analysts. Major weapons programs undergo elaborate design reviews, testing, milestone decisions, audits, congressional oversight, and independent evaluation. Serious technical problems are seldom unknown. The records of troubled programs such as the F-35, Littoral Combat Ship, KC-46 tanker, and Ford-class aircraft carrier are filled with warnings, test failures, cost increases, delayed capabilities, and unresolved deficiencies.

Why, then, have so many consequential problems failed to stop or substantially redirect these programs while intervention could still have prevented enormous downstream costs? The easy explanation is corruption or incompetence. Neither is adequate. The defense establishment contains a great many highly competent people, and criticism is hardly suppressed. Engineers report deficiencies. Test organizations document failures. GAO publishes critical reports. Congress holds hearings. Journalists investigate. Outside experts object. The criticism is often technically sound and eventually vindicated.

The deeper problem is that important negative information does not adequately influence program progress. Weapons programs advance through consequential decisions: development, design maturity, production, appropriations, operational testing, acceptance, and full-rate production. These are the gates that determine a program’s fate. A program does not have to answer every criticism successfully in order to continue. It has to pass the next gate. How seriously troubled programs repeatedly manage to do so is the subject of this article.

The Remarkable Professor Postol

Ted Postol first became widely known as a missile-defense critic after the 1991 Gulf War. The Patriot missile emerged from the conflict as a technological hero. Initial Army claims credited it with extraordinarily high success rates against Iraqi Scuds, and President George H. W. Bush hailed its performance as evidence that missile defense worked. Postol was skeptical. Working with George Lewis of MIT and examining television footage of Patriot engagements, he concluded that there was little evidence Patriot had actually destroyed incoming Scud warheads. In congressional testimony in 1992, he estimated that the intercept rate could have been below 10 percent and possibly zero. A subsequent House investigation found little evidence that Patriot had successfully intercepted more than a few Scuds.

This was not criticism from an outsider unfamiliar with military technology. Postol had analyzed the MX missile for Congress’s Office of Technology Assessment and served in the Pentagon as a scientific adviser to the Chief of Naval Operations, working on ballistic missile defense, strategic weapons, countermeasures, and advanced sensors. He later joined Stanford’s Center for International Security and Arms Control before moving to MIT. By credentials and experience, Postol belonged to the technical defense establishment he was challenging.

Yet a striking feature of Postol’s subsequent career is how isolated his criticism often became. The United States possesses an enormous community of scientists and engineers capable of evaluating missile-defense claims, supplemented by national laboratories, federally funded research centers, defense contractors, intelligence organizations, universities, and government test agencies. Nevertheless, when Postol challenged important claims about missile-defense performance, remarkably few specialists from inside this establishment publicly joined him.

The pattern became particularly clear in the controversy over National Missile Defense and the ability of an interceptor to distinguish a warhead from decoys during midcourse flight. This was fundamental: in space, lightweight decoys can travel alongside a warhead without the atmospheric effects that might help distinguish them. An interceptor could therefore possess exquisite accuracy and still fail if it could not determine which object to hit.

A 1997 flight experiment collecting sensor data on a mock warhead and decoys became the center of an extended dispute involving TRW, Boeing, the Missile Defense Agency, MIT Lincoln Laboratory, and engineer Nira Schwartz, who had raised concerns about the treatment of the test data. Postol argued that the results showed simple decoys could not reliably be distinguished from the mock warhead and that analyses supporting contrary conclusions were scientifically unsound. He carried his objections to Congress and the White House and became embroiled in disputes over classified information and the adequacy of independent technical review.

The controversy was not merely an academic dispute about missile-defense technology. The National Missile Defense program evolved into the Ground-based Midcourse Defense (GMD) system, intended to defend the United States against long-range ballistic missiles. Postol’s objection went directly to a fundamental requirement of that mission: successful interception depended upon reliably discriminating the incoming warhead from accompanying decoys.

Yet the unresolved discrimination problem did not prevent deployment. In 2002, President George W. Bush directed that an initial missile-defense capability be fielded beginning in 2004. GAO subsequently reported that GMD had not been tested under unscripted, operationally realistic conditions and that its accelerated schedule left only limited opportunity to characterize performance before initial fielding. The Director of Operational Test and Evaluation had identified target discrimination as a principal concern and concluded that the existing test program was inadequate to produce credible estimates of GMD system performance. Nevertheless, GMD interceptors began entering their operational silos in 2004.

First GMD interceptor lowered into its silo at Fort Greely, Alaska, July 2004 — ready or not?

The GMD experience therefore presents two related puzzles. Why did Postol find so few technically qualified allies willing to carry the discrimination objection outside the institutions responsible for the program? And why was the system deployed before such a fundamental capability had been convincingly demonstrated?

The asymmetry confronting Postol was formidable. He was disputing technical claims made by organizations possessing vastly greater institutional resources than an individual professor. Program officials, contractors, government laboratories, advisory panels, and other credentialed specialists could contradict him. To a member of Congress or journalist unable to reproduce the technical analysis independently, the controversy could appear to pit one obstinate professor against an impressive consensus of expert authority. But counting experts is not the same thing as weighing evidence.

Many specialists defending a weapons program work for organizations that design, build, test, or manage it, or are indirectly associated with such organizations. That does not establish bad faith or mean that their technical judgments are wrong. It does mean that an apparent expert consensus cannot automatically be treated as a collection of independent observations. The institutional structure whose performance is being questioned may supply much of the expertise invoked to validate it.

The incentives facing potential dissenters were correspondingly asymmetric. An engineer could question a technical conclusion internally and allow it to proceed through ordinary review. Publicly siding with an outside critic against one’s own program, laboratory, contractor, or agency was different. It could mean challenging colleagues, supervisors, institutional commitments, and technical judgments upon which enormous expenditures already depended.

None of this requires conspiracy or conscious misrepresentation. Technical specialists may defer to managers, reviewers may conclude that deficiencies can be mitigated, contractors may believe their designs can be improved, and officials may accept remaining uncertainty. Individually, each position can be defensible. Collectively, they can create a formidable institutional consensus against the person arguing that the underlying premise is wrong.

Postol’s isolation points toward a deeper systemic failure. The defense establishment does not lack technically competent criticism. It has difficulty converting that criticism into the effective feedback necessary for sound development program management. The remarkable fact is not only that Ted Postol kept speaking. It is that neither his criticism nor the broader technical uncertainty surrounding GMD determined the decision that mattered: whether the system would be deployed.

The Silence of the Insiders: Criticism Without Control

The scarcity of Postols does not mean that major defense programs operate without internal criticism. Troubled programs routinely generate extensive records of technical deficiencies, failed tests, cost overruns, schedule delays, reliability problems, and unmet requirements. Engineers report problems. Test organizations document failures. Review teams investigate. GAO publishes critical assessments. Congress holds hearings. The puzzle is not the absence of criticism. It is the inability of criticism to control program outcomes.

Part of the explanation lies in fragmented responsibility. An engineer who identifies a serious deficiency may have neither the authority nor responsibility to decide whether the program should continue. The engineer reports the problem. A technical organization evaluates it. A contractor proposes a corrective action. Program management assesses cost and schedule effects. A review board classifies the deficiency. Senior officials determine whether the remaining risk is acceptable. Each participant performs a rational and institutionally appropriate function. Yet nowhere in this sequence must anyone answer the larger question: Does this problem mean that the program should stop?

This is a kind of fractional sanity: individually rational actions can combine to produce an irrational collective outcome. The engineer has reported the problem, the contractor proposed a solution, the review board evaluated it, and the program manager balanced technical, cost, and schedule considerations. Responsibility for the outcome is distributed across the system.

The process also produces procedural absolution. Once an objection has been documented, investigated, reviewed, and formally dispositioned, the organization can demonstrate that it took the problem seriously. The procedure becomes evidence of institutional responsibility even when it has little effect on the program’s trajectory.

Criticism and control are not the same thing. A thermostat does not regulate a furnace because it accurately reports the temperature. It regulates the furnace because its information is connected to a mechanism capable of turning the furnace off. An acquisition system can likewise contain excellent sensors—engineers, testers, auditors, GAO, congressional oversight—while possessing weak negative feedback if their findings are not coupled to decisions capable of stopping or redirecting the program.

The sheer quantity of criticism can even obscure this distinction. A troubled program may accumulate hundreds of deficiencies, each becoming a separate subject for investigation, rebuttal, mitigation, retesting, or future correction. Critics naturally assume that the cumulative weight strengthens their case. Institutionally, the opposite can occur.

What matters, therefore, is where criticism acquires leverage. A defense program must pass development approvals, design reviews, appropriations, production authorizations, operational testing, acceptance, and full-rate production. These are the key points at which criticism can halt or redirect a program.

Program advocates therefore do not need to refute every objection. They need to survive them until the next consequential decision. While critics contest numerous technical issues, the program need only secure passage through a comparatively small number of gates.

None of this requires suppressing criticism. The criticism may remain in the record, accurately stated and officially acknowledged. It simply ceases to be dispositive. The decisive question is not whether a problem has been identified. It is whether that problem has the power to keep the next gate closed.

Candidate Show-Stoppers That Did Not Stop the Show

GMD was hardly an isolated case. The same pattern appears across the U.S. defense establishment. Major programs have encountered deficiencies fundamental enough to raise the question of whether further commitment was prudent. Yet the programs continued.

Calling these deficiencies candidate show-stoppers does not mean that each should have caused cancellation. The relevant question is whether the acquisition system had determined in advance which failures were serious enough to prevent passage through the next major commitment gate.

Stopping at a gate does not necessarily mean cancelling a program. Usually it should mean something less dramatic: do not make the next major commitment until the specified problem has been corrected and the required capability demonstrated. Development can continue, redesigns can be made, and testing repeated. Cancellation becomes appropriate only when the deficiency cannot be corrected at acceptable cost or within an acceptable time.

ALT_TEXT

These examples differ technically and operationally, and some deficiencies were eventually corrected. What they share is more important: each program encountered a problem that could reasonably have raised a fundamental question about readiness to proceed, yet institutional commitment continued while the deficiency migrated downstream. That migration changes the problem. Before production, an inadequate design is principally an engineering problem. After production begins, the same deficiency becomes an engineering problem plus a retrofit, cost, scheduling, and operational problem. Once equipment has been delivered and organizations built around it, reversal becomes harder still.

GMD provides an especially revealing example because it connects the problem directly to Postol’s criticism. Discrimination between a warhead and plausible decoys was not a peripheral capability; it was fundamental to successful midcourse interception. Yet the program proceeded toward deployment without operationally realistic testing demonstrating that capability. The unresolved technical question migrated across the deployment gate. What might have been a prerequisite for fielding instead became a limitation to be investigated and improved after the system had been deployed.

The examples therefore pose a more important question than whether any particular deficiency justified cancellation: What, exactly, would have stopped these programs from proceeding to the next stage? If the answer is determined only after a deficiency appears, requirements are inherently vulnerable to reinterpretation. A capability described as essential when a program is justified can become non-essential when failure threatens advancement. A prerequisite can become a future upgrade. An unacceptable test result can become acceptable risk.

A genuine show-stopper works differently. Its consequence is established before the result is known: If X has not been demonstrated by gate Y, the program does not proceed beyond gate Y. The defense acquisition system already contains numerous points at which such a rule could operate. The question is whether those gates retain sufficient stopping power.

The Gates of Defense Program Fate

The gates discussed here are not a vague metaphor. For major weapons programs, they exist as formal milestones, technical reviews, testing decisions, production authorizations, congressional funding, and government acceptance. Each provides an opportunity to determine whether a program is sufficiently mature to justify further commitment.

The Department of Defense acquisition framework has changed repeatedly, and contemporary programs can follow different pathways. But the traditional Major Capability Acquisition pathway illustrates the basic structure: a Materiel Development Decision initiates consideration of potential solutions; Milestone A can authorize technology maturation and risk reduction; Milestone B normally begins engineering and manufacturing development; and Milestone C permits movement into production and deployment. Operational testing and the Full-Rate Production Decision provide additional opportunities to assess readiness for large-scale procurement.

ALT_TEXT

Formal milestones are not the only consequential gates. Design reviews, appropriations, contract awards, successive production lots, operational declarations, and acceptance of delivered equipment can also increase commitment. Their common characteristic is simple: each can increase the cost of saying no later.

A gate, however, is only as strong as the criteria controlling passage through it. If an unmet requirement can be waived, an unsuccessful test deferred, an immature capability accepted for later improvement, or a missing component installed after delivery, the gate remains administratively real while becoming functionally porous.

Nor are all gates equally important. The most consequential are those that substantially alter the cost of reversal. The important measure of acquisition discipline is therefore not how many reviews a program undergoes or how much documentation they produce. It is whether identifiable conditions exist under which the reviewing authority will actually refuse permission to proceed. A genuine decision gate requires a genuine stop condition.

The F-35 provides an unusually consequential demonstration of how these gates can lose their controlling function. The table below shows the key decision gates through which the program progressed.

ALT_TEXT

The F-35: How a Program Became Unstoppable

The F-35 provides an extraordinary example of what happens when program commitment overwhelms the gates intended to control it. As the most expensive weapons program in Department of Defense history, intended to produce thousands of aircraft for U.S. and allied forces, its scale made mistakes exceptionally consequential and remediation very costly.

A central F-35 program decision was concurrency: development, testing, and production substantially overlapped. In principle, concurrency promised faster fielding. In practice, aircraft entered production while testing was still discovering changes that otherwise could have been incorporated before manufacturing. Every aircraft produced before the design stabilized increased the inventory potentially requiring later modification.

Concurrency therefore altered the acquisition system’s feedback architecture. Testing traditionally supplies information before production so that deficiencies can prevent an immature design from being replicated. Under heavy concurrency, testing increasingly identified problems in aircraft already built or ordered. Potential reasons not to produce became requirements to retrofit later.

Successive low-rate production lots compounded the problem. Each added batch of aircraft increased potential remediation costs while deepening commitments by suppliers, military units, allies, and operational planners. The formal Full-Rate Production decision remained years away while the arguably premature decision to manufacture the aircraft in large numbers was being made incrementally.

F-35s at Hill Air Force base – too many produced too soon?

By the time that formal decision arrived in March 2024, more than two decades after program initiation, manufacturing had already operated at or near full rate for years and hundreds of aircraft had been delivered. A gate intended to determine readiness for large-scale production thus arrived after large-scale production was an accomplished fact.

The F-35’s Autonomic Logistics Information System (ALIS) illustrates how a potentially fundamental problem became a downstream remediation issue. ALIS supported maintenance, supply, deployment, mission planning, and other functions essential to operating the aircraft. Yet the Marine Corps declared the F-35B operational in 2015 without comprehensive deployability testing of ALIS. GAO subsequently warned of serious infrastructure and operational deficiencies. Rather than preventing the operational milestone, ALIS became a continuing remediation project. Problems persisted for years, and the Pentagon eventually began replacing the system with ODIN. A capability described as critical to F-35 operations had not stopped the program from proceeding.

Gun accuracy provides a simpler example. Operational testing reported that the internally mounted F-35A gun produced unacceptable accuracy and failed to meet its specification. Yet the failure became another deficiency to investigate and correct rather than a stop condition. One can reasonably argue that gun accuracy should not determine the fate of an aircraft whose principal capabilities lie elsewhere. But that is precisely the gating question: Was acceptable gun accuracy a mandatory capability that had to be demonstrated before a specified gate, or wasn’t it?

If failure to meet a requirement does not prevent the next commitment, the requirement has little power to constrain the program. The same question applies to ALIS and other major deficiencies. The issue is not our retrospective judgment about whether any one problem justified halting or cancelling the program. It is whether the acquisition system had established beforehand which failures would prevent further commitment.

The F-35’s history shows how a program can progressively convert uncertainty into commitment. Production precedes completed testing; testing discovers deficiencies after aircraft exist; deficiencies become retrofit and remediation programs; and successive production and operational decisions make reversal progressively more expensive. The F-35 did not become unstoppable because its problems disappeared. It became unstoppable because the program kept passing approval milestones despite persistent serious problems.

The Defense Program Protection Playbook

The F-35 reveals a broader institutional strategy. Program critics and advocates concentrate on different objectives. Critics focus on technical deficiencies because deficiencies are what their investigations uncover. Program advocates focus on approval decisions because those decisions determine whether programs continue.

This creates a powerful asymmetry. Critics assume that accumulating objections strengthens their case. But a long deficiency list can diffuse opposition. Each problem generates its own point-counterpoint: a component is being redesigned; software will improve; reliability is trending upward; another test is scheduled. Numerous objections can thus be reframed as the ordinary problems of a complex development program. Meanwhile, the next program approval date approaches.

Program advocates need not prove every criticism wrong. They need only establish that the remaining problems do not justify withholding the next approval. The process follows a recurring pattern:

Fragment the criticism. Treat potentially fundamental objections as separate deficiencies requiring individual disposition.

Process the objections. Investigate, document, mitigate, and schedule corrective actions. Criticism is absorbed rather than suppressed.

Run out the clock. Technical debate continues while the next program approval decision approaches.

Soften the criteria. A mandatory capability becomes a remediation objective; demonstrated performance becomes expected performance; failure becomes accepted risk for a planned fix.

Pass the gate. Additional money is spent, contracts executed, equipment produced, and organizations committed.

Move the problems downstream. Deficiencies that might have delayed approval become problems to correct in the program already approved.

Then this cycle begins again from a stronger position of program momentum.

This strategy has emerged from institutional incentives. Program offices exist to deliver programs. Contractors fulfill contracts. Military services plan around anticipated capabilities. Congressional districts acquire jobs, and allied governments make commitments. Each participant has rational reasons to keep moving.

Each approval gate crossed therefore changes the next contest. More money becomes sunk cost, more organizations depend upon the program, and more alternatives are foreclosed. The acquisition process develops a commitment ratchet: every successive gate becomes easier to cross.

This explains how highly critical test findings, GAO reports, congressional hearings, and engineering assessments can coexist with continued program growth. The criticism may be entirely accurate, but it may not control the decision that matters. The lesson for reformers is counterintuitive: effective criticism is not principally a problem of quantity. It is a problem of leverage. A hundred valid objections dispersed across a program may matter less than a single failed criterion that must be satisfied before the next gate can open.

Restoring Effective Defense Program Approval Practices

If program sustainment depends upon passing successive decision gates, the reform strategy is straightforward: make the gates harder to pass when important requirements have not been met. This requires no new layer of defense bureaucracy. The acquisition system already contains milestone reviews, design reviews, production decisions, operational testing, appropriations, and acceptance decisions. What it needs are stronger stop conditions at the most consequential gates.

Congress has the power to add rigor to these decision points. Through authorization and appropriations legislation, it can condition advancement or expenditure on specified findings, require independent testing before major production commitments, and prescribe the authority required to waive mandatory criteria. Congress need not adjudicate technical performance itself; it can strengthen the rules governing the decisions at which commitment becomes difficult to reverse.

For major programs, a small number of show-stopper criteria should therefore be established before the relevant decision. They should be objectively verifiable, tied explicitly to a gate, and determined before anyone knows whether the program will satisfy them: If criterion X has not been demonstrated, the program does not pass gate Y.

Critics should then concentrate on defending these criteria rather than maximizing the number of deficiencies they identify. A show-stopper must remain a show-stopper after the program fails it. A mandatory requirement should not quietly become an objective, a required demonstration become projected performance, or a failed test automatically become downstream remediation.

Hard approval criteria need not be absolute. Exceptional circumstances may justify proceeding despite a failed criterion, but the burden should fall heavily on those seeking the exception. They should identify the failure, explain why advancement is necessary, assess the technical and operational risks and downstream remediation costs, and obtain independent technical review. For the most consequential gates, Congress could require notification or approval before an exception takes effect.

Most importantly, the exception decision should have names attached to it. The public record should identify the officials requesting and authorizing an exception, together with the date and justification. Classified performance details can remain classified; secrecy about weapons capabilities does not require secrecy about the decision architecture. The purpose is not to punish officials whose reasonable judgments later prove wrong. It is to prevent responsibility from dissolving into committees and procedures and to force fundamental problems to be corrected while commitments remain reversible.

The proposed reform strategy requires neither more reports nor more critics. It requires fewer, harder criteria attached to decisions that matter. The acquisition system already has gates. The task is to restore their intended controlling function.

Conclusion

Ted Postol’s career is a cautionary tale, but not primarily because he may have been right about missile-defense systems. A defense establishment containing enormous technical expertise should not depend upon exceptional individuals willing to fight prolonged public battles to give serious technical objections institutional attention.

The problem is not a shortage of criticism. Troubled weapons programs accumulate warnings, failed tests, audits, engineering deficiencies, and congressional investigations. Defense critics have made many good plays. They have exposed problems, challenged official claims, and often been vindicated by subsequent events. But the establishment has been winning the government approval games and accumulating a record of costly mismanaged programs.

The score that matters is not the number of deficiencies identified or arguments won. It is whether a troubled program receives the next development approval, appropriation, production authorization, operational declaration, or acceptance decision. Program advocates need not defeat every criticism. They just need to pass the current approval gate. Once they do, additional commitments make the next gate harder to close.

Defense program critics need to adopt a corresponding strategy. Instead of trying to address every deficiency, critics should concentrate on the critical requirements that must be satisfied before the next consequential commitment. The battle is not over how many things are wrong with a program. It is over which things are sufficiently important that the program cannot proceed until they are put right.

Congress can reinforce that strategy by strengthening the existing gates: requiring independent verification of critical criteria, imposing a heavy burden of justification on exceptions, and making responsibility for overrides explicit. The objective is not more oversight for its own sake. It is to reconnect negative technical feedback to decisions that control program outcomes.

The defense establishment does not need more Ted Postols fighting lonely battles after programs have acquired enormous momentum. It needs an acquisition system in which technically valid negative feedback has leverage before programs become unstoppable. The current dysfunctional defense program sustainment strategy is to pass through approval gates by weakening or circumventing their criteria. The appropriate reform strategy is to strengthen those gates and restore their necessary controlling function.

Ted Postol’s long campaign may not have achieved the results he sought, but his remarkable isolation exposes a deeper problem: a defense establishment capable of generating abundant internal technical criticism while too often disregarding that criticism in a manner detrimental to U.S. national interests.

 

Print Friendly, PDF & Email

Leave a Reply

Your email address will not be published. Required fields are marked *