Level 4 — Upper Intermediate · Argument & Evidence Basics

Lesson 30: Evidence, Claims, Data & Interpretation

How to separate data from observation, observation from interpretation, and interpretation from inference — so you can judge whether a study's evidence actually matches its claim, its sample supports its scope, and its conclusion is genuinely earned.

upper-intermediate 95 min evidence-evaluation data-and-interpretation research-reading scope-and-generalization

Learning Objectives

By the end of this lesson, you should be able to:

  • Separate data/observation, interpretation, inference, and conclusion as four distinct layers, rather than reading them as one undifferentiated claim
  • Judge a piece of evidence on three questions — is it relevant, is it sufficient, and is it reliable — rather than accepting it because it exists
  • Match a study's claimed scope to its actual sample, catching generalizations that outrun the population actually studied
  • Read the support-language family precisely — consistent with, supports, contradicts, calls into question — each signalling a different evidential weight
  • Tell 'absence of evidence' (we didn't find support) apart from 'evidence of absence' (we found support for nothing being there) — one of the most consequential distinctions in research reading

Introduction

Lesson 26 built claim → reason → evidence. Lesson 29 built cause → evidence → alternative explanation → qualification. Today's lesson takes evidence itself apart, layer by layer: data, observation, interpretation, inference, claim, conclusion. English nonfiction rarely announces "this is my opinion" outright — instead, an author moves quietly from data to pattern to interpretation to inference to conclusion, and a careless reader absorbs all of it as one seamless fact. Learning to separate these layers is what lets you judge a research paper, a business report, or a news article on its actual merits, rather than on how confidently it's written.

What Is Evidence?

Information used to support a claim

Evidence can take many forms: numerical data, experimental results, survey results, observations, historical records, documents, quotations, expert analysis, case studies, multiple studies.

Employee productivity increased by 18% after training. — this is a data point/result. Only once an author writes This suggests that training improved productivity does that number start functioning as evidence for a claim.

Data Is Not Automatically Evidence

Collected information only becomes evidence once it's used to support something

Data = collected information. Evidence = information used to support a claim.

500 employees participated in the study. — this is data. If the claim is "training improved productivity," the bare fact that 500 employees participated doesn't, on its own, support that claim at all — participation numbers and productivity outcomes are two different pieces of information.

Observation, Interpretation, Inference, and Conclusion

Four layers, easy to blur into one

Observation: Processing time decreased from 20 minutes to 15 minutes. — a directly observed change.

Interpretation: The new system appears to have improved efficiency. — the meaning the author assigns to that observation. Notice appears to signalling the interpretation's own uncertainty.

Inference: Several independent studies found similar improvements (evidence) → the intervention is likely to be effective (inference — a logically supported conclusion, built from the evidence).

Conclusion: Overall, the evidence supports continued use of the intervention. — the argument's final position.

The Full Chain

Every arrow in this chain needs its own justification

DATA/OBSERVATION → PATTERN → INTERPRETATION → INFERENCE → CLAIM → CONCLUSION.

None of these arrows are automatic. A careful reader checks the justification behind each step individually, rather than accepting the whole chain because its first link (the data) looked solid.

Worked Example: Why Timing Alone Isn't Enough

Sales increased by 25% after the company introduced the new website

Data: sales increased by 25%. Temporal relationship: the increase happened after the website launch. Possible interpretation: the website may have contributed. Stronger claim: the website caused the increase — but this last claim needs additional evidence the passage hasn't yet supplied.

This is before → after proving only that A happened before B — not that A caused B (Lesson 29's central warning, worth repeating here: advertising campaigns, price reductions, seasonal demand, a competitor's failure, a new product, or a broader economic change could all explain the same 25% increase just as well).

Evidence Must Match the Claim

Evidence can be real and still support the wrong claim

Claim: The software is faster. Evidence offered: Employees say they like the software.

This evidence doesn't match the claim. Employees liking it is satisfaction evidence, not speed evidence — and a text that slides from one to the other without noticing is making an unearned leap.

Three Questions for Any Piece of Evidence

Testing a piece of evidence before accepting it

1Is the evidence relevant to this specific claim?
2Is the evidence sufficient to support a claim this size?
3Is the evidence reliable — trustworthy in how it was produced?

Relevant but insufficient, or reliable but irrelevant

Relevant but insufficient: One employee reports that the software saved time. Relevant: yes. Sufficient to prove a company-wide effect? No — one person's experience is not the whole organization.

Reliable but irrelevant: A highly reliable survey shows employees prefer blue interfaces, offered as evidence for the claim blue interfaces increase productivity. The survey itself may be reliable, but its relevance to a productivity claim is weak — preference and productivity are not the same thing.

Anecdote and Case Study

Two useful but limited forms of evidence

Anecdote — an individual story or example: One manager reported that automation dramatically improved his team's performance. Useful? Yes. Strong general evidence? Usually not.

Case study — a specific organization, person, or event, examined in depth. Useful for understanding mechanism and context in detail, but its findings' generalizability may be limited beyond that one case.

Sample Size and Representativeness

Bigger is not automatically better evidence

A study of 5,000 employees found… vs. a study of 8 employees found… — generally, a larger sample can support stronger generalizability. But a large sample does not automatically mean perfect evidence: study design, sampling method, and measurement quality all matter just as much.

Representative sample: if the population is 1 million workers and the study covers 10,000, the key question is are those 10,000 representative? If yes, findings may generalize reasonably well. If no, even a large sample can carry real bias.

Selection Bias and the Self-Report Problem

Who was asked, and how they answered, both shape the evidence

Selection bias: A company surveys only its highest-performing employees and finds high satisfaction. Can this generalize to all employees? Probably not — the sample was chosen in a way that skews the result before any analysis even begins.

Self-reported data: Employees reported that they were more productive. Reported productivity and objectively measured productivity are different things — self-reports can be shaped by memory error, perception, social desirability, or expectation, even when they're offered honestly.

Objective measurement: Average processing time fell from 12 minutes to 8 minutes is more directly measurable — but why it fell still requires additional reasoning, not just the measurement itself.

Correlation, Reverse Causation, and Bidirectional Relationships

An association can run in more than one direction at once

Employees who received more training had higher productivity. Association: training ↔ productivity. Possible explanations: training caused productivity; already-productive employees received more training; experienced employees received more training; some other factor affected both.

Reverse causation: organizations might train employees because they're already high performers — meaning productivity → training, not only training → productivity.

Bidirectional relationship: Better communication improves teamwork, while stronger teamwork also improves communication — a genuine feedback relationship, where X influences Y and Y influences X at the same time. (This builds directly on Lesson 24's and 29's causal-direction skills.)

An Evidence Hierarchy

Tier Typical evidence
Stronger systematic reviews · multiple independent studies · replicated experiments · large, high-quality datasets · well-designed controlled studies
Moderate/contextual individual studies · longitudinal studies · observational research · expert analysis · case studies
Weaker for generalization anecdotes · personal experience · isolated examples · unsupported assertions

This is a reading guide, not an absolute ranking — the research question and study design still matter, echoing Lesson 26's evidence hierarchy from a research-specific angle.

Framing the Data: According To, The Data Show, Suggest, Demonstrate

Four common ways to introduce a finding, at four different strengths

According to a 2025 survey, 62% of employees preferred hybrid work. — source: the survey; ask who says this?

The data show that processing time declined. — ask: which data? how were they collected? over what period? which population?

The findings suggest that training improved performance.suggest signals interpretation, offered with caution.

The results demonstrate a clear relationship between the variables. — a stronger claim; ask whether the results actually justify demonstrate, rather than trusting the reporting verb on its own.

"Consistent With" and "Supports"

Fitting the evidence is not the same as proving the claim

Consistent with: The findings are consistent with the hypothesis. — the findings don't contradict the hypothesis; but consistent with is not proves — other explanations may fit the same evidence equally well.

Supports: The evidence supports the conclusion that training improved performance. — the evidence increases the justification for the conclusion; again, support is not absolute proof.

Strongly supports: The evidence strongly supports the hypothesis — carries greater evidential weight than a bare supports.

Weaker and Negative Support Language

A whole family of phrases for evidence that doesn't fully back a claim

Provides little support for: The findings provide little support for the proposed explanation. — weak support; not necessarily proof the explanation is false.

Provides no support for: The evidence provides no support for the claim. — doesn't help justify the claim; again, not proof of the opposite.

Contradicts: The new evidence contradicts the earlier explanation. — stronger: evidence actively points against a previous claim.

Challenges: The findings challenge the assumption that automation always reduces costs. — the assumption is put under real pressure, not necessarily fully disproved.

Calls into question: The results call into question the reliability of the earlier study. — creates serious doubt.

Raises questions about: The findings raise questions about the policy's effectiveness. — a weaker cousin of calls into question: increased uncertainty, not yet a strong challenge.

"Fails To Demonstrate" and "Fails To Find"

A study's failure to detect something is not proof that it isn't there

The study fails to demonstrate a causal relationship. — the study didn't establish causation; this does not automatically mean no causal relationship exists.

The study failed to find a significant relationship. — the study didn't detect one; this doesn't necessarily mean the relationship doesn't exist. Possible reasons: sample size, measurement, study design, or a true absence — any of these could explain a failed detection.

"No Significant Difference" — A Technical Trap

'Significant' can carry a precise statistical meaning

The study found no significant difference between the groups.

In research contexts, significant often has a specific statistical meaning. Don't automatically read this as no difference whatsoever — it more precisely means no statistically significant difference was detected, which leaves room for a real, if small or undetected, difference to still exist.

Evidence of Absence vs. Absence of Evidence

One of the most consequential distinctions in research reading

Absence of evidence: We did not find evidence that X occurred. — evidence simply wasn't found.

Evidence of absence: The data provide evidence that X did not occur. — a stronger claim: the evidence actively supports X's non-occurrence.

Never treat these as identical. The investigation found no evidence of fraud does not necessarily mean fraud definitely did not occur — it means the investigation, as conducted, didn't turn up supporting evidence. That gap between the two readings can matter enormously, in journalism as much as in research.

Generalization and Scope Matching

A study's scope and a claim's scope have to actually match

A study of 50 software developers found that AI tools improved coding speed. Can we say "AI tools improve everyone's productivity"? No — scope mismatch: evidence covers 50 software developers; the claim reaches for all workers worldwide. A better-matched version: Among the developers studied, AI tools were associated with higher coding speed.

Scope Markers

Among, in this sample, under these conditions, in similar settings

Among: Among experienced employees, the intervention was effective. — don't generalize automatically to inexperienced employees.

In this sample: In this sample, productivity increased after training. — scope explicitly limited to the group studied.

Under these conditions: Under these conditions, automation reduced processing time. — don't assume automation always reduces processing time everywhere.

In similar settings: The findings may apply to organizations operating in similar settings. — again, limited, conditional generalization.

External Validity, and Internal vs. External Reasoning

Two separate questions a study's findings raise

External validity — how far a study's findings can reasonably generalize to other populations or contexts. A study conducted in a large German manufacturing company doesn't automatically apply to a small logistics company in Bangladesh — context differences matter.

Internal question: did the study establish the relationship within its own design? External question: can we generalize the result beyond the study? A study can be internally strong while still being externally limited — these are genuinely separate judgments.

Triangulation and Replication

Two ways evidence can become more convincing over time

Triangulation — when multiple kinds of evidence (survey, interview, transaction data, observational data) all point toward the same conclusion, confidence may reasonably increase.

Replication — when independent studies repeatedly find similar results (Study A → same result, Study B → same result, Study C → same result), this is far more convincing than one isolated study.

Across studies: Across several studies, the intervention was associated with improved performance — signals a broader evidence base than a single trial. Consistently: Results were consistently positive across multiple trials — strengthens confidence, though consistent findings still aren't automatic proof of causation.

Mixed and Conflicting Evidence

Don't quietly round 'mixed' up to 'supports'

Mixed evidence: The evidence is mixed means some evidence supports, some does not — never summarize this as simply "the evidence supports."

Conflicting evidence: Studies have produced conflicting evidence means the results genuinely disagree — an author reporting this will often go on to discuss methodological differences, different populations, measurement choices, context, or sample size as possible explanations for the disagreement.

Preliminary, Limited, Robust, and Compelling Evidence

Four adjectives, four different confidence levels

Preliminary: Preliminary evidence suggests that the intervention may be effective. — early evidence, not a final conclusion.

Limited: There is limited evidence that the policy improves productivity. — evidence exists but is small, insufficient, or weak in some respect.

Robust: Robust evidence supports the effectiveness of the intervention. — evidence considered strong and resistant to reasonable methodological concerns.

Compelling: There is compelling evidence that the policy improved outcomes. — a strong author stance; still worth asking compelling according to what, exactly?

A Full Passage, Fully Decoded

Read once, slowly, before checking the sentence-by-sentence breakdown

A logistics company introduced an AI-assisted scheduling system and subsequently reported a 22 percent reduction in average vehicle waiting time. At first glance, this result appears to support the view that automated scheduling improves operational efficiency. However, the company also changed its staffing arrangements and redesigned several workflow procedures during the same period. Moreover, the evaluation was based on data from a single terminal over a six-month period. The findings therefore provide some evidence that the new system contributed to improved efficiency, but they do not establish that the technology was solely responsible for the reduction. Further studies across multiple terminals would be needed to determine whether the results can be generalized more broadly.

Sentence 1 — data: a 22 percent reduction, an observed result. Sentence 2 — interpretation: "appears to support" — cautious, signalled by appears. Sentence 3 — alternative explanations: staffing changes and workflow redesign — automation alone can't easily be isolated. Sentence 4 — scope limitation: one terminal, six months — external validity is limited. Sentence 5 — qualified conclusion: "provide some evidence" (not prove), then "do not establish solely responsible" — a carefully qualified position, closing with an explicit call for further evidence.

The passage's reasoning, as a map

DATA (22% reduction) → INITIAL INTERPRETATION (supports automation benefit) → ALTERNATIVE EXPLANATIONS (staffing + workflow changes) → SCOPE LIMITATION (one terminal, six months) → QUALIFIED CONCLUSION (automation probably contributed) → LIMIT (sole causation not established) → NEXT EVIDENCE NEEDED (multiple terminals, longer studies).

What Advanced Reading Looks Like

Compare these two readings of the same passage

Basic reader: "AI scheduling reduced waiting time by 22%."

Advanced reader: "A 22% reduction was observed after implementation, but because other operational changes occurred and the evidence came from one terminal over six months, the study supports — but does not establish — a causal effect attributable solely to the AI system."

The second reading is today's target — not because it's more complicated for its own sake, but because it's the version that's actually true to what the passage supports.

A Nine-Step Master Framework for Reading Evidence

Working through any research-based paragraph

1What is the data?
2What pattern is observed?
3What does the author interpret from it?
4What claim is being made?
5What evidence supports that claim?
6What alternative explanations exist?
7What are the study's limitations?
8How far can the result be generalized?
9How certain is the final conclusion?

Vocabulary in Context

representativeadjective

accurately reflecting the characteristics of a larger group or population (প্রতিনিধিত্বমূলক)

“A representative sample makes a study's findings more likely to generalize beyond the people actually studied.”

biasnoun

a systematic distortion in how a sample or study is put together, which skews its results (পক্ষপাত/পক্ষপাতিত্ব)

“Surveying only top performers introduces selection bias into the results.”

replicateverb

to repeat a study independently to see whether it produces the same result (পুনরাবৃত্তি/পুনরায় সম্পাদন করা)

“A finding that has been replicated across several independent studies is generally more trustworthy.”

triangulationnoun

using several different kinds or sources of evidence to check whether they point toward the same conclusion (বহুমুখী যাচাই)

“Triangulation across survey data, interviews, and transaction records strengthened the study's conclusion.”

preliminaryadjective

happening at an early stage, before a fuller or final result is available (প্রাথমিক)

“Preliminary evidence suggests the intervention may help, but a final conclusion would require more research.”

robustadjective

strong and able to withstand reasonable challenges or scrutiny (মজবুত/সুদৃঢ়)

“Robust evidence tends to hold up even when researchers test it with different methods.”

generalizeverb

to apply a finding from a specific sample or study to a broader population or context (সাধারণীকরণ করা)

“A study of 50 developers cannot automatically be generalized to all workers.”

external validitynoun phrase

the degree to which a study's findings can reasonably apply beyond the specific sample or context it examined (বহিঃস্থ বৈধতা)

“A study's external validity is limited when its sample comes from only one unusual context.”

anecdotenoun

a single, informal example or story, generally weaker as evidence than systematic data (একক ঘটনা/কাহিনি)

“One manager's anecdote about automation is useful context, but not strong general evidence.”

scopenoun

the range or extent of people, situations, or conditions that a claim or study actually covers (পরিধি/সীমা)

“A claim's scope should match the scope of the evidence actually collected for it.”

Guided Reading Practice

Read this passage once for its overall claim, then go back and label each sentence as data, interpretation, alternative explanation, scope limitation, or qualified conclusion before checking the notes below.

A regional port authority piloted a new predictive-maintenance system for its container cranes at one terminal. Over the following year, unplanned crane downtime fell by 35 percent compared with the previous year. Officials described the result as evidence that predictive maintenance had improved crane reliability. However, the terminal had also replaced several aging cranes with newer equipment during the same period, and the pilot covered only a single terminal out of the port's five. The available data therefore provide some support for the system's value, but they do not establish how much of the improvement came from predictive maintenance rather than equipment replacement, or whether similar results would occur at the port's other terminals.

Data: a 35 percent fall in unplanned downtime. Interpretation: officials describe this as evidence of improved reliability. Alternative explanation: aging cranes were also replaced during the same period — a competing explanation for the same result. Scope limitation: one terminal out of five — a clear external- validity concern. Qualified conclusion: "provide some support…but do not establish" — exactly the same careful, two-part qualification pattern as this lesson's worked AI-scheduling passage, applied to a different domain.

Golden Rule

Golden Rule

Evidence does not speak for itself — always separate what happened from what the author thinks it means, and from how far that meaning can reasonably travel.

Lesson Summary

Today's lesson pulled apart a chain that casual reading usually compresses into one impression: data, observation, interpretation, inference, claim, and conclusion. You practised testing evidence on three questions — relevant, sufficient, reliable — and catching the gap between a study's actual sample and the broader claim resting on top of it. You learned to read the support-language family precisely (consistent with, supports, contradicts, calls into question) rather than treating them as interchangeable, and to separate absence of evidence from the much stronger evidence of absence — a distinction that changes how a "no evidence found" sentence should actually be read. Every one of these skills serves one goal: judging whether a conclusion has genuinely earned its place, or has simply been written with enough confidence to sound like it has.

The habit to carry forward

Whenever a passage reports a result, run it through the chain: what was observed? what does the author think it means? what evidence backs that interpretation? what else could explain it? and how far can this reasonably generalize? That five-part question is this whole lesson, compressed into one habit.

Practice: Test What You've Learned

Work through every question yourself before checking anything.

Before you start

Before answering, decide for yourself which layer — data, interpretation, inference, or conclusion — each sentence belongs to; mixing up two layers is usually where a wrong answer comes from.

Part A — Identify the Layer

Label each sentence: D = Data/Observation, I = Interpretation, E = Evidence-based inference, C = Conclusion.

  1. Average processing time fell from 14 minutes to 10 minutes.
  2. This reduction suggests that the new workflow improved efficiency.
  3. Taken together, the findings support continued investment in process automation.
  4. The improvement was observed in 85 percent of the participating departments.

Part B — Evidence Strength

  1. One manager reported that the new system was highly effective. What kind of evidence is this?
  2. Five independent studies reported similar results. Why is this stronger?
  3. The study found no evidence that the intervention caused harm. Does this mean the intervention is definitely harmless?

Part C — Scope

  1. A study of 100 experienced software developers found that AI tools increased coding speed. Can this evidence support the claim "AI tools increase productivity for all workers"? Explain why or why not.

Part D — Deep Analysis

Read this paragraph for questions 9-15: A port introduced an automated cargo-documentation system in January. By June, the average time required to process documentation had fallen by 30 percent. Management interpreted the improvement as evidence that automation had increased operational efficiency. However, the port had also introduced additional staff training and revised several documentation procedures during the same period. Furthermore, the evaluation covered only one operational unit. The findings therefore provide some support for the effectiveness of the automated system, but they do not establish that automation alone caused the improvement or that the results would necessarily apply to the entire port.

  1. What is the data?
  2. What is the observation?
  3. What is management's interpretation?
  4. What alternative explanations does the passage raise?
  5. What is the scope limitation?
  6. What is the author's final position?
  7. Name two specific things the evidence does not establish.

Part E — Absence of Evidence vs. Evidence of Absence

  1. A safety audit found no evidence that the equipment was faulty. Does this mean the equipment is definitely not faulty? Explain your reasoning.
  2. Rewrite the audit's finding as a genuine "evidence of absence" claim, and explain exactly how that version would differ from the original.

With evidence, claims, data, and interpretation now separated into their own distinct layers, you're ready to apply this same discipline to any research paper, business report, or serious nonfiction passage you read next.