LearnAnyLanguage.com

How can I set up an experiment to investigate second language acquisition?

Evidence-based answer · Last updated · How it’s made

The short answer

Setting up a solid second language acquisition (SLA) experiment means picking a specific, well-defined linguistic target, deciding whether you'll test in a controlled lab-like setting or a real classroom (each has trade-offs), using a pretest/treatment/posttest structure with a control group, and choosing measures that actually capture the kind of knowledge you care about (quick, unmonitored production vs. careful grammar tests tap different things). The sources describe many concrete design choices — artificial vs. natural language input, eye-tracking, delayed posttests, standardized proficiency tests — that let you trade off experimental control against real-world relevance depending on your question.

Answered for language learning in general. No specific language was set for this question, so the findings are drawn from general second-language research rather than one language’s literature.

Lab control vs. real-world relevance trade-off

Researchers face a basic trade-off: a tightly controlled lab study lets you isolate exactly one variable (like a type of feedback or an artificial grammar rule) with confidence, but findings may not transfer to messy real classrooms; a classroom study is more realistic but has more uncontrolled variables (student motivation, teacher style, mixed prior knowledge) that can muddy the results.1,2

In plain terms

'Internal validity' means how confident you can be that your treatment, and nothing else, caused the result. 'Ecological validity' means how well the finding generalizes to real-life learning outside the experiment.

Artificial or semi-artificial languages let you control prior knowledge

Many SLA experiments use invented mini-languages, or natural languages modified so participants have zero prior exposure, specifically to guarantee that nobody already knows the grammar rule being taught. Some researchers run a 'twin' design: the same experiment once with an artificial language (high control) and once with a real language (high realism), hoping the two sets of results agree.1,5,7

Example

For example, one researcher taught English speakers a version of Finnish stripped of some features, while another built entirely made-up verbs and word orders so that no participant could have learned the target rule beforehand.

Pretest–treatment–posttest with a control group is the standard skeleth

The classic design used across the studies is: test learners before any teaching (pretest), give one group the instructional treatment while a control group gets none or a different treatment, then test again immediately after (posttest) and sometimes again weeks later (delayed posttest) to see if gains stick.6,9,12,13

Example

In one study on English relative clauses, learners were randomly split into two treatment groups and a control group, pretested on their existing knowledge, given different types of instruction, and then compared on posttest improvement.

Delayed posttests reveal whether learning actually lasts

Several studies added a delayed posttest (weeks after training) in addition to an immediate posttest, because gains measured right after teaching can fade or, in some cases, even continue to grow — so testing only right after training can be misleading about durability.12

Example

In one grammar study, one instructed group scored well right after training but a different group actually caught up and matched it only on the delayed test given ten weeks later.

Pretests are needed to rule out prior knowledge as a hidden variable

When using a real language as the experimental material, almost all studies had to pretest participants and exclude anyone who already knew the target grammar structure, because uneven prior knowledge among participants can distort what looks like a training effect.1

Choice of test format determines what kind of knowledge you're measuring

Different test formats (grammar judgment tests, elicited imitation, free speech, written tests) tap different kinds of knowledge — quick, spontaneous production tasks are thought to better reveal 'implicit' knowledge (automatic, unconscious know-how), while careful, monitored tasks favor 'explicit' knowledge (conscious rule knowledge) — so the measurement tool you pick can change your conclusions even with identical teaching.4,11,12

In plain terms

'Implicit knowledge' is language ability you can use automatically without consciously thinking about the rule, like knowing a sentence sounds wrong without being able to say why. 'Explicit knowledge' is being able to state the rule itself.

Eye-tracking and ERP brain measures can capture real-time processing

Beyond behavioral tests and scores, some SLA experiments use eye-tracking (recording where and how long someone looks at text) or ERP (a brain-response measure) to see moment-by-moment how learners process language, which can reveal effects that a simple accuracy score misses.3,4

In plain terms

ERP stands for event-related potentials, a way of measuring the brain's electrical response to a word or sentence as it's being processed.

Standardized proficiency tests improve comparability

Using an externally standardized, independently validated proficiency test (rather than a test the researcher invented just for that study) makes results easier to compare across studies and more directly tied to real communicative ability.3,9

Meta-analyses require coding many design features to compare studies fairly

When trying to combine results across many separate experiments (a meta-analysis), researchers have to carefully record details like learners' proficiency level, treatment duration, and target skill (speaking, grammar, vocabulary) because studies vary so much in these features that raw scores aren't directly comparable.8

In plain terms

A meta-analysis is a study that statistically combines the results of many separate experiments to look for an overall pattern.

Quasi-experimental classroom designs use intact groups

When true random assignment isn't possible in a school setting, researchers often use 'quasi-experimental' designs comparing whole pre-existing classes rather than individually randomized learners, accepting some loss of control in exchange for practicality.9,10,13

In plain terms

Quasi-experimental means the study has treatment and comparison groups like a true experiment, but participants weren't randomly assigned to them — often because you're working with already-formed classes.

What to do with this

  • Start by narrowing your question to one specific, well-defined linguistic target (a grammar structure, a set of vocabulary words, a pronunciation feature) rather than 'language learning' broadly — this makes both instruction and measurement manageable.
  • Decide early whether you need lab-level control (useful if you want to isolate one mechanism, e.g. by using an artificial or unfamiliar language so no one already knows the target rule) or classroom realism (useful if you want findings that generalize to real teaching, accepting more uncontrolled variables).
  • Always pretest participants on the target structure so you can exclude or account for people who already know it, then randomly assign to treatment and control groups where possible, or use matched intact classes if random assignment isn't feasible.
  • Measure with more than one instrument if you can: a controlled task (like a grammar judgment test) for explicit knowledge and a spontaneous production task (free speech or writing) for implicit, automatic knowledge, since these can show different results from the same treatment.
  • Add a delayed posttest weeks later, not just an immediate one, to see whether any effect actually lasts.
  • If possible, anchor your outcome measure to an independent, standardized proficiency test rather than a homemade one, so your results are easier to compare to other research.

Now pick something to do it with

Our resource list is an overview of the apps, courses and immersion material worth a look, whatever you're working on.

See our resources

Worth knowing

  • The sources are drawn from academic SLA methodology papers and individual studies rather than a single unified guide to experimental design, so this report synthesizes design principles that appear repeatedly rather than quoting one authoritative 'how-to.
  • ' Several sources (mindfulness, gestures, glossing, extensive reading) are individual studies illustrating specific design choices rather than general methodology papers, so they are used here as examples of design elements rather than as comprehensive guidance.
  • The sources also disagree somewhat implicitly on how much lab control is worth sacrificing for real-world relevance, and this is a genuinely unresolved trade-off in the field, not a settled answer.

Want a deeper literature dive?

There’s enough published research here to go wider than the usual pass. Available on the Pro plan.

See plans

  1. [1]Jan H. Hulstijn. SECOND LANGUAGE ACQUISITION RESEARCH IN THE LABORATORY. Studies in Second Language Acquisition 1997. doi.org/10.1017/s0272263197002015
  2. [2]Doing SLA Research with Implications for the Classroom. Language learning and language teaching 2019. doi.org/10.1075/lllt.52
  3. [3]Stefano Rastelli, John W. Schwieter. From the laboratory to the classroom: Rethinking ERP methods in second language acquisition.. Acta Psychologica 2026. doi.org/10.1016/j.actpsy.2026.107407
  4. [4]Cécile Laval, Harriet Lowe. The use of eye-tracking in experimental approaches in second language acquisition research: the primary effects of Processing Instruction in the acquisition of the French imperfect. Journal of French Language Studies 2020. doi.org/10.1017/S0959269520000046
  5. [5]Simon Kirby, Hannah Cornish, Kenny Smith. Cumulative cultural evolution in the laboratory: An experimental approach to the origins of structure in human language. Proceedings of the National Academy of Sciences 2008. doi.org/10.1073/pnas.0707835105
  6. [6]Catherine J. Doughty. Second Language Instruction Does Make a Difference. Studies in Second Language Acquisition 1991. doi.org/10.1017/s0272263100010287
  7. [7]J. Culbertson, Kathryn D. Schuler. Artificial Language Learning in Children. Annual Review of Linguistics 2019. doi.org/10.1146/ANNUREV-LINGUISTICS-011718-012329
  8. [8]Wei‐Chen Lin, Hung–Tzu Huang, Hsien‐Chin Liou. The Effects of Text-Based SCMC on SLA: A Meta Analysis. Language learning & technology 2013. doi.org/10.64152/10125/44327
  9. [9]María J. de la Fuente, Carola Goldenberg. Understanding the role of the first language (L1) in instructed second language acquisition (ISLA): Effects of using a principled approach to L1 in the beginner foreign language classroom. Language Teaching Research 2020. doi.org/10.1177/1362168820921882
  10. [10]Geòrgia Pujadas, Carmen Muñoz. Extensive viewing of captioned and subtitled TV series: a study of L2 vocabulary learning by adolescents. Language Learning Journal 2019. doi.org/10.1080/09571736.2019.1616806
  11. [11]Rod Ellis, Shawn Loewen, Rosemary Erlam. IMPLICIT AND EXPLICIT CORRECTIVE FEEDBACK AND THE ACQUISITION OF L2 GRAMMAR. Studies in Second Language Acquisition 2006. doi.org/10.1017/s0272263106060141
  12. [12]Broszkiewicz, Anna. The Effect of Focused Communication Tasks on Instructed Acquisition of English Past Counterfactual Conditionals. Studies in Second Language Learning and Teaching 2011. https://eric.ed.gov/?id=EJ1136449
  13. [13]Namhee Suk. The Effects of Extensive Reading on Reading Comprehension, Reading Rate, and Vocabulary Acquisition. Reading Research Quarterly 2016. doi.org/10.1002/rrq.152