🧠 Psychology · Research Methods

Research tricks that make methodology click

Study designs, variables, validity, and statistics β€” mastered.

πŸ“Š Research

Memory tricks

Proven mnemonics — fast to learn, hard to forget.

Correlation vs Causation
Correlation β‰  Causation: ice cream and drowning both rise in summer
Correlation vs Causation
Two things happening together doesn't mean one causes the other
Hot weather causes both ice cream sales and drowning rates to rise β€” neither causes the other. To establish causation: need an experiment with random assignment.
Statistical Significance
p < 0.05: less than 5% chance the result occurred by chance
Statistical Significance
The p-value tells you if a result is likely to be random
p < 0.05 = less than 5% probability the result is due to chance. Does NOT mean the effect is large β€” just unlikely to be random. Effect size tells you the magnitude.
Research Threats
RAVEN: Random assignment, Attrition, Validity, Experimenter bias, Null hypothesis
Research Threats
Five threats to the validity of a study
Random assignment eliminates pre-existing differences. Attrition (dropout) can bias results. Internal validity: did IV cause DV? Experimenter bias: researcher inadvertently influences results. Null hypothesis: default 'no effect' claim.
R
Random assignment
A
Attrition β€” dropout
V
Validity β€” internal and external
E
Experimenter bias
N
Null hypothesis
Reliability and Validity
Reliability vs Validity: consistency vs accuracy. Can be reliable but not valid.
Reliability and Validity
Two essential qualities of any good measurement tool
Reliability: gives same result each time (consistent). Validity: measures what it claims to measure (accurate). A scale that always reads 5 lbs too heavy is reliable but not valid. Both are required for a good study.
Experimental Design
Experimental design: independent variable manipulated, dependent variable measured, extraneous variables controlled
Experimental Design
The structure of a true experiment β€” the only way to establish causation
Random assignment: participants randomly placed in experimental or control group β€” controls for pre-existing differences. Control group: doesn't receive treatment β€” provides baseline. Experimental group: receives the IV manipulation. Double-blind: neither participants nor researchers know who's in which group β€” prevents bias.
Effect Size
Effect size: Cohen's d. Small=0.2, Medium=0.5, Large=0.8. More meaningful than p-value alone.
Effect Size
How big is the effect β€” the practical significance question
Statistical significance (p-value) only tells you the result probably isn't random β€” not how important it is. A huge study can find a tiny, meaningless effect at p<0.001. Effect size measures the magnitude. Cohen's d = (mean₁ - meanβ‚‚)/pooled SD. Always report effect size alongside p-value.
Sampling Methods
Sampling methods: random (every member has equal chance), stratified (proportional subgroups), convenience (whoever is available)
Sampling Methods
How researchers select participants β€” affects generalizability
Simple random: each person equally likely to be selected β€” best for generalizability. Stratified: divide population into subgroups (strata), randomly sample from each β€” ensures representation. Cluster: randomly select groups then sample within. Convenience: whoever is available β€” biased, poor generalizability (WEIRD problem: Western, Educated, Industrialized, Rich, Democratic).
Research Methods Overview
Naturalistic observation: observe in natural setting, no manipulation. Strength: ecological validity.
Research Methods Overview
The main research designs and their key trade-offs
Naturalistic observation: high ecological validity, no cause-effect. Case study: rich detail, poor generalizability. Survey: large samples quickly, self-report bias. Correlational: shows relationships, no causation. Experimental: only method establishing causation, may lack ecological validity. Choose method based on research question.
Operational Definitions
Operational definition: precisely how a variable is measured. 'Intelligence' measured as 'IQ score on Wechsler.'
Operational Definitions
Turning abstract concepts into measurable variables
Operational definition: specifies the exact procedures used to measure or manipulate a variable. 'Stress' is abstract β€” measured as cortisol level, heart rate, or score on perceived stress scale. Good operational definitions: reliable (consistent), valid (measures what it claims), practical. Allows replication.
Longitudinal vs Cross-Sectional
Longitudinal study: same people over time. Cross-sectional: different age groups at one time. Each has limitations.
Longitudinal vs Cross-Sectional
Two ways to study development β€” each with different flaws
Longitudinal: follow same people over years/decades. Strength: sees actual change. Weaknesses: dropout (attrition), time-consuming, expensive, cohort effects. Cross-sectional: compare different age groups at same time. Strength: quick, no attrition. Weakness: cohort effects (different generations, not just age).
Blind Procedures
Single-blind: participants don't know condition. Double-blind: neither participants nor researchers know. Reduces expectancy effects.
Blind Procedures
Controlling for expectation effects in research
Demand characteristics: participants guess the study's purpose and change behavior accordingly. Experimenter bias: researcher unconsciously treats groups differently or interprets results based on expectations. Single-blind eliminates demand characteristics. Double-blind eliminates both. Placebo effect: inert treatment produces real changes because of expectations.
Research Designs
EXCEL β€” Experimental, Cross-sectional, Experience sampling, Longitudinal, Lab vs. field
Five major research design types and when to use each
Research design choice determines what conclusions you can draw β€” causation requires experiments
Experimental: random assignment to conditions β†’ can infer causation. Quasi-experimental: no random assignment β€” groups naturally differ. Cross-sectional: different age groups at one time β€” fast, but cohort effects confound. Longitudinal: same people over time β€” shows true development but time-consuming and attrition. Case study: deep individual analysis β€” rich data, poor generalizability. Naturalistic observation: behavior in natural settings β€” ecological validity but no control.
Experimental
Random assignment β†’ only design that proves causation
Longitudinal
Same people over time β†’ shows developmental change
Cross-sectional
Different ages at once β†’ cohort effects are a problem
Statistical Significance
P-VALE β€” P-value, Variance, Alpha level, Level of significance, Effect size
The key concepts for interpreting whether a finding is real or due to chance
Statistical significance (p < .05) does not mean practically important β€” effect size tells you magnitude
P-value: probability of getting results at least this extreme if the null hypothesis is true. p < .05 = less than 5% chance of a false positive (Type I error). Alpha level (Ξ±): threshold set before study (usually .05). Type I error (false positive): rejecting true null. Type II error (false negative): failing to reject false null. Effect size (Cohen's d): how large is the difference β€” small (.2), medium (.5), large (.8). A huge study can make trivial effects statistically significant.
Type I error
False positive β€” seeing effect that doesn't exist (Ξ± = .05)
Type II error
False negative β€” missing real effect (Ξ², related to power)
Effect size
Cohen's d β€” how big is the difference, practically speaking
Sampling Methods
RSCCP β€” Random, Stratified, Cluster, Convenience, Purposive
Five sampling strategies and their tradeoffs between representativeness and feasibility
Random sampling is the gold standard for generalizability β€” most real studies use convenience samples
Random sampling: every person in population has equal chance β€” best for generalizability. Stratified random: sample proportionally from subgroups β€” ensures minority groups represented. Cluster: randomly select groups (schools, hospitals), then sample within. Convenience: whoever is available β€” cheap but biased (WEIRD: Western, Educated, Industrialized, Rich, Democratic). Purposive: deliberately select particular individuals β€” used in qualitative research. Sample size affects statistical power.
Random
Most representative β€” everyone has equal chance
Stratified
Ensures proportional representation of subgroups
WEIRD
Most psychology studies use Western college students β€” poor generalizability
Research Ethics
BIRD β€” Beneficence, Informed consent, Risk minimization, Debriefing
Four core ethical principles in psychological research
APA ethics guidelines emerged from historical abuses β€” Milgram and Zimbardo changed research ethics forever
Beneficence: research must benefit society and minimize harm. Informed consent: participants must understand and voluntarily agree to procedures (deception requires IRB approval and debriefing). Confidentiality: protect participant data and identity. Right to withdraw: can leave at any time without penalty. Debriefing: explain true purpose after deception. Milgram obedience study and Stanford Prison Experiment sparked ethics reform β€” IRBs (Institutional Review Boards) now required for all human research.
IRB
Institutional Review Board β€” must approve all human research
Deception
Allowed only if necessary and followed by debriefing
Debriefing
Explain true purpose β€” required after deception studies
🎓 Common Exam Questions
Q: What is the difference between reliability and validity in psychological measurement?
A: Reliability: consistency of measurement β€” does the test give the same result across time (test-retest reliability), across items (internal consistency β€” Cronbach's alpha), across raters (inter-rater reliability)? A reliable test gives consistent results. Validity: does the test measure what it claims to measure? Face validity: looks like it measures the construct. Content validity: covers all aspects of the construct. Criterion validity: correlates with other measures of the same construct (concurrent) or predicts future outcomes (predictive). Construct validity: overall evidence that the test measures the theoretical construct. Key relationship: a test can be reliable without being valid (consistent but measuring the wrong thing) but cannot be valid without being reliable.
Q: Explain internal validity, external validity, and the threats to each.
A: Internal validity: confidence that the independent variable caused the change in the dependent variable β€” not some confound. Threats: selection bias (groups differ at baseline), history (other events occur during study), maturation (participants change over time), testing effects (practice improves performance), regression to the mean, demand characteristics (participants guess hypothesis), experimenter bias. Controls: random assignment, double-blind procedures, control groups, standardized procedures. External validity: generalizability of findings to other people, places, and times. Threats: convenience samples (WEIRD), artificial lab settings, demand characteristics. Tradeoff: high internal validity (lab experiment) often means lower external validity (artificial setting). Replication across diverse samples and settings builds external validity.
Q: What is the scientific method and how does hypothesis testing work in psychology?
A: Scientific method: (1) Observation and question. (2) Literature review. (3) Hypothesis β€” specific, falsifiable prediction (if-then format). (4) Research design selection. (5) Data collection. (6) Analysis and interpretation. (7) Report and peer review. (8) Replication. Null hypothesis (H0): no effect or relationship exists. Alternative hypothesis (H1): effect or relationship exists. Statistical testing determines probability of results given null hypothesis is true. p < .05 convention: less than 5% chance of false positive. Replication crisis: many classic psychology findings fail to replicate β€” prompted pre-registration (register hypothesis and methods before data collection) and open science movement. Publication bias: positive results are published; null results are not β€” inflates apparent effect sizes.
Q: What is correlation and why doesn't it imply causation?
A: Correlation measures the strength and direction of the linear relationship between two variables. Pearson's r ranges from -1.0 (perfect negative) to +1.0 (perfect positive). Moderate correlations in psychology: r = .3-.5 is typical for complex behavior. Correlation does not imply causation because of: (1) Directionality problem β€” we don't know which variable causes which. (2) Third variable problem (confound) β€” a third unmeasured variable causes both. Example: ice cream sales correlate with drowning rates β€” both caused by hot weather. Only experimental manipulation with random assignment can establish causation. Correlation is useful for: prediction, establishing that a relationship exists, measuring associations that cannot be experimentally manipulated (genetics, personality).
Q: Compare experimental and correlational research designs in terms of control and generalizability.
A: Experimental design: researcher manipulates the independent variable, randomly assigns participants to conditions, controls extraneous variables. Can establish causation. Weaknesses: artificial lab settings limit generalizability (external validity); many variables of interest cannot be manipulated ethically (trauma, poverty, genetics). Correlational design: measures naturally occurring variables without manipulation. Cannot establish causation but can study variables that cannot be manipulated. Good external validity if using representative samples. Quasi-experimental: compares naturally occurring groups (e.g., people with and without depression) β€” no random assignment, so cannot rule out selection effects. The choice between designs involves weighing the need for causal inference against ethical and practical constraints.