0:00 / 0:00

Reliability and Validity


Reliability - the extent to which a measurement is consistent

Example - an IQ test should give the same result each time it is given to an individual

Validity - the extent to which a measurement reflects the true state of something

Example - an IQ test predicts real-world performance on tasks that require intellectual ability


Reliability is a necessary but not sufficient condition for validity:
  • Something can be reliable but not valid - the tape measures above are individually reliable. Each time I measure an object with one of the tape measures, it gives me the same result. Despite this, the bottom tape measure is invalid - it is stretched out and the inches are longer than a true inch
  • Something cannot be valid if it is not reliable - if every time you step on a scale it gives you a different result, it is not reliable. It cannot be valid

PAGE BREAK

Reliability

Test-retest reliability - give the test or take the measurement multiple times and see if you get the same result. Only appropriate if we do not expect change over time.

Example - an IQ or personality test

Inter-rater reliability - have multiple people rate/score something and compare their results. Appropriate if there is subjectivity in the scoring.

Example - counting the number of aggressive behaviours witnessed in a video of children playing

Internal consistency (or split-half) reliability - randomly select half of the items on a test and compare the result to the other half of the items. Appropriate if there are a large number of items.

Example - a personality test with 200 items should give you the same score no matter which 100 items you select

PAGE BREAK

Validity

Internal validity is the extent to which we can establish cause and effect relationships between variables.
Internal validity is improved by random assignment

Threats to Internal Validity
  • Spontaneous recovery/remission - change without an apparent external reason
  • Maturation - growth of the person
  • Measurement - act of being measured changes how we behave
  • Secular drift - gradual changes in culture
  • History effects - effects of events that happen in larger society
  • Regression to the mean - people who represent extremes of the distribution tend to drift back towards the mean in later tests
  • Instrument effects - poor measurement tools
  • Selection effects - if groups were not well-chosen
  • Attrition/Experimental mortality - if there is a systematic reason people drop out of the study over time

PAGE BREAK

External validity is the extent to which we can generalize our results to the population we're interested in.
External validity is improved by random selection

Threats to External Validity
  • Experimental conditions do not reflect real world
  • Selection criteria are too restrictive
  • Situational effects - presence of lab conditions changes the outcome

Practice: Reliability and Validity

You and a coworker take a personality test. It tells you that you're an introvert who relies on their feelings to make decisions. It tells your coworker that they're an extrovert who relies on their thoughts to make decisions. This is really strange to you, because the two of you have very similar work styles. What is missing in this personality test?