Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Sampling and populations and experiments, oh my

Defining Statistics and Biostatistics

Before we approach more rigorous ideas of probability and statistics, it can be helpful to understand how we conceptualize statistical (and biostatistical) practices.

Let’s first look at a list of different definitions for the field of statistics:

  1. Wikipedia: Statistics is the discipline that concerns the collection, organization, analysis, interpretation, and presentation of data.

  2. American Statistical Association: The science of learning from data, and of measuring, controlling, and communicating uncertainty.

  3. Bartholomew in this Presidential Address to the Royal Statistical Society summarized the myriad different definitions proposed, some of which are below. ( Bartholomew, D. J. (1995). What is Statistics? Journal of the Royal Statistical Society. Series A (Statistics in Society), 158(1), 1–20.)

  4. Grenander and Miller: The object of statistics is information. The objective of statistics is the understanding of information contained in data.

  5. Egon Pearson: The study of the collective characters of populations

  6. Kendall: The science of collectives and group properties

  7. S.S. Stevens: A straightforward discipline designed to amplify the power of common sense in the discernment of order amid complexity.

  8. E.J. Goncourt: Statistics is the premier inexact science.

There are many different ways to define the field of statistics, and, as expected, there are many different ways to define Biostatistics. However, definition of Biostatistics benefits from having a large intersection with the larger field of statistics.

  1. University of Washington:The branch of statistics that deals with data related to living organisms.

  2. Wikipedia The branch of statistics that applies statistical methods to a wide range of topics in the biological sciences, with a focus on clinical medicine and public health applications

  3. Berger and Matthews: Biostatistics is the discipline concerned with how we ought to make decisions when analyzing biomedical data. It is the evolving discipline concerned with formulating explicit rules to compensate both for the fallibility of human intuition in general and for biases in study design in particular.

Links that supported the above:

  1. https://academic.oup.com/jrsssa/article/158/1/1/7106950

  2. https://pmc.ncbi.nlm.nih.gov/articles/PMC3190464/

  3. https://www.biostat.washington.edu/about/biostatistics

  4. https://magazine.amstat.org/blog/2022/10/01/what-is-statistics/

Population, Sample, Experiment

The goal of many, if not all, of statistics is to develop knowledge and conclusions about a larger collection of individuals (interpreted broadly to mean more than just about humans). The typical presentation is that statisticians wish to draw conclusions about the set of all individuals, called the target population by collecting information about a smaller set of like individuals contained inside that population, called a sample.

To be clear, a target population is the set of individuals that you wish to draw conclusions about. The act of selecting individuals from this population is called sampling.

That said, there are important subtleties about the above process that are worth exploring. 1. Does your target population of interest exist as a static entity over time? 2. Do all individuals have a positive probability of being selected (i.e. sampled)? 3. Do all individuals have the same probability of selection?

If the target population you wish to study actually exists in space and time then this population is further described as a real population. A real population is a finite, definite, countable collection of individuals that actually exist in our perceived reality at this time.

Ex: What's the deal with "Real" Populations
Suppose you wish to study the impact of a new vaccine technology on its effectiveness to provide an immune response to influenza. The population that the technology was designed for was college students aged 18 to 24. Then the "Real population" is all individuals enrolled in a college who are aged 18 to 24. If you collect a smaller set of college students and conduct the study you may aim to conclude that the immune response in this smaller "sample" applies all college students aged 18-24. But is this correct? If you collected all possible college students aged 18-24 (assuming you have infinite resources) then is there any uncertainty in the immune response? Discuss.

Smaller than a real population, a statistical population is the set of individuals that can feasibly be collected and studied. The subtlety between a real and statistical population is that a statistical population admits that a positive probability cannot be assigned to every individual. Due to time or resources or complexity, we might not really be able to collect everyone. A statistical population is therefore a limited real population because of limitations in the sampling process.

A super-population is a hypothetical (usually infinite) population, or data-generating process, that can produce real populations. You can think of the real population that exists now as a single realization from a super-population. For example, let’s suppose that we wanted to better understand how students understand the derivative (I am a mathematician, after all). Further suppose that we collected absolutely everyone in our real population. If we summarize our findings are they final? Is there some uncertainty still left over? If you, as the scientist, feel like there is still uncertainty left over then you’re thinking about the super population. Even though you collected the entire real population, that real population may be changing over time: growing, shrinking, the individuals may be morphing, so that the real population you collected at time tt looks a bit different now at time t+t+.

Much of this discussion is owed to great works like https://www.sciencedirect.com/science/article/pii/S0028377024000791

Sampling

We stated before that sampling is a strategy for selecting individuals from a target population such that learning about this sample tells us something about the target population.

For example, suppose we are interested in better understanding who is most at risk for hospitalization due to influenza. Our target population are those infected with influenza, and our sample might select a smaller set of individuals who are infected: not all of them. With a sample in hand, we can measure different characteristics of that set of collected individuals and claim that the estimates in our sample tell us something about the larger target population.

But how do we pick who is in our sample?

Simple random sampling

Approach

Simple random sampling (SRS) assigns the same probability to every individual in our target (often statistical) population. An individual from the population is picked and placed in the sample.

Pros A (true) simple random sample will provide you accurate measurements of any variables of interest. In addition, a SRS will also supply you with individuals who have other attributes that are representative of your target population. For example, in the influenza (flu) example above, if we use a SRS and assign the same probability of choosing individuals sick with the flu then if 60% of those who are sick with the flu are under the age of three then our sample will likely have 60% of individuals under three.

Cons That said, a true SRS isn’t usually feasible. We often never have access to every member of a population and so we can never really assign an equal probability to selecting everyone. In addition, because we assign the same probability to selecting an individual then it may be very difficult to have collected individuals with rare (but maybe important) attributes. For example, in the flu example, if a very small number of flu patients are sick because of an infection with an antibiotic resistant bacteria then we are likely to miss selecting these patients in our sample.

Stratified sampling

Approach

Given a population where every individual has a yet-to-be-measured variable of interest yy and a known attribute xx that can take one of KK potential values, select n/Kn/K individuals with x=1x=1 using a SRS; then select n/Kn/K individuals with x=2x=2, and so on until you get to the KthK^{\text{th}} group with x=Kx=K, selecting n/Kn/K individuals.

Pros A stratified sample allows you to equally represent, in your sample, KK different groups that are likely heterogeneous in nature across their KK groups but homogeneous within each group. Therefore, stratified sampling is good for when you want to understand your variable-of-interest among different groups, often called strata. This is particularly effective when there exist very small strata in the population that are important.

Cons That said, if you estimate your variable-of-interest across all strata then it will no longer be representative of the entire population.

Why?

Consider our variable of interest to be a preference for pepsi vs coca-cola in the University population. We decide to conduct a stratified sampling procedure across freshman, sophomores, juniors, and seniors, selecting 10 from each grade.

Then our overall estimate will assume that each grade is equally represented in our overall target population of the University. However, this is usually untrue. The freshman class is typically largest and represents 50% of the University. However in our stratified sample they are only represented at 10/40=25% of the entire population: they are under-represented. Likewise, seniors are usually the smallest group at 10% but our sample assumes that they are 25% of the population: they are over-represented.

Cluster sampling

Approach

Given a target population, variable-of-interest, yy , and attribute xx with (again) KK different categories, cluster sampling selects from the KK different categories and then samples all individuals within that category.

A typical setup for cluster sampling is spatial sampling. Our variable of interest may be the incidence of west-nile virus among students at Lehigh University. Rather than collecting a SRS of all 8,000 we will instead select 3 dorm residences at random and then collect our variable-of-interest from every single member who resides in those three dormitories.

Pros Cluster sampling can be cheap and convenient, and quick (especially if you need to travel to one of many locations for data collection).

Cons However, when you conduct a cluster sampling approach you typically underestimate the variability (i.e. how spread out) in your variable-of-interest yy. This is because individuals within a cluster tend to be more similar than those outside the cluster.

Classroom exercise:

The questions that we discussed in class:

  1. Why does each group have a different estimate and which one is correct?

  2. Where is the randomness and uncertainty “coming from”?

  3. If we combined all our data would there still be uncertainty about another 30 students?

A main story-line for statistics is how to infer population level characteristics from a smaller, finite sample.

Homework / Exercises

  1. Population or Sample?: Suppose you decide to characterize the number of days from diagnosis to all-cause mortality (death from any cause) of a rare disease. There exist just 100 (unfortunate) individuals with this disease. Your experiment follows all one hundred individuals for two years and all individuals experience death within this time period. What is the population for your experiment? What is the sample for your experiment? In this case, the sample happens to be all individuals. Then, is there any randomness left? Are you estimating the number of days from diagnosis to death or are you discovering the population parameters in this population? Please explain.

  2. Run your own experiment: Please run your own experiment and collect at minimum ten data points that contain three pieces of information each. Organize this information: (a) create a folder on your local computer called “./Plots_and_probability/” and (b) a folder called “./Plots_and_probability/experiment/”. In the experiment folder, please

    1. Add a file called “summary.txt” that outlines your experiment.

    2. Add a file called “data_collection.csv” that includes the data that you collected with descriptive column names.

    3. Add a file called “data_dictionary.txt”. This file will describe what every column means. What types of data are allowed in each column? Are there any irregularities or are there minimum/maximums associated with these columns?

    4. What’s the population that your experiment is extracting data from?

  3. Book problem (common cold): Please complete exercise 1.12 on page 78 of the book titled Introductory Statistics for the Biomedical and Life Sciences.

  4. In the wild: Please read the 2010 paper titled “Transcatheter Aortic-Valve Implantation for Aortic Stenosis in Patients Who Cannot Undergo Surgery”. The link for this (free) article is here Leon et al. (2010). Answer the following questions about this statistical experiment.

    1. Please describe the population that this experiment was meant to draw conclusions about

    2. Briefly describe the core experiment (how many people were recruited and how? what happened after recruitment? what was measured about these patients? etc).

    3. In Table 1, the authors present summary statistics. Can you comment on the purpose of presenting this table?

    4. The plot of a story is the specific connected chain of events that lead from the beginning to the end of a story. In scientific papers, the plot is typically pretty straightforward: but not the narrative. The narrative is (usually) the approach for telling the plot. Please take one paragraph to read over this paper and describe elements that were persuasive, elements that felt rhetorical in nature, used to convince the reader, and please also add one thing that was difficult for you to grasp.

References
  1. Leon, M. B., Smith, C. R., Mack, M., Miller, D. C., Moses, J. W., Svensson, L. G., Tuzcu, E. M., Webb, J. G., Fontana, G. P., Makkar, R. R., Brown, D. L., Block, P. C., Guyton, R. A., Pichard, A. D., Bavaria, J. E., Herrmann, H. C., Douglas, P. S., Petersen, J. L., Akin, J. J., … Pocock, S. (2010). Transcatheter Aortic-Valve Implantation for Aortic Stenosis in Patients Who Cannot Undergo Surgery. New England Journal of Medicine, 363(17), 1597–1607. 10.1056/nejmoa1008232