Sampling and populations and experiments, oh my
Table of Contents¶
Defining Statistics and Biostatistics¶
Before we approach more rigorous ideas of probability and statistics, it can be helpful to understand how we conceptualize statistical (and biostatistical) practices.
Let’s first look at a list of different definitions for the field of statistics:
Wikipedia: Statistics is the discipline that concerns the collection, organization, analysis, interpretation, and presentation of data.
American Statistical Association: The science of learning from data, and of measuring, controlling, and communicating uncertainty.
Bartholomew in this Presidential Address to the Royal Statistical Society summarized the myriad different definitions proposed, some of which are below. ( Bartholomew, D. J. (1995). What is Statistics? Journal of the Royal Statistical Society. Series A (Statistics in Society), 158(1), 1–20.)
Grenander and Miller: The object of statistics is information. The objective of statistics is the understanding of information contained in data.
Egon Pearson: The study of the collective characters of populations
Kendall: The science of collectives and group properties
S.S. Stevens: A straightforward discipline designed to amplify the power of common sense in the discernment of order amid complexity.
E.J. Goncourt: Statistics is the premier inexact science.
There are many different ways to define the field of statistics, and, as expected, there are many different ways to define Biostatistics. However, definition of Biostatistics benefits from having a large intersection with the larger field of statistics.
University of Washington:The branch of statistics that deals with data related to living organisms.
Wikipedia The branch of statistics that applies statistical methods to a wide range of topics in the biological sciences, with a focus on clinical medicine and public health applications
Berger and Matthews: Biostatistics is the discipline concerned with how we ought to make decisions when analyzing biomedical data. It is the evolving discipline concerned with formulating explicit rules to compensate both for the fallibility of human intuition in general and for biases in study design in particular.
Links that supported the above:
Population, Sample, Experiment¶
The goal of many, if not all, of statistics is to develop knowledge and conclusions about a larger collection of individuals (interpreted broadly to mean more than just about humans). The typical presentation is that statisticians wish to draw conclusions about the set of all individuals, called the target population by collecting information about a smaller set of like individuals contained inside that population, called a sample.
To be clear, a target population is the set of individuals that you wish to draw conclusions about. The act of selecting individuals from this population is called sampling.
That said, there are important subtleties about the above process that are worth exploring. 1. Does your target population of interest exist as a static entity over time? 2. Do all individuals have a positive probability of being selected (i.e. sampled)? 3. Do all individuals have the same probability of selection?
If the target population you wish to study actually exists in space and time then this population is further described as a real population. A real population is a finite, definite, countable collection of individuals that actually exist in our perceived reality at this time.
Ex: What's the deal with "Real" Populations
Smaller than a real population, a statistical population is the set of individuals that can feasibly be collected and studied. The subtlety between a real and statistical population is that a statistical population admits that a positive probability cannot be assigned to every individual. Due to time or resources or complexity, we might not really be able to collect everyone. A statistical population is therefore a limited real population because of limitations in the sampling process.
A super-population is a hypothetical (usually infinite) population, or data-generating process, that can produce real populations. You can think of the real population that exists now as a single realization from a super-population. For example, let’s suppose that we wanted to better understand how students understand the derivative (I am a mathematician, after all). Further suppose that we collected absolutely everyone in our real population. If we summarize our findings are they final? Is there some uncertainty still left over? If you, as the scientist, feel like there is still uncertainty left over then you’re thinking about the super population. Even though you collected the entire real population, that real population may be changing over time: growing, shrinking, the individuals may be morphing, so that the real population you collected at time looks a bit different now at time .
Much of this discussion is owed to great works like
https://
Sampling¶
We stated before that sampling is a strategy for selecting individuals from a target population such that learning about this sample tells us something about the target population.
For example, suppose we are interested in better understanding who is most at risk for hospitalization due to influenza. Our target population are those infected with influenza, and our sample might select a smaller set of individuals who are infected: not all of them. With a sample in hand, we can measure different characteristics of that set of collected individuals and claim that the estimates in our sample tell us something about the larger target population.
But how do we pick who is in our sample?
Simple random sampling¶
Approach
Simple random sampling (SRS) assigns the same probability to every individual in our target (often statistical) population. An individual from the population is picked and placed in the sample.
Pros A (true) simple random sample will provide you accurate measurements of any variables of interest. In addition, a SRS will also supply you with individuals who have other attributes that are representative of your target population. For example, in the influenza (flu) example above, if we use a SRS and assign the same probability of choosing individuals sick with the flu then if 60% of those who are sick with the flu are under the age of three then our sample will likely have 60% of individuals under three.
Cons That said, a true SRS isn’t usually feasible. We often never have access to every member of a population and so we can never really assign an equal probability to selecting everyone. In addition, because we assign the same probability to selecting an individual then it may be very difficult to have collected individuals with rare (but maybe important) attributes. For example, in the flu example, if a very small number of flu patients are sick because of an infection with an antibiotic resistant bacteria then we are likely to miss selecting these patients in our sample.
Stratified sampling¶
Approach
Given a population where every individual has a yet-to-be-measured variable of interest and a known attribute that can take one of potential values, select individuals with using a SRS; then select individuals with , and so on until you get to the group with , selecting individuals.
Pros A stratified sample allows you to equally represent, in your sample, different groups that are likely heterogeneous in nature across their groups but homogeneous within each group. Therefore, stratified sampling is good for when you want to understand your variable-of-interest among different groups, often called strata. This is particularly effective when there exist very small strata in the population that are important.
Cons That said, if you estimate your variable-of-interest across all strata then it will no longer be representative of the entire population.
Why?
Consider our variable of interest to be a preference for pepsi vs coca-cola in the University population. We decide to conduct a stratified sampling procedure across freshman, sophomores, juniors, and seniors, selecting 10 from each grade.
Then our overall estimate will assume that each grade is equally represented in our overall target population of the University. However, this is usually untrue. The freshman class is typically largest and represents 50% of the University. However in our stratified sample they are only represented at 10/40=25% of the entire population: they are under-represented. Likewise, seniors are usually the smallest group at 10% but our sample assumes that they are 25% of the population: they are over-represented.
Cluster sampling¶
Approach
Given a target population, variable-of-interest, , and attribute with (again) different categories, cluster sampling selects from the different categories and then samples all individuals within that category.
A typical setup for cluster sampling is spatial sampling. Our variable of interest may be the incidence of west-nile virus among students at Lehigh University. Rather than collecting a SRS of all 8,000 we will instead select 3 dorm residences at random and then collect our variable-of-interest from every single member who resides in those three dormitories.
Pros Cluster sampling can be cheap and convenient, and quick (especially if you need to travel to one of many locations for data collection).
Cons However, when you conduct a cluster sampling approach you typically underestimate the variability (i.e. how spread out) in your variable-of-interest . This is because individuals within a cluster tend to be more similar than those outside the cluster.
Classroom exercise:¶
The questions that we discussed in class:
Why does each group have a different estimate and which one is correct?
Where is the randomness and uncertainty “coming from”?
If we combined all our data would there still be uncertainty about another 30 students?
A main story-line for statistics is how to infer population level characteristics from a smaller, finite sample.
Fun blunders on sampling¶
https://
Homework / Exercises¶
Population or Sample?: Suppose you decide to characterize the number of days from diagnosis to all-cause mortality (death from any cause) of a rare disease. There exist just 100 (unfortunate) individuals with this disease. Your experiment follows all one hundred individuals for two years and all individuals experience death within this time period. What is the population for your experiment? What is the sample for your experiment? In this case, the sample happens to be all individuals. Then, is there any randomness left? Are you estimating the number of days from diagnosis to death or are you discovering the population parameters in this population? Please explain.
Run your own experiment: Please run your own experiment and collect at minimum ten data points that contain three pieces of information each. Organize this information: (a) create a folder on your local computer called “./Plots_and_probability/” and (b) a folder called “./Plots_and_probability/experiment/”. In the experiment folder, please
Add a file called “summary.txt” that outlines your experiment.
Add a file called “data_collection.csv” that includes the data that you collected with descriptive column names.
Add a file called “data_dictionary.txt”. This file will describe what every column means. What types of data are allowed in each column? Are there any irregularities or are there minimum/maximums associated with these columns?
What’s the population that your experiment is extracting data from?
Book problem (common cold): Please complete exercise 1.12 on page 78 of the book titled Introductory Statistics for the Biomedical and Life Sciences.
In the wild: Please read the 2010 paper titled “Transcatheter Aortic-Valve Implantation for Aortic Stenosis in Patients Who Cannot Undergo Surgery”. The link for this (free) article is here Leon et al. (2010). Answer the following questions about this statistical experiment.
Please describe the population that this experiment was meant to draw conclusions about
Briefly describe the core experiment (how many people were recruited and how? what happened after recruitment? what was measured about these patients? etc).
In Table 1, the authors present summary statistics. Can you comment on the purpose of presenting this table?
The plot of a story is the specific connected chain of events that lead from the beginning to the end of a story. In scientific papers, the plot is typically pretty straightforward: but not the narrative. The narrative is (usually) the approach for telling the plot. Please take one paragraph to read over this paper and describe elements that were persuasive, elements that felt rhetorical in nature, used to convince the reader, and please also add one thing that was difficult for you to grasp.
- Leon, M. B., Smith, C. R., Mack, M., Miller, D. C., Moses, J. W., Svensson, L. G., Tuzcu, E. M., Webb, J. G., Fontana, G. P., Makkar, R. R., Brown, D. L., Block, P. C., Guyton, R. A., Pichard, A. D., Bavaria, J. E., Herrmann, H. C., Douglas, P. S., Petersen, J. L., Akin, J. J., … Pocock, S. (2010). Transcatheter Aortic-Valve Implantation for Aortic Stenosis in Patients Who Cannot Undergo Surgery. New England Journal of Medicine, 363(17), 1597–1607. 10.1056/nejmoa1008232