Data Summarization: Quick summaries

Nicky Wakim

Quick summaries

  • A good first step when working with data is to explore the dataset, variables, and observations
    • Does every variable have the expected number of observations?
    • Are there any missing values?
    • Are there any weird values that don’t make sense?
  • We want to quickly browse the data and get summary statistics for each variable

 

  • There are a few functions in R that give you automatic summaries of your data:
    • summary() in base R
    • skim() in skimr
    • get_summary_stats() in rstatix

What does summary() do?

  • summary() will show:
    • For categorical data: the counts in each category
    • For numeric data: the minimum, 1st quartile, median, mean, 3rd quartile, maximum, and count for NA’s

 

  • No package needed: it’s built into base R, so it always worksks
  • Downside: no standard deviation, and the output is harder to read/scroll through for datasets with many columns
  • Often what I use when I’m feeling super lazy, but not the best for a first look at a whole dataset

What does summary() output look like?

Input is a data frame, output is a list of summary statistics for each variable

summary(hrs_00)
      hhid        pn               id        diab     
 559839 :   8   010:1944   Length   :2728   No :1993  
 559734 :   7   020: 784   N.unique :1717   Yes: 735  
 554789 :   7              N.blank  :   0             
 555578 :   7              Min.nchar:  10             
 557721 :   7              Max.nchar:  10             
 555053 :   6                                         
 (Other):2686                                         
                                 degree       birthyr        birthmo      
 High school diploma                :512   Min.   :1922   Min.   : 1.000  
 Master's degree                    :487   1st Qu.:1958   1st Qu.: 3.750  
 Associate's degree                 :411   Median :1967   Median : 7.000  
 None                               :384   Mean   :1963   Mean   : 6.631  
 Professional degree (Ph.D./M.D./JD):194   3rd Qu.:1970   3rd Qu.:10.000  
 (Other)                            :270   Max.   :1990   Max.   :12.000  
 NAs                                :470                                  
   birthdate               proxy                 coupled         sex      
 Min.   :-12314.0   Respondent:2721   Not a couple HH:1093   Female:1568  
 1st Qu.:  -480.8   Proxy     :   7   Couple HH      :1635   Male  :1160  
 Median :  2602.0                                                         
 Mean   :  1460.0                                                         
 3rd Qu.:  3757.0                                                         
 Max.   : 11215.0                                                         
                                                                          
     age_mo         age_yr             ed                     race_original 
 Min.   : 382   Min.   : 34.00   Min.   : 0    White/Caucasian       :1000  
 1st Qu.: 635   1st Qu.: 52.00   1st Qu.:12    Black/African American: 884  
 Median : 671   Median : 55.00   Median :13    Other                 : 358  
 Mean   : 711   Mean   : 58.82   Mean   :13    NAs                   : 486  
 3rd Qu.: 778   3rd Qu.: 64.00   3rd Qu.:16                                 
 Max.   :1200   Max.   :100.00   Max.   :17                                 
                                 NAs    :507                                
 smoke_ever smoke_now  drink          height             srh     
 No :2267   No :1484   No : 854   Min.   :1.041   Fair     :743  
 Yes: 459   Yes:1243   Yes:1874   1st Qu.:1.626   Good     :870  
 NAs:   2   NAs:   1              Median :1.676   Very Good:647  
                                  Mean   :1.689   Excellent:277  
                                  3rd Qu.:1.765   Poor     :187  
                                  Max.   :2.083   NAs      :  4  
                                  NAs    :55                     
         act_vig       bp       cancer      lung       hrt        strk     
 Never       :1443   No :1287   No :2457   No :2568   No :2355   No :2535  
 >1 per week : 586   Yes:1433   Yes: 271   Yes: 160   Yes: 373   Yes: 193  
 1 per week  : 270   NAs:   8                                              
 1-3 per week: 275                                                         
 Every day   : 145                                                         
 NAs         :   9                                                         
                                                                           
 psych      sleep       arth        cond_count         cesd      
 No :2188   Yes: 534   No :1722   Min.   :0.000   Min.   :0.000  
 Yes: 540   No :2194   Yes:1006   1st Qu.:1.000   1st Qu.:0.000  
                                  Median :2.000   Median :1.000  
                                  Mean   :1.733   Mean   :1.791  
                                  3rd Qu.:3.000   3rd Qu.:3.000  
                                  Max.   :7.000   Max.   :8.000  
                                                  NAs    :7      
     income       
 Min.   :      0  
 1st Qu.:  19852  
 Median :  51462  
 Mean   :  91813  
 3rd Qu.: 116500  
 Max.   :1800012  
                  

You can input a specific variable:

summary(hrs_00$age_yr)
1
Use $ to select a specific variable from a data frame
   Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
  34.00   52.00   55.00   58.82   64.00  100.00 

What does skim() do?

  • skim() will show a few tables:
    • Data summary: number of rows, number of columns, and types of variables
    • Variable type summaries:
      • For categorical data: the counts in each category, the number of unique values, and the most common value
      • For numeric data: the minimum, 1st quartile, median, mean, 3rd quartile, maximum, standard deviation, and count for NA’s

 

  • Best of the three for a first look at a whole dataset: nicely formatted, includes standard deviation, and separates variables by type automatically
  • Downside: the output is a special skim_df object, so it’s not as

What does skim() output look like?

skim(hrs_00)
Data summary
Name hrs_00
Number of rows 2728
Number of columns 32
_______________________
Column type frequency:
character 1
factor 21
numeric 10
________________________
Group variables None

Variable type: character

skim_variable n_missing complete_rate min max empty n_unique whitespace
id 0 1 10 10 0 1717 0

Variable type: factor

skim_variable n_missing complete_rate ordered n_unique top_counts
hhid 0 1.00 FALSE 1468 559: 8, 559: 7, 554: 7, 555: 7
pn 0 1.00 FALSE 2 010: 1944, 020: 784
diab 0 1.00 FALSE 2 No: 1993, Yes: 735
degree 470 0.83 FALSE 7 Hig: 512, Mas: 487, Ass: 411, Non: 384
proxy 0 1.00 FALSE 2 Res: 2721, Pro: 7
coupled 0 1.00 FALSE 2 Cou: 1635, Not: 1093
sex 0 1.00 FALSE 2 Fem: 1568, Mal: 1160
race_original 486 0.82 FALSE 3 Whi: 1000, Bla: 884, Oth: 358
smoke_ever 2 1.00 FALSE 2 No: 2267, Yes: 459
smoke_now 1 1.00 FALSE 2 No: 1484, Yes: 1243
drink 0 1.00 FALSE 2 Yes: 1874, No: 854
srh 4 1.00 FALSE 5 Goo: 870, Fai: 743, Ver: 647, Exc: 277
act_vig 9 1.00 FALSE 5 Nev: 1443, >1 : 586, 1-3: 275, 1 p: 270
bp 8 1.00 FALSE 2 Yes: 1433, No: 1287
cancer 0 1.00 FALSE 2 No: 2457, Yes: 271
lung 0 1.00 FALSE 2 No: 2568, Yes: 160
hrt 0 1.00 FALSE 2 No: 2355, Yes: 373
strk 0 1.00 FALSE 2 No: 2535, Yes: 193
psych 0 1.00 FALSE 2 No: 2188, Yes: 540
sleep 0 1.00 FALSE 2 No: 2194, Yes: 534
arth 0 1.00 FALSE 2 No: 1722, Yes: 1006

Variable type: numeric

skim_variable n_missing complete_rate mean sd p0 p25 p50 p75 p100 hist
birthyr 0 1.00 1963.48 9.41 1922.00 1958.00 1967.00 1970.00 1990.00 ▁▁▃▇▁
birthmo 0 1.00 6.63 3.48 1.00 3.75 7.00 10.00 12.00 ▇▅▅▆▇
birthdate 0 1.00 1460.04 3422.19 -12314.00 -480.75 2602.00 3757.00 11215.00 ▁▂▃▇▁
age_mo 0 1.00 711.05 113.91 382.00 635.00 671.00 778.00 1200.00 ▁▇▃▁▁
age_yr 0 1.00 58.82 9.50 34.00 52.00 55.00 64.00 100.00 ▁▇▃▁▁
ed 507 0.81 13.00 3.40 0.00 12.00 13.00 16.00 17.00 ▁▁▁▆▇
height 55 0.98 1.69 0.11 1.04 1.63 1.68 1.77 2.08 ▁▁▇▇▁
cond_count 0 1.00 1.73 1.41 0.00 1.00 2.00 3.00 7.00 ▇▃▅▁▁
cesd 7 1.00 1.79 2.13 0.00 0.00 1.00 3.00 8.00 ▇▂▁▁▁
income 0 1.00 91812.78 128588.35 0.00 19852.00 51461.87 116500.00 1800012.00 ▇▁▁▁▁

What does get_summary_stats() do?

  • get_summary_stats() will show a table of summary statistics for each numeric variable

 

  • Only works for numeric variables: categorical variables are silently dropped
  • Returns a data frame (a tibble), which makes it easy to use the results in further code, tables, or plots: this is its biggest advantage over summary() and skim()

What does get_summary_stats() output look like?

hrs_00 |>
  get_summary_stats()
# A tibble: 10 × 13
   variable       n      min    max median      q1     q3     iqr     mad   mean
   <fct>      <dbl>    <dbl>  <dbl>  <dbl>   <dbl>  <dbl>   <dbl>   <dbl>  <dbl>
 1 birthyr     2728   1.92e3 1.99e3 1.97e3  1.96e3 1.97e3 1.2 e+1 5.93e+0 1.96e3
 2 birthmo     2728   1   e0 1.2 e1 7   e0  3.75e0 1   e1 6.25e+0 4.45e+0 6.63e0
 3 birthdate   2728  -1.23e4 1.12e4 2.60e3 -4.81e2 3.76e3 4.24e+3 2.26e+3 1.46e3
 4 age_mo      2728   3.82e2 1.2 e3 6.71e2  6.35e2 7.78e2 1.43e+2 7.56e+1 7.11e2
 5 age_yr      2728   3.4 e1 1   e2 5.5 e1  5.2 e1 6.4 e1 1.2 e+1 5.93e+0 5.88e1
 6 ed          2221   0      1.7 e1 1.3 e1  1.2 e1 1.6 e1 4   e+0 2.96e+0 1.30e1
 7 height      2673   1.04e0 2.08e0 1.68e0  1.63e0 1.76e0 1.4 e-1 1.13e-1 1.69e0
 8 cond_count  2728   0      7   e0 2   e0  1   e0 3   e0 2   e+0 1.48e+0 1.73e0
 9 cesd        2721   0      8   e0 1   e0  0      3   e0 3   e+0 1.48e+0 1.79e0
10 income      2728   0      1.80e6 5.15e4  1.99e4 1.16e5 9.66e+4 5.58e+4 9.18e4
# ℹ 3 more variables: sd <dbl>, se <dbl>, ci <dbl>

Flexibility of get_summary_stats(): show certain variables

  • You can use the same syntax as the select() in dplyr to keep or remove specific columns
    • Use variable_name to keep and -variable_name to remove
hrs_00 |>
  get_summary_stats(-birthmo)
# A tibble: 9 × 13
  variable       n       min    max median      q1     q3     iqr     mad   mean
  <fct>      <dbl>     <dbl>  <dbl>  <dbl>   <dbl>  <dbl>   <dbl>   <dbl>  <dbl>
1 birthyr     2728   1922    1.99e3 1.97e3  1.96e3 1.97e3 1.2 e+1 5.93e+0 1.96e3
2 birthdate   2728 -12314    1.12e4 2.60e3 -4.81e2 3.76e3 4.24e+3 2.26e+3 1.46e3
3 age_mo      2728    382    1.2 e3 6.71e2  6.35e2 7.78e2 1.43e+2 7.56e+1 7.11e2
4 age_yr      2728     34    1   e2 5.5 e1  5.2 e1 6.4 e1 1.2 e+1 5.93e+0 5.88e1
5 ed          2221      0    1.7 e1 1.3 e1  1.2 e1 1.6 e1 4   e+0 2.96e+0 1.30e1
6 height      2673      1.04 2.08e0 1.68e0  1.63e0 1.76e0 1.4 e-1 1.13e-1 1.69e0
7 cond_count  2728      0    7   e0 2   e0  1   e0 3   e0 2   e+0 1.48e+0 1.73e0
8 cesd        2721      0    8   e0 1   e0  0      3   e0 3   e+0 1.48e+0 1.79e0
9 income      2728      0    1.80e6 5.15e4  1.99e4 1.16e5 9.66e+4 5.58e+4 9.18e4
# ℹ 3 more variables: sd <dbl>, se <dbl>, ci <dbl>

Flexibility of get_summary_stats(): show certain summary statistics

  • Use type = to specify which summary statistics you want to see
    • Options include: "common", "full", "mean_sd", "mean_ci", "median_iqr", and more (see ?get_summary_stats for the full list)
hrs_00 |>
  get_summary_stats(type = "common")
# A tibble: 10 × 10
   variable       n     min    max median     iqr   mean      sd      se      ci
   <fct>      <dbl>   <dbl>  <dbl>  <dbl>   <dbl>  <dbl>   <dbl>   <dbl>   <dbl>
 1 birthyr     2728  1.92e3 1.99e3 1.97e3 1.2 e+1 1.96e3 9.40e+0 1.8 e-1 3.53e-1
 2 birthmo     2728  1   e0 1.2 e1 7   e0 6.25e+0 6.63e0 3.48e+0 6.7 e-2 1.31e-1
 3 birthdate   2728 -1.23e4 1.12e4 2.60e3 4.24e+3 1.46e3 3.42e+3 6.55e+1 1.28e+2
 4 age_mo      2728  3.82e2 1.2 e3 6.71e2 1.43e+2 7.11e2 1.14e+2 2.18e+0 4.28e+0
 5 age_yr      2728  3.4 e1 1   e2 5.5 e1 1.2 e+1 5.88e1 9.5 e+0 1.82e-1 3.57e-1
 6 ed          2221  0      1.7 e1 1.3 e1 4   e+0 1.30e1 3.4 e+0 7.2 e-2 1.41e-1
 7 height      2673  1.04e0 2.08e0 1.68e0 1.4 e-1 1.69e0 1.09e-1 2   e-3 4   e-3
 8 cond_count  2728  0      7   e0 2   e0 2   e+0 1.73e0 1.41e+0 2.7 e-2 5.3 e-2
 9 cesd        2721  0      8   e0 1   e0 3   e+0 1.79e0 2.13e+0 4.1 e-2 8   e-2
10 income      2728  0      1.80e6 5.15e4 9.66e+4 9.18e4 1.29e+5 2.46e+3 4.83e+3

Which one should I use?

Function Best for Watch out for
summary() -A fast, no-setup check of a small dataset

-No standard deviation

-Output can be hard to read for many columns

skim() -A first, fairly thorough look at a whole dataset (categorical + numeric)

-Requires skimr

-Output isn’t a plain data frame

get_summary_stats()

-Numeric summaries you want to reuse (e.g., in a table or plot)

-Grouped summaries (to be covered later)

-Requires rstatix

-Numeric variables only

 

  • In practice, you’ll often start with skim() to explore
  • Then use get_summary_stats() if you have certain numeric variables and statistics you actually need