Loading…
Card 0/28
28 cards
Keep studying on Mneva
You’ve explored three public decks. Create a free account to keep studying unlimited cards and save your progress.
Free forever. No credit card needed.
Frequency table
A frequency table lists each category and the number of observations in it. The frequencies must sum to the total number of observations, $n$.
Frequency
Frequency is the number of times a particular category or data value occurs in a data set.
How is a frequency table constructed from raw categorical data?
List each distinct category, then count how many observations fall in each category. Check the table by verifying that all frequencies sum to $n$.
How should a frequency table be checked for internal consistency?
The frequency total should equal the sample size $n$, and the relative-frequency total should be approximately $1$. If cumulative relative frequencies are used, they should increase down the table and end near $1$.
Relative frequency
Relative frequency is the proportion of observations in a category: $\text{relative frequency}=f/n$, where $f$ is the category frequency and $n$ is the total number of observations. It may be expressed as a fraction, decimal, or percent.
How is a relative frequency found for a category containing 40 of 100 observations?
Use $f/n=40/100=0.40$, which is $40\%$. The same calculation applies to any category.
What should the relative frequencies in a complete frequency table sum to?
They should sum to $1$, or equivalently $100\%$. Small deviations can occur if individual decimal relative frequencies were rounded.
Cumulative relative frequency
Cumulative relative frequency is the sum of the relative frequencies up to and including a given row. It represents the proportion of observations in that category or any preceding category, when the categories have a meaningful order.
What does the final cumulative relative frequency represent?
The final cumulative relative frequency represents all observations, so it should equal $1.00$ or $100\%$, apart from rounding error.
Bar chart
A bar chart displays categorical data with rectangular bars whose lengths or heights are proportional to the represented values. One axis lists categories, while the other shows frequency, relative frequency, or another measured quantity.
When can categories in a bar chart be arranged in any order?
If the categories have no natural order, such as smartphone brands or favorite foods, they can be arranged in any order. If categories have an inherent order, preserving it usually improves interpretation.
How does a bar chart differ from a histogram?
A bar chart is designed for discrete categories and generally has separated bars whose order may be arbitrary. A histogram represents quantitative data grouped into numerical intervals, so adjacent bars typically touch.
Pareto chart
A Pareto chart is a bar chart whose categories are ordered from greatest to least frequency or incidence. This arrangement makes the most common categories immediately visible.
Why can truncating the baseline of a bar chart be misleading?
Bar lengths are visually compared, so starting the axis above zero can exaggerate differences between similar values. A zero baseline is generally appropriate unless a justified scale, such as temperature or a logarithmic axis, makes it unsuitable.
Grouped (clustered) bar chart
A grouped bar chart places two or more color-coded bars beside one another for each category. It is used to compare multiple groups or measured variables across the same categories.
Stacked bar chart
A stacked bar chart divides each bar into segments representing subcategories, so the total bar length shows the combined value. It also shows each subcategory's contribution to the total.
Pie chart
A pie chart displays the relative frequencies of categories as sectors of a circle. Each sector's area and central angle are proportional to the category's share of the whole.
How is the central angle of a pie-chart sector calculated?
Multiply the category's relative frequency by $360^\circ$: $\text{angle}=\text{relative frequency}\times360^\circ$. Equivalently, use $\text{angle}=(f/n)360^\circ$.
What must be true of the categories represented in a pie chart?
The categories should be mutually exclusive and collectively exhaustive, so each observation belongs to exactly one category and the sectors represent the entire sample.
How does a pie chart differ from a bar chart?
Both can display categorical frequencies or relative frequencies. A pie chart emphasizes parts of a whole through sector sizes, whereas a bar chart makes comparisons among category values easier.
Two-way table
A two-way table displays counts for two categorical variables. The rows represent the categories of one variable, the columns represent the categories of the other, and each interior cell gives the count for a combination of categories.
Joint frequency
A joint frequency is the count in an interior cell of a two-way table. It represents observations having a particular category of the row variable and a particular category of the column variable.
Marginal frequency
A marginal frequency is a row or column total in a two-way table. It gives the frequency for one variable without regard to the category of the other variable.
What does the grand total in a two-way table represent?
The grand total is the sum of all interior cell counts and represents the total number of observations in the data set.
How is a joint relative frequency calculated from a two-way table?
Divide the relevant cell count by the grand total: $\text{joint relative frequency}=\text{cell count}/n$. It represents the proportion of all observations in that combination of categories.
How is a conditional relative frequency calculated from a two-way table?
Divide a cell count by the appropriate row or column total, depending on the condition. For example, the proportion in column category $B$ among row category $A$ is $\text{count}(A\text{ and }B)/\text{row total for }A$.
What is the difference between row and column relative frequencies in a two-way table?
Row relative frequencies use each row total as the denominator, so each row sums to $1$. Column relative frequencies use each column total as the denominator, so each column sums to $1$.
How can a two-way table be used to compare categorical distributions?
Calculate conditional relative frequencies within the relevant groups and compare the resulting proportions. Differences between the conditional distributions indicate that the variables may be associated.
Free forever. No credit card needed.
Ready to study AP Statistics 1.3-1.4: Representing Categorical Data?
Free forever. No credit card needed.