Missing is not zero: a small pandas support-time lab
Prepared with AI assistance and linked primary sources. Examples are illustrative unless stated otherwise.
A missing measurement means you do not know the value; zero is a recorded value. Before calculating an average, count the rows, count the available measurements and explain exclusions. A six-ticket example makes the difference visible without a large dataset or a complicated dashboard.
Ask a question before choosing a chart
This evergreen lesson fills the October 3 learning slot and was first published on October 5, 2026. It is a fundamentals exercise using current pandas documentation, not an announcement of a new analytics feature. All ticket records below are fictional and should not be interpreted as SpotHub support performance.
The question is: what was the average recorded first-response time for this small sample? Keep the word recorded in your notes. It prevents the result from silently becoming a claim about every ticket. Decide that the unit is minutes and that the sample contains exactly six tickets before calculating or comparing anything.
Create six rows and predict the answer
Use ticket IDs T1 through T6 with response_minutes values 4, 8, missing, 0, 12 and missing respectively. In a Python environment with pandas installed, create a DataFrame from those two lists, using None for the two missing values. Print the table and data types so the input is visible alongside the analysis.
Work through the expected result by hand: the four recorded values total 24 minutes, so their average is 6 minutes. There are six tickets, four recorded measurements and two missing measurements. If you replace the missing entries with zero, the same total divided by six becomes 4 minutes. That lower figure answers a different question.
Inspect what pandas is counting
The pandas summary-statistics tutorial demonstrates column aggregations, while the missing-data guide explains isna and notna. In your notebook, compare len(df), df['response_minutes'].count(), df['response_minutes'].isna().sum() and df['response_minutes'].mean(). For this numeric example, expect 6, 4, 2 and 6.0 respectively.
Place the observed output next to the manual calculation. If your result differs, inspect the input values and their types before changing the expected answer. For instance, a literal word such as 'missing' is not the same input as None. Keep the original dataset intact while investigating, so you can explain what changed and why.
Source: pandas tutorial: Calculate summary statisticspandas guide: Working with missing data
Compare a tempting but unsupported cleanup
Create a separate calculation using df['response_minutes'].fillna(0).mean(). It should return 4.0 for this example. Present it as a deliberate what-if experiment, not a correction. You have assumed the two unrecorded responses took no time. Nothing in the sample supports that assumption, even though the code runs and produces a neat number.
Keep T4's recorded zero in the main calculation unless the data owner explains that zero is a placeholder in this system. Ask what the field means, when it is populated, and whether an unanswered ticket remains blank. Different answers may call for different metrics. An analyst should resolve the definition before deciding which rows to remove or replace.
Write the result with its denominator
A defensible summary is: among four tickets with recorded response times, the mean was six minutes; two of the six tickets had no recorded time. Include the sample definition and unit next to the result. Avoid a headline that simply says the support team responds in six minutes, because this dataset does not establish that claim.
As a second check, calculate the median of the recorded values: sorted values are 0, 4, 8 and 12, so the median is also 6. Matching mean and median here does not establish that the sample is representative. Add a much slower fictional seventh ticket as an extension and explain how both summaries change.
Turn the notebook into useful evidence
Save the six input rows, the handwritten expectation, the pandas commands, observed outputs and your interpretation. Record the Python and pandas versions used, because reproducibility includes the environment. A reviewer should be able to rerun the notebook and distinguish the main analysis from the intentionally misleading zero-fill comparison without asking you to reconstruct the steps.
This tiny sample cannot reveal a team's service quality, explain why measurements are absent, or justify a staffing decision. State those limits, then propose collecting the missingness reason and a larger defined sample. The related Data Analytics + AI path can help you extend the exercise into a documented analysis where data definitions matter as much as chart design.
Sources and further reading
- pandas tutorial: Calculate summary statistics · checked 2026-10-05
- pandas guide: Working with missing data · checked 2026-10-05
Spot an error? Email info@spothub.in with the article link and correction.
