DevOps

Follow one request: a beginner log-investigation exercise

By SPOTHUB · · 4 min read

Prepared with AI assistance and linked primary sources. Examples are illustrative unless stated otherwise.

Start with a reproducible symptom, then follow the same request identifier through a small sequence of events. Record where the expected sequence stops, compare it with a successful request and test one explanation. Logs help establish what was recorded; they do not automatically prove why a failure happened.

Turn a vague failure into a concrete question

This evergreen observability lesson fills the October 8 slot and was first published on October 9, 2026. The example is fictional: a practice enquiry page reports a temporary submission error. Use synthetic records in a local exercise. No real customer enquiry, contact detail or production incident is required to learn the investigation method.

Write the question before looking through output: did this request fail during input validation, storage or notification? Note the approximate time and the visible symptom. Avoid beginning with an assumed cause such as the database is down. A useful investigation narrows uncertainty step by step and keeps observations separate from explanations.

Give events a consistent shape

OpenTelemetry's logging documentation distinguishes a consistent schema from merely encoding arbitrary text as JSON. For our practice records, use timestamp, level, request_id, event, outcome and duration_ms with agreed meanings and types. Keep duration_ms numeric, timestamps in one documented timezone, and the request identifier stable across events for the same operation.

Choose event names such as enquiry.received, enquiry.validated and enquiry.storage_completed. Explain which event indicates completion and whether an error record names a component or an operation. A field named status with several incompatible meanings makes filtering less reliable. The exercise needs only a few explicit fields, not a dump of every variable in memory.

Source: OpenTelemetry documentation: Logs

Build two short fictional timelines

Create request demo-A with received at 10:00:00.000Z, validated at 10:00:00.010Z, storage_completed at 10:00:00.040Z and response_sent at 10:00:00.045Z with outcome success. Then create request demo-B with received at 10:00:01.000Z, validated at 10:00:01.010Z and storage_failed at 10:00:03.010Z with outcome timeout.

Add a response_sent event for demo-B at 10:00:03.015Z with outcome temporary_error. Mix the two requests' records together in one sample file. Your first task is to recover each sequence by request_id and timestamp. Keep the successful and failed timelines side by side rather than reading the mixed file as if every neighbouring line belonged to one user.

State what the evidence supports

For demo-B, the sample supports that validation completed and a storage attempt recorded a timeout before the application returned a temporary error. It does not establish whether the cause was a connection problem, lock contention, an overloaded dependency or a deliberately short timeout. Put those possibilities in a hypotheses column rather than in the conclusion.

The example also does not prove whether a timed-out write eventually committed. That distinction matters before suggesting a retry that might duplicate an enquiry. In a real system, check persistence and the operation's retry design with the responsible owner. In this lab, record the uncertainty and define the extra event or test you would need to resolve it.

Use traces when the operation crosses components

OpenTelemetry describes traces as a way to follow a request's path through an application, with spans representing units of work. If a practice request involves an API and another service, a trace can help relate their operations. A hand-written request_id in this exercise is a correlation convention; it is not automatically a complete distributed trace.

For an extension, sketch an API operation containing validation and storage spans. Mark the interval associated with the observed timeout and identify which detail remains missing. You do not need to install a monitoring platform to draw this model. The goal is to decide what evidence would help, before adding instrumentation that produces more output without answering the question.

Source: OpenTelemetry documentation: Traces

Finish with a bounded incident note

Write a short note containing the symptom, the selected request, the observed event sequence, the current hypothesis and a proposed verification step. If you later run a controlled fault-injection experiment, clearly label the failure as intentional. Repeat the successful path after a change so that a fix does not merely replace one failure with another.

Keep contact details, passwords, access tokens and unnecessary payloads out of shared logs. Choose synthetic examples for portfolio work. This exercise demonstrates investigation and communication, not a production reliability certification. In the related Cloud + DevOps path, extend it with documented retention, access controls and a test that verifies expected events for both success and failure.

Sources and further reading

Spot an error? Email info@spothub.in with the article link and correction.

← All articles
Find My IT Career Path