Back to list

Blog

Debugging Production Issues from Your Phone

When production breaks at dinner time, your phone becomes your debugging interface. Here's how to diagnose and fix issues without rushing back to your laptop.

Published Tags: Getting started / debugging / production / incident-response
Blog

Follow product and engineering updates from this channel.

Browse category

It's 7 PM on a Friday. You're at dinner with friends when your phone buzzes with a PagerDuty alert: error rate on your payment service has spiked to 15%. Your laptop is at home. Your teammates are offline. The incident is yours to handle.

This scenario is why Tactic Remote exists. This tutorial walks through a realistic production debugging workflow using only your iPhone, demonstrating how to investigate, diagnose, and fix production issues remotely.

Scenario Setup

For this tutorial, we'll work through a realistic scenario:

  • Service: Node.js payment processing API
  • Symptom: 500 errors on the /api/payments/process endpoint
  • Impact: 15% of payment attempts failing
  • Environment: Production logs available via SSH, codebase on your Mac

Your Mac is on at home, running the Tactic Remote companion app. You connect from your iPhone and begin the investigation.

Phase 1: Triage (5 minutes)

The first goal is understanding the scope and urgency. Start a Claude Code session pointed at your project directory:

We have a production incident. The /api/payments/process endpoint
is returning 500 errors for ~15% of requests. This started approximately
30 minutes ago.

Help me investigate:
1. Check the git log for any recent deployments (last 24 hours)
2. Look at the payment processing code in src/api/payments/
3. Check if there were any dependency updates recently

Claude Code will examine the code and recent changes. While it works, open your monitoring dashboard in your phone's browser to check:

  • Is the error rate stable, increasing, or decreasing?
  • Are all instances affected or just some?
  • Is there a pattern in the failing requests?

What Claude Code might find: A commit from 3 hours ago updated the Stripe SDK dependency from v12.1.0 to v13.0.0, which included breaking API changes.

Phase 2: Diagnosis (10 minutes)

With a hypothesis (the Stripe SDK update), dig deeper:

Compare the Stripe SDK v12 and v13 changelog.
Check our payment processing code for any API calls that changed
between versions. Focus on the payment intent creation and
confirmation flow in src/api/payments/processor.ts.

Claude Code will analyze the code and identify specific incompatibilities. A typical finding:

Found the issue. In Stripe SDK v13, the `payment_intents.create()` method
changed its error handling behavior:

- v12: Returns { error: { type, message } } for card declines
- v13: Throws StripeCardError exception for card declines

Our code in processor.ts:45-52 checks for `result.error` but doesn't
catch the new exception type. When a card is declined, the uncaught
exception propagates as a 500 error.

This explains the 15% error rate — approximately 15% of payment
attempts involve card declines (expired cards, insufficient funds, etc.)
that now throw uncaught exceptions.

This diagnosis is specific, testable, and explains the observed behavior. From your phone, you can verify it makes sense before proceeding.

Phase 3: Fix (10 minutes)

Now implement the fix. You have two options:

Option A: Quick Rollback

If the situation is urgent and the previous version was stable:

Revert the Stripe SDK update. Change package.json to pin
stripe at v12.1.0. Create a commit with message:
"hotfix: revert Stripe SDK to v12 (fixes payment 500 errors)"

Approval step: Claude Code will request approval to modify package.json and run npm install. Review the changes in the approval card and approve.

Option B: Forward Fix

If you want to fix the code to work with v13 (avoiding a revert that loses other improvements):

Fix the payment processing code to handle both the legacy error
response format and the new exception-based error handling in
Stripe SDK v13.

Specifically:
1. Wrap the payment_intents.create() call in a try/catch
2. Catch StripeCardError and convert it to our standard error format
3. Keep the existing result.error check as a fallback
4. Add a test case for the card decline scenario

Commit with: "hotfix: handle Stripe v13 exception-based card decline errors"

Claude Code will make the changes and run the test. Watch the terminal view for the test results.

Approval flow: You'll likely see multiple approval requests:

  1. Modifying src/api/payments/processor.ts — Approve (this is the fix)
  2. Modifying src/api/payments/__tests__/processor.test.ts — Approve (adding test)
  3. Running npm test — Approve (verifying the fix)

Phase 4: Deploy (5 minutes)

With the fix committed and tested:

Push this commit to the main branch.
Then trigger the deployment pipeline by running our deploy script:
./scripts/deploy.sh production

Critical approval: Claude Code will request approval for the deploy command. This is the highest-risk action in the workflow. Review carefully:

  • Verify the commit message matches your fix
  • Confirm the deploy target is correct (production)
  • Ensure tests passed in the previous step

After approval, Claude Code executes the deployment. Monitor the terminal output for deployment progress.

Phase 5: Verify (5 minutes)

After deployment completes:

Check the deployment status. Then help me verify:
1. Run a test request against the production payment endpoint
2. Check if the error rate has started decreasing
3. Look for any new error patterns in the last 5 minutes

Simultaneously, check your monitoring dashboard on your phone. The error rate should begin dropping within minutes of the deploy completing.

Post-Incident: Write the Report

While the incident is fresh in your mind (still at dinner, between courses):

Write a post-incident report based on our investigation:

Include:
- Timeline of events
- Root cause (Stripe SDK v13 breaking change)
- Impact (15% payment failure rate for ~1 hour)
- Resolution (code fix to handle new exception type)
- Action items to prevent recurrence

Format as a markdown document in docs/incidents/

This gives your team a complete incident report before Monday morning. Your dinner companions barely noticed the interruption.

Tips for Effective Remote Debugging

Start broad, then narrow. Don't jump to hypotheses immediately. Let Claude Code survey the landscape (recent changes, error patterns, system state) before focusing on a specific cause.

Use approval gates as checkpoints. Each approval request is a natural checkpoint where you review what Claude Code is doing. Use these to verify the investigation is on the right track.

Separate diagnosis from fix. Get a clear diagnosis and confirm it makes sense before implementing any changes. This prevents "fix chasing" where you make changes without understanding the root cause.

Communicate while debugging. If your team has an incident channel, post updates from your phone:

  • "Investigating payment 500 errors. Suspect Stripe SDK update."
  • "Confirmed: SDK v13 changed error handling for card declines. Deploying fix."
  • "Fix deployed. Error rate returning to normal."

This keeps stakeholders informed and prevents duplicate investigation.

Know your limits. Some production issues require infrastructure access, database queries, or system-level debugging that can't be done through Claude Code. If you hit a wall, escalate — don't spend an hour from your phone on something that needs 5 minutes at a workstation.

Building Your Incident Toolkit

Prepare for incidents before they happen:

  1. Pre-stage a production debugging session with access to your deployment scripts and monitoring tools.
  2. Create incident response prompts in your prompt library that Claude Code can follow.
  3. Test the workflow during a non-critical situation so you're familiar with the flow when pressure is real.
  4. Ensure Cloudflare Tunnel access so you can debug from anywhere, not just your home network.

Production incidents don't wait for convenient timing. Tactic Remote ensures that wherever you are, you have the tools to respond effectively.

Try Tactic Remote

Control your coding Agents from your phone

Connect to Claude Code, Codex, and other Agents on your Mac, Windows, or Linux computer. Check progress and send the next instruction from iPhone or iPad.