Questions across responsibility, technical craft, culture, and dilemmas. Click any question to expand the answer.
| Role | Responsibility | Meetings |
|---|---|---|
| Product Manager | Defines acceptance criteria, signs off on what "done" means | Review specs before development and upon changes that require redefinition. |
| Designer (UX/UI) | Provides expected behavior and edge case intent for UI flows | Design review - catch ambiguities before development begins and upon changes. |
| BE & FE Developers | Implements functionality, owns unit tests, shares integration coverage | During sprint - understand API contracts, dependencies, risk areas; bug triage; clarifying bug steps; reviewing testability. |
| DevOps / SRE | Owns test environments, pipelines, observability | CI/CD setup - align on test gates, environment stability, and monitoring. |
| QA | Designs, executes, and maintains test strategy for functional, regression, and non-functional tests | All phases (daily) - sync on task status per member. |
| Suite Type | Purpose | When Used |
|---|---|---|
| Unit Tests | Validate individual functions/modules in isolation | Continuously, during development |
| Integration Tests | Validate that components interact correctly | After unit testing, before E2E |
| End-to-End Tests | Simulate full user flows through the live system | Pre-release and on schedule |
| Regression Tests | Ensure existing functionality isn't broken by new changes | Every release cycle |
| Smoke Tests | Quick pass/fail check that a build is stable | After every deployment |
| Performance/Load Tests | Measure response time, throughput, stability under load | Before major releases or infrastructure changes |
| Security Tests | Identify vulnerabilities - injection, auth bypass, exposure | Pre-release and on schedule |
| Acceptance Tests | Validate the product meets business requirements | Pre-release, often with stakeholders |
| Visual Regression Tests | Catch unintended UI changes via screenshot diffs | After UI-touching changes |
| Exploratory Testing | Unscripted investigation guided by tester intuition | Throughout, especially before major releases |
| Category | Tools |
|---|---|
| UI/E2E & Cross-Browser | Playwright, Cypress, Selenium |
| API | Postman, pytest |
| Mobile | Appium |
| Real device (cloud) | BrowserStack |
| Performance/Load | k6, JMeter, Locust |
| Unit/Integration | pytest |
| Execution & CI | GitLab CI, GitHub Actions, Azure Pipelines |
| Reporting | Allure, Grafana |
| Alerting | Slack |
| Security | OWASP ZAP, manual boundary testing |
| DB | SQL queries to validate data integrity |
| Accessibility | axe DevTools, Lighthouse |
| AI-assisted | Cursor, ChatGPT, Claude (workflow-integrated via Skills and MCPs: test case generation from specs, log analysis, coverage gap analysis) |
| LLM Evaluation | LangSmith |
Best fit to company tech stack from TestRail, Qase, Zephyr, Xray, Tuskr or other compatible test management system.
data-testid attributes, well-documented versioned APIBrowser DevTools, Wireshark, Charles Proxy
| Tool | Purpose |
|---|---|
| Chrome DevTools | DOM inspection, network, console errors, performance profiling, storage |
| Playwright / Cypress | Trace viewer, step-through test replay, screenshot on failure |
| Postman / pytest | Isolate and replay API calls independent of the UI |
| Allure | Structured test reports with failure context and history |
| Sentry / Datadog | Production error tracking, stack traces, user session replay |
| Grafana | Dashboards for error rates, latency, and anomaly detection |
| Lighthouse | Performance, accessibility, and SEO audit in one report |
| axe DevTools | Accessibility violation scanner |
| BrowserStack | Real device/browser testing without local setup |
| Bug Class | Description & How to Test |
|---|---|
| Race conditions | Two operations close in time produce unexpected state. Test by: simulate concurrent users hitting the same endpoint simultaneously using k6, targeting shared resources like inventory counts, balances, booking slots. |
| Off-by-one errors | Boundaries behave wrong at the exact limit. Test by: always test at N-1, N, and N+1 for any numerical boundary - pagination, character caps, retry counts, date ranges. |
| Timezone and DST bugs | Logic that works in one timezone silently breaks in another. Test by: run with server and client in different timezones, around DST transitions, end of month/year, and midnight boundaries. |
| Floating point precision | Decimal arithmetic produces imperceptibly wrong results that compound. Test by: financial/measurement calculations with values that can't be represented in binary (0.1 + 0.2, currency rounding at scale). |
| State mutation bugs | A shared object is modified in place and affects unrelated parts. Test by: run operations in sequence and verify earlier operations don't contaminate later ones - caching layers and session state. |
| Silent failures | An operation fails but returns 200 with no error signal. Test by: inject failures at the dependency level (DB down, third-party timeout) and verify the response accurately reflects the failure. |
| Permission escalation | Lower-permission roles access higher-permission resources via direct URL or modified API. Test by: authenticate as each role and attempt every protected endpoint directly; manipulate resource IDs in API calls. |
| Locale and encoding | Non-ASCII characters break display, storage, or sort order. Test by: use "Ångström", "日本語", "Ñoño", Arabic and Hebrew RTL text; verify round-trip integrity from input to storage to display. |
| Eventual consistency lag | Read immediately after write returns stale data. Test by: add read-after-write assertions, run tests with artificial replication delay, verify system returns consistent data or signals pending. |
| Cascading default values | Missing/null value silently replaced by a default at multiple layers. Test by: deliberately omit optional fields and trace what value actually reaches the database. |
| Silent data truncation | Input longer than DB column length silently cut with no error. Test by: submit at max length, max+1, and with multi-byte characters (emoji, CJK). |
| Idempotency violations | Same operation twice creates duplicate records or double charge. Test by: rapid double-click, simulate retries with network throttling - verify exactly one outcome regardless of how many times the request arrives. |
I'm in requirement reviews asking "how will we test this?" before implementation starts. I'm in design conversations catching edge cases before they're built in. By the time code reaches formal testing, most of the ambiguity is already resolved.
Developers own unit tests for their own code. QA owns the E2E and regression suites that validate the system as users experience it. Integration coverage sits in between and belongs to whoever has the most context. If developers aren't writing meaningful unit tests, my E2E suite ends up doing work it shouldn't have to.
Quality becomes shared when the definition of done includes it. Done means tested, observable, and ready to ship - not "dev complete and thrown over the wall." That single shared definition changes how developers write tickets, how they hand off features, and how they respond to bug reports.
Bug triage is the clearest test of the relationship. In a good team, it's a joint conversation: QA brings evidence and impact, dev brings implementation context, PM brings priority. Nobody is defending territory. In a broken team, bugs are accusations and triage is negotiation.
The goal isn't a QA department that enforces quality. It's a team where quality is built into how everyone works, and QA is the function that keeps the signal clear.
data-testid attributes.| Metric | What It Measures |
|---|---|
| Defect escape rate | Bugs found in production vs. caught pre-release - the primary signal of QA effectiveness |
| Test automation coverage | % of regression suite automated - tracks progress and identifies manual testing debt |
| Test cycle time | How long a full regression run takes - slower cycles mean slower releases |
| Mean time to detection | How quickly QA catches a bug after introduction - lower is better |
| Defect reopen rate | Measures bug report quality and fix verification thoroughness |
| Test flakiness rate | % of test runs that fail non-deterministically - high flakiness erodes trust in the suite |
The goal is to make sure it's a real decision, not just deadline pressure defaulting into the assumption that QA is optional when things get tight.
Either way: document what I found, when I escalated, who made the call, and what the stated trade-off was. If it ships and breaks, that record matters.
I weigh three things: how bad it is when it does hit, the real user volume behind that rate (1/50 on 50k weekly users is 1,000 people affected - not a rounding error), and how reproducible it is. A bug I can't reliably reproduce is one I can't confidently fix or regression-test.
Works as designed. But I push back if:
None of these are bugs against the written spec. The design is just incomplete or unsafe. My job is to surface those gaps and advocate for the user even when the spec doesn't.
This is the core challenge of QA for AI systems. A UI either renders correctly or it doesn't. An API either returns 200 or it doesn't. But when the output is a generated summary, a recommended action, or a test case drafted by an agent, "correct" isn't binary.
The frame I use: turn subjective quality into a measurable loop. Collect examples of outputs that a domain expert marks good and bad, then define criteria that don't ask "is this right?" but "does this stay within an acceptable range across runs? Does it degrade on edge cases?"