← Back to the field notes

PLANNING APP / REVIEW

What a passing checklist missed.

A fresh review found four practical problems the first checklist had missed. The interesting part wasn't the score. It was what each problem could mean in everyday use.

AT THE TIME: TRIED, FIXED, CHECKED AGAINBy App Wizzud

A passing checklist wasn't the same as finished.

This was a planning app used by a small business. Before I called the update done, every existing automated check was passing. Those checks were useful, but they didn't cover every situation.

A separate AI reviewer started fresh, tried a few less obvious things, and found four practical problems.

One passing test answers one question. It doesn't mean you asked every important question.

Finding a defect was only half the job.

The reviewer could show that a save had failed or that an unusual value produced the wrong result. It couldn't decide what the app should do next. That depended on what kind of mistake would be most dangerous in the real workflow.

For this planner, a convincing schedule that was never safely saved was worse than no schedule at all. The app therefore had to stop, preserve the existing file, and make recovery explicit. The same rule applied to questionable quantities and dates: reject them instead of quietly turning them into something plausible. Those were decisions about trust, not merely code fixes.

The four things we missed.

The reviewer used safe copies, never the live planner. All four problems were in the earlier version and were fixed later.

01 / A failed save could still make the app look ready.

If the app couldn't save its first planner, it still showed starter information as if everything were fine. Showing an error wasn't enough; the app also had to stop pretending it was ready.

02 / A file arriving at exactly the wrong moment could be replaced.

If a real planner appeared while the app was preparing its first file, the app could overwrite it. Weird timing, serious result.

03 / An invisible character could turn “1000” into “1.”

A hidden character inside a number made the app accept only the digits before it. The input looked almost normal, but the answer was wrong.

04 / Some unusual dates were judged incorrectly.

One impossible date got through, while another perfectly valid date was rejected. These were odd examples, not proof that everyday schedules had already been harmed.

One round of fixes wasn't enough.

The first changes addressed all four problems. Then the fresh reviewer looked again and found one smaller rule that was still too loose.

We tightened that rule and checked the whole path again. The important part wasn't the final number. The app got fixed, questioned again, and improved one more time.

Then we tried it with the real planner.

After making a verified backup, I opened the real planner, checked a few familiar things, closed the app, and opened it again. The planner file was exactly the same before and after.

I kept this deliberately boring. I didn't try to break the live planner. I only wanted to answer one question: could the updated app open the existing information without changing it?

Could I get this far without a programming background?

In this case, I got further than I would have guessed. AI helped me build the app and widened the search for failure. I still had to decide which failures were unacceptable and what the safe response should be.

My knowledge of this workflow helped me make those decisions. I had to understand the consequences, insist on evidence, and keep asking, “Could this look correct while being wrong?” This maintenance result describes what I managed with this app; it does not establish what someone else could achieve in a different setting.

The reviewer found ways the app could fail. Knowing the work determined how it had to fail safely.

What supports this note?

The original work records show 31 distinct tests passing before the problems were found and 39 passing after the final fixes.

The fresh review was performed by a separate AI reviewer working from the same version of the app. It was not an outside company or human audit.

The real-data check was limited and read-only. It showed that the existing planner stayed unchanged during that check; it did not prove that every future action would be safe.

This is one app and one round of maintenance. It does not prove that this approach finds every problem or that building software is always better or cheaper than buying it. Identifying details and private business data have been left out.