Docs
Flows

Testing and Quality

Test your Flows and ensure quality with automated grading.

Testing and Quality

Building a great AI Flow is an iterative process. Unlike traditional software, the same input can produce different output every time the model runs - so the only way to know if your Flow is actually working is to test it. FormWise gives you a full set of testing and evaluation features in the Quality Check tab of your Flow editor.

Plan feature

The AI Quality Checker is included with the Pro and Agency plans (and AppSumo tier 3 and above). On other plans the Quality Check tab shows a lock and an upgrade prompt.

Why testing matters: small wording changes in your prompt, a new model version, or a tweak to your Knowledgebases can all affect output quality in ways you would not catch with a single preview run. Tests catch regressions before your users do.

[Screenshot: The Quality Check tab inside a Flow, showing rubric rules on the left and test cases on the right]


Preview testing

The quickest way to sanity-check your Flow is with the built-in live preview:

  1. Open your Flow and go to the Design tab.
  2. The live preview sits to the right of the canvas. Drag its edge to make it wider.
  3. Fill in sample inputs (or type a message for chatbot Flows) and send them.
  4. Review the output to see if it matches your expectations.

The preview always runs your latest draft, so edits on the canvas show up on your next run.

This is great for quick, hands-on checks while you are building. You can run the preview as many times as you want and adjust your workflow between runs. Preview runs also appear in Execution History so you can dig into the per-node trace afterward.


Test cases

Test cases are the inputs you want to check against. Each test case is paired with the Flow's current workflow and rubric - when you run the evaluator, FormWise executes the workflow for every test case and grades each output.

What a test case looks like

A test case has:

  • A name - so you can recognize it in the list ("Empty URL," "Long product description," "User asks for a refund").
  • Inputs - for form Flows, a value for every form question. For chatbot Flows, a single message string.
  • Optional expected behavior - a short note describing what a good output should look like. This is fed to the grader so it knows what to check for.

You do not need a hard-coded "expected output" string. FormWise uses rubric-based grading (see below) - you describe the qualities a good answer should have, and the grader judges every output against those qualities.

Creating test cases manually

  1. Go to the Quality Check tab.
  2. Click Add Test Case.
  3. Give it a Test Case Name so you can find it later.
  4. Fill in the input fields with sample values (for chatbot Flows, the message to send).
  5. Optionally, add Expected Behavior - a sentence or two describing the ideal response.
  6. Click Done.

You can create as many test cases as you need. Good test coverage means checking a variety of inputs - not just the happy path. Include edge cases: empty fields, weirdly long input, off-topic requests, things you would not expect a real user to send.

Auto-generating test cases

Do not want to write every test case by hand? FormWise can generate them for you:

  1. In the Quality Check tab, click Generate Tests.
  2. Choose how many test cases you want.
  3. The AI reads your Flow's configuration - input fields, system instructions, and workflow - and produces a set of realistic test cases.
  4. Review the generated tests and edit or remove any that do not fit your use case.

Auto-generated tests are a great starting point, especially when you are setting up the Quality tab for the first time. Always review them and add edge cases the AI might have missed.


Rubric rules

A rubric rule is a criterion the grader uses to score each output. Each rule has:

  • A rule name - a short label for the rule.
  • A pass condition (assertion) - the rule itself, written in plain language. For example: "The response always includes a citation," or "The tone is friendly and never apologetic."
  • A severity level - Critical, Major, or Minor. Severity sets how sure the grader must be before it passes an output: a Critical rule needs a more confident pass than a Minor one.

You can write rubric rules by hand with Add Rule, or click Generate Rules and describe what you want to evaluate in plain language ("never use the word 'delve', always include citations, maintain a professional tone"). FormWise turns that into a structured rule set you can edit.

Rubric grading is faster and more flexible than exact-match testing - the grader can recognize that two differently worded responses both meet the same criterion.

[Screenshot: The rubric rules section with three example rules and their severity badges]


Running the evaluator

Once you have at least one test case and one rubric rule:

  1. Click Run Evaluator in the Quality Check tab.
  2. FormWise runs your current workflow against every test case.
  3. Each output is graded against every rubric rule.
  4. The result is recorded as a new Execution - a numbered run you can come back to later.

While the evaluator is running, you will see a live progress indicator. A typical suite of 5--10 test cases takes anywhere from 30 seconds to a few minutes, depending on the models and tools your workflow uses.

Each evaluator run executes your Flow once per test case and grades every output, so keep your test suite focused to keep runs fast.


Reading the results

Every evaluator run produces an Execution entry in the history list. Click into one to see:

  • Pass / fail summary - how many test cases passed all rules, and how many had at least one failure.
  • Failed rules - which rules failed and on which test cases.
  • Per-test-case detail - the input, the actual output, and the grader's reasoning for any failures.
  • Full execution trace - click View run on any test case to open the same per-node waterfall view used in Execution History. This shows exactly what each node received, what it produced, and where time was spent.

Comparing one execution to the next is how you measure progress. If Execution #4 fails three rules and Execution #5 only fails one, your most recent change moved things in the right direction.

[Screenshot: An execution detail panel showing the pass/fail breakdown by rubric rule]


Turning failures into fixes

A failed test case is information, not a verdict. FormWise can use the failures to suggest concrete improvements:

  1. Open a failed Execution from the history list.
  2. Pick a failed rule to see the AI Analysis for This Rule and the suggested fix.
  3. Each suggestion targets a specific rule and explains what to change in the prompt or workflow configuration.
  4. Review the Suggested Revision, then click Apply Fix to Model to rewrite the target node's prompt. A diff preview shows you exactly what is changing before you confirm. When a fix touches several nodes, step through them with Apply & Next Node.

You can also pull insights from failures manually:

  • If a rule fails on most test cases, the rule might be too strict, or your system instruction is missing something important. Try rewording the instruction or adding a few-shot example.
  • If a rule fails on one specific test case, that input might be an edge case worth handling explicitly.
  • If failures cluster around a specific node, that node's prompt is a good place to start refining.

If your Flow uses a Knowledgebase, check that the relevant knowledge is actually in it and is being retrieved - the per-node trace will show you what context the model received.

Improve automatically

Below the evaluator, Improve automatically starts a background run that tries fixes to your prompts and step settings against your rules, keeps only the changes that help, and checks the best one on test cases it never tuned on. Pick a run size (each shows its number of runs and estimated credit spend), then click Start. You can leave the page while it runs.

When it finds a better version, click Apply to draft to apply it to your draft. Nothing goes live until you publish. The run needs enough test cases to split into tuning and checking sets; if you don't have enough, the panel tells you what to add.


When to run tests

You do not need to run the evaluator on every save. A few good moments:

  • Before publishing a version. Always run the evaluator before promoting a draft to production. See Versioning.
  • After big changes. Switched models? Rewrote a system instruction? Added a node? Run the suite to make sure nothing broke.
  • After updating your Knowledgebases. New documents in a Knowledgebase can change retrieved context and shift outputs.
  • Periodically on a schedule. Model providers update their models occasionally. Re-running your suite catches drift before users notice.

For one-off bug reports from real users, jump straight to Execution History - that is where you can replay actual production runs and inspect what happened.


Best practices

  • Test with a variety of inputs. Short, long, edge cases, unusual requests. Your users will surprise you.
  • Keep rules independent. Each rubric rule should test one thing. Bundling multiple checks into a single rule makes failures harder to act on.
  • Use Critical sparingly. Reserve Critical severity for behaviors that would genuinely make the Flow unusable - safety, factual accuracy, brand-breaking tone.
  • Re-run after changes. Every time you update your system instruction, switch models, or adjust your workflow, run the suite again.
  • Start testing early. You do not need a perfect Flow before you start. Early tests help you spot problems before they compound.

Next steps

Happy with your test results? Learn how to save your progress with Versioning so you can track changes and roll back if needed. To debug real user runs after you publish, see Execution History.

On this page