FOUNDER FIELD GUIDE

How to evaluate AI tools for your startup.

Evaluate one real workflow against your current process before adopting an AI tool. Define acceptable outputs, test with permitted representative inputs, record failures and review effort, then make a bounded decision.

1. Define the decision before trying the tool

“Should we use AI?” is too broad to test. A more useful question is: “Can this tool draft a particular type of support reply that a teammate can review within our quality and time constraints?” Name the owner and the decision deadline.

Write the alternatives: adopt for the bounded use case, continue evaluating, keep the present workflow, or clarify a missing requirement. An attractive demonstration does not resolve those choices.

2. Write the acceptance criteria first

Choose criteria that reflect the job, not a generic model benchmark. Specify hard failures separately from preferences. An unapproved disclosure, unsupported factual claim, or incorrect action may rule out adoption even when other outputs read well.

A scorecard to adapt to your workflow
CriterionRecord for each testDecision rule to set before testing
Task correctnessExpected result, actual result, and material omissionsWhich failures are unacceptable?
Evidence and uncertaintyWhether claims have support and missing inputs are acknowledgedWhich claims require verification?
Human reviewReview time and substantive correctionsHow much review work is acceptable?
Cost and latencyObserved cost and elapsed time per completed task, including retriesWhat budget and response time fit the workflow?
Data and permissionsApproved inputs, access scope, and processing restrictionsWhich conditions must hold before any live use?

Keep the underlying observations. A single combined score can hide a hard failure behind several strong results.

3. Use a small, representative test set

Include ordinary work, ambiguous inputs, missing context, and cases where the right answer is to ask for clarification. Use only data you are authorized to process. Synthetic inputs can check behavior, but they do not by themselves establish quality on actual customer work.

Record the tool version, relevant settings, input, output, test date, and reviewer. Keep the evaluation bounded enough that someone can inspect every result. Select the cases before seeing the outputs to avoid reporting only the successes.

4. Compare against the current workflow

Run comparable tasks through the existing process and the candidate workflow. Include the time spent preparing inputs, checking claims, correcting output, and recovering from failures. An answer that arrives quickly can still create more work overall.

Illustrative example: if a generated reply takes less time to draft but requires a lengthy policy check, record both parts. Do not describe drafting time alone as total time saved.

5. Decide what the evidence supports

Separate observed results from expectations about wider deployment. A small test can support another bounded trial; it cannot establish reliability across every customer, language, or task.

  • Act: authorize the next specific step within the tested scope.
  • Monitor: retain the current approach and name a trigger for reconsideration.
  • Ignore: close this option because it does not fit the present question.
  • Clarify: resolve the missing requirement before choosing.

6. Preserve a revisit trigger

Save the tested scope, failures, unresolved risks, and decision in a decision journal. Choose a review date or condition, such as a product version change, a measured failure, or a new business requirement. Keep later results separate from the original record.

Feedsion is designed to keep evidence, company context, and revisits together. It does not replace the evaluation itself. Explore a labeled sample to see that workflow.

TAKE A LOOK INSIDE

Your next question is a good place to start.

Explore the sample