What Is a Flavor Harmony Score?

Most cooking tools answer the question "what do you want to make?" The question you can actually answer at 5:45pm with the fridge open is "what do I have?" — and the honest follow-up is "will any of this taste good together?"

The Flavor Harmony Score exists to answer that second question, before you commit to cooking anything.

The short version

It is a number from 0 to 10 that updates live as you enter ingredients. It estimates how naturally your ingredients work together, drawn from our Food Intelligence Layer — which today covers 195 core ingredients and how they pair across 35 cuisine styles.

A high score means your set sits close to combinations that appear again and again in established cooking traditions. A low score does not mean "bad" — it means you are further from well-trodden ground, and you may need a bridging ingredient to pull it together.

Why a score at all?

Because the failure mode of AI recipe generation is confident plausibility.

Ask any general-purpose chatbot for a recipe using chicken, blue cheese and banana and it will write you one. It will look like a recipe. It will have steps and quantities and a serving suggestion. What it will not do is tell you that those three ingredients have almost nothing in common, and that you are about to spend forty minutes finding that out.

Scoring first changes the order of operations. You see the number, you see the profile, and you get one concrete suggestion — all before a single word of the recipe is written.

What goes into the number

Three things:

Pairing strength. Every ingredient in the graph carries weighted connections to the ingredients it works with. Chicken and lemon are strongly connected. Chicken and banana are not. The score aggregates the strength of the connections between everything you have entered.

Cuisine coherence. Pairings are tagged with the cuisines they are authentic to. A set that is strong within a single tradition scores better than one drawn from several unrelated traditions — that is what makes the difference between a coherent dish and a list of ingredients.

Set size and balance. Two ingredients that pair beautifully is a promising start, not a meal. The score accounts for whether the set is developed enough to build a dish around.

The other two things you see

The score on its own would be a verdict without a remedy. Two companions make it actionable.

The Flavor Profile label names the character of what you are building in plain language — "Bright, Mediterranean", "Umami-forward, Japanese", "Rich, Classical French". It tells you what kind of dish you are heading toward while you can still change direction.

The One Best Addition names the single ingredient most likely to lift the dish. One suggestion, not forty-seven. Adding a well-chosen bridging ingredient is often the difference between a set that scores a 6 and the same set scoring an 8.5 — and it is usually something already in the cupboard.

What the score is not

Worth being direct about the limits, because they are real.

It is not a guarantee. It is a probability signal drawn from how ingredients have historically been combined. Cooking has room for combinations that no graph would predict, and some of the best dishes come from ignoring exactly this kind of advice. A low score is information, not a prohibition.

It is not derived from flavor chemistry. The graph was built by querying a large language model for ingredient pairings tagged by cuisine, then reviewed by a person. It reflects documented culinary practice, not molecular compound analysis. That is a meaningful distinction and worth stating plainly: this is a structured, reviewed model of how people actually cook — not a laboratory result.

It does not know your taste. The score reflects general culinary affinity. If you dislike cilantro, a set full of cilantro scores well and remains wrong for you. That is what diner profiles are for — preferences and off-limits ingredients are applied as hard constraints before generation, separately from the score.

Why it is a graph rather than a prompt

This is the part that matters technically.

A prompt is instructions you send to a model each time. A graph is a structured dataset that is queried, versioned, and cached — it exists independently of any one model, and it improves as ingredients are added and pairing weights are tuned against real use.

That difference is why the score can appear while you type rather than after a generation round trip, why it stays consistent between sessions, and why a cuisine selection actually re-ranks the suggestions instead of decorating the request.

Trying it

Enter two or three things you actually have. Watch the number. Add whatever it suggests and watch it move.

The fastest way to understand what the score is measuring is to feed it a deliberately terrible combination and see how far it drops — and then see which single ingredient it reaches for to rescue it.

Score your fridge tonight.

Free. No credit card. 10 recipes a month, every one with flavor science built in.

Generate your first recipe free

← All articles