GingerOSHow we test

Stop taking accuracy claims on faith.

We test GingerOS the way you’d audit a classification: real CBP rulings, answers sealed, one round of questions, full-code matches only. Here is every step.

Test planSealed
  • Tariff editionOne HTSUS revision, frozen for the whole run
  • Code lengthAgreed up front. A matching heading is not a match
  • Test casesListed before any result is seen
  • QuestionsOne round of answers, then it has to decide
  • What is gradedOnly what a user would see on screen
  • Pass lineSet by a person before the scores come in

Any change after sealing is logged, with the reason.

A Test plan window listing six rules fixed before a run: tariff edition, code length, test cases, questions, what is graded and pass line. Each rule’s padlock closes in turn, then the Draft tag changes to Sealed.

Start from real CBP rulings. Check every one.

A ruling only becomes a test case after it passes the checks you’d run yourself and a second reviewer agrees.

  • Pulled from CROSS in a fixed date range, with source and date saved. A defined pool, not a random sample of all trade.
  • The case list is fixed before any result, so a hard case can’t be swapped out for being hard.
  • A code that still exists isn’t enough. Its wording in the frozen tariff must still fit the ruling’s facts.
Case intake

Candidates

1 accepted · 1 on hold · 2 left out · 1 escalated

Mounting bracket

Four checks

  • One product, as importedPassed
  • Deciding facts stated in the rulingMaterial never stated
  • Not revoked or modifiedPassed
  • Code still fits the frozen HTSUSPassed
DecisionOn hold

The ruling gives a code but never says what the bracket is made of. We go back to the full text. If it isn’t there, we don’t fill it in from the answer.

Second reviewer: agrees

Illustration with sample data: a Case intake window lists five made-up candidate rulings, such as a hose clamp and a brass valve, each tagged Accepted, On hold, Left out or Escalated. Selecting one shows its four checks, any failing check, the decision and the reason.

It gets the product. Never the answer.

GingerOS gets only the description of the goods. The code and reasoning stay sealed until grading.

  • Facts are quoted word for word. Specs in the analysis can count; “therefore it falls in heading X” never does.
  • Each saved ruling is fingerprinted. If its text changes, every case built on it is reviewed again.
  • CBP’s descriptions are tidier than real invoice lines, so this step is easier than your desk. That’s one reason acceptance uses your own data.
Practice ruling · made-up product

Sent to GingerOSSealed until grading

Description of merchandise

The item is a stainless steel ring with an outside diameter of 18 mm.

The wall is 1 mm thick. The ring has no surface coating.

Law and analysis

Holding

Illustration with sample data: a practice ruling for a made-up stainless steel ring, with the description of merchandise highlighted as sent to GingerOS and the law and analysis and holding blacked out as sealed until grading.

Take out one deciding fact. Does it ask?

A good broker asks before guessing. We remove one fact that decides the code and grade whether the first question goes after it.

Missing-fact test · stainless steel ring

Made-up ruleOutside diameter up to 25 mm is Code A. Over 25 mm is Code B.

Full case

Stainless steel ring, 18 mm

  • Stainless steel
  • Outside diameter 18 mm
  • Wall 1 mm thick
  • No surface coating
What GingerOS gets

Stainless steel ring 18 mm

  • Stainless steel
  • Outside diameter 18 mm
  • Wall 1 mm thick
  • No surface coating

The diameter appeared twice, so both come out. Nothing else changes.

Which first question gets credit?

GingerOSIs the outside diameter over 10 mm?
Test answererYes.
Doesn’t cover it

Too broad. A ring of 18 mm and one of 32 mm both answer yes, so the code is still open.

Illustration with sample data: a missing-fact test on a made-up stainless steel ring, with the 18 mm diameter struck out of the case GingerOS gets. Four first questions can be picked, each showing the reply and a verdict, Covers the gap or Doesn’t cover it.

A gap only counts if

  • Only one fact is missing

    Remove only that fact, never a sentence that also holds the surface finish.

  • It really changes the code

    18 mm and 32 mm lead to different codes, and the tariff text backs both.

  • Putting it back is enough

    Put it back and nothing decisive is still unknown.

  • No clue is left behind

    Leave “1.8 cm across” in the name and the test is broken.

One round of questions. Then it decides.

Every case stops at the same line, so none gets extra tries.

The test answerer’s rules

  • Answers only what was asked: “Over 10 mm?” gets “Yes”, never “18 mm”.
  • Anything not in the case file gets “I don’t know”.
  • Every answer is quoted from the case file, so an 80°C operating rating never becomes an 80°C maximum.

One round is a rule of this test, not a promise that real products need only one question.

Test run · stainless steel ring
GingerOSBefore I classify it: what is the ring’s outside diameter?
Test answerer18 mm.
GingerOSFinal classificationCode AOutside diameter 18 mm, under the 25 mm line.

Run stops here and is graded

Every possible first reply

  • Asks for the missing factAnswered once. The next reply is graded.
  • Gives a code without askingStops. Code graded, question check fails.
  • Asks about something elseStops. Question check fails.
  • Asks again after the answerNo third round. Counts as not finishing.

Illustration with sample data: a test run chat where GingerOS asks for the made-up ring’s outside diameter, the test answerer replies 18 mm, and GingerOS gives a final classification of Code A before the run stops to be graded.

Only the full code counts. Every miss stays in.

A matching heading isn’t a match, and errored runs are never quietly dropped.

  • Good wording can’t make up for a wrong code. The code is graded on its own.
  • Full-facts and missing-fact tests are reported apart, with planned, run and excluded counts.
Grading sheet

Answer key8471.30.0100

GingerOS returnedCounted as
8471.30.0100MatchExact match
8471300100MatchDots and spaces don’t matter
8471.30MissStopped at 6 digits
Three options, none pickedMissNo single final code
Timed outErrorStays in the planned total, listed apart

Two accuracy numbers, never one

On runs that finished
Correct codesRuns that finished normally
No final code counts as a miss
Against the whole plan
Correct codesAll planned cases, minus proven bad tests
Errors and unfinished runs stay in

Illustration with sample data: a grading sheet with the example answer key 8471.30.0100 and five example replies, two counted as Match, two as Miss and a timed-out run as Error, and two accuracy formulas with no numbers filled in.

Right code, wrong reason? The reason fails.

We read the explanation like an auditor: facts traced to the user, legal scope respected, citations looked up.

  • A fact the user gives later can’t rescue an earlier guess.
  • An “Other” line is read with the heading text and notes above it.
  • Short is fine: no set length or full tariff text, just reasons that hold up.
Report review · stainless steel ring

Final classificationCode A

The report saysOur check
The outside diameter is 18 mm, within the 25 mm limit for Code A.PassThe deciding fact is tied to the code.
Polished finish, rated for continuous outdoor use.FailNobody said that. An invented fact fails even when the code is right.
Code B is ruled out because it only covers rings over 25 mm.PassThe rejected alternative has a reason you can follow.

Illustration with sample data: a report review table for the made-up ring, final classification Code A, checking three sentences of the report. Two pass, and the sentence claiming a polished finish rated for outdoor use is framed as an invented fact and fails.

Every run is kept. A person signs off.

Every link, from the source ruling to the sign-off, can be opened and checked. A rerun never erases a miss.

Run record
  • Test planSealed before the run
  • Case listFixed before any result
  • Stainless steel ring
    • ConversationEvery message, as the user saw it
    • ResultThe final code
    • FingerprintsShows nothing was edited later
  • Grading
    • Each checkVerdict and the quoted words
    • ReportTo share
    • SpreadsheetEvery case, one row

Attempts on this case

Attempt 1Timed out. Kept.

Attempt 2Rerun under the plan’s rule. Used.

Illustration with sample data: a Run record file tree with the sealed test plan, the case list, the stainless steel ring case folder and a grading folder. Attempt 1 timed out and is kept, and attempt 2 was rerun under the plan’s rule and used.

Sign-off
  • Test cases qualified
  • Run followed the plan
  • Scores can be recomputed
  • Disputes settled or listed
  • Pass line set before the run

Signed by

Our test runs
Our test lead
Your acceptance
Your expert

No result is accepted until a person signs here.

A Sign-off window with five ticked items: test cases qualified, run followed the plan, scores can be recomputed, disputes settled or listed, and pass line set before the run. Below them, our test lead signs for our test runs and your expert for your acceptance.

  • Consistency gets its own planned repeat runs. Being consistently wrong is still wrong.
  • A flawed case, once a person confirms it, is fixed for every run it touched, not only the misses.
  • AI-assisted notes are labeled. Only a person signs.

The 14 checks, in plain words.

Each one says when it applies, what passes, what fails and where it stops. Open any of them.

Grading checklist

Every check gets one of four verdicts

  • PassIt applies, and the evidence shows it met the bar.
  • FailIt applies, with quoted evidence of a real problem.
  • Doesn’t applyThe reply never triggered it, such as no ruling cited.
  • Can’t tellIt applies, but the evidence isn’t enough to judge.

Only Pass counts as a pass. Pass rate = Passes ÷ (Passes + Fails), always shown with the Doesn’t apply and Can’t tell counts.

The product

Did it classify what was actually described?

Classifies the article as it actually arrives
Applies when
The reply judges whether the goods are a whole article, a part, a set, or unassembled.
Passes
The thing classified matches the input. A part isn’t treated as a whole machine, and goods imported unassembled aren’t described as assembled. Classifying them as the complete article under GRI 2(a) is fine.
Fails
It swaps in a different article or condition and classifies that instead.
Limits
It doesn’t have to restate every detail of the condition. Whether the law was applied correctly is a separate check.
Never changes or invents a product fact
Applies when
The reply states a material, construction, function, spec or use, or explains a technical term or unit conversion.
Passes
Every product fact comes from what the user had said by then, or follows directly from it. Technical explanations and conversions are correct.
Fails
It states as known a fact nobody gave, contradicts a fact that was given, or uses a wrong conversion that misleads the classification or the user.
Limits
Options inside a question aren’t claims. A later answer can’t justify an earlier guess. Reasonable rounding is fine. A disputed technical meaning with no reliable source is marked Can’t tell.

The questions

Did it ask the right thing, in a way people can answer?

The first question goes after the missing factMissing-fact tests only
Applies when
The first reply in a valid missing-fact test.
Passes
It asks for the missing fact, and a truthful answer to just that question settles the code. Other questions in the same batch are fine, and so is a range or threshold that settles it.
Fails
It asks only about other things, asks for “more information” in general, classifies without asking, or asks something too broad to settle the code.
Limits
No keyword matching. Mentioning the fact, assuming it, or asking the user to pick a code doesn’t count as asking. It is judged on the first questions only, never on what came later.
Someone with the spec sheet can understand the question
Applies when
Any question to the user.
Passes
It asks for product facts in words the person holding the spec sheet can follow. Technical terms come with context or are standard for that product.
Fails
The user would need to understand classification law, or pick a tariff code, to know what to answer.
Limits
There is no banned-word list. Mentioning a GRI or a tariff term doesn’t fail on its own, and legal terms are fine in the final explanation.
The question is specific and matters for the code
Applies when
Any question to the user.
Passes
It is clear which product fact is wanted and why it bears on classification, and it can be answered directly.
Fails
It is irrelevant, has no clear target, or hands the legal call to the user, such as asking which part gives the essential character.
Limits
A reasonable question the case file can’t answer gets “I don’t know”, not a fail. Several questions at once are fine. If whether a question was needed is genuinely disputed, it is marked Can’t tell or sent back for case review.
It doesn’t ask again for what it already has
Applies when
Any question to the user.
Passes
No repeat requests for facts already given clearly. Confirming a real ambiguity, a conflict, or a different level of detail is fine.
Fails
It asks again for a fact the user already stated, or repeats a question the user already said they can’t answer.
Limits
The repeated question still gets answered once. Facts the system was never shown can’t count as repeats.
Answer choices let the user tell the truth
Applies when
A question offers choices.
Passes
The choices are clear and don’t push the user to make something up. Single choice, multiple choice or free text, whichever the question needs.
Fails
The choices mislead, mix product facts with classification conclusions, or leave no way to answer truthfully.
Limits
No need for one option per candidate code, or for perfectly exclusive options. If free text is available, a missing “Other” or “Don’t know” button isn’t a fail.

The explanation

Would the reasoning hold up in your file?

The explanation links the deciding facts to the code
Applies when
A final code is given and an explanation is part of the agreed task.
Passes
A short explanation shows why these facts lead to this code. On a close call, the deciding reason is named.
Fails
The explanation is empty, or it is only a code, boilerplate, or a conclusion with no link to the facts.
Limits
No required length, no full tariff text at every level, no fixed list of factors. It is not compared with the source ruling’s own reasoning. If we failed to capture the explanation, that is Can’t tell, not a fail.
Rejected alternatives get a reason you can follow
Applies when
The explanation openly rules out a serious competing classification.
Passes
The reader can see why the main alternative doesn’t fit. One shared reason or a clear comparison is enough.
Fails
A main alternative is listed as ruled out with no understandable basis.
Limits
It doesn’t need to list every candidate or give a reason at every branch. If no alternatives are discussed, it doesn’t apply.
What it says about the code matches the code’s legal scope
Applies when
The explanation describes what the chosen code covers.
Passes
No real conflict with the heading text, the levels above it, and the section and chapter notes.
Fails
It says the code covers goods, materials or technology its legal text excludes, or denies a condition the text requires.
Limits
The wording needn’t match. “Other” is read with the heading and notes above it, and legal scope isn’t the everyday meaning of the heading’s words. Matching the answer key doesn’t earn an automatic pass.
Any cited ruling exists and is quoted fairly
Applies when
The reply cites a specific ruling or court decision.
Passes
Trusted sources supplied for the run confirm the ruling and what it is said to hold.
Fails
Trusted sources show a wrong number, a different product, or a real misstatement.
Limits
No citation means it doesn’t apply; citing isn’t required. If a citation can’t be verified, it is Can’t tell. Not finding a ruling doesn’t prove it was made up.
The legal reasoning it shows has no real error
Applies when
The reply uses or paraphrases a GRI, a tariff provision or a note.
Passes
The number, meaning, conditions and application fit the applicable law. A correct explanation without a rule number is fine.
Fails
A real misapplication: mixing up GRI 3(a) and 3(b), treating several candidate headings as if all applied at once, ignoring an exclusion note, misreading which parts provision takes priority, or applying HS rules to an export control (ECCN) question.
Limits
Only what is shown is judged. Loose wording isn’t a fail when the comparison and the standard are right. Citing GRI 1 or GRI 3 at the subheading level without naming GRI 6 isn’t a fail by itself. With no authoritative material to check against, it is Can’t tell.

Finishing the job

Did it land on one valid code?

The final code is valid and full length
Applies when
A final classification is given.
Passes
It is in the agreed code system, at the agreed length, and valid in the frozen tariff edition.
Fails
No code, the wrong format or length, or a code that didn’t exist on the edition date.
Limits
A valid code isn’t necessarily the right code; that is graded separately. The tariff edition never changes mid-test.
Once it has the facts, it finishes
Applies when
The first reply in a full-facts test, or the reply after the missing fact is supplied.
Passes
It gives a final classification without asking for more.
Fails
It still asks for more on a valid case that has enough facts, or ends without a code.
Limits
An optional “want more help?” isn’t a follow-up question. System errors are recorded separately. If another fact really is missing, the test case is reviewed rather than an answer invented.

A Grading checklist window: the four verdicts (Pass, Fail, Doesn’t apply, Can’t tell), then 14 checks in four groups: the product, the questions, the explanation and finishing the job. Each check opens to show when it applies, what passes, what fails and its limits.

Questions about how we test

How do you test an AI classifier’s accuracy?

We build test cases from real CBP rulings, give GingerOS only the product description, and seal the ruling’s code until grading. Only an exact match of the full code at the agreed length counts. We also remove one deciding fact from some cases to see whether GingerOS asks for it, and we review every explanation against 14 written checks.

Why use CBP CROSS rulings as test cases?

They are public, official, and state both the product facts and the classification. But a ruling is only a candidate. Each one must describe one article, state every deciding fact, still be in force, and carry a code that is valid in the tariff edition we froze. A second reviewer agrees before it goes in.

Do you give credit when the first six digits match?

No. The required length is set before the run. If the test calls for 10 digits, a code that matches only the first six is a miss. Dots and spaces are ignored. Listing the right code among several options without choosing one is also a miss.

What happens to a ruling that has been revoked or modified?

It is put on hold or left out, with the reason written down. We never move the old answer to a similar-looking code in the current tariff.

Does GingerOS see the ruling’s answer during the test?

No. It receives only the description of the goods, quoted from the ruling. The code, the legal analysis and the holding stay sealed and are used only for grading.

Why test what happens when a fact is missing?

Real product data is often incomplete. A system that guesses a missing material or dimension can produce a confident wrong code. We check whether its first question goes after the fact that decides the code, the way a good broker would ask the importer.

How should we compare AI classification accuracy claims?

Ask three things of any accuracy number: was it measured at 6 digits or the full 10, on whose test cases, and who checked the answers. A 6-digit match is far easier than the 10-digit code an entry needs, and most published figures are self-reported. A fair test looks like the one on this page: real CBP rulings, answers sealed, full codes only, then the same test on your own parts.

Where is your accuracy score?

This page explains the method, not a headline number. A score only means something next to the cases it was measured on. For your products, the same method runs on your own parts during acceptance: your experts confirm the answer key, and the pass line is agreed before anything runs.

Does the same method work for export control classification?

The checks are written to keep the two apart: applying HS rules to an export control question is a fail. An export control test needs its own answer key and legal basis, fixed before it runs.

Who decides whether a result passes?

A person. AI helps pull facts and flag problems, and those notes are labeled AI-assisted. On our own runs, our test lead qualifies the test cases, settles disputes and signs off against a pass line set before the run. When the test runs on your parts during acceptance, your expert signs.

Run the test on your own parts.

Your experts confirm the answer key. The pass line is agreed before anything runs.

Our Risk-Free Promise

Full refund.

If you’re not satisfied before MVP acceptance, get a full refund of your deposit.

Click Me!

See if you qualify