02 · v0.4.0 · Published 3 August 2026

Case 02 — AI-Native Inventory: Image-to-Unit Identification, Condition Capture and Per-Stay Reconciliation

Point the camera at the shelf. Four seconds later, the count exists — an AI capability embedded in Casa Admin's Add Inventory module, identifying items, quantity, condition and location from a single unstaged photo.

Public, redacted render. Guest identifiers, payment references and document images are withheld.

1. Business Problem

A serviced property carries a large and heterogeneous item count per room — glassware, crockery, kitchen and bar items, electricals, linen, furnishings, consumables, and the long tail of small objects that are individually cheap and collectively material. One space in this case, the Stores at 17VC, carried 37 items on record.

Two operations depend on knowing exactly what is in each space:

a. Initial setup. Establishing the baseline — what this space contains, in what quantity, in what condition — when a room is commissioned or re-set.

b. Pre-arrival setup verification. Confirming, before the guest is handed the key, that the space has been set to standard and matches the base. This is the handover of custody: it fixes what the guest was given, and every later claim rests on it.

c. Per-stay checkout count. Re-establishing that baseline at guest departure and identifying the delta, in both quantity and condition.

(a) is a recording task. (b) and (c) are comparisons. The distinction looks pedantic and is not — it is the whole of Section 7.1.

Done manually, all three are slow. The checkout count is the binding constraint: re-inventorying a room took hours, and that cost lands at precisely the wrong moment in the operating cycle — inside the turnover window between one guest leaving and the next arriving, when time is least available and error is most expensive.

The failure modes are structural rather than effort-related:

  • It is a listing problem disguised as a counting problem. The counter must first recall the full expected item set, then verify each one. Absence is far harder to notice than presence — the missing item is the one nobody looks at.
  • Naming is unstable. Two staff inventorying the same shelf produce two different lists for the same objects. Without a controlled vocabulary the counts are not comparable across rooms or across stays, which makes loss attribution and trend analysis impossible.
  • Condition is not captured at all. Manual counts record quantity. Whether a glass is clean, chipped or cracked is either omitted or recorded as an unstructured note, which means damage cannot be attributed to a stay.
  • No defensible record. A count that exists only as a signed sheet cannot settle a guest dispute, cannot be audited, and cannot be re-checked after the fact.
  • It does not scale. Hours per room × rooms × stays is the whole problem in one line. Three properties are in scope — 17VC, BF16 and 520.

◇ Scale of the problem (to confirm): approximately 120–200 line items per full apartment across all spaces; 2–3 hours for a full manual re-inventory at checkout; 3 properties, 8–12 lettable spaces, 25–40 stays per month across the portfolio. At 2.5 hours per checkout count, that is 60–100 staff-hours per month spent counting.


2. AI Solution

Point the camera at the shelf. Four seconds later, the count exists.

An AI capability was built directly into the Casa Admin Add Inventory module. The operator selects a property and a specific room or space, opens the camera inside the app, photographs what is there, and submits the frame for analysis. A multimodal model — Gemini 2.5 Flash, one call per image — identifies each distinct object, counts units, assesses condition, and infers spatial location within the frame. The structured result is presented as an editable table for human review before it is committed to the inventory record.

Three properties of the design carry the case:

  1. Condition is captured at the same instant as quantity. Every detected line carries a condition grade. This converts checkout from a counting exercise into a genuine reconciliation: at departure the same capture runs and the condition grade is compared against the stored baseline for clean, damaged, missing.
  2. The system checks itself against what it already holds. Detected items are matched against the existing inventory for that space, and an item already on record is flagged rather than silently duplicated.
  3. Nothing is written without a human. The result screen is explicitly headed Review and edit items before saving. Identification is AI-executed; the commit is human.

A Manual tab sits alongside the AI tab throughout, preserving the unassisted path.

◇ Invocation: analysis is triggered by the explicit Analyze Image with AI action; the round trip completes in approximately four seconds. The captured frame is retained in object storage against the inventory event as audit evidence.


The Story

One shelf in the Stores at 17VC, captured and counted. Five screens, four seconds.

Figure 1 — Choosing where the count belongs.

Figure 1. The Add Inventory screen opens with a quick-select row across the three properties — 17VC, BF16, 520 — and a link to select a specific room. Below it, the AI / Manual toggle. Why it matters: the first thing the system establishes is not what is being counted but where it belongs. An item count without a location is not inventory; it is a list. Every downstream comparison — baseline against checkout, room against room, stay against stay — depends on this being fixed before the camera opens.

Figure 2 — Down to the space.

Figure 2. Property resolves to 17VC (lmc_ro); room resolves to Stores. Why it matters: the granularity is the space, not the property. This is the unit at which a baseline is meaningful and at which a housekeeping delta is actionable.

Figure 3 — The capture, in the conditions that actually exist.

Figure 3. The in-app camera modal, live preview, Cancel and Capture. The frame shows a shelf of glassware — mugs and tumblers inverted and stacked in overlapping rows, metal jiggers at the right edge, and below the shelf line a microwave and a cardboard carton in near-darkness.

Why it matters — and this is the figure to look at twice. The frame is bad by every conventional standard: low light, oblique angle, motion blur, heavy occlusion between near-identical transparent objects, two depth planes in one shot, and a partly cut-off item at the frame edge. It was not staged. No grid, no labels turned to the camera, no lighting rig. This is a member of staff holding a phone at arm's length in a store room, which is the only condition the system will ever actually operate in. A system that needs a clean frame does not survive contact with a turnover window.

It is also, as Section 7.1 sets out, the exact source of the principal unsolved problem in this deployment. The tolerance that makes first ingestion work is the same tolerance that makes checkout comparison unreliable.

Figure 4 — Submission.

Figure 4. The captured frame in preview, the Analyze Image with AI action, and beneath it the collapsed panel: Existing Inventory (37 items). Why it matters: the existing record is loaded and present at the moment of analysis, not consulted afterwards. The new capture is read against what the space is already known to hold. That adjacency is what makes the deduplication in Figure 5 possible.

Figure 5 — The count, four seconds later.

Figure 5. Inventory Items (5) — Review and edit items before saving.

Item Name Quantity Condition Location Notes
Clear glass mugs 10 excellent on shelf All mugs are stacked neatly
Clear glass tumblers 6 excellent on shelf Stacked in two rows
Stainless steel jiggers 2 excellent on shelf Stacked together
Microwave oven 1 good below shelf Standard size microwave
Cardboard box *(flagged: Exists (Qty: 1))* 1 good below shelf Closed box, contents unknown

Five rows out of that frame. Four things in this table are worth more than the count itself:

  • It separated two visually similar classes. Mugs and tumblers are both clear glass, both inverted, both stacked, both overlapping. It resolved them as distinct classes at 10 and 6 — a distinction a hurried human counter routinely collapses.
  • It resolved depth as a spatial attribute. on shelf versus below shelf, from a single monocular frame. The location column is not free text a person typed; it is inferred structure.
  • It declined to guess. The carton is recorded as Closed box, contents unknown. The model had every opportunity to invent a plausible contents list and did not. In an inventory system, calibrated refusal is worth more than a confident answer — a fabricated line item propagates into every subsequent checkout comparison as a permanent phantom discrepancy.
  • It caught its own duplicate. The carton carries an orange Exists (Qty: 1) badge against the 37 items already on record for that space. This is the normalisation layer surfacing in the interface: the operator is told you already have this, and decides whether it is the same box or a second one. Without this, repeated captures of the same shelf inflate the baseline until it is worthless.

And then the header line does the governance work: Review and edit items before saving. Five rows, each individually editable, none committed. The machine produced the list; a person owns the record.

The rest of the loop. The same capture runs twice more: once before the next guest arrives, to verify the space has been set to standard, and once at their departure. Quantity is differenced against the baseline; condition is compared grade against grade. The question the system answers is not only is the count short but did this stay damage anything — and it answers it against a stored, image-backed baseline rather than against memory.

That is the design. What The Story demonstrates is the first of those three operations. Section 7.1 records where the other two currently fall short in practice, and why the difference between recording and comparing is the whole of the problem.


Add Inventory screen with property quick-select and AI/Manual toggle
Figure 1 — Choosing where the count belongsThe Add Inventory screen opens with a quick-select row across the three properties — 17VC, BF16, 520 — and a link to select a specific room. Below it, the AI / Manual toggle.
Room/space selector resolving to 17VC, Stores
Figure 2 — Down to the spaceProperty resolves to 17VC (lmc_ro); room resolves to Stores.
In-app camera modal showing an unstaged, low-light shelf of glassware
Figure 3 — The capture, in the conditions that actually existThe in-app camera modal, live preview, Cancel and Capture. The frame shows a shelf of glassware — mugs and tumblers inverted and stacked in overlapping rows, metal jiggers at the right edge, and below the shelf line a microwave and a cardboard carton in near-darkness. Not staged.
Captured frame in preview with Analyze Image with AI action and Existing Inventory panel
Figure 4 — SubmissionThe captured frame in preview, the Analyze Image with AI action, and beneath it the collapsed panel: Existing Inventory (37 items).
AI result review table: 5 detected items with quantity, condition, location, notes, and an Exists duplicate flag
Figure 5 — The count, four seconds laterInventory Items (5) — Review and edit items before saving. Five detected items with the Exists (Qty: 1) duplicate flag on the cardboard box.

4. Implementation Journey

◇ Build. A single developer, working through Lovable, over approximately 48 hours of directed development effort excluding the underlying inventory system. Full economics in Section 7 and in Document control.

◇ Pilot. Single space — the Stores at 17VC, 37 items on record — used as the proving ground. One shelf, one frame, five items, four seconds is the run documented in The Story.

◇ Scale-up. Sequenced by space rather than by property: all spaces at 17VC, then BF16, then 520. Rationale — the frame-registration problem in Section 7.1 is a per-space problem and is best solved once at one property before it is replicated across three.

Change management — the part that is not a technology question. The housekeeping role at turnover changes from enumerating to adjudicating. A staff member who previously produced a list now reviews one. This is a different skill with a different failure mode: fatigue produced the errors in the old process; over-trust produces them in the new one. A well-formed table of five rows with green condition badges is persuasive, and the reviewer's job is to catch what is missing — which the table, by construction, cannot show.

◇ Training implication. HK staff need to be trained on framing and coverage, not on the software. The interface is a camera button; the skill is knowing that a verification shot must include everything and must match the baseline — a requirement that does not apply when they are simply recording a new space. That training does not currently exist and is a precondition for either verification leg to be trustworthy (Section 7.1). The training is not "how to use the app." It is "why this photograph is different from the last one you took with the same button."

Governance. No agent holds this role, so the eight gates do not apply (Section 12). The governance surface is the commit gate, the override log, and image retention (Section 13).


5. Business Benefits

Productivity. ◇ The checkout count moves from hours per room to minutes — approximately 4 seconds of inference per frame plus operator review time. On ◇ 6–10 frames per full apartment, the machine portion of a full room count is under a minute. The reviewing time is the remaining cost and is the number to measure.

Quality — the benefit that did not exist before. Condition is captured as structured data at the moment of the count. Manual counting never produced this, which means damage was previously either unattributable or contested. It is now a graded field against a timestamped, image-backed baseline.

Defensibility. ◇ A manual count produces a number. This produces a number with the photograph that generated it, retained against the stay. In a guest dispute the difference between an assertion and a record is the entire argument.

Comparability. ◇ Counts are now comparable across spaces, properties and stays, because naming is normalised against the existing record rather than reinvented by each counter. This makes portfolio-level questions answerable for the first time — which item classes attrit fastest, which spaces lose most, whether a given item is worth restocking.

Loss visibility. ◇ Slow attrition of low-value items was previously invisible, not because it was hidden but because nobody re-counted often enough to see the trend. Frequent cheap counting makes a slow leak visible; expensive counting cannot, at any level of diligence.

Employee experience. ◇ The most disliked task in the turnover window is materially reduced. This matters for retention in a role where turnover-window pressure is the main source of stress.

What is deliberately not claimed. No headcount reduction. The count was never a full-time role; it was a task inside a role. Framing this as labour saving would misstate both the benefit and the economics.


6. KPIs

Framed on completeness and reliability rather than cost reduction, consistent with Case JD-01.

Measure Baseline (manual) Observed / proposed Status
Inference time per frame Hours per room ~4 seconds Observed, Figs. 4→5
Items resolved per frame n/a 5 from one unstaged frame Observed, Fig. 5
Condition captured Not captured 100% of detected lines Observed, Fig. 5
Duplicate detection against existing record None Flagged in-line, pre-commit Observed, Fig. 5
Identification accuracy, first pass n/a ◇ target ≥ 90% of lines accepted unedited To measure
Lines requiring human correction n/a ◇ target ≤ 10% To measure
Frames per space to full coverage n/a ◇ 3–5 small space; 6–10 apartment To measure — see note
Coverage completeness (items on record not seen in any frame) Unknown ◇ target ≤ 2% To measure
Spurious delta rate — pre-arrival (B) n/a ◇ currently material — see §7.1 To measure — priority
Spurious delta rate — checkout (C) n/a ◇ currently material — see §7.1 To measure — priority
Pre-arrival verifications completed before key handover ◇ informal ◇ target 100%, recorded To measure
Operator review time per frame n/a ◇ 20–40 s To measure
Full-room count, end to end ◇ 2–3 hours ◇ target < 15 minutes To measure
Cost per full-room count ◇ 2.5 staff-hours ◇ ~₹1 inference + review time See Section 7

The two numbers that matter most are both missing. Frames per space is the multiplier without which "four seconds" means nothing. Spurious delta rate is the number that decides whether the verification legs are usable at all — and it must be logged separately for pre-arrival and checkout, because they will have different rates and different causes. Both are cheap to instrument and neither can be reconstructed later.


Chain completion audit

4 complete · 2 soft · 3 not done · 5 unverified · 36% chain completion

#StepStatusEvidence
1Inference time per frameComplete~4 seconds. Observed, Figs. 4→5.
2Items resolved per frameComplete5 from one unstaged frame. Observed, Fig. 5.
3Condition capturedComplete100% of detected lines. Observed, Fig. 5.
4Duplicate detection against existing recordCompleteFlagged in-line, pre-commit. Observed, Fig. 5.
5Identification accuracy, first passUnverifiedTarget proposed at ≥90% of lines accepted unedited. To measure.
6Lines requiring human correctionUnverifiedTarget proposed at ≤10%. To measure.
7Frames per space to full coverageNot doneProposed 3–5 small space, 6–10 apartment. To measure — the missing multiplier on the four-second figure.
8Coverage completeness (items on record not seen in any frame)UnverifiedTarget proposed at ≤2%. To measure.
9Spurious delta rate — pre-arrival (B)Not doneCurrently material per §7.1. To measure — priority.
10Spurious delta rate — checkout (C)Not doneCurrently material per §7.1. To measure — priority.
11Pre-arrival verifications completed before key handoverUnverifiedInformal today; target 100%, recorded. To measure.
12Operator review time per frameUnverifiedProposed 20–40 s. To measure.
13Full-room count, end to endSoftBaseline 2–3 hours manual; target <15 minutes. To measure.
14Cost per full-room countSoftBaseline 2.5 staff-hours; proposed ~₹1 inference + review time. See Section 7/Economics.

7. Challenges

7.1 The frame-registration gap — the principal unsolved problem

Three operations, not one. The system is used at three distinct moments, and they are not the same kind of task:

Operation Kind of task Status
A Ingestion — first capture of a space; the baseline is created Recording Works. Documented in The Story.
B Pre-arrival setup verification — the room is set to standard and the inventory is verified against the base before the guest is handed the key Comparison This is where the gap bites.
C Checkout count — the space is re-captured at departure and differenced Comparison Inherits the gap, and inherits any error from B.

This is not a problem with A. At ingestion there is no reference. The system records what it sees, an operator reviews it, and whatever is captured becomes the baseline. Tolerance for a bad frame is high precisely because a bad frame simply produces a slightly different baseline, which is then the baseline. Nothing can be wrong because nothing is being checked.

The gap is in B and C — the verification legs. A verification is a comparison, and a comparison requires the two observations to be commensurable. Nothing in the current design enforces that. The interface treats every capture identically: same camera, same button, same tolerance. It does not know, and does not signal, that a capture taken at pre-arrival or at checkout is being asked to do a fundamentally harder job than the capture that created the baseline.

B is the one that matters most, and it is the easier one to overlook. The pre-arrival verification is the chain-of-custody handover: it fixes the state the guest is being given. Every claim later made against that guest — a missing tumbler, a chipped mug — rests on B having been correct. A spurious result at B is worse than a spurious result at C, because C at least fails visibly and in the operator's presence; B fails silently and then licenses a charge weeks later against the wrong person. If the reference is wrong, the delta is wrong, and the delta is what reaches the guest.

In practice the housekeeping team does not reproduce the baseline frame at either B or C. The same shelf is photographed at a different distance, a different angle, a different crop, under different light, at a different time of day. The operator is not being careless — nothing in the interface tells them the frame has to match, and there is nothing on screen to match it against. The result is that the item sets differ because the frame changed, not because the inventory changed. The system reports a delta that is an artefact of photography.

The specific spurious results this produces:

Frame variance Spurious result
Tighter crop than baseline Items outside the new frame read as missing
Wider crop than baseline Items newly included read as added — or as duplicates
Partial occlusion of a stack 8 of 10 mugs visible reads as quantity 8 — a two-unit loss that did not happen
Different angle A class resolves differently — tumblers read as mugs, or one class splits into two
Different lighting or blur Condition grade shifts; the same glass reads excellent in one frame and good in another, manufacturing a damage event
Different zone photographed entirely Whole item groups appear and disappear between stays

Why this is the serious failure, and not merely an annoyance. The errors are asymmetric in consequence:

  • False positives destroy trust faster than false negatives. If the system tells the HK team three times that two mugs are missing and they are not, the team stops believing the fourth alert — which is the real one. An alerting system that cries wolf is worse than no alerting system, because it consumes attention and then fails at the moment it matters.
  • A spurious damage flag is guest-facing. A false missing is an internal irritation. A false damaged that reaches a guest as a charge or a claim is a reputational and commercial event, and it is unrecoverable — the apology does not undo it.
  • It is invisible to the review gate. The human gate in Figure 5 catches a wrong row. It cannot catch a right row that should not have been compared, because the operator has no view of the baseline frame at the moment of review.
  • It compounds. If a spurious delta is committed, the baseline itself drifts, and the next comparison is against a corrupted reference.
  • It breaks the chain of custody at its first link. An error at B, the pre-arrival verification, is not merely one bad reading — it silently redefines what the guest was given. Every subsequent claim against that stay inherits the error and is unfalsifiable, because the only record of the pre-arrival state is the flawed capture. A wrong answer at C is an argument. A wrong answer at B is a wrong premise, and no amount of care downstream recovers from it.

Restating the underlying principle: the system currently instruments what was seen. Verification requires what was looked at. Those are different things, and the gap between them is where every spurious result in this deployment lives.

And restating the scope, because it is easy to misread this section as a criticism of the run in The Story: the capability demonstrated in Figures 1–5 is sound. Operation A — building the baseline from an unstaged frame — works, and works in bad conditions. What has not been built is the discipline that makes B and C comparisons rather than fresh observations. The system today is an excellent instrument for recording inventory and an unreliable one for verifying it. Those are two different products, and only the first one has been built.

◇ Proposed remediation, in order of leverage:

  1. Ghosted baseline overlay in the camera modal. Store the baseline frame per zone; display it semi-transparently in the live preview so the operator aligns the shot before capturing. This is the highest-leverage fix by a wide margin — it addresses the cause rather than the symptom, requires no model change, and turns an unstated expectation into a visible one.
  2. Capture zones, not rooms. Define named capture positions per space — Stores / upper shelf, Stores / under-counter — each with a fixed baseline frame. Coverage then becomes checkable: the system knows how many zones a space has and can refuse to compute a delta until all are captured.
  3. Three-state item status, not two. Present / not present / not seen. An item absent from the frame is not seen, which is not the same claim as missing. Only items inside a declared, captured zone are eligible to be marked missing. This single change eliminates the largest class of spurious result.
  4. Per-item visibility and confidence flags. Ask the model to report occlusion and partial visibility per line. Low-confidence or partially-occluded lines are excluded from delta generation and routed to human check instead.
  5. Condition changes require corroboration. A condition downgrade should not be actionable on a single wide frame. Require either a second frame, a close-up, or explicit human confirmation before a damage event is recorded. Condition is the field with the highest consequence and the highest sensitivity to lighting — it should have the highest evidentiary bar, and currently has the same bar as everything else.
  6. Never auto-charge. No delta becomes a guest charge or a damage claim without human adjudication, and the adjudicator sees both frames side by side.
  7. Instrument the spurious rate as a first-class KPI. Every delta an operator dismisses is a labelled false positive. Log it. It is the only way to know whether remediations 1–5 worked. Log it separately for B and C — they will have different rates and different causes, and a blended number will hide both.
  8. Treat the pre-arrival verification as a distinct, named operation in the interface. It is currently indistinguishable from any other capture. It should have its own entry point, its own completion state, and its own record — this space was verified to standard, by this person, at this time, against these frames — because that record is the thing a guest charge ultimately rests on. Consider requiring it to pass before a space can be marked ready for arrival.

◇ Sequencing recommendation. Items 1 and 3 together address the majority of the problem and can ship independently of the rest. Items 2 and 7 are the foundation for measuring whether that is true. Item 8 is what makes B auditable at all.

Interim operating rule. Until at least 1, 3, 7 and 8 are in place: both verification legs are advisory. A delta at B is a prompt for the housekeeping team to look again before the guest arrives — the cheapest possible moment to correct an error, and a use the system is already good enough for. A delta at C is a prompt for a human to check, never a finding, and never the basis of a guest charge. Used this way the system is useful today and harmless if wrong, which is the correct posture for a comparison capability that has not yet been made commensurable.

7.2 Coverage, not accuracy, is the underlying risk

Related to 7.1 but distinct, and unlike 7.1 it applies to ingestion as well. The dangerous failure is not a misidentified item — the reviewer catches that, and the review gate is designed for it. It is the item outside the frame, which is never flagged at all and silently shrinks the baseline. Accuracy is self-correcting through the human gate; coverage is not. Nothing currently tells an operator that a space is incompletely captured.

7.3 Cost at pilot versus cost at production

The pilot is cheap because it is small. Per-image inference cost × frames per space × spaces × stays is a consumption line that scales with adoption. The full arithmetic, with the actual build inputs, is in Document control → Economics. The headline: inference is negligible and will remain so. The one-time build dominates by three orders of magnitude, and the real ongoing cost is operator review time, which is not on any invoice.

7.4 Model retirement risk

Gemini 2.5 Flash is a retiring model. Google has set the 2.5 family for retirement in October 2026, and the 2.0 family was already shut down in June 2026. This deployment has a dated dependency. ◇ Mitigation: abstract the inference call behind an internal interface now, while there is one call site, and treat the model as a swappable component. Re-baseline the accuracy expectation on the successor model before the cutover rather than after — a model change is a silent behaviour change in a system where the output is committed to a database.

7.5 The closed box

Contents unknown is correct behaviour and an open operational question. An inventory system that faithfully records opaque containers as single units has a blind spot it is honest about but has not solved. Proposal: opaque containers get a manual one-time contents declaration at first sighting, after which the container is treated as a single auditable unit with known contents.

7.6 Over-trust at the review gate

Discussed in Section 4. The reviewer's job is to catch what is missing, which the table by construction cannot show. Partial mitigation: display the count of items on record for the space that were not seen in this capture, adjacent to the results table. This turns an absence into something visible.

7.7 Vocabulary drift

Clear glass mugs today; glass mugs or coffee mugs on a later capture. The Exists badge handles the cases it catches; the cases it misses fragment the baseline into near-duplicate classes. Proposal: the first human-accepted name for an item class becomes canonical for that space, and subsequent captures are matched against the canonical list rather than free-text compared.


8. Lessons Learned

◇ 1. The recognition was never the hard part. Gemini 2.5 Flash read an unstaged, badly lit, motion-blurred frame and separated two near-identical classes of transparent glassware, resolved depth, and declined to guess at a closed box. The engineering that made this a system was elsewhere entirely: scoping before capture, loading the existing record adjacent to analysis, and putting an editable table between the model and the database.

◇ 2. Recording and verifying are different problems, and one interface was built for both. This is the most transferable lesson in the case. Ingestion is tolerant — it records what it sees, and whatever it sees becomes true. Verification is a comparison and demands that two observations be commensurable. The same camera button serves both, which means the harder job inherited the easier job's assumptions. Anyone replicating this should design the verification capture first and derive the ingestion capture from it, not the other way round — the reverse order is what produced Section 7.1.

◇ 2a. The pre-arrival verification was the leg nobody thought about. Attention went to ingestion, because it is the visible build, and to checkout, because it is where money is at stake. The pre-arrival check sat between them and was treated as a lesser version of ingestion. It is not: it is the chain-of-custody handover, and it is the leg on which every guest-facing claim silently depends.

◇ 3. The residual human task is the one to instrument. Automating identification left the human holding the camera. Framing turned out to be the step the whole system is most sensitive to, and it is the step that received no design attention, no guidance, and no measurement.

◇ 4. Calibrated refusal is a feature and should be selected for. Closed box, contents unknown is worth more than a plausible invented list, because a fabricated line becomes a permanent phantom discrepancy in every subsequent comparison.

◇ 5. Keeping the abandoned approach as a fallback was correct. The Manual tab is the reason this system has a working degraded mode when a frame fails or the API is down.

◇ 6. The free accuracy dataset was nearly thrown away. Every human correction at the review gate is a labelled example, generated at zero marginal cost by normal operation. It only exists if the pre-edit machine output is retained, and that decision has to be made before the data starts flowing, not when someone asks for accuracy figures.

◇ 7. A build like this is now measured in days, not quarters, and that changes which problems are worth solving. At roughly $3,500 all-in, the question is no longer whether the business case clears a capital threshold. It is whether the operational discipline exists to use the thing — which is exactly what Section 7.1 is about.


Findings

F1 The recognition was never the hard part

Gemini 2.5 Flash read an unstaged, badly lit, motion-blurred frame and separated two near-identical classes of transparent glassware, resolved depth, and declined to guess at a closed box. The engineering that made this a system was elsewhere entirely: scoping before capture, loading the existing record adjacent to analysis, and putting an editable table between the model and the database.

F2 Recording and verifying are different problems, and one interface was built for bothBlocking

This is the most transferable lesson in the case. Ingestion is tolerant — it records what it sees, and whatever it sees becomes true. Verification is a comparison and demands that two observations be commensurable. The same camera button serves both, which means the harder job inherited the easier job's assumptions. Anyone replicating this should design the verification capture first and derive the ingestion capture from it, not the other way round — the reverse order is what produced Section 7.1.

F3 The pre-arrival verification was the leg nobody thought aboutBlocking

Attention went to ingestion, because it is the visible build, and to checkout, because it is where money is at stake. The pre-arrival check sat between them and was treated as a lesser version of ingestion. It is not: it is the chain-of-custody handover, and it is the leg on which every guest-facing claim silently depends.

F4 The residual human task is the one to instrument

Automating identification left the human holding the camera. Framing turned out to be the step the whole system is most sensitive to, and it is the step that received no design attention, no guidance, and no measurement.

F5 Calibrated refusal is a feature and should be selected for

Closed box, contents unknown is worth more than a plausible invented list, because a fabricated line becomes a permanent phantom discrepancy in every subsequent comparison.

F6 Keeping the abandoned approach as a fallback was correct

The Manual tab is the reason this system has a working degraded mode when a frame fails or the API is down.

F7 The free accuracy dataset was nearly thrown awayBlocking

Every human correction at the review gate is a labelled example, generated at zero marginal cost by normal operation. It only exists if the pre-edit machine output is retained, and that decision has to be made before the data starts flowing, not when someone asks for accuracy figures.

F8 A build like this is now measured in days, not quarters, and that changes which problems are worth solving

At roughly $3,500 all-in, the question is no longer whether the business case clears a capital threshold. It is whether the operational discipline exists to use the thing — which is exactly what Section 7.1 is about.

9. Replicability

Yes, conditionally — and the conditions are organisational rather than technical.

◇ What a replicating organisation must already have:

  1. A location hierarchy that is fixed before capture — property → space, and ideally property → space → zone. Without it, counts are not comparable and none of the rest works.
  2. An existing inventory record per space to reconcile against. The Exists badge, and therefore the whole anti-duplication property, depends on there being something to check against.
  3. Acceptance of a human commit gate on every write. An organisation that wants unattended writing to its inventory system should not build this.
  4. A staff practice that can absorb a framing discipline. This is the binding constraint, and it is the one most likely to be underestimated. The technology transfers in days; the capture discipline does not.
  5. A clear-eyed separation of recording from verifying. An organisation that wants only a register — what do we own, where is it — can deploy this as-is and get full value today. An organisation that wants verification against a reference state is buying a harder product and should read Section 7.1 as a specification of what still has to be built.

◇ What transfers directly. The pattern — scope, capture, single multimodal call with structured output, reconcile against existing record, human-gated commit — is domain-general. It applies to warehouse spot-counts, retail shelf audits, facilities and asset registers, laboratory consumables, and site-safety equipment checks. Nothing about it is hospitality-specific.

◇ What does not transfer. The tolerance for a four-second latency and a wide frame suits a low-value, high-count, low-consequence item mix. High-value serialised assets need identity, not class-and-count — a different problem requiring tags or serial capture, not this design.

◇ Cost of replication. Materially lower than the ~$3,500 in this case, because the abandoned paths and the design iteration are now documented. A second implementation is ◇ 15–20 hours plus the underlying inventory system.


10. Future Roadmap

Near term — close the gap in Section 7.1. This precedes everything else on this list.

  1. ◇ Ghosted baseline overlay in the camera modal
  2. ◇ Named capture zones per space, with coverage completion checks
  3. ◇ Three-state item status: present / not present / not seen
  4. ◇ Spurious-delta logging as a first-class KPI, split by leg (pre-arrival / checkout)
  5. ◇ Retention of pre-edit machine output as the accuracy dataset
  6. Pre-arrival verification as a distinct, named operation with its own entry point, completion state and record — and ◇ a gate on marking a space ready for arrival

Medium term

  1. ◇ Delta screen with baseline and current frames displayed side by side for adjudication — used at both legs
  2. ◇ Condition-change corroboration rule — no damage event on a single wide frame
  3. ◇ Inference call abstracted behind an internal interface ahead of the Gemini 2.5 retirement
  4. ◇ Consumable replenishment triggered by count thresholds
  5. ◇ Opaque-container contents declaration at first sighting

Longer term

  1. ◇ Linkage to the reservation and billing chain documented in Case JD-01, so an adjudicated checkout delta becomes a charge line. Note the ordering: this must not ship before items 1–8. A guest charge is only as sound as the pre-arrival verification it is measured against — wiring a system with a known spurious-result rate and an unaudited reference state directly to guest billing is the specific thing not to do.
  2. ◇ Portfolio analytics — attrition rates by item class, by space, by property
  3. Agent-initiated capture. An apeople prompting for a count at turnover rather than a human remembering to run one. This is the threshold at which this case becomes an apeople case and the eight gates of EAIS-03 begin to apply. It is not one today, and this document should not be read as implying otherwise.

Roadmap

  • Near term: Ghosted baseline overlay in the camera modal
  • Near term: Named capture zones per space, with coverage completion checks
  • Near term: Three-state item status: present / not present / not seen
  • Near term: Spurious-delta logging as a first-class KPI, split by leg (pre-arrival / checkout)
  • Near term: Retention of pre-edit machine output as the accuracy dataset
  • Near term: Pre-arrival verification as a distinct, named operation with its own entry point, completion state and record — and a gate on marking a space ready for arrival
  • Medium term: Delta screen with baseline and current frames displayed side by side for adjudication — used at both legs
  • Medium term: Condition-change corroboration rule — no damage event on a single wide frame
  • Medium term: Inference call abstracted behind an internal interface ahead of the Gemini 2.5 retirement
  • Medium term: Consumable replenishment triggered by count thresholds
  • Medium term: Opaque-container contents declaration at first sighting
  • Longer term: Linkage to the reservation and billing chain documented in Case JD-01, so an adjudicated checkout delta becomes a charge line. Must not ship before items 1-8.
  • Longer term: Portfolio analytics — attrition rates by item class, by space, by property
  • Longer term: Agent-initiated capture. This is the threshold at which this case becomes an apeople case and the eight gates of EAIS-03 begin to apply.

Open items

  1. Case ID renumber — JD-02 → Case 02 / UC02. Propagate to Ravi's seed data before the public route ships.
  2. §7.1 remediation sequencing — confirm items 1, 3, 7 and 8 as the near-term build, and confirm the interim rule that both verification legs are advisory only. 2a. Pre-arrival verification as a named operation — decide whether it gates "ready for arrival." This is the chain-of-custody decision and should be taken deliberately, not inherited.
  3. Retention of pre-edit machine output — decide and implement before scale-up. It cannot be backfilled.
  4. Frames per space — measure. The missing multiplier on the four-second figure.
  5. Spurious delta rate — instrument. The number that decides whether the checkout half is usable.
  6. Manual-first sequence — confirm §3 what was tried and abandoned is factually correct.
  7. Image retention and access policy — 90-day proposal to be confirmed; guest-belongings exposure to be settled explicitly.
  8. Gemini 2.5 retirement — abstract the inference call ahead of the October 2026 cutover.
  9. EAIS-01 stage names — substitute into §11 once fixed.
  10. Verification screenshots — both the pre-arrival check and the checkout comparison are described but not evidenced. Needed for The Story; the pre-arrival one is the more important of the two.
  11. Redaction map — internal property identifiers and admin UI appear in all five figures.

3. AI Approach and Deployment Choices

The harness

The model contributes one capability: it looks at a photograph and says what is in it. Everything that turns that into an inventory system sits in the harness around it.

Harness layer What it does Source
Scoping Property → room/space fixed before capture Figs. 1–2
Capture In-app camera with live preview; also file upload and clipboard paste (Ctrl+V / Cmd+V) Figs. 1, 3
Submission Analyze Image with AI; ~4 s round trip Fig. 4
Identification Single Gemini 2.5 Flash call returning name, quantity, condition, location, notes Fig. 5
Reconciliation Detected items matched against existing inventory for that space; duplicates badged Exists (Qty: n) Fig. 5
Human gate Editable table; explicit review before saving; per-row edit action Fig. 5
Persistence Commit to the space's inventory record on Supabase
Fallback Manual tab retained throughout Figs. 1, 2, 4
Differencing Pre-arrival and checkout captures compared to baseline on quantity and condition ◇ — built; not yet evidenced by screenshot
Frame registration Nothing enforces that a verification capture is commensurable with the baseline it is compared against Absent — §7.1
Audit Captured frame retained against the inventory event

Three harness decisions did the work.

Scope before capture. Location is not a field the model guesses; it is a constraint the interface fixes first. This single sequencing choice is why the output is comparable across rooms and across stays.

The existing record is in the room at analysis time. Loading the 37 known items adjacent to the analysis is what makes the Exists badge possible. The alternative — analyse first, reconcile later in a batch job — produces the same information hours later, to nobody, at a point where the operator has left the store room and cannot resolve the ambiguity by looking.

The commit is human and the interface says so. Not a confirmation dialog, but an editable table with a heading that states the contract.

◇ Retention of the pre-edit machine output. Recommended and not believed to be implemented. The delta between what the model returned and what the operator saved is a labelled accuracy dataset generated free by normal operation. Capturing it costs one additional column; not capturing it means the accuracy figures in Section 6 can never be filled retrospectively. This is the single cheapest improvement available and it should be made before scale-up, not after.

AI-executed versus AI-assisted

Step Mode
Property and space scoping Human
Image capture and framing Human
Item identification, unit count, condition grade, spatial location AI-executed
Reconciliation against existing inventory AI-executed, surfaced for human decision
Row-level correction and commit Human
Pre-arrival verification capture and framing Human
Pre-arrival and checkout delta computation AI-executed
Delta adjudication, damage attribution, guest charge Human ◇

The load is shifted, not removed. The human no longer enumerates; the human adjudicates. That is a job-content change and is reported as one in Section 4. It is also, per Section 7.1, where the residual risk now sits: the human's remaining task — framing the shot — turns out to be the task the system is most sensitive to and gives the least guidance on.

On-premise versus cloud

◇ Fully cloud. Application built and hosted on Lovable with Supabase as the system of record; inference via the Gemini API; captured images in cloud object storage.

◇ The decisive factors, and what the choice cost. No regulatory constraint binds the item data itself — a count of glassware carries no personal data and no residency obligation. Latency at four seconds is comfortably inside the operator's tolerance. The property has no infrastructure, no operations staff, and no appetite for either, so the on-premise option would have required creating a capability rather than using one. Capital expenditure is nil and time-to-first-run was days rather than months.

◇ What it bought and what it costs. It bought speed of build (Section 7) and zero fixed cost. It costs a consumption line that scales with adoption rather than sitting flat, and — the material one — it places photographs of the interior of guest-occupied rooms in third-party cloud storage. Room images may incidentally capture guest belongings, documents or the guests themselves. That is a privacy exposure the item data alone does not carry, and it is the reason image retention policy (Section 13) is a live decision rather than a formality.

Contrast with Case JD-01, which is deliberately hybrid: JD-01 holds credentials and acts across five systems, so the harness sits on-premise. This case holds no credentials and acts on nothing. Different risk, different shape. The comparison is instructive: deployment shape should follow what the system can do if it goes wrong, not what technology it uses.

Training and adaptation method

◇ Off-the-shelf, prompt-engineered, with the existing item list supplied in-context. No fine-tuning, no custom training, no labelled dataset. Gemini 2.5 Flash is called as-is with a structured-output instruction defining the return schema — name, quantity, condition, location, notes — and the space's current item list supplied alongside the image so the model can flag matches against it.

◇ What was tried first and abandoned. The manual entry form was built first, and it still ships — the Manual tab visible in every screenshot is its residue. It was the observed cost of that path in real use, not a hypothesis about AI, that motivated the vision approach. The abandoned approach was retained as the fallback rather than deleted, which is why the system has a working degraded mode. ◇ Confirm the sequence — if the AI path was designed first and Manual added as fallback, this paragraph inverts and should be rewritten.

Why no fine-tuning. ◇ The item classes are ordinary household and hospitality objects that a general-purpose multimodal model already recognises. Fine-tuning would buy accuracy on the property's specific SKUs at the cost of a labelling programme, a training pipeline, and a model that must be retrained whenever the property buys something new. For a portfolio of three properties the arithmetic does not close. It would begin to close at portfolio scale, or if the item mix were genuinely unusual.


11. EAIS Stage

Field Value
Stage entered from Stage IIa — a single source of truth existed. The inventory record for each space was already in Casa Admin (37 items on record for the Stores at 17VC), with a fixed property→space hierarchy.
Stage reached Stage III — AI-assisted execution embedded in an operational workflow, human-gated at commit. Not Stage IV: no agent holds the role, no autonomous trigger, no delegated authority.
Prerequisites that had to be true first A canonical location hierarchy (property → room/space) and an existing inventory record per space. Both are Stage IIa single-source-of-truth dependencies. Without them there is nothing to reconcile against and the Exists badge cannot exist.
The move this case demonstrates ◇ Stage IIa is not a preliminary to AI work; it is the substrate that makes AI work reconcilable. The most useful behaviour in Figure 5 — self-checking against 37 known items — is a Stage IIa dividend, not a model capability.

Stage numbering follows EAIS-01. Stage names for I–VI are pending from the EAIS-01 SSOT and should be substituted here once fixed — this also resolves the open item on FIG. 1 of the patent drawings.


12. Actor Class

Field Value
Actor hpeople, operating an AI capability embedded in Casa Admin
The line — tool used, or actor governed? A tool that people use. No agent holds this role. There is no aJD, no Constitution, no Soul, no persistent identity, no memory across runs. A human scopes, captures, reviews and commits every time.
Contrast with Case JD-01 JD-01 documents an apeople — Jeni De — holding a role and executing a chain across five systems from a single instruction. This case documents hpeople using a capability. The two cases sit at opposite sides of the tool/actor line and are deliberately paired for that reason.

Case ID consequence. Because no apeople holds the role, the JD track prefix does not apply. This case is numbered Case 02, paper ID EAIS-CASE-02, slug UC02. This affects the Supabase uc_case.slug in the seed data already delivered to Ravi — one line, but change it before the public route is built at /usecase/.

Roadmap item 13, agent-initiated capture, is precisely the point at which this becomes an apeople case and the eight gates start to apply.


13. Governance

Not assessed against the eight gates of EAIS-03, because no apeople holds this role (field 12). The governance questions that do apply, with proposed answers:

Question ◇ Proposed
Who may commit an inventory change? Authenticated Casa Admin users with property access. Every commit carries user, timestamp, space and source frame.
Is a human override of the AI result logged, with the original machine output retained? Recommended, believed not implemented. Should be closed before scale-up — see §3 and Lesson 6.
Who owns the canonical item vocabulary? Property operations. First human-accepted name for a class becomes canonical for that space (§7.7).
Are captured images retained as audit evidence, and for how long? Retained against the inventory event. Proposed retention: 90 days for routine captures; indefinite where a delta was adjudicated or a guest charge raised.
Who may view images that may contain guest belongings? Restricted. Property operations and admin only. Not exposed on any public route. This is the sharpest privacy question in the case and should be settled explicitly rather than inherited from the default table permissions.
What happens on a disputed checkout delta — who adjudicates? A named human, viewing baseline and current frames side by side. No delta becomes a guest charge automatically (§7.1, remediation 6).
Who is accountable for a wrong count that reaches a guest? The committing user, not the model. The interface states the contract — review and edit items before saving — and the commit is a human act.
What is the off-switch? The Manual tab. If the AI path is disabled or the API is unavailable, the workflow degrades to manual entry rather than stopping.

Governance gates

GateStatusDetail
Who may commit an inventory change?HeldAuthenticated Casa Admin users with property access. Every commit carries user, timestamp, space and source frame.
Is a human override of the AI result logged, with the original machine output retained?Not heldRecommended, believed not implemented. Should be closed before scale-up — see §3 and Lesson 6.
Who owns the canonical item vocabulary?PartialProperty operations. First human-accepted name for a class becomes canonical for that space (§7.7).
Are captured images retained as audit evidence, and for how long?PartialRetained against the inventory event. Proposed retention: 90 days for routine captures; indefinite where a delta was adjudicated or a guest charge raised.
Who may view images that may contain guest belongings?HeldRestricted. Property operations and admin only. Not exposed on any public route. This is the sharpest privacy question in the case and should be settled explicitly rather than inherited from the default table permissions.
What happens on a disputed checkout delta — who adjudicates?PartialA named human, viewing baseline and current frames side by side. No delta becomes a guest charge automatically (§7.1, remediation 6).
Who is accountable for a wrong count that reaches a guest?HeldThe committing user, not the model. The interface states the contract — review and edit items before saving — and the commit is a human act.
What is the off-switch?HeldThe Manual tab. If the AI path is disabled or the API is unavailable, the workflow degrades to manual entry rather than stopping.

Figure manifest

Fig. File Shows
1 fig-01_add-inventory-property-select.png Add Inventory; property quick-select 17VC / BF16 / 520; AI–Manual toggle
2 fig-02_room-space-select-ai-manual-toggle.png Property 17VC (lmc_ro), room Stores
3 fig-03_camera-capture-modal.png In-app camera, live preview, unstaged shelf of glassware
4 fig-04_image-preview-analyze.png Captured frame, Analyze Image with AI, Existing Inventory (37 items)
5 fig-05_ai-result-review-before-save.png Five detected items with quantity, condition, location, notes; Exists (Qty: 1) duplicate flag; review-before-save gate
◇ 6 pending Pre-arrival verification — space set to standard, checked against base
◇ 7 pending Checkout comparison — baseline vs current, delta surfaced

Open items

  • Case ID renumber — JD-02 -> Case 02 / UC02. Propagate to Ravi's seed data before the public route ships.
  • Section 7.1 remediation sequencing — confirm items 1, 3, 7 and 8 as the near-term build, and confirm the interim rule that both verification legs are advisory only.
  • Pre-arrival verification as a named operation — decide whether it gates "ready for arrival." This is the chain-of-custody decision and should be taken deliberately, not inherited.
  • Retention of pre-edit machine output — decide and implement before scale-up. It cannot be backfilled.
  • Frames per space — measure. The missing multiplier on the four-second figure.
  • Spurious delta rate — instrument. The number that decides whether the checkout half is usable.
  • Manual-first sequence — confirm section 3 'what was tried and abandoned' is factually correct.
  • Image retention and access policy — 90-day proposal to be confirmed; guest-belongings exposure to be settled explicitly.
  • Gemini 2.5 retirement — abstract the inference call ahead of the October 2026 cutover.
  • EAIS-01 stage names — substitute into section 11 once fixed.
  • Verification screenshots — both the pre-arrival check and the checkout comparison are described but not evidenced. Needed for The Story; the pre-arrival one is the more important of the two.
  • Redaction map — internal property identifiers and admin UI appear in all five figures. No screenshot image files have been supplied to this pipeline yet either.

Document control

Case
02
Version
0.4.0
Series
Enterprise AI Series
Track
Use Cases / Case studies
Author
Alok Sinha
Operating mode
AI-assisted — human-scoped, AI-executed identification, human-committed.
Run date
3 August 2026
Classification
Review before external circulation. Figures show internal property identifiers (17VC, BF16, 520, lmc_ro) and an internal admin interface. Room photographs may incidentally capture guest belongings. Redact identifiers before FICCI submission or publication on the public /usecase/ route. Redaction map to be prepared. No screenshot image files have been supplied to this pipeline yet.
Primary objective
Count completeness and condition capture; turnover-window time recovery — not cost reduction.

Deployment

  • Interfacecloud: Casa Admin, mobile web — built on Lovable
  • Modelcloud: Gemini 2.5 Flash via API, one call per image
  • Application and datacloud: Supabase
  • Image storagecloud: Cloud object storage, retained against the inventory event

Economics

Setup tokens
9,800
Setup cost
≈ $3,540
Tokens per call
2,5002,500
Cost per call
≈ $0.00
Harness
9,800 Lovable credits + 48 developer man-hours (~$3,540 all-in one-time build). Per-scan inference ~2,500 in / 400 out tokens, ~$0.002/scan (Gemini 2.5 Flash). See Document control -> Economics for full breakdown.

Changelog

  • v0.4.0 · 3 August 2026Section 7.1 restructured around three distinct operations: (A) ingestion, (B) pre-arrival setup verification, (C) checkout count. The gap is located on B and C -- the verification legs -- and expressly not on A. B named as the chain-of-custody handover and the more consequential leg; fifth asymmetry added (an error at B is a wrong premise, not a wrong reading). Remediation item 8 added. Interim advisory rule extended to both legs. Distinction propagated through sections 1, 2, 3, 4, 6, 7.2, 8, 9, 10, figure manifest and open items.
  • v0.3.0 · 3 August 2026Complete draft. All placeholders replaced with proposed text under the diamond convention. Economics computed from actual build inputs (9,800 Lovable credits, 48 developer man-hours) and Gemini 2.5 Flash published rates verified 3 Aug 2026. New section 7.1 documents the frame-registration gap at checkout as the principal unsolved problem, with seven-item remediation ranked by leverage and an interim advisory-only rule. Model retirement risk added (7.4). Lessons, roadmap, replicability and governance rewritten around the gap.
  • v0.2.0 · 3 August 2026Rewritten against five production screenshots. The Story written in full. Case renumbered JD-02 -> Case 02 (UC02) on the actor-class finding. Figure manifest added.
  • v0.1.0 · 3 August 2026Skeleton against TPL-B in canonical order.