Weekly Update · July 14-21, 2026

This week, PortMind turned review work into a usable product.

An invited reviewer can now inspect a port image, record a decision, and save the result. The next step is to finish the first checked set and take it into customer conversations.

What Moved

Three concrete steps moved the project forward.

01 · SHIPReviewer appInvites, instructions, image inspection, decisions, and saved notes now work in one flow.
02 · PACKAGEFirst reference setThe immediate deliverable is a small, checked set of port images that we can show and test.
03 · ALIGNTwo callsCounsel today; Simon Savard at the Port of Montreal tomorrow.

The focus is execution: finish the first set, learn from it, and use the result to guide the next customer conversation.

Roadmap

Build the port-data standard, then turn it into products.

01 · CollectPort datasetAdd time-stamped images and events from each port we serve.
02 · LabelPort assetsMark tankers, trucks, vessels, and hard cases with reviewer checks.
03 · BenchmarkPort tasksCompare systems on detection, classification, and bounding boxes.
04 · ReleasePublic + privateShare a public benchmark; sell private evaluations and labeling.
05 · ApplyPort toolsUse the history to build monitoring, alerts, and forecasts.

The open release is a defined benchmark and supporting tooling; private customer work stays private. The first benchmark starts with five systems and expands as review capacity grows.

Evaluation

Measure every signal, with a human in the loop.

Established benchmark72 hard-case rows
historical comparison
LabelerTruckContainer truck
Human reference
72 rows
ReferenceReference
Codex
72 rows
59.7%50.7%
Grok 4.5
72 rows
77.6%67.2%

Agreement against the human reference. Useful as a first-pass signal; human review remains required.

This week’s open-model pass24 calibration rows
agreement vs Human 1
ModelTruckContainer truck
Llama 3.2 Vision95.7%82.6%
LLaVA 1.5100%82.6%
Moondream 3.1100%78.3%
Llama 4 Scout26.1%91.3%
Mistral Small 3.1100%87.0%
Second human labeler9 / 16 known-answer controls

56.3% scene agreement; consistent, but failed qualification. Retrain before promotion.

Open-model scores use the same 24-row packet as Human 2; binary scores use 23 completed reference rows. Report-only calibration results, not production accuracy or benchmark truth.

Human Review

The reviewer app now turns an image into a saved, checkable decision.

Screenshots from the PortMind reviewer workspace. This is the app itself, not the onboarding video shown in the previous update. Next: complete the first checked set.

The Review Loop

Each image follows the same path from invite to usable record.

  1. InviteSend a reviewer a scoped link to the assigned work.
  2. Set the rulesExplain what to look for and what a good decision includes.
  3. InspectUse the original image and grid view to check difficult scenes.
  4. RecordSave the decision, confidence, image quality, and a short note.
  5. PackageMove checked records into the reference set and customer demo.

Next Conversations

Use the next two calls to remove the next blockers.

Today · Counsel

Choose the structure

Leave with a clear project structure, ownership, and the next workstream to execute.

Tomorrow · Port of Montreal

Validate the first use case

Ask Simon Savard which port workflow is worth testing first and who should evaluate it.

Role reference: Port of Montreal Innovation Hub. Portrait: Startupfest public speaker profile.

Next Seven Days

Finish the first set, then use it to make the next sale.

  • Complete the reviewer flow. Open the first set and get every assigned image reviewed.
  • Summarize the results. Record where reviewers agree, disagree, or need clearer rules.
  • Prepare the customer demo. Show the checked set and one practical port use case.
  • Close the loop on both calls. Leave with owners, dates, and one next test.