Skip to content
Menu
Applied AI 6 min read ·

Cameras on the dock: choosing between Jetson, Hailo and a GPU server

Bizmap engineering team
The stance
Camera count, privacy and who maintains the hardware decide where vision runs; the model is the easy part.
A question about this?Ask the team that wrote this. An engineer, not a ticket system, replies.Talk to an engineer →

A camera on a loading dock can tell you when a truck docks, how long a bay stays open, whether someone has walked into a forklift lane, or how many pallets went through a door. The model that detects a truck or a person is now the easy part. The harder decision is where the model runs, because that choice settles cost, privacy, reliability and who has to maintain it for the next five years.

Vision AI is in build at Bizmap. It does not run at a client yet, so this is not a deployment report and there are no benchmark figures in it. It is the decision guide we use when a dock, a line or a yard asks which hardware to plan for.

Four places the model can run

On the camera. Some cameras carry their own inference chip; Sony's IMX500 sensor is one example. The camera sends events or metadata rather than video. Nothing else to install, very little bandwidth, and footage need not leave the device. The trade-off is that the models must be small and are tied to that camera's toolchain.

On an edge box near the cameras. A small computer with an accelerator takes a handful of RTSP streams over the local network and runs detection and tracking. NVIDIA's Jetson Orin modules, with the DeepStream video pipeline, are the common choice. Hailo's accelerators (the Hailo-8, for example) attach to a host such as an industrial PC and are designed for low power draw.

On a GPU server on site. One machine with a data-centre or workstation GPU takes every camera stream in the building. It runs larger models, more streams per box, and can also be used to retrain.

In the cloud. Streams or frames go to a cloud GPU. It is the easiest to start and the hardest to justify for continuous video, for reasons below.

The questions that decide it

Can the footage leave the premises?

Ask this first, because it removes options. Many warehouses and plants will not allow video of their staff to leave the site, and for some customers or sectors it is a contract term. If the answer is no, the cloud is out and the choice is between the camera, edge boxes and an on-site server. Our default for vision is that detection happens on the camera or on site, and only events (a timestamp, a zone, a class, a snapshot if agreed) are sent to the business system.

How many cameras, and how spread out?

A single dock with a few cameras is an edge-box problem. A site with dozens of cameras across docks, yard and production starts to favour a central server, because one well-managed machine is easier to look after than many small ones. Cameras spread over buildings a long way apart push back towards edge boxes, because carrying many video streams across a site network is its own project.

How fast does someone need to know?

A safety event, a person in a forklift lane or a door left open on a cold store, needs to reach someone within seconds, and should not depend on an internet link. That argues for inference on site. A daily report of dock dwell times can tolerate delay and could be processed anywhere.

How big does the model need to be?

Counting pallets or detecting trucks and people works with compact detection models that run comfortably on edge hardware. Reading damaged labels, classifying many product types or tracking across several overlapping cameras needs larger models and more memory. Jetson runs standard CUDA and TensorRT, so most models can be deployed with modest conversion work. Hailo requires compiling the model with its own toolchain, which supports a narrower set of model operations; models that compile run efficiently, and models that do not need changing. A GPU server takes almost anything.

Where will it sit, and what power is there?

A dock is dusty, often unheated or hot, and sometimes vibrates. Edge boxes need enclosures suitable for that environment, a power supply and somewhere to mount them. A GPU server needs a proper room with cooling and a UPS. Power draw matters on a remote site or a solar-backed one, and it is where low-power accelerators earn their place.

Who maintains it?

This is the question buyers ask last and should ask first. Every edge box is a computer with an operating system, drivers, a model version and a network connection to keep patched. Ten boxes means ten of everything. A central server concentrates the maintenance but also the risk: when it is down, every camera is blind. Either way, someone must own updates, monitoring and spares.

A side-by-side view

Factor On the camera Edge box (Jetson or Hailo) GPU server on site Cloud
Footage leaves the site No No No Yes
Model size Small Small to medium Large Large
Cameras per unit One A few Many Many
Works without internet Yes Yes Yes No
Power and space Minimal Low, needs enclosures A room, cooling, UPS None on site
Maintenance Firmware per camera A fleet of small computers One machine, one point of failure Provider plus bandwidth
Changing the model Constrained Moderate (Hailo needs a compile step) Easy Easy

The table describes the shape of each option, not measured performance. Vendor figures for throughput and power draw exist for each device and are worth reading, but they are measured under the vendor's conditions; the only numbers we will quote to a client are ones measured on their cameras and their footage.

Why the cloud rarely wins for continuous video

Continuous video is heavy. Uploading it costs bandwidth every minute of the day, a link failure stops detection, and the footage leaves the site. The cloud is a good place for training, for occasional batch analysis of recorded clips, and for the dashboards that summarise events. It is rarely the right place for the inference itself on a dock that runs all day.

What the model should do, and what it should not

Whichever hardware runs it, the system raises events and a person acts on them. A truck docked at bay four, a door open longer than the agreed limit, a person in a restricted zone: each becomes a record in the business system that a supervisor sees and closes. The model does not stop a conveyor, post a goods receipt or discipline anyone.

Two working rules follow from how we deliver and govern AI:

  • Agree the number first. Events caught, false alarms per shift, minutes of dwell time recovered: whatever the measure is, it is baselined before anything is built.
  • Keep an eval set from the site. A few hundred labelled frames from the client's own cameras, in their light and their weather, is what the model must pass before it goes live and after every model change. A model that scores well on public datasets and has never seen your dock at dusk has not been tested.

A short decision path

  1. If footage cannot leave the site, rule out the cloud.
  2. Count the cameras and map where they are.
  3. A few cameras in one place, small models: an edge box, or cameras with built-in inference if the model fits.
  4. Many cameras on one site, larger models, or a plan to retrain: a GPU server on site, with a plan for when it fails.
  5. Low power available, or a model that compiles cleanly to the accelerator: consider Hailo; otherwise Jetson for flexibility.
  6. Whatever you choose, name the person who maintains it before you buy it.

If this is your situation

If you have cameras on a dock, a line or a yard and want to know whether vision AI fits, start with an AI fit check: one workflow, a baseline and a build plan, with a supervisor at the checkpoint. Read how we deliver and govern AI, then tell us about your site.

Start a conversation

Want this for your operation?

Tell us which case looked closest to yours.