Model training

Custom Computer Vision Model Training Services

We train detection, segmentation and classification models on your own images, then prove the accuracy condition by condition before anything ships.
  • Annotation and QA included
  • Scores reported per class
  • Air-gapped training on request
run-07 · site-safety-detector (sample)
Illustrative run
2,260 images140 near-duplicates removed
Train 1,582Val 339Test 339

Labelled instances per class

  • person4,200
  • hard hat3,100
  • hi-vis vest2,700
  • forklift380

forklift: under 1,000 examples, collect more

1/4Images are de-duplicated, balanced and split before anyone labels them. Classes with too few examples are flagged early.

How it works

Training is a loop, not a one-off run

Every run tells us which class to collect more data for. That is the work.

  • 150+ models delivered
  • Proof of concept in 4–6 weeks
  • Best checkpoint kept and versioned
  1. Stage 1

    Curate the dataset

    Near-duplicates are removed, classes are balanced and the train, validation and test split is fixed before any labelling starts.
    Clean, split dataset
  2. Stage 2

    Annotate and review

    Boxes, masks or keypoints are drawn to written guidelines, then a second annotator checks and corrects every frame.
    Reviewed ground truthCVATRoboflow
  3. Stage 3

    Train from a pre-trained checkpoint

    Transfer learning cuts how much of your data we need. Architecture and hyperparameters are benchmarked rather than guessed.
    Tracked training runsPyTorchMLflowWeights & Biases
  4. Stage 4

    Evaluate per class and condition

    A held-out test set gives a score for every class and capture condition, so a good average can't hide a class the model keeps missing.
    AP per class, failure cases
  5. Stage 5

    Mine hard cases and retrain

    Frames the model gets wrong are collected, labelled and fed into the next run. The same loop runs again once the model is live.
    Next run, measurably better

Architecture

The architecture follows the question

We shortlist by what the decision needs and what your hardware can run, then benchmark on your data.

  • YOLO family for real-time detection
  • U-Net and Mask R-CNN for exact boundaries
  • Vision transformers for context-heavy scenes
  • Vision-language models where text and image mix
How we compare candidate models
Flowchart branching from an inspection goal to object detection with YOLO or Faster R-CNN, image segmentation with U-Net or Mask R-CNN, classification models such as ResNet and EfficientNet, and a multimodal system combining images, drone, thermal and metadata
A model-selection flowchart from our property inspection guide; the same logic applies to any visual task.

Data

Accuracy is a data problem first

Most gains come from better labels and the frames nobody thought to collect.

  • Your own frames

    Captured from your cameras, including the awkward ones.
  • Annotation and QA

    Written guidelines, then a second reviewer on every frame.
  • Synthetic rare cases

    Generated images for defects too rare or risky to film.
  • Active learning

    The model flags what it finds confusing; those get labelled next.

Optimisation

Trained to fit the device it runs on

A model is only finished when it hits the frame rate on your hardware.

  • Quantisation and pruning to cut size
  • Knowledge distillation into a smaller model
  • TensorRT, ONNX and Core ML export
  • Measured on the target board, not a workstation
TensorRT throughput measured in our Jetson tutorial
ModelOrin NX, 20 WAGX Orin, 50 W
YOLOv8 Nano145 FPS420 FPS
YOLOv10 Nano160 FPS465 FPS
YOLO-World Small38 FPS115 FPS

Figures from our YOLO on Jetson tutorial. Your numbers depend on input size and camera count.

Getting started

Start with the data you already have

Training can run in your own cloud account or fully air-gapped on your servers. See private AI deployment and how we handle IP.

  1. Step 1: Data review

    30 minutes

    We look at a sample of your images and say whether they can support the task yet.

    • NDA on request
  2. Step 2: First trained model

    4–6 weeks

    Annotation, a training run and a test report per class against the agreed target.

    • Target set before we start
  3. Step 3: Retraining in production

    Ongoing

    New failure cases are labelled and folded into scheduled retraining runs.

    • You own the weights

FAQ

Questions, answered

What teams ask before handing over a dataset.

Need the whole build, not just training? See custom computer vision models.

  • Yes. We specialize in computer vision model retraining services to improve computer vision model accuracy as your production data evolves. We identify accuracy bottlenecks and re-train using optimized hyperparameters and targeted data augmentation.

Book a strategy session

Talk to an AI engineer about your project

Tell us what you want to automate. The first call is a 30-minute working session with an engineer, not a sales pitch.

  • Send the form, it takes 2 minutes
  • We reply within 1 business day, under NDA if you need it
  • A 30-minute call to scope feasibility and next steps

Tell us about your project

By submitting, you agree to our Privacy Policy. We never share your details.