Back to stuffs
Computer Vision

YOLO Object Detection & Architecture

•4 min read
#computer vision#deep learning

png

I recently took a deep dive into the YOLO (You Only Look Once) architecture. Originally introduced by Joseph Redmon in 2015, I think it completely changed the game for object detection. Instead of treating detection as a classification task on multiple regions, YOLO does everything in a single pass, making it fast and efficient.


What I Learned About YOLO’s Architecture

The biggest takeaway? YOLO is fast. Instead of scanning an image multiple times (like R-CNN-based models), YOLO looks at the image once, divides it into a grid, and predicts bounding boxes and class probabilities at the same time. Perfect for real-time applications.

Key Things I Found Out:

  1. Single-stage detection: Unlike older models that use multiple passes, YOLO does it all at once.

  2. Grid-based prediction: The image is split into a grid (e.g., 7x7 in YOLOv1), and each cell predicts objects within it.

  3. Regression-based output: Instead of generating region proposals, YOLO directly predicts bounding boxes and class labels.

  4. Backbone evolution: The first YOLO used Darknet, but newer versions bring in CSPDarknet and EfficientNet.

png


How It Has Evolved Over Time

YOLOv1

It was back in 2015 that Joseph Redmon and his team introduced YOLOv1. Unlike traditional object detection models, which relied on region proposals (like R-CNN and Faster R-CNN), YOLO did something radical: it treated object detection as a single regression problem.

The idea was simple—split an image into an S×S grid, and for each cell, predict bounding boxes and class probabilities. Since everything was done in a single pass, YOLO was insanely fast, achieving real-time detection. However, it struggled with small objects and overlapping ones. But the foundation was laid.

png

YOLOv2 (YOLO9000)

The next iteration, YOLOv2, came in 2016. This one improved on its predecessor in multiple ways:

  • Batch Normalization: Reduced overfitting and stabilized training.

  • Anchor Boxes: Inspired by Faster R-CNN, this helped with detecting objects of varying sizes.

  • High-Resolution Classifier: Pre-trained at higher resolutions for better feature extraction.

  • Multi-Scale Training: Trained on images of different sizes to improve robustness.

png

And then there was YOLO9000, which could detect over 9000 classes using WordTree, a hierarchy of object classes. This was a game-changer, though still lacking in accuracy compared to two-stage detectors. More objects, better performance.

YOLOv3

By 2018, YOLOv3 arrived with several refinements:

  • Darknet-53 as the backbone, replacing Darknet-19, improving feature extraction.

  • Multi-Scale Predictions: Used three different scales for detecting small, medium, and large objects.

  • Binary Cross-Entropy Loss for Multi-Label Classification.

png

YOLOv4

Then came YOLOv4, released in 2020 by Alexey Bochkovskiy (since Redmon had stepped away from computer vision due to ethical concerns). This version was packed with engineering tricks to boost performance:

  • CSPDarknet53: A more efficient backbone.

  • Mish Activation: A smoother activation function for better gradient flow.

  • Spatial Pyramid Pooling (SPP) and Path Aggregation Network (PANet): Enhanced feature fusion.

  • Mosaic Augmentation: Combined multiple images to improve robustness.

png

YOLOv5

YOLOv5 was released by Ultralytics and it wasn’t an official continuation of YOLO but quickly became one of the most used versions because it was incredibly lightweight and easy to use. Key improvements are:

  • Implemented in PyTorch instead of Darknet.

  • Smaller and faster, with different model sizes (YOLOv5s, m, l, x).

  • AutoAnchor Optimization: Automatically select the best anchor boxes for a given dataset. Improved detection accuracy and ensures the model adapts well to different datasets without requiring manual tuning.

png

YOLOv6:

YOLOv6 was introduced by Meituan, focusing on industrial applications. It refined previous ideas while adding:

  • Reparameterization: A technique that fuses multiple layers into a single layer during inference, making the model faster without affecting accuracy.

  • Efficient Backbone Networks: Designed to improve computational efficiency while retaining detection accuracy, making YOLOv6 more lightweight and scalable.

png

YOLOv7

Released in 2022, YOLOv7 added:

  • Extends Efficient Layer Aggregation Networks (ELAN): Improves feature reuse and gradient flow, making training deeper networks more efficient.

  • Model scaling without sacrificing speed: YOLOv7 optimizes scaling to improve accuracy while maintaining real-time performance.

png

YOLOv8

Then came YOLOv8, developed by Ultralytics in 2023. It built upon the strengths of previous YOLO versions while introducing major improvements:

  • Enhanced Backbone and Neck Architecture: Improved feature extraction and multi-scale feature fusion.

  • Instance Segmentation: YOLOv8 is not just an object detector but can also segment objects within images.

Better Edge Deployment: Optimized for deployment on mobile and edge device.

png