Maciej Płoński
Streszczenie
Celem artykułu jest empiryczna ocena użyteczności wybranych architektur głębokich sieci neuronowych w zadaniu wykrywania broni na obrazach z monitoringu wizyjnego. Analizie poddano detektory jednoetapowe YOLOv5 oraz dwuetapowy Faster R-CNN, trenowane w identycznym reżimie na ograniczonym zbiorze danych. Eksperyment wykorzystuje zbiór z platformy Kaggle obejmujący 4000 obrazów przedstawiających broń palną i noże. Modele uczono z zastosowaniem transfer learningu na wagach wstępnie wytrenowanych na COCO, a skuteczność oceniano z użyciem mAP, precyzji, czułości (recall) oraz liczby klatek na sekundę (FPS). Uzyskane wyniki wskazują, że rodzina YOLO zapewnia korzystny kompromis między dokładnością a szybkością, umożliwiając przetwarzanie strumienia wideo w czasie zbliżonym do rzeczywistego.
Słowa kluczowe: sztuczna inteligencja, głębokie sieci neuronowe, detekcja broni, YOLO, Faster R-CNN, monitoring wizyjny
Abstract
The purpose of this paper is to conduct an empirical assessment of the usefulness of selected deep neural network architectures in detection of weapons in CCTV images. The analysis compared single-stage YOLOv5 detectors and two-stage Faster R-CNN detectors, each trained under identical conditions on a limited dataset. The experiment used a dataset from the Kaggle platform comprising 4,000 images of firearms and knives. The models were trained with transfer learning on COCO pre-trained weights, and their performance was assessed in terms of mAP, precision, recall and frames per second (FPS). The results obtained indicate that the YOLO family offers a favourable trade-off between accuracy and speed, enabling video stream processing in near real time.
Keywords: artificial intelligence, deep neural networks, weapons detection, YOLO, Faster R-CNN, CCTV monitoring
Identyfikator DOI: https://doi.org/10.34836/pk.2026.325.7