Multi-target detection and distance measurement integrated system based on YOLOv5

By optimizing image quality, dynamically adapting parameters, and accelerating computation through the YOLOv5 integrated system, the problems of high hardware cost, high error rate, and insufficient real-time performance of existing traffic detection and ranging systems are solved. This enables efficient and low-cost multi-target detection and ranging, meeting the needs of real-time monitoring and data analysis.

CN120928329AInactive Publication Date: 2025-11-11ANT-MAN TECH (SHENZHEN) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511447111.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2025-11-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing traffic detection and ranging systems require separate deployment of different devices, resulting in high hardware costs and complex system integration. They use fixed target sizes and device parameters to calculate distances, leading to high error rates. They are not adapted to different target types and scene changes, and the model inference process has not been optimized, failing to meet real-time monitoring requirements. Furthermore, the detection and ranging results do not generate structured data.

Method used

A multi-target detection and ranging integrated system based on YOLOv5 is adopted, including hardware and software modules. The hardware module optimizes image quality, uses NVIDIA Jetson AGX Xavier for computation, supports TensorRT acceleration, and a storage module is used for data storage; the software module achieves integrated detection and ranging by dynamically adapting parameters and combining a target type-actual size mapping library, generating structured data.

Benefits of technology

It achieves the goal of eliminating the need for additional ranging equipment, reducing hardware costs by more than 40%, controlling the ranging error rate to within 5%, increasing inference speed by 2-5 times, meeting real-time monitoring needs, and generating structured data to support traffic flow statistics and dangerous target early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120928329A_ABST
    Figure CN120928329A_ABST
Patent Text Reader

Abstract

The invention discloses a YOLOv5-based multi-target detection and distance measurement integrated system, which belongs to the technical field of detection and distance measurement systems and comprises a hardware module and a software module. Wherein the hardware module comprises a detection module, a calculation module and a storage module; the software module is used for detecting and ranging a plurality of targets in a traffic road section and judging the relative position relation between the targets, and the detection and ranging process comprises parameter configuration, data acquisition, reasoning calculation and multi-dimensional result output; the system realizes the integrated design of detection and distance measurement through a software algorithm, does not need to additionally deploy distance measurement equipment, improves the distance measurement accuracy through dynamic parameter adaptation and error correction, meets the requirements of traffic monitoring and other real-time scenes after the system is subjected to TensorRT acceleration, improves the real-time performance of detection and distance measurement, supports the docking of a data analysis platform, and improves the real-time performance of detection and distance measurement. The method can be applied to secondary applications such as traffic flow statistics and dangerous target early warning, and the system value is expanded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of detection and ranging system technology, specifically to an integrated multi-target detection and ranging system based on YOLOv5. Background Technology

[0002] In the field of traffic management, vehicle flow detection and ranging provide the data foundation for traffic flow analysis and safety control. Among current target detection technologies, YOLO series models such as YOLOv5 have been widely used for target recognition in image and video scenarios. However, existing solutions suffer from functional limitations, only achieving target classification and bounding box annotation, and failing to simultaneously acquire distance information between the target and the detection device. For example, a conventional YOLOv5 detection system in traffic scenarios can only identify targets such as vehicles and pedestrians. If target distance needs to be obtained, dedicated ranging equipment such as LiDAR must be deployed, resulting in high system hardware costs and increased deployment complexity. In addition, existing detection solutions that combine ranging functions mostly use fixed parameters to calculate distance, without considering differences in target size under different scenarios (such as the actual width of different vehicle models) and changes in device parameters (such as camera focal length adjustment), resulting in low ranging accuracy. At the same time, most solutions are not optimized for inference efficiency, and are prone to stuttering when processing real-time video streams, making it difficult to meet the needs of scenarios with high real-time requirements such as traffic monitoring and autonomous driving assistance.

[0003] For example, the patent with announcement number CN108830131B discloses a traffic target detection and ranging method based on deep learning, which includes: an algorithm model training step, selecting image region samples near vehicles, adding samples to the image, changing the pixels of the added samples to make them images with equal intensity, and finally performing deep learning training to obtain the model; an algorithm model prediction step, based on the trained model, selecting image region samples near vehicles, normalizing the contrast of the image region samples, and then performing deep convolutional neural network prediction to output image data; using deep learning methods to unify target recognition and ranging under a single network model, completing the detection and ranging of traffic targets, and achieving a balance between performance and efficiency in the system.

[0004] For example, patent CN104485002B discloses a traffic flow detection system based on PSD (Power Segment Detection). This system uses a PSD ranging module mounted on a bracket above the road to measure distances. Traffic flow is determined based on the change curve of the distance measured by the PSD; a shorter distance indicates the passage of a vehicle. The vehicle type is determined based on the magnitude of the distance change. The vehicle direction is determined by the time difference before and after a vehicle passes the PSD ranging module, and the vehicle speed is calculated. Finally, the measured traffic flow, vehicle type, traffic direction, and vehicle speed are wirelessly transmitted to a traffic flow management platform. This traffic flow detection system is simple to assemble and easy to operate.

[0005] For example, patent CN117079252A discloses a traffic sign detection and ranging method based on an improved YOLOv5, including: an improved YOLOv5 feature extraction backbone network, using Ghost instead of standard convolution, incorporating the top-level feature vector of the feature pyramid fusion structure into the EVC structure, adopting the Alpha-CIOU bounding box regression loss function, loading the network model, and combining it with a D435i depth camera in real-time detection within the application scenario to detect and recognize traffic sign image information and measure depth information within the scene. The image information detection and recognition results and depth information are fused to obtain the final detection and ranging results of the system. The traffic sign detection and ranging network model has improved generalization ability for multi-scale traffic sign detection and recognition, and at the same time, it is combined with a depth camera to measure the depth distance information between the traffic sign and the traffic sign, expanding the application function of traffic sign detection in real-world scenarios.

[0006] However, some existing detection and ranging systems require separate deployment of different devices for detection and ranging, resulting in high hardware costs and difficulty in system integration. They use fixed target size and device parameters to calculate distance, which does not adapt to different target types and scene changes. The error rate often exceeds 20%, and there is no optimization for the model inference process. When processing 1080P video streams, the frame rate is often below 15 FPS, which cannot meet the needs of real-time monitoring. The detection and ranging results are only displayed in the form of image annotations, without generating structured data, making it difficult to connect to subsequent data analysis systems.

[0007] To address the aforementioned issues, there is an urgent need for innovative design based on the existing detection and ranging system. Summary of the Invention

[0008] The purpose of this invention is to provide an integrated multi-target detection and ranging system based on YOLOv5, to address the problems mentioned in the background art. These existing detection and ranging systems require separate deployment of different devices for detection and ranging, resulting in high hardware costs, significant system integration difficulties, and the use of fixed target sizes and device parameters to calculate distances. This fails to adapt to different target types and scene variations, often leading to error rates exceeding 20%. Furthermore, the system lacks optimization for the model inference process, resulting in frame rates often below 15 FPS when processing 1080P video streams, failing to meet real-time monitoring requirements. The detection and ranging results are also only displayed as image annotations, without generating structured data, making it difficult to integrate with subsequent data analysis systems.

[0009] To achieve the above objectives, the present invention provides the following technical solution: a multi-target detection and ranging integrated system based on YOLOv5, comprising a hardware module and a software module.

[0010] The hardware modules include: The detection module is used to optimize image quality in low-light and backlight scenes in traffic sections, balance the brightness of the image, provide high-quality raw data for traffic target detection and ranging, and calibrate the detection focal length to adapt to different installation distance requirements. The computing module is responsible for model inference and data computation. It is equipped with NVIDIA Jetson AGX Xavier and supports TensorRT acceleration, providing computing power for multi-target detection and ranging, and processing a large amount of target detection data in parallel. The storage module is used for temporary storage of detection results and video streams. It can quickly read and write data and supports connection to cloud storage services. Detection data is stored synchronously in the local area and the cloud to achieve long-term data backup and management. The software module is used to detect and measure the distance of multiple targets in traffic sections, monitor people and vehicles in the traffic section area in real time, and determine the relative positional relationship between targets through the distance measurement function. The detection and distance measurement process includes parameter configuration, data collection, inference calculation, and multi-dimensional result output.

[0011] Preferably, in the parameter configuration, the user selects the application scenario through a web interface, and the system automatically loads preset parameters, specifically including: The detection parameters include the confidence threshold, the NMS IOU threshold, and the maximum number of targets to be detected. The confidence threshold for traffic scenarios is 0.4, the confidence threshold for indoor scenarios is 0.3, the default NMS IOU threshold is 0.45, and the default number of targets to be detected is 1000.

[0012] The ranging parameters are loaded, the target type and actual size mapping library are loaded, and the current camera focal length is input through the camera calibration tool. The default focal length pixel is 1500.

[0013] The dynamic adaptation mechanism for detection and ranging parameters includes establishing a target type-actual size mapping library, pre-setting differentiated actual widths for different targets such as pedestrians, vehicles, and buses, and supporting users to add custom target parameters; it also provides a camera focal length calibration interface to dynamically adjust focal length parameters according to the device's installation height and angle.

[0014] Preferably, the data acquisition adopts a multi-source acquisition method, with the detection terminal acquiring video streams or images in real time and transmitting them to the computing unit via the RTSP protocol, supporting the import of local video files and image files.

[0015] Deploy diverse detection terminals to ensure comprehensive coverage of targets, specifically as follows: Fixed monitoring equipment is deployed on both sides of the road to collect frontal and side images of vehicles and pedestrians at long distances; it is deployed at high points at intersections to support 360° rotation and zoom to track fast-moving targets; and it is deployed directly above the lane to collect rear features of vehicles, with supplementary lighting to ensure clear imaging at night; mobile acquisition equipment is installed on camera terminals at moving targets to collect dynamic views of the road surface.

[0016] Preferably, the data is preprocessed after acquisition, as follows: The image was resized to 640×640, and letterboxes were used to fill the image to prevent distortion.

[0017] Data normalization converts pixel values ​​from 0-255 to 0.0-1.0 and then converts them to Tensor format; Data augmentation, for static images, can enable random flipping and brightness adjustment enhancement strategies to improve the model's generalization ability.

[0018] Inference computation integrates detection and ranging into a single process, including target detection inference and dynamic distance calculation.

[0019] Preferably, the target detection inference specifically includes: Load the pre-trained YOLOv5 TensorRT engine and initialize the model through the DetectMultiBackend interface.

[0020] The model takes preprocessed image data as input and outputs the target bounding box, confidence score, and class ID.

[0021] Non-maximum suppression is performed to remove redundant bounding boxes with an overlap greater than the IOU threshold, while retaining the best detection results.

[0022] Preferably, the dynamic distance calculation specifically involves: Extract the pixel width x2-x1 of each target bounding box, where x1 and x2 are the x-coordinates of the left and right vertices of the bounding box.

[0023] The preset actual width is retrieved based on the target category ID. If the category is not in the mapping library, the distance is marked as N / A.

[0024] Distance is calculated using the similar triangle distance measurement formula: Distance = (focal length × actual target width) / target pixel width.

[0025] Preferably, the multi-dimensional result output specifically includes: Image annotation involves using the Annotator tool to draw bounding boxes on the original image, labeling the target name, confidence level, and distance, and using color to distinguish the target category.

[0026] Structured data storage generates CSV format result files, with fields including image, video frame name, target name, confidence level, distance, and detection time.

[0027] Real-time display and push: Displays annotated images and video streams using OpenCV, and supports exiting by pressing the 'q' key; it can also be configured with HTTP requests to push detection results to the backend management platform in real time.

[0028] Preferably, the optimization of the key algorithms of the software module specifically includes: TensorRT model acceleration converts YOLOv5 PyTorch models into the TensorRT engine, improving inference speed.

[0029] Layer fusion combines convolutional, activation, and batch normalization layers into a single computational unit, reducing data transmission overhead.

[0030] Precision quantization supports FP16 half-precision inference, improving inference speed by 2-3 times while maintaining less than 3% precision loss.

[0031] Dynamic batch processing automatically adjusts the batch size based on the amount of input data, balancing inference speed and memory usage.

[0032] Preferably, in the ranging error correction, when the target pixel width is less than 10, a multi-frame fusion strategy is enabled to take the average distance of 3 consecutive frames to reduce measurement fluctuations.

[0033] Compared with the prior art, the beneficial effects of the present invention are: Integrated design: Through software algorithms, detection and ranging are integrated. It can not only detect multiple targets such as people and vehicles in the monitoring area in real time, but also determine the relative positional relationship between targets through the ranging function, and judge violations such as speeding and driving in the wrong direction. No additional ranging equipment is required, reducing hardware costs by more than 40%.

[0034] High-precision ranging: Through dynamic parameter adaptation and error correction, the pixel width of the bounding box of each traffic target is extracted. The preset actual width is queried according to the target category ID, and the data in the mapping library is retrieved. If the detected target is not in the database, the target parameter error is corrected by the similar triangle ranging formula, so that the ranging error rate is controlled within 5%, which is better than the existing fixed parameter scheme.

[0035] Real-time performance improvement: TensorRT can improve the inference speed of YOLOv5 by 2-5 times through technologies such as layer fusion, precision calibration, and automatic kernel optimization. After TensorRT acceleration, the detection frame rate of 1080P video stream reaches more than 30 FPS, which meets the needs of real-time scenarios such as traffic monitoring.

[0036] High reusability: The structured output results generate CSV format result files with fields including image and video frame names, target names, confidence scores, distances, and detection times. It supports pushing results to the backend platform via an HTTP interface and can be used for secondary applications such as traffic flow statistics and dangerous target warnings, thereby expanding the value of the system. Attached Figure Description

[0037] Figure 1 This is a diagram showing the system module composition of the present invention.

[0038] Figure 2 This is a flowchart of the traffic multi-source target detection and ranging process of the present invention.

[0039] Figure 3 This is a size mapping library chart for the present invention. Detailed Implementation

[0040] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0041] Example 1: Please refer to Figure 1 The present invention provides the following technical solution: a multi-target detection and ranging integrated system based on YOLOv5, comprising a hardware module and a software module.

[0042] The hardware modules include: The detection module is used to optimize image quality in low-light and backlight scenes in traffic sections, balance the brightness of the image, provide high-quality raw data for traffic target detection and ranging, and calibrate the detection focal length to adapt to different installation distance requirements. The detection module employs a 2-megapixel high-definition camera with a built-in ISP image enhancement chip. The camera lens focal length can be calibrated via software (range 8-12mm). By strategically deploying the camera acquisition module along traffic routes, it comprehensively captures and detects pedestrians and vehicles. The 2-megapixel high-definition camera, a core component of the detection terminal, plays a crucial role. It supports 1080P resolution and can stably output images at a frame rate of 30FPS, capturing smooth and clear images, providing high-quality raw data for subsequent target detection and ranging. The camera's built-in ISP image enhancement chip can perform noise reduction and brightening in low-light environments, ensuring that objects in the image remain clearly discernible, such as in scenes like a parking lot at night or a dimly lit alley. Ordinary cameras cannot capture such clear images in these conditions. Images captured by traditional cameras often suffer from high noise and blurred details. In backlit scenes, such as road surveillance at sunset, traditional cameras tend to underexpose the subject and overexpose the background. ISP chips, however, can balance the brightness of the image through dynamic range adjustment and other technologies, ensuring that foreground vehicles, pedestrians, and the background are all clearly presented. This greatly expands the camera's applicable scenarios and improves the system's detection capabilities under complex lighting conditions. When the system is used for long-distance detection scenarios, such as monitoring vehicles on highways, setting the focal length to 12mm allows the camera to capture details of distant vehicles more clearly, such as license plate numbers and vehicle models. This meets the requirements for accurate identification of target objects under different installation distances, enabling the system to perform at its best in various complex real-world application scenarios.

[0043] The computing module is responsible for model inference and data computation. It is equipped with NVIDIA Jetson AGX Xavier and supports TensorRT acceleration, providing computing power for multi-target detection and ranging, and processing a large amount of target detection data in parallel. The computation module, combined with TensorRT acceleration technology, further enhances the computational efficiency of the NVIDIA Jetson AGX Xavier. TensorRT can optimize deep learning models, reducing computational load and memory consumption, and significantly improving inference speed. In practical applications, the inference speed of the YOLOv5 model optimized by TensorRT can be increased by several times or even tens of times. Taking multi-source target detection in traffic segments as an example, what used to take a long time to complete the synchronous detection and analysis of multi-source targets can now be completed instantly with the help of TensorRT acceleration, improving the system's efficiency in detecting and ranging multi-source targets.

[0044] The storage module is used for temporary storage of detection results and video streams. It can quickly read and write data and supports connection to cloud storage services. Detection data is stored synchronously in the local area and the cloud to achieve long-term data backup and management. The storage module is locally configured with a 128GB SSD. When the system is used for traffic segment detection, it immediately stores the relevant traffic target detection results and video streams to the local SSD during use, ensuring that the data is not lost due to temporary network failures or other problems. This data can be retrieved at any time for subsequent event analysis and processing, providing accurate information support for event handling. The storage module supports connection to cloud storage services, realizing long-term data backup and more convenient management. Cloud storage has unlimited storage space and high reliability. Even if the local storage fails, the data can be safely saved in the cloud.

[0045] Please see Figure 2 The software module is used for the detection and ranging of multiple targets in traffic segments. It monitors and detects people and vehicles in the traffic segment area in real time, and determines the relative positional relationship between targets through the ranging function. The detection and ranging process includes parameter configuration, data acquisition, inference calculation, and multi-dimensional result output, specifically including the following steps: S1. Based on the traffic environment of the application installed in the detection module, select the application scenario in the system, automatically load the preset parameters under the scenario, and collect multi-source targets of the traffic segment.

[0046] In the traffic scene described above, low light and backlight conditions frequently occur during data acquisition in S1, which seriously affect the quality of image acquisition. The detection module integrates advanced image enhancement algorithms. In low light environments, the algorithm analyzes and adjusts the pixel values ​​of the image to improve the overall brightness of the image. At the same time, it uses noise reduction technology to suppress noise that may be introduced due to increased brightness, clearly presenting the outlines and details of objects in dark areas. In backlight scenes, dynamic range adjustment technology is used to balance the brightness of overly bright and overly dark areas in the image, ensuring that information from different lighting areas is completely preserved. For example, when shooting a car in backlight, not only can the car body be clearly displayed, but the license plate number can also be clearly identified.

[0047] S2. The collected data is preprocessed by using letterbox filling to scale the collected image data to a uniform size and fill the blank areas at the edges of the images. Then, the pixels of the images are converted to unify the data of different images and reduce the differences between the data. Finally, the image features are adjusted through data augmentation so that the model can quickly identify the location and category of the image target.

[0048] The data augmentation function in S2 mentioned above includes horizontal or vertical flipping of the image, which helps the model better understand the essential features of the target, rather than relying solely on features in a specific direction. At the same time, it can simulate images under different lighting conditions, enabling the model to adapt to various complex lighting environments. In actual traffic scenarios, whether under strong light or in low light, the model can accurately detect and measure distances.

[0049] S3. Load the pre-trained YOLOv5 TensorRT engine, input the pre-processed image data into it, perform feature extraction and analysis on the image, and output key information such as target bounding box, confidence score, and class ID. The system performs non-maximum suppression operation and calculates the IOU value between each bounding box and other bounding boxes to obtain the optimal detection result.

[0050] The pre-trained YOLOv5 model in S3 has been trained on a large amount of image data and has learned rich target features. It can quickly and accurately identify various targets. The system calculates the IOU value between each bounding box and other bounding boxes, removes redundant bounding boxes with IOU values ​​higher than a set threshold, and retains only the bounding box with the highest score that most accurately represents the target location.

[0051] S4. Based on the optimal detection results, extract the pixel width of each target bounding box, query the preset actual width in the mapping library according to the target category ID, and quickly match the target to the corresponding preset value in the mapping library. If the traffic target is not within the mapping library, the target parameter error is corrected by the similar triangle distance measurement formula, and the distance of each detected target is calculated in real time.

[0052] The mapping library in S4 mentioned above includes target type and actual size mapping libraries. For different targets such as pedestrians, vehicles, and buses, the mapping library presets different actual widths. For example, the average width of pedestrians is set to a reasonable numerical range, and common family cars also have corresponding standard width preset values. At the same time, in order to meet the special needs of different users, the system also supports users to add custom target parameters. For example, in some special traffic scenarios, there may be some non-standard sized vehicles or special objects. Users can add the parameters of these targets to the mapping library according to the actual situation, enriching the internal data of the mapping library, so that the preset values ​​corresponding to traffic targets can be quickly found in the mapping library.

[0053] S5. The system uses the Annotator tool to draw bounding boxes on the original image, clearly marking the detected targets. It generates a CSV format result file from the detection data and displays the annotated image and video stream in real time using OpenCV. Users can observe the target detection and ranging results in the traffic scene on the screen in real time.

[0054] The S5 system is configured with HTTP requests to push detection results to the backend management platform in real time. The backend management platform can centrally manage and analyze a large number of detection results, and realize advanced functions such as traffic flow statistics and dangerous target warning. For example, when a vehicle is detected to be speeding or driving against traffic, the backend management platform can issue an alarm in a timely manner and notify relevant departments to handle the situation, providing strong support for traffic management.

[0055] In this embodiment, the user selects the application scenario through a web interface during parameter configuration, and the system automatically loads preset parameters, specifically including: detection parameters, which include a confidence threshold, an NMS IOU threshold, and a maximum number of detected targets. The confidence threshold for traffic scenarios is 0.4, the confidence threshold for indoor scenarios is 0.3, the default NMS IOU threshold is 0.45, and the default maximum number of detected targets is 1000; ranging parameters, such as... Figure 3 As shown, load the target type and actual size mapping library, and input the current camera focal length through the camera calibration tool. The focal length pixel is 1500 by default.

[0056] It should be noted that the dynamic adaptation mechanism for detection and ranging parameters in this application includes establishing a target type-actual size mapping library, pre-setting differentiated actual widths for different targets such as pedestrians, vehicles and buses, and supporting users to add custom target parameters; at the same time, it provides a camera focal length calibration interface to dynamically adjust the focal length parameters according to the device installation height and angle.

[0057] In practical implementation, scenario-based parameter configuration is tailored to different scenarios such as traffic monitoring, park security, and indoor inspection, with preset differentiated inference parameters. For example, the confidence threshold is 0.4 for traffic scenarios and 0.3 for indoor scenarios. Users can switch scenario modes with one click through a visual interface, reducing operational complexity. Diverse detection terminals are deployed to ensure comprehensive coverage of targets, as follows: fixed monitoring equipment is deployed on both sides of the road to collect frontal and side images of distant vehicles and pedestrians; it is deployed at high points at intersections, supporting 360° rotation and zoom to track fast-moving targets; and it is deployed directly above the lane to collect vehicle rear features, with supplementary lighting to ensure clear nighttime imaging; and mobile acquisition equipment is installed on camera terminals at moving targets to collect dynamic road surface perspectives.

[0058] In this embodiment, data acquisition adopts a multi-source acquisition method. The detection terminal acquires video streams or images in real time and transmits them to the computing unit via the RTSP protocol. It also supports the import of local video files and image files.

[0059] In this embodiment, data acquisition is followed by preprocessing, specifically as follows: image resizing to 640×640, using letterbox padding to avoid image distortion; data normalization, converting pixel values ​​from 0-255 to 0.0-1.0 and then to Tensor format; data augmentation, enabling random flipping and brightness adjustment enhancement strategies for static images to improve model generalization ability; and inference computation, integrating detection and ranging, including object detection inference and dynamic distance calculation.

[0060] It should be noted that the implementation logic of letterbox filling in this application includes: calculating the scaling factor, finding the ratio of the width and height of the original image to the target size, and taking the smaller ratio as the uniform scaling factor. For example, for an original image of 1280×720, the width scaling factor = 640 / 1280 = 0.5, the height scaling factor = 640 / 720≈0.888, and finally taking 0.5; scaling the image proportionally, scaling the original image with the above factors to obtain an image with a width of 640 and a height of 360, keeping the target ratio unchanged; filling the blank area, filling the top and bottom of the scaled image with black pixels to complete it to 640×640, ensuring that the input size is uniform; and preserving the original shape of the target to the greatest extent.

[0061] It should be noted that the data augmentation strategies in this application include: random flipping, including horizontal and vertical flipping, such as flipping the car on the left to the right, to avoid the model's reliance on the bias of the target being on the left side of the image; brightness adjustment, randomly increasing and decreasing the image brightness to simulate daytime, evening, bright light, and shadow scenes, to avoid the model's reliance on fixed lighting conditions; and other supplementary strategies, such as random contrast adjustment, hue shift, random cropping (simulating the target being in different positions in the image), and Gaussian noise addition (to enhance the model's anti-interference ability).

[0062] In this embodiment, the object detection inference is specifically as follows: load the pre-trained YOLOv5 TensorRT engine and initialize the model through the DetectMultiBackend interface; the model takes pre-processed image data as input and outputs the target bounding box, confidence score, and class ID; perform non-maximum suppression to remove redundant bounding boxes with overlap higher than the IOU threshold and retain the best detection result.

[0063] It should be noted that, compared to PyTorch's native inference, TensorRT in this application can improve the inference speed of YOLOv5 by 2-5 times through techniques such as layer fusion, precision calibration, and automatic kernel optimization, thus meeting real-time requirements. It should be noted that the bounding boxes after NMS are based on the coordinates of a 640×640 image. If it is necessary to map them back to the original image size, the scaling factor and offset when filling the letterbox need to be recorded and restored through inverse operation. If multiple images need to be processed simultaneously, it is necessary to ensure that the TensorRT engine supports dynamic batch processing when exporting.

[0064] In this embodiment, the dynamic distance calculation is specifically as follows: extract the pixel width x2-x1 of each target bounding box, where x1 and x2 are the horizontal coordinates of the left and right vertices of the bounding box; query the preset actual width according to the target category ID; if the category is not in the mapping library, mark the distance as N / A; calculate the distance using the similar triangle ranging formula: distance = (focal length × actual target width) / target pixel width.

[0065] It should be noted that the accuracy of dynamic distance calculation in this application is affected by a variety of factors, including: (1) deviation of the actual width of the target, such as calculating the width of a wide-body vehicle as that of an ordinary vehicle, which leads to a smaller distance measurement. The optimization solution is to subdivide the categories, such as distinguishing between sedans and SUVs, and to use the target height as an auxiliary calibration; (2) camera focal length error, which directly leads to a proportional deviation in distance due to inaccurate focal length. The optimization solution is to periodically calibrate the camera using a checkerboard pattern and update the intrinsic parameter matrix; (3) target tilt angle, which means that the target is not perpendicular to the camera, such as a car driving on the side, where the pixel width is smaller than the actual width, which leads to a larger distance measurement. The optimization solution is to prioritize detecting frontal targets and correct the width using a 3D pose estimation algorithm; (4) image resolution, which means that the pixel width of the letterbox after scaling deviates from the original image. The optimization solution is to calculate the pixel width uniformly under the original image size.

[0066] In this embodiment, the multi-dimensional result output specifically includes: image annotation, using the Annotator tool to draw bounding boxes on the original image, annotating the target name, confidence level, and distance, with annotation colors distinguished by target category; structured data storage, generating a CSV format result file, with fields including image, video frame name, target name, confidence level, distance, and detection time; real-time display and push: displaying the annotated image and video stream via OpenCV, supporting exit by pressing the 'q' key; and the ability to configure HTTP requests to push the detection results to the backend management platform in real time.

[0067] It should be noted that, in addition to image annotation, the multi-dimensional result output system in this application automatically generates CSV structured data containing target name, confidence level, distance, detection time, and image path, and supports pushing the results to the backend platform via HTTP interface to realize secondary analysis and application of detection data.

[0068] In this embodiment, the key algorithm optimizations of the software module specifically include: TensorRT model acceleration, converting the YOLOv5PyTorch model to the TensorRT engine to improve inference speed; layer fusion, merging convolution, activation, and batch normalization layers into a single computational unit to reduce data transmission overhead; precision quantization, supporting FP16 half-precision inference, improving inference speed by 2-3 times while maintaining a precision loss of less than 3%; dynamic batch processing, automatically adjusting the batch size based on the amount of input data to balance inference speed and memory usage; and in ranging error correction, when the target pixel width is less than 10, enabling a multi-frame fusion strategy to take the average distance of three consecutive frames to reduce measurement fluctuations.

[0069] It should be noted that the TensorRT-accelerated real-time inference architecture in this application uses the YOLOv5 TensorRT engine to replace the native PyTorch model. By optimizing the inference speed through model quantization and layer fusion, the frame rate of 1080P video streams is increased to more than 30 FPS in a GPU (NVIDIA Tesla T4) environment, which meets the real-time requirements.

[0070] Example 2: To better understand the above technical solution, the following will provide a detailed explanation of the specific implementation method using actual examples. To provide efficient and accurate solutions for various practical application scenarios, a multi-target detection and ranging integrated system based on YOLOv5 is proposed, including hardware and software modules, as detailed below: The detection module optimizes image quality in low-light and backlight scenarios in traffic areas, balances image brightness, provides high-quality raw data for traffic target detection and ranging, and calibrates the detection focal length to adapt to different installation distance requirements. The computing module is responsible for model inference and data computation, equipped with NVIDIA Jetson AGX Xavier, supports TensorRT acceleration, provides computing power for multi-target detection and ranging, and processes large amounts of target detection data in parallel. The storage module temporarily stores detection results and video streams, enabling fast data read and write, and supports connection to cloud storage services, allowing detection data to be stored synchronously locally and in the cloud, achieving long-term data backup and management.

[0071] The software module is used to detect and measure the distance of multiple targets in traffic sections, monitor people and vehicles in the traffic section area in real time, and determine the relative positional relationship between targets through the distance measurement function. The detection and distance measurement process includes parameter configuration, data collection, inference calculation, and multi-dimensional result output.

[0072] In the software module, the parameter configuration stage plays a crucial guiding role. Through a simple and intuitive web interface, users can easily select the required application scenario. Taking traffic monitoring as an example, when the user selects the scenario on the web interface, the system will respond quickly and automatically load a series of preset parameters. Regarding detection parameters, the confidence threshold is set to 0.4, which means that only when the model's confidence in recognizing the target object is higher than 0.4 will it be judged as a valid detection result. In actual traffic scenes, if the confidence in detecting a vehicle is lower than 0.4, the system will ignore the detection to avoid misjudgments caused by low-confidence detection.

[0073] Regarding distance measurement parameters, the system loads a target type and actual size mapping library. For pedestrians, the system presets their actual width to a reasonable average value; for vehicles, it presets corresponding actual widths based on different vehicle models. Users can also easily customize and add parameters for special targets through the interface, inputting the actual size information of a specific vehicle model, which the system will then include in the mapping library for accurate distance calculation in subsequent detection. Furthermore, users need to input the current camera focal length using the camera calibration tool. The default focal length pixel value is 1500. This parameter plays a crucial role in accurate distance measurement, dynamically adjusting based on factors such as the actual installation height and angle of the device to ensure accuracy.

[0074] The inference and computation module can be regarded as the intelligent engine of the system. It closely integrates target detection inference and dynamic distance calculation, realizes the integrated function of multi-target detection and ranging, and endows the system with powerful intelligent analysis capabilities. During the object detection inference process, the system first loads the pre-trained YOLOv5 TensorRT engine and initializes the model through the DetectMultiBackend interface to prepare for subsequent detection tasks. When the pre-processed image data is input into the model, the model quickly analyzes and judges each region in the image and outputs the target bounding box, confidence score, and category ID. In an image of a city street, the model can accurately identify different targets such as vehicles, pedestrians, and traffic lights, and mark their positions in the image. It also provides the confidence score for each target recognition and the ID representing the target category. Subsequently, the system removes redundant bounding boxes with an overlap higher than the IOU threshold from the numerous bounding boxes output by the model. For example, when multiple bounding boxes surround the same vehicle, non-maximum suppression will retain the bounding box that most accurately represents the vehicle's position and size based on the IOU threshold, and remove other overlapping boxes, thereby obtaining the optimal detection result. In terms of dynamic distance calculation, the system extracts the pixel width x2 - x1 of each target bounding box, where x1 and x2 are the x-coordinates of the left and right vertices of the bounding box. This pixel width information is one of the key data for distance calculation. Next, the system queries the preset actual width based on the target category ID. This actual width information comes from the previously loaded target type and actual size mapping library. If the target category is found in the mapping library, the system can accurately obtain its actual width and then calculate the distance using the similar triangle distance measurement formula: distance = (focal length × actual target width) / target pixel width. However, if the category is not in the mapping library, the system will mark the distance as N / A, prompting the user that the distance of the target cannot be accurately calculated and further target parameter information is needed. Through this dynamic distance calculation process, the system can obtain the distance information of each detected target in real time and accurately, providing important data support for subsequent applications.

[0075] The multi-dimensional results output module is a showcase for the system to display detection and ranging results. It presents the processed data to users intuitively in various ways, meeting the needs of data display and application in different scenarios. In terms of image annotation, the system uses the Annotator tool to draw meticulously on the original image. For each detected target, a bounding box is drawn to clearly outline the target's contour. At the same time, the target name, confidence level, and distance information are labeled next to the bounding box. Different types of targets are labeled with different colored bounding boxes. In traffic monitoring footage, vehicles may be labeled with yellow bounding boxes, while pedestrians may be labeled with green bounding boxes. This allows users to distinguish different types of targets at a glance. Through this intuitive image annotation method, users can quickly understand the relevant information of each target in the image. Structured data storage generates result files in CSV format from the detection results. Fields in the file include image or video frame name, target name, confidence level, distance, and detection time. In traffic flow statistical analysis, researchers can easily obtain data such as the number of vehicles detected, vehicle type, and vehicle distance in different time periods by reading these CSV files, providing structured data support for traffic flow analysis and prediction. The real-time display and push functionality further enhances the system's usability. Using OpenCV, the system can display annotated images and video streams in real time, allowing users to see the detection and ranging results on the monitoring screen. Users can exit the display interface at any time by pressing the 'q' key, making operation simple and convenient. Furthermore, the system supports configuring HTTP requests to push detection results to the backend management platform in real time. In the intelligent security system, the backend management platform can promptly detect anomalies and issue warnings based on the received real-time detection results, achieving real-time monitoring and management of the monitored area and greatly improving the system's application value and efficiency.

[0076] In multi-target detection and ranging integrated systems, YOLOv5, as a target detection model based on PyTorch, requires conversion from a PyTorch model to the TensorRT engine as a key step to achieve acceleration in practical applications. The YOLOv5 model contains numerous convolutional layers, activation function layers, and batch normalization layers. TensorRT intelligently merges these adjacent and combinable layers into a single computational unit. This reduces the overhead of intermediate data transfer and multiple kernel calls during computation, significantly improving computational efficiency. In a module containing multiple consecutive convolutional, activation, and batch normalization operations, TensorRT's fusion allows data to be processed continuously within a single kernel, avoiding frequent data transfers between memory and GPU memory, thus significantly improving inference speed. In practical tests, for processing a complex video stream, the YOLOv5 model accelerated by TensorRT showed a several-fold increase in inference speed compared to the unaccelerated version, enabling faster output of detection results. This provides more timely data support for subsequent dynamic distance calculations and multi-dimensional result output, allowing the system to better meet the real-time requirements of application scenarios.

[0077] Layer fusion technology is a key means to improve the operating efficiency of the system. In the YOLOv5 model, convolution, activation, batch normalization and other layers are the basic components of building the model. They are executed sequentially in the forward propagation process of the model to complete the feature extraction and transformation of the input data.

[0078] Layer fusion technology breaks this conventional model by merging adjacent layers into a single computational unit. Taking convolutional and batch normalization layers as examples, in the traditional approach, after convolutional layer computation, data needs to be transferred to the batch normalization layer for normalization, involving data storage and retrieval operations. With layer fusion, however, convolution and batch normalization operations are completed in a unified kernel, eliminating the need for additional data transfer between different layers and allowing for continuous processing within the same computational unit. This significantly reduces data transfer overhead and improves computational efficiency. From a hardware resource utilization perspective, layer fusion enables GPUs and other computing devices to utilize their resources more efficiently, reducing idle time and thus improving the overall system efficiency. In practical applications, the YOLOv5 model optimized with layer fusion can complete object detection tasks faster when processing large-scale image data, providing strong support for the system's real-time performance and efficiency.

[0079] The system supports P16 half-precision inference, which is an important way to implement precision quantization. In deep learning models, FP32 (32-bit floating-point numbers) is usually used to represent the model weights and data in the calculation process. Although this high-precision representation can ensure the accuracy of the model, it also brings a large amount of computation and memory usage. FP16 half-precision inference uses 16-bit floating-point numbers to represent data, reducing the number of data bits by half compared to FP32. This change brings significant advantages. During computation, the reduced number of data bits correspondingly decreases the computational load, resulting in a substantial improvement in inference speed. According to actual test data, the system's inference speed can be increased by 2-3 times when using FP16 half-precision inference. However, this reduction in precision may have some impact on model accuracy. To ensure improved inference speed while keeping the precision loss within an acceptable range, the system underwent careful optimization and adjustment. Through extensive experiments and data analysis, it was found that as long as the quantization strategy and parameters are selected appropriately, the precision loss can be controlled within less than 3%. When detecting images in a large number of different scenes, comparing the detection results under FP32 and FP16 precision, it was found that although there are some differences in the detection of some subtle targets, the overall detection accuracy and recall rate do not change significantly and can still meet the needs of most practical application scenarios. Therefore, FP16 half-precision inference achieves a good balance between accuracy and speed, providing strong support for the system to operate efficiently under resource constraints.

[0080] In multi-target detection and ranging integrated systems, dynamic batch processing technology automatically adjusts the batch size according to changes in the amount of input data, thereby achieving a clever balance between inference speed and memory usage. In the deep learning inference process, batch processing is a common optimization method. It combines multiple input data into a batch for processing, which can make full use of the parallel computing capabilities of computing devices (such as GPUs) and improve computing efficiency. Dynamic batch processing technology can flexibly adjust the batch size based on the real-time input data volume. When the input data volume is small, it appropriately reduces the batch size to avoid wasting computing resources and memory space by filling in too much invalid data. When the input data volume is large, it increases the batch size to fully utilize the parallel computing capabilities of the computing device and improve inference speed. In real-time traffic monitoring scenarios, the number of vehicles and pedestrians varies at different times, causing the amount of video stream data input to the system to change. During the late-night hours when traffic is low, dynamic batch processing technology automatically reduces the batch size to adapt to the processing needs of small amounts of data and reduce memory usage. During peak traffic hours, when the number of vehicles and pedestrians increases significantly, dynamic batch processing technology increases the batch size, utilizing the parallel computing capabilities of the GPU to quickly process large amounts of data, ensuring that the system can detect and analyze complex scenes in real time. Through this dynamic adjustment of the batch size, the system can maintain high inference efficiency and reasonable memory usage under different data load conditions, providing a strong guarantee for the efficient execution of multi-target detection and ranging tasks.

[0081] In the actual operation of the multi-target detection and ranging integrated system, when the target pixel width is less than 10, the measurement fluctuation problem faced by the system is more prominent. By comprehensively analyzing and processing multiple consecutive frames of images, the fluctuation of the measurement results can be reduced, and the stability and accuracy of the ranging can be improved. In traffic monitoring scenarios, when pedestrians or small vehicles enter the monitoring screen from a distance, their pixel width in the image may be less than 10. Due to the influence of environmental noise, changes in lighting, and the movement of the target itself, the single measurement results of these targets often have large fluctuations, making it difficult to accurately reflect the actual distance of the target. To address this issue, the system automatically employs a multi-frame fusion strategy. It continuously acquires three frames of images, extracts the pixel width of the target's bounding box in each frame, and calculates the target's distance in each frame using the similar triangle ranging formula. Since noise and interference between frames are random, averaging the distances from these three frames effectively mitigates the impact of these random factors. For example, assuming a pedestrian target with a pixel width less than 10 is measured at a certain moment, the distance calculated in the first frame is 50 meters, the distance in the second frame is 55 meters due to momentary light reflection interference, and the distance in the third frame is 48 meters under normal conditions. Through the multi-frame fusion strategy, the average of these three distances (50+55+48) / 3 ≈ 51 meters. This result is closer to the target's true distance than a single-frame measurement, effectively reducing measurement fluctuations and improving the accuracy and reliability of ranging.

[0082] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "connected" and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0083] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A multi-target detection and ranging integrated system based on YOLOv5, characterized in that, Includes hardware modules and software modules; The hardware modules include: The detection module is used to optimize image quality in low-light and backlight scenes in traffic sections, balance the brightness of the image, provide high-quality raw data for traffic target detection and ranging, and calibrate the detection focal length to adapt to different installation distance requirements. The computing module is responsible for model inference and data computation. It is equipped with NVIDIA Jetson AGX Xavier and supports TensorRT acceleration, providing computing power for multi-target detection and ranging, and processing a large amount of target detection data in parallel. The storage module is used for temporary storage of detection results and video streams. It can quickly read and write data and supports connection to cloud storage services. Detection data is stored synchronously in the local area and the cloud to achieve long-term data backup and management. The software module is used to detect and measure the distance of multiple targets in traffic sections, monitor people and vehicles in the traffic section area in real time, and determine the relative positional relationship between targets through the distance measurement function. The detection and distance measurement process includes parameter configuration, data collection, inference calculation, and multi-dimensional result output.

2. The multi-target detection and ranging integrated system based on YOLOv5 according to claim 1, characterized in that: In the parameter configuration, users select application scenarios through a web interface, and the system automatically loads preset parameters to ensure efficient system operation, specifically including: The detection parameters include the confidence threshold, the NMS IOU threshold, and the maximum number of detection targets. The confidence threshold for traffic scenarios is 0.4, the confidence threshold for indoor scenarios is 0.3, the default NMS IOU threshold is 0.45, and the default maximum number of detection targets is 1000. The ranging parameters are loaded, the target type and actual size mapping library are loaded, and the current camera focal length is input through the camera calibration tool. The default focal length pixel is 1500. The dynamic adaptation mechanism for detection and ranging parameters includes establishing a target type-actual size mapping library, pre-setting differentiated actual widths for different targets such as pedestrians, vehicles, and buses, and supporting users to add custom target parameters; it also provides a camera focal length calibration interface to dynamically adjust focal length parameters according to the device's installation height and angle.

3. The multi-target detection and ranging integrated system based on YOLOv5 according to claim 1, characterized in that: The data acquisition adopts a multi-source acquisition method, with the detection terminal acquiring video streams or images in real time and transmitting them to the computing unit via the RTSP protocol. It also supports the import of local video and image files. Deploy diverse detection terminals to ensure comprehensive coverage of the target area, as detailed below: Fixed monitoring equipment is deployed on both sides of the road to collect frontal and side images of vehicles and pedestrians at a distance. It is also deployed at a high point at the intersection to support 360° rotation and zoom to track fast-moving targets. It is also deployed directly above the lane to collect the rear features of vehicles and is equipped with supplementary lighting to ensure clear imaging at night. Mobile data acquisition equipment, a camera terminal installed on a moving target, captures dynamic views of the road surface.

4. The multi-target detection and ranging integrated system based on YOLOv5 according to claim 3, characterized in that: The data is preprocessed after acquisition, as follows: The image was resized to 640×640, and letterbox padding was used to avoid image distortion. Data normalization converts pixel values ​​from 0-255 to 0.0-1.0 and then converts them to Tensor format; Data augmentation, for static images, can enable random flipping and brightness adjustment enhancement strategies to improve the model's generalization ability.

5. The multi-target detection and ranging integrated system based on YOLOv5 according to claim 1, characterized in that: The inference calculation integrates detection and ranging into one, including target detection inference and dynamic distance calculation.

6. The multi-target detection and ranging integrated system based on YOLOv5 according to claim 5, characterized in that: The target detection inference is specifically as follows: Load the pre-trained YOLOv5 TensorRT engine and initialize the model through the DetectMultiBackend interface; The model takes preprocessed image data as input and outputs the target bounding box, confidence score, and class ID. Non-maximum suppression is performed to remove redundant bounding boxes with an overlap greater than the IOU threshold, while retaining the best detection results.

7. The multi-target detection and ranging integrated system based on YOLOv5 according to claim 5, characterized in that: The dynamic distance calculation specifically involves: Extract the pixel width x2-x1 of each target bounding box, where x1 and x2 are the x-coordinates of the left and right vertices of the bounding box; Query the preset actual width based on the target category ID. If the category is not in the mapping library, mark the distance as N / A. Distance is calculated using the similar triangle distance measurement formula: Distance = (focal length × actual target width) / target pixel width.

8. The multi-target detection and ranging integrated system based on YOLOv5 according to claim 1, characterized in that: The multi-dimensional result output specifically includes: Image annotation: Use the Annotator tool to draw bounding boxes on the original image, and annotate the target name, confidence level, and distance. The annotation colors are distinguished according to the target category. Structured data storage generates CSV format result files, with fields including image, video frame name, target name, confidence score, distance, and detection time; Real-time display and push: Displays annotated images and video streams using OpenCV, and supports exiting by pressing the 'q' key; it can also be configured with HTTP requests to push detection results to the backend management platform in real time.

9. The multi-target detection and ranging integrated system based on YOLOv5 according to claim 1, characterized in that: The key algorithm optimizations for the software module specifically include: TensorRT model acceleration: Convert YOLOv5 PyTorch models to the TensorRT engine to improve inference speed. Layer fusion combines convolutional, activation, and batch normalization layers into a single computational unit, reducing data transmission overhead. Precision quantization supports FP16 half-precision inference, improving inference speed by 2-3 times while maintaining less than 3% precision loss; Dynamic batch processing automatically adjusts the batch size based on the amount of input data, balancing inference speed and memory usage; Ranging error correction addresses pixel width deviations caused by lens distortion by generating distortion coefficients using camera calibration tools to correct the bounding box pixel width.

10. The multi-target detection and ranging integrated system based on YOLOv5 according to claim 9, characterized in that: In the ranging error correction, when the target pixel width is less than 10, a multi-frame fusion strategy is enabled to take the average distance of 3 consecutive frames to reduce measurement fluctuations.

Citation Information

Patent Citations

  • A PSD-based traffic flow detection system

    CN104485002B

  • Traffic Target Detection and Ranging Methods Based on Deep Learning

    CN108830131B

  • Traffic sign detection and distance measurement method based on improved YOLOv5

    CN117079252A

  • Traffic condition comprehensive monitoring method based on YOLO

    CN114067564A

  • Distance measurement method and device based on target width and electronic equipment

    CN120233349A