Real-time multi-target detection system and method based on deep learning

Through deep learning multi-scale feature extraction and fusion, improved YOLOv5 model and Kalman filter algorithm, the complex environmental interference and detection time-consuming and labor-consuming problems in multi-object detection are solved, and fast and accurate multi-object detection and tracking are achieved.

CN120339907AInactive Publication Date: 2025-07-18ZHEJIANG COLLEGE OF CONSTR
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510400164.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art faces complex environmental interference in multi-object detection, such as occlusion, messy background, target deformation, etc., which affects the detection accuracy. Multi-object detection requires time-consuming and labor-intensive testing and has poor detection performance for a few categories.

Method used

Real-time multi-object detection method based on deep learning is adopted, including video stream preprocessing, multi-scale feature extraction, feature fusion and attention mechanism, improved YOLOv5 object detection model and Kalman filter algorithm, to generate the target's motion trajectory and post-processing to output the detection results.

Benefits of technology

Improves the comprehensiveness and accuracy of multi-object detection. The YOLOv5 model provides fast detection, and the Kalman filter ensures the stability of real-time tracking, and is suitable for complex backgrounds and target-intensive scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339907A_ABST
    Figure CN120339907A_ABST
Patent Text Reader

Abstract

The invention discloses a real-time multi-target detection system and method based on deep learning, and relates to the technical field of deep learning, and the method comprises the steps: collecting a video stream and an image sequence of a to-be-detected scene, and carrying out the preprocessing of the video stream and the image sequence; multi-scale feature extraction is carried out on the preprocessed input data by using a deep learning model, a multi-scale feature pyramid is constructed, and each layer of feature map represents target information of different scales; fusing the feature maps of different scales, introducing an attention mechanism, and performing weighting processing on the features of the target area; using an improved YOLOv5 target detection model to carry out target detection on the fused feature map, and outputting a bounding box and a category probability of a target; performing real-time tracking on the detected target by adopting a Kalman filter algorithm to generate a motion track of the target; carrying out post-processing on the detection result, wherein the post-processing comprises non-maximum suppression, target association and trajectory smoothing; and finally outputting a detection result and a tracking trajectory of the target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and specifically relates to a real-time multi-object detection system and method based on deep learning. Background Art

[0002] In practical applications, object detection faces complex environmental interferences, such as occlusion, cluttered background, object deformation, etc., which seriously affect the detection accuracy. When there is an occlusion relationship between multiple objects, how to accurately detect and identify each object is a challenge. Multi-object detection requires precise object annotation for a large amount of image or video data, which is time-consuming and laborious, and the accuracy of the annotation directly affects the training effect of the model; in addition, the detection performance of the model for a small number of categories is poor.

[0003] The rapid development of deep learning technology provides strong support for real-time multi-object detection. Models such as convolutional neural networks (CNNs) can automatically extract features in images, and have higher expressive power and robustness compared to traditional manually designed features, greatly improving the accuracy of object detection. The computing power of graphics processing units (GPUs) has been continuously improved, significantly accelerating the training and inference speed of deep learning models, and meeting the speed requirements for real-time detection.

[0004] In the fields of security monitoring, autonomous driving, intelligent transportation, unmanned aerial vehicles, robots, etc., the demand for real-time multi-object detection is increasing day by day, and these fields require the system to be able to quickly and accurately detect and identify multiple objects in images or videos. Summary of the Invention

[0005] To solve the above technical problems, a real-time multi-object detection system and method based on deep learning are provided. This technical solution solves the problems that the above object detection faces complex environmental interferences, such as occlusion, cluttered background, object deformation, etc., which seriously affect the detection accuracy. When there is an occlusion relationship between multiple objects, how to accurately detect and identify each object is a challenge; multi-object detection requires precise object annotation for a large amount of image or video data, which is time-consuming and laborious, and the accuracy of the annotation directly affects the training effect of the model; the detection performance of the model for a small number of categories is poor.

[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:

[0007] A real-time multi-object detection method based on deep learning, including:

[0008] Collect the video stream and image sequence of the scene to be detected, and perform preprocessing on it, including denoising, normalization, size adjustment, and enhancement operations;

[0009] Use a deep learning model to perform multi-scale feature extraction on the preprocessed input data and construct a multi-scale feature pyramid, where each layer of feature map represents target information at different scales;

[0010] Fuse the feature maps at different scales, introduce an attention mechanism, and perform weighted processing on the features of the target area;

[0011] Use an improved YOLOv5 object detection model to perform object detection on the fused feature map, and output the bounding box and class probability of the object; use the Kalman filter algorithm to perform real-time tracking on the detected object and generate the motion trajectory of the object;

[0012] Perform post-processing on the detection results, including non-maximum suppression, object association, and trajectory smoothing; finally output the detection results and tracking trajectories of the objects.

[0013] Preferably, the video stream and image sequence of the scene to be detected are collected and preprocessed, including denoising, normalization, size adjustment, and enhancement operations, specifically including:

[0014] Enhance the contrast of the image by adjusting the pixel distribution of the image; use Gaussian filtering and median filtering methods to remove noise in the image; adjust the brightness, contrast, and saturation of the image;

[0015] Adjust the pixel values of the image to a distribution with a mean of zero and a standard deviation of one; make the covariance matrix between different pixels close to the identity matrix to remove the correlation between data;

[0016] Realize data enhancement through histogram equalization, contrast stretching, color space conversion, and color space conversion.

[0017] Preferably, the use of a deep learning model to perform multi-scale feature extraction on the preprocessed input data and construct a multi-scale feature pyramid, where each layer of feature map represents target information at different scales specifically includes:

[0018] Select a pre-trained deep learning model as the basic network, which has a multi-layer structure, and each layer corresponds to different scales and semantic information;

[0019] Extract feature maps from different stages of the basic network; perform upsampling and downsampling operations on the extracted feature maps to generate feature maps at different scales;

[0020] Start from the deep feature map, sequentially upsample and fuse the shallow feature maps;

[0021] In the top-down process, introduce lateral connections to fuse the high-semantic feature maps in the deep layer with the high-resolution feature maps in the shallow layer;

[0022] Through the above process, feature maps of different scales are generated, forming a pyramid structure, and each layer of the feature map corresponds to a different target detection scale.

[0023] Preferably, the fusion of feature maps of different scales and the introduction of an attention mechanism to weight the features of the target region specifically include:

[0024] Fuse the feature maps of different scales, and the methods include concatenation, addition, and weighted averaging;

[0025] Apply the attention mechanism to the fused feature map, and enhance the feature representation of the target region through the calculation of channel attention and spatial attention;

[0026] Through backpropagation and optimization algorithms, train the model to learn the optimal feature fusion and attention weights to maximize the performance of target detection.

[0027] Preferably, using the improved YOLOv5 target detection model to perform target detection on the fused feature map and output the bounding box and class probability of the target specifically include:

[0028] Use the backbone network to enhance the feature extraction ability; introduce the attention mechanism, perform adaptive spatial feature fusion, and improve the model's perception ability of the target region;

[0029] Introduce multi-scale prediction layers at different stages of the model to detect targets of different sizes; use the feature pyramid network to enhance the expression of multi-scale features;

[0030] Adopt focal loss to alleviate the problem of class imbalance, introduce smooth L1 loss, and improve the accuracy of bounding box regression;

[0031] Use data augmentation techniques to augment the training data; adopt transfer learning, initialize with a model pre-trained on a dataset, and fine-tune on a specific task dataset

[0032] Input the preprocessed and feature-fused multi-scale feature maps into the improved YOLOv5 target detection model; the model gradually extracts high-level features in the image through multiple convolutional layers and attention modules;

[0033] The model outputs multiple candidate bounding boxes, each bounding box containing the coordinate information and confidence of the target; classify the targets within each bounding box and output the probability distribution of each predefined class.

[0034] Preferably, the use of the Kalman filter algorithm to perform real-time tracking on the detected targets and generate the motion trajectories of the targets specifically includes:

[0035] Initialize a Kalman filter for each detected target, setting the initial state and covariance matrix; predict the state of the target in the next frame according to the target's motion model; update the covariance matrix to reflect the uncertainty of the prediction.

[0036] After detecting the target in the current frame, update the state estimate of the Kalman filter using the observation data; calculate the residual between the observation and the prediction, and adjust the state estimate to reduce the error;

[0037] Use the Hungarian algorithm to associate the detection results of the current frame with the previous trajectories to ensure continuous tracking of each target;

[0038] For the unmatched detection results, regard them as new targets and initialize new Kalman filters;

[0039] Maintain the trajectory of each target, recording its historical positions and states; when the target is not detected for a long time, mark it as lost and terminate the tracking.

[0040] Preferably, the real-time tracking of the detected targets using the Kalman filter algorithm to generate the motion trajectories of the targets specifically includes:

[0041] Predict the state of the target in the next frame according to the target's motion model, and the prediction formula is:

[0042]

[0043] In the formula, is the predicted state of the target in the next frame; F is the state transition matrix, which is used to describe the relationship between the target state and time; is the state estimate value of the target in the previous frame; B is the control matrix, u k-1 is the control input at the previous moment;

[0044] Update the covariance matrix to reflect the uncertainty of the prediction, and the formula is:

[0045]

[0046] In the formula, is the predicted covariance matrix, is the covariance matrix of the previous frame, Q is the covariance matrix of the process noise, which is used to describe the uncertainty in the target motion model, and F is the state transition matrix, which is used to describe the relationship between the target state and time.

[0047] Preferably, the post-processing of the detection results includes non-maximum suppression, target association, and trajectory smoothing; the final output of the detection results and tracking trajectories of the targets specifically includes:

[0048] Calculate the confidence scores of all detected bounding boxes, sort them from high to low, and check the bounding boxes one by one. If the intersection over union (IoU) with the previously retained bounding boxes exceeds the threshold, discard the current bounding box;

[0049] Associate the detection results of the current frame with the tracks of the previous frame to ensure the continuity of tracking; use the Hungarian algorithm to calculate the matching cost between the detected bounding boxes and the existing tracks, and associate the bounding boxes with the tracks according to the principle of minimum cost matching;

[0050] Use moving average to smooth the tracks, perform weighted average on the historical positions of the tracks to reduce mutations;

[0051] Output the filtered detected bounding boxes and class probabilities; output the unique identifier, track points, and motion state of the target; draw the detected bounding boxes, class labels, and tracks on the video frame for users to view.

[0052] Furthermore, a real-time multi-object detection system based on deep learning is proposed to implement the real-time multi-object detection method based on deep learning as described above, including:

[0053] Data acquisition and preprocessing module: used to collect video streams and image sequences of the scene to be detected and perform preprocessing on them, including denoising, normalization, size adjustment, and enhancement operations;

[0054] Multi-scale feature extraction module: use a deep learning model to perform multi-scale feature extraction on the preprocessed input data, construct a multi-scale feature pyramid, where each layer of feature map represents target information at different scales;

[0055] Feature fusion and attention mechanism module: fuse feature maps at different scales, and at the same time introduce an attention mechanism to weight the features of the target area and enhance the expression ability of target features;

[0056] Object detection and tracking module: use an improved YOLOv5 object detection model to perform object detection on the fused feature map and output the bounding boxes and class probabilities of the targets; use the Kalman filter algorithm to perform real-time tracking on the detected targets and generate the motion tracks of the targets;

[0057] Post-processing and result output module: perform post-processing on the detection results, including non-maximum suppression, object association, and track smoothing; finally output the detection results and tracking tracks of the targets.

[0058] Optionally, the multi-scale feature extraction module specifically includes:

[0059] The multi-scale feature pyramid extracts feature maps at different scales through a deep residual network and a lightweight network, and realizes the alignment and fusion of feature maps through upsampling and downsampling operations;

[0060] In the top - down process, horizontal connections are introduced, and convolutional operations are performed on the shallow feature maps to adjust their number of channels to match the size of the deep feature maps; the adjusted shallow feature maps are added and concatenated pixel - by - pixel with the upsampled deep feature maps to fuse the high - semantic feature maps in the deep layer with the high - resolution feature maps in the shallow layer.

[0061] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0062] The present invention proposes to construct a multi - scale feature pyramid, which can simultaneously capture target information at different scales, enabling the model to effectively detect targets of different sizes and shapes, effectively improving the comprehensiveness and accuracy of detection; YOLOv5, as an improved object detection model, has a fast detection speed and high detection accuracy; the Kalman filter can track the detected targets in real time, generate the motion trajectories of the targets, and continuously correct the position and state estimates of the targets through two steps of prediction and update, ensuring the real - time performance and stability of tracking, so that the entire system can achieve real - time multi - target detection and tracking; the deep - learning - based framework makes this method have good scalability, and users can further train and optimize the model according to needs to adapt to specific object detection and tracking tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 is a flowchart of a real - time multi - target detection method based on deep learning;

[0064] Figure 2 is an internal framework diagram of a real - time multi - target detection system based on deep learning. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0065] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious variations.

[0066] Referring to Figure 1 as shown, a real - time multi - target detection method based on deep learning includes:

[0067] Collect video streams and image sequences of the scene to be detected and pre - process them, including denoising, normalization, size adjustment, and enhancement operations;

[0068] Use a deep - learning model to perform multi - scale feature extraction on the pre - processed input data to construct a multi - scale feature pyramid, where each layer of feature maps represents target information at different scales;

[0069] Fuse feature maps of different scales, introduce an attention mechanism, and weight the features of the target area;

[0070] Use the improved YOLOv5 object detection model to perform object detection on the fused feature map, and output the bounding box and class probability of the object; Use the Kalman filter algorithm to track the detected object in real time and generate the motion trajectory of the object;

[0071] Post-process the detection results, including non-maximum suppression, object association, and trajectory smoothing; Finally, output the detection results and tracking trajectories of the objects.

[0072] This solution can extract features of different scales in the image by using convolutional layers and pooling layers at different levels through a deep learning model; The low-level feature map contains more detailed information and is suitable for detecting smaller objects; The high-level feature map has a larger receptive field and can capture more extensive object context information, making it suitable for detecting larger objects; Combine feature maps at different levels to form a feature pyramid, where each layer of the feature map corresponds to object information at different scales. In this way, the model can simultaneously use features at different levels to detect objects of different sizes and improve the detection ability for multi-scale objects.

[0073] It should be noted that for feature map fusion and the attention mechanism: Fusing feature maps of different scales can integrate feature information at different levels, enabling the model to more comprehensively understand the features of the object. For example, combining the detailed information in the low-level feature map with the semantic information in the high-level feature map can more accurately locate and identify the object;

[0074] The purpose of introducing the attention mechanism is to enable the model to automatically learn the important features of the target area, weight these features, highlight the key information, and suppress irrelevant information. For example, by calculating the importance weights of each position in the feature map, the model can pay more attention to the features in the area where the object is located, thereby improving the detection accuracy of the object, especially in scenarios with complex backgrounds or dense objects.

[0075] The improved YOLOv5 object detection model: YOLOv5 itself is a deep learning-based object detection model with the characteristics of fast speed and high accuracy. Improving it may include adjusting the network structure, optimizing the loss function, adding additional modules or layers, etc., to further improve the detection performance of the model. For example, it may introduce more efficient feature extraction modules, improved anchor box mechanisms, etc., to enable the model to better adapt to multi-object detection tasks.

[0076] Kalman Filter Algorithm: The Kalman filter is a recursive least squares algorithm that can estimate the state of an object (such as position, velocity, etc.) in real time based on the object's motion model and observation data. In object detection, the Kalman filter is used to track the detected objects, and the motion trajectory of the object can be generated. It continuously corrects the position and state estimation of the object through two steps: prediction and update. Even when the object is occluded or the detection is lost, it can maintain stable tracking of the object, thereby improving the accuracy and stability of tracking.

[0077] The post-processing steps include:

[0078] Non-Maximum Suppression (NMS): In object detection, there may be multiple detection boxes covering the same object. Non-maximum suppression compares the confidence scores of different detection boxes, retains the detection box with the highest confidence, and suppresses other duplicate detection boxes, thereby reducing false detections and duplicate detections and obtaining more accurate object positions.

[0079] Object Association: In multi-object tracking, it is necessary to correctly associate the objects in different frames to form continuous trajectories. Object association can calculate the similarity of features such as the positions and appearances of objects in different frames, and match the objects with the highest similarity, thereby achieving continuous tracking of the objects.

[0080] Trajectory Smoothing: Due to noise and jitter in the detection and tracking process, the object trajectory may not be smooth. Through trajectory smoothing techniques such as moving average filtering, the motion trajectory of the object can be smoothed, making the trajectory more real and stable, and improving the readability and reliability of the tracking results.

[0081] Refer to Figure 2 As shown, the real-time multi-object detection system based on deep learning includes:

[0082] Data Acquisition and Preprocessing Module: Used to acquire the video stream and image sequence of the scene to be detected and preprocess them, including denoising, normalization, size adjustment, and enhancement operations;

[0083] Multi-Scale Feature Extraction Module: Use a deep learning model to perform multi-scale feature extraction on the preprocessed input data and construct a multi-scale feature pyramid, where each layer of feature map represents object information at different scales;

[0084] Feature Fusion and Attention Mechanism Module: Fuse feature maps at different scales, and at the same time introduce an attention mechanism to weight the features of the target area and enhance the expression ability of the target features;

[0085] Object Detection and Tracking Module: Use an improved YOLOv5 object detection model to perform object detection on the fused feature map, and output the bounding boxes and class probabilities of the objects; adopt the Kalman filter algorithm to perform real-time tracking on the detected objects and generate the motion trajectories of the objects;

[0086] Post-processing and Result Output Module: Perform post-processing on the detection results, including non-maximum suppression, object association, and trajectory smoothing; finally output the detection results and tracking trajectories of the objects.

[0087] The construction of the multi-scale feature pyramid in the multi-scale feature extraction module includes:

[0088] Deep Residual Network and Lightweight Network: The deep residual network (such as ResNet) solves the problem of gradient disappearance during the training of deep networks by introducing residual connections, enabling the network to learn and extract features more effectively. The lightweight network aims to reduce the computational amount and number of parameters of the network, improve the running efficiency of the model, and at the same time try to maintain the performance of the model. In the multi-scale feature extraction module, combining these two network structures can improve the running speed of the model while ensuring the feature extraction effect, which is suitable for real-time multi-object detection tasks.

[0089] Upsampling and Downsampling Operations: Downsampling operations (such as max pooling, strided convolution, etc.) can reduce the size of the feature map, expand the receptive field, and extract higher-level and more abstract semantic features, which are suitable for detecting larger objects; upsampling operations (such as transposed convolution, interpolation, etc.) can increase the size of the feature map, restore the detailed information, and combine with the downsampled feature map to achieve the alignment and fusion of features at different scales, enabling the model to detect objects of different sizes by using features at different levels simultaneously.

[0090] Top-down Feature Fusion Process:

[0091] Introduce Lateral Connections: During the top-down process, when transmitting information from the deep feature map to the shallow feature map, introducing lateral connections can effectively combine the high-resolution detailed information in the shallow feature map with the high-semantic information in the deep feature map. This can make up for the possible loss of detailed information when detecting solely relying on the deep feature map and improve the detection ability for small objects.

[0092] Adjust the Number of Channels through Convolution Operations: Adjust the number of channels of the shallow feature map through convolution operations to match the number of channels of the deep feature map for subsequent fusion operations. This step ensures the dimensional consistency of feature maps at different levels, enabling fusion operations such as pixel-by-pixel addition or concatenation.

[0093] Pixel-by-pixel addition and splicing fusion: Adding the adjusted shallow feature map and the upsampled deep feature map pixel by pixel can enable the deep fusion of the information of the two feature maps, highlighting the feature information of the target. The splicing operation can combine the two feature maps in the channel dimension, retaining more feature details. Through the fusion of these two methods, the semantic information of the deep feature map and the spatial information of the shallow feature map can be fully utilized, improving the model's positioning and recognition accuracy of the target.

[0094] In summary, the advantages of the present invention are as follows: By constructing a multi-scale feature pyramid, it is possible to simultaneously capture target information at different scales, enabling the model to effectively detect targets of different sizes and shapes, effectively improving the comprehensiveness and accuracy of detection. Introducing an attention mechanism to weight the features of the target area can highlight the key features of the target and suppress the interference of irrelevant information such as the background, further improving the detection accuracy of the target, especially in scenarios with complex backgrounds or dense targets; As an improved object detection model, YOLOv5 has a fast detection speed and high detection accuracy. It adopts efficient network structures such as CSPNet, reducing the computational amount. At the same time, through strategies such as adaptive training, the model can quickly and accurately process input data to meet the requirements of real-time detection; The Kalman filter can track the detected targets in real time, generating the motion trajectories of the targets. Through two steps of prediction and update, the position and state estimation of the targets are continuously corrected, ensuring the real-time and stability of tracking, enabling the entire system to achieve real-time multi-object detection and tracking; This method can be applied to different scenarios, such as traffic monitoring, security monitoring, intelligent driving, etc., and can effectively detect and track various types of targets. Users can select different model configurations and parameters according to actual needs and flexibly apply them to different scenarios and tasks.

[0095] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art of this industry should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection required by the present invention is defined by the appended claims and their equivalents.

Claims

1. A real-time multi-object detection method based on deep learning, characterized in that, Including: Collecting the video stream and image sequence of the scene to be detected, and preprocessing them, including denoising, normalization, size adjustment and enhancement operations; Using a deep learning model to perform multi-scale feature extraction on the preprocessed input data, constructing a multi-scale feature pyramid, where each layer of feature map represents target information at different scales; Fusing the feature maps at different scales, introducing an attention mechanism, and weighting the features of the target region; Using an improved YOLOv5 object detection model to perform object detection on the fused feature map, outputting the bounding box and class probability of the target; using the Kalman filter algorithm to perform real-time tracking on the detected target, generating the motion trajectory of the target; Performing post-processing on the detection results, including non-maximum suppression, target association and trajectory smoothing; finally outputting the detection results and tracking trajectories of the target.

2. The real-time multi-object detection method based on deep learning according to claim 1, wherein, The collecting the video stream and image sequence of the scene to be detected, and preprocessing them, including denoising, normalization, size adjustment and enhancement operations specifically includes: Enhancing the contrast of the image by adjusting the pixel distribution of the image; removing the noise in the image by using Gaussian filtering and median filtering methods; adjusting the brightness, contrast and saturation of the image; Adjusting the pixel values of the image to a distribution with a mean of zero and a standard deviation of one; making the covariance matrix between different pixels close to the identity matrix to remove the correlation between data; Realizing data augmentation through histogram equalization, contrast stretching, color space conversion, color space conversion.

3. The real-time multi-object detection method based on deep learning according to claim 2, characterized in that The using a deep learning model to perform multi-scale feature extraction on the preprocessed input data, constructing a multi-scale feature pyramid, where each layer of feature map represents target information at different scales specifically includes: Selecting a pre-trained deep learning model as the base network, which has a multi-layer structure, and each layer corresponds to different scales and semantic information; Extracting feature maps from different stages of the base network; performing upsampling and downsampling operations on the extracted feature maps to generate feature maps at different scales; Starting from the deep feature map, sequentially upsampling and fusing the shallow feature maps; In the top-down process, introducing lateral connections to fuse the deep high-semantic feature map with the shallow high-resolution feature map; Through the above process, generating feature maps at different scales, forming a pyramid structure, and each layer of feature map corresponds to a different object detection scale.

4. The real-time multi-object detection method based on deep learning according to claim 3, characterized in that The fusing the feature maps at different scales, introducing an attention mechanism, and weighting the features of the target region specifically includes: Fusing the feature maps at different scales, and the methods include concatenation, addition and weighted average; Applying an attention mechanism to the fused feature map, and enhancing the feature representation of the target region through the calculation of channel attention and spatial attention; Training the model through backpropagation and optimization algorithms to learn the best feature fusion and attention weights to maximize the performance of object detection.

5. The real-time multi-object detection method based on deep learning according to claim 4, wherein, The using an improved YOLOv5 object detection model to perform object detection on the fused feature map, outputting the bounding box and class probability of the target specifically includes: Use a backbone network to enhance the feature extraction ability; introduce an attention mechanism for adaptive spatial feature fusion to improve the model's perception ability of the target area; Introduce multi-scale prediction layers at different stages of the model to detect targets of different sizes; use a feature pyramid network to enhance the expression of multi-scale features; Adopt focal loss to alleviate the class imbalance problem, and introduce smooth L1 loss to improve the accuracy of bounding box regression; Use data augmentation techniques to augment the training data; adopt transfer learning, initialize with a model pre-trained on a dataset, and fine-tune on a specific task dataset Input the multi-scale feature maps after preprocessing and feature fusion into an improved YOLOv5 object detection model; the model gradually extracts high-level features in the image through multiple convolutional layers and attention modules; The model outputs multiple candidate bounding boxes, each bounding box containing the coordinate information and confidence of the target; classify the targets within each bounding box and output the probability distribution of each predefined class.

6. The real-time multi-object detection method based on deep learning according to claim 5, characterized in that, The use of the Kalman filter algorithm to perform real-time tracking on the detected targets and generate the motion trajectories of the targets specifically includes: Initialize a Kalman filter for each detected target, set the initial state and covariance matrix; predict the state of the target in the next frame according to the motion model of the target; update the covariance matrix to reflect the uncertainty of the prediction; After detecting a target in the current frame, use the observation data to update the state estimate of the Kalman filter; calculate the residual between the observation and the prediction, and adjust the state estimate to reduce the error; Use the Hungarian algorithm to associate the detection results of the current frame with the previous trajectories to ensure continuous tracking of each target; For the unmatched detection results, regard them as new targets and initialize new Kalman filters; Maintain the trajectory of each target, record its historical positions and states; when a target is not detected for a long time, mark it as lost and terminate the tracking.

7. The real-time multi-object detection method based on deep learning according to claim 6, characterized in that, The use of the Kalman filter algorithm to perform real-time tracking on the detected targets and generate the motion trajectories of the targets specifically includes: Predict the state of the target in the next frame according to the motion model of the target, and the prediction formula is: In the formula, is the predicted target state of the next frame; F is the state transition matrix, which is used to describe the relationship between the target states over time; is the state estimation value of the target in the previous frame; B is the control matrix, and u k-1 is the control input at the previous moment; Update the covariance matrix to reflect the uncertainty of the prediction, and the formula is: wherein, is the predicted covariance matrix, is the covariance matrix of the previous frame, Q is the covariance matrix of the process noise, which is used to describe the uncertainty in the target motion model, and F is the state transition matrix, which is used to describe the relationship between the target states over time.

8. The real-time multi-object detection method based on deep learning according to claim 7, characterized in that The post-processing of the detection results includes non-maximum suppression, target association, and trajectory smoothing; the final output of the detection results and tracking trajectories of the targets specifically includes: Calculate the confidence scores of all detection boxes, sort them from high to low according to the confidence scores, check the detection boxes one by one, and discard the current detection box if the intersection over union with the reserved detection boxes exceeds the threshold; Associate the detection results of the current frame with the trajectories of the previous frame to ensure the continuity of tracking; use the Hungarian algorithm to calculate the matching cost between the detection boxes and the existing trajectories, and associate the detection boxes with the trajectories according to the principle of minimum cost matching; Use a moving average to smooth the trajectories, perform a weighted average on the historical positions of the trajectories to reduce mutations; Output the filtered detection boxes and class probabilities; output the unique identifier, trajectory points, and motion states of the targets; draw the detection boxes, class labels, and trajectories on the video frames for users to view.

9. A real-time multi-object detection system based on deep learning, characterized in that, A method for implementing a real-time multi-object detection method based on deep learning as described in claims 1-8, comprising: A data acquisition and preprocessing module: used to collect video streams and image sequences of the scene to be detected, and perform preprocessing on them, including denoising, normalization, size adjustment, and enhancement operations; A multi-scale feature extraction module: uses a deep learning model to perform multi-scale feature extraction on the preprocessed input data, constructs a multi-scale feature pyramid, where each layer of feature map represents target information at different scales; A feature fusion and attention mechanism module: fuses feature maps of different scales, and at the same time introduces an attention mechanism to perform weighted processing on the features of the target area to enhance the expression ability of target features; A target detection and tracking module: uses an improved YOLOv5 target detection model to perform target detection on the fused feature map, and outputs the bounding boxes and class probabilities of the targets; uses the Kalman filter algorithm to perform real-time tracking on the detected targets to generate the motion trajectories of the targets; A post-processing and result output module: performs post-processing on the detection results, including non-maximum suppression, target association, and trajectory smoothing; finally outputs the detection results and tracking trajectories of the targets.

10. The real-time multi-object detection system based on deep learning according to claim 9, characterized in that, The multi-scale feature extraction module specifically includes: The multi-scale feature pyramid extracts feature maps of different scales through a deep residual network and a lightweight network, and realizes the alignment and fusion of feature maps through upsampling and downsampling operations; In the top-down process, lateral connections are introduced to perform convolution operations on the shallow feature maps to adjust their number of channels to match the size of the deep feature maps; the adjusted shallow feature maps are added and concatenated pixel by pixel with the upsampled deep feature maps to realize the fusion of the high-semantic feature maps in the deep layer and the high-resolution feature maps in the shallow layer.

Citation Information

Cited By

  • Anomaly detection method and system based on RGB-D fusion and dual-time flow feature learning

    CN121074851A

  • Offshore multi-target tracking detection method for unmanned ship

    CN121305346A

  • Power transmission line mechanized operation target detection method and device based on deep learning

    CN121415325A

  • Road user trajectory extraction and behavior labeling method integrating detection and tracking functions

    CN121527733A

  • Road user trajectory extraction and behavior labeling method fusing detection and tracking functions

    CN121527733B