A multi-target regional dynamic vehicle pedestrian identification method and system
By employing a multi-target area dynamic vehicle and pedestrian recognition method, which utilizes preprocessing of multi-camera image data, target detection, and spatiotemporal correlation, the problems of duplicate recognition and missed tracking in multi-camera target recognition and tracking are solved, achieving more efficient and accurate target monitoring.
Patent Information
- Application Number
- CN202510158035.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-02-13
AI Technical Summary
Existing multi-camera target recognition and tracking technologies are prone to problems such as repeated target recognition and missed tracking in highly dynamic environments and complex backgrounds, especially due to spatiotemporal correlation difficulties caused by differences in camera perspective, position, and time synchronization.
A multi-target area dynamic vehicle and pedestrian recognition method is adopted. Image data from multiple cameras is acquired, preprocessed, and target detection is performed. Kalman filtering or particle filtering algorithms are used to predict the target's trajectory. Spatiotemporal correlation is performed through joint data association methods, and the target trajectory is updated in real time by combining data fusion models and trajectory optimization algorithms.
It improves the coverage and accuracy of target recognition, overcomes the blind spots of a single camera view, provides more comprehensive monitoring capabilities, and ensures efficient target identification and tracking in complex environments.
Smart Images

Figure CN120071300B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of multi-camera systems, and particularly relates to a multi-target region dynamic vehicle and pedestrian identification method and system. BACKGROUND
[0002] With the rapid development of intelligent transportation systems, public safety monitoring, and unmanned driving, multi-camera target identification and tracking technology has been widely applied in various scenarios. By obtaining information of target objects (such as vehicles, pedestrians, etc.) through multiple cameras, more comprehensive and accurate target state and location data can be provided. However, existing multi-camera target identification and tracking technology still faces several challenges, especially in high dynamic environments, complex backgrounds, and target occlusion.
[0003] In the process of multi-camera target identification and tracking, due to the differences in camera view angle, position, and time synchronization, spatio-temporal correlation is a complex and error-prone link. Traditional methods often rely on simple spatial distance and temporal difference for correlation, lacking in-depth understanding of target motion trajectories and appearance features. Such methods are prone to produce mismatching or missed tracking phenomena when facing fast-moving targets, occlusions, or frequent camera switching, resulting in target loss or repeated identification. SUMMARY
[0004] The purpose of the embodiments of the present application is to propose a multi-target region dynamic vehicle and pedestrian identification method and system to solve the technical problems of target repeated identification and missed tracking.
[0005] To solve the above technical problems, the embodiments of the present application provide a multi-target region dynamic vehicle and pedestrian identification method, which adopts the following technical solutions:
[0006] A multi-target region dynamic vehicle and pedestrian identification method, comprising the following steps:
[0007] Obtaining image data of a target region collected by multiple cameras;
[0008] Preprocessing the image data to extract features of the target region;
[0009] Performing target detection on the image data to identify vehicle and pedestrian targets in the image data and determine the positions and features of the targets;
[0010] Preliminary positioning of each target and prediction of the motion trajectory of the target through a Kalman filter algorithm or a particle filter algorithm;
[0011] Based on a joint data correlation method of the positions, features, and motion trajectories of the targets, spatio-temporal correlation of the targets identified by multiple cameras is performed.
[0012] In a possible implementation, after the step of spatio-temporally correlating the targets recognized by the plurality of cameras based on the joint data correlation method of the positions, features and motion trajectories of the targets, the method further includes:
[0013] The target trajectory is updated in real time through the data fusion model and the trajectory optimization algorithm.
[0014] In a possible implementation, the step of updating the target trajectory in real time through the data fusion model and the trajectory optimization algorithm includes:
[0015] Different weights are assigned to the targets provided by each camera;
[0016] According to the weighted fusion algorithm, the targets of each camera are fused according to their weights, the weighted average position of the target is calculated, and the final target position is generated;
[0017] Based on the target position, the target trajectory is fitted using the least squares method;
[0018] The position, speed and direction of the target are updated in real time using the target trajectory.
[0019] In a possible implementation, after the step of updating the position, speed and direction of the target in real time using the target trajectory, the method further includes:
[0020] Abnormal data in the target tracking process are detected;
[0021] The abnormal data are corrected using the Kalman filter or the robust estimation algorithm.
[0022] In a possible implementation, the step of detecting the targets in the image data, identifying the vehicle and pedestrian targets in the image data, and determining the positions and features of the targets includes:
[0023] The targets are detected and classified to identify whether they are vehicles or pedestrians, and to distinguish different types of targets;
[0024] The positions of each identified target are determined;
[0025] The features of each target are extracted using a convolutional neural network;
[0026] The categories, positions and features of the detected targets are output.
[0027] In a possible implementation, the step of preliminarily positioning each target and predicting the motion trajectory of the target through the Kalman filter algorithm or the particle filter algorithm includes:
[0028] initializing a state variable for each detected target, the state variable including position and velocity information of the target;
[0029] performing trajectory prediction according to the state variable;
[0030] updating the state variable of the target, and performing trajectory correction if there is a large deviation between the actual position and the predicted position of the target in the current frame.
[0031] In a possible implementation, the step of associating the targets recognized by the multiple cameras in time and space based on the joint data association method of the positions, features, and motion trajectories of the targets specifically includes:
[0032] sorting the targets according to timestamps;
[0033] matching the targets in different cameras in the same time period according to the timestamps, and performing time sequence alignment;
[0034] converting the targets of different cameras to a unified coordinate system;
[0035] performing spatial association according to the coordinate system of the targets.
[0036] To solve the above technical problems, the embodiments of the present application further provide a multi-target region dynamic vehicle and pedestrian identification system, which adopts the technical solutions as follows:
[0037] A multi-target region dynamic vehicle and pedestrian identification system includes:
[0038] An acquisition module is configured to acquire image data of a target region collected by multiple cameras.
[0039] A processing module is configured to pre-process the image data and extract features of the target region.
[0040] A detection module is configured to detect targets in the image data, identify vehicle and pedestrian targets in the image data, and determine positions and features of the targets.
[0041] A prediction module is configured to preliminarily locate each target and predict a motion trajectory of the target by using a Kalman filtering algorithm or a particle filtering algorithm.
[0042] An association module is configured to associate targets recognized by the multiple cameras in time and space based on a joint data association method of positions, features, and motion trajectories of the targets.
[0043] To solve the above technical problems, the embodiments of the present application further provide a computer device, which adopts the technical solutions as follows:
[0044] A computer device comprises a memory and a processor, the memory stores computer readable instructions, and the processor executes the computer readable instructions to implement the steps of the multi-target area dynamic vehicle and pedestrian identification method.
[0045] To solve the above technical problems, the embodiment of the application further provides a computer readable storage medium, which adopts the technical scheme as follows:
[0046] A computer readable storage medium stores computer readable instructions, and the processor executes the computer readable instructions to implement the steps of the multi-target area dynamic vehicle and pedestrian identification method.
[0047] Compared with the prior art, the embodiment of the application has the following beneficial effects:
[0048] The multi-target area dynamic vehicle and pedestrian identification method disclosed in the application comprises the following steps: acquiring image data of a target area collected by multiple cameras; preprocessing the image data to extract features of the target area; performing target detection on the image data to identify vehicle and pedestrian targets in the image data and determine positions and features of the targets; preliminarily positioning each target and predicting a motion trajectory of the target through a Kalman filtering algorithm or a particle filtering algorithm; and performing spatio-temporal correlation on the targets identified by the multiple cameras based on a joint data correlation method of the positions, features and motion trajectories of the targets. The application defines the acquisition, preprocessing, target detection and tracking of multi-camera image data and a spatio-temporal correlation method, so as to ensure that the system can efficiently identify and track targets in a wide area. Through cooperation of the multiple cameras, the system can overcome the problem of a blind area of a single camera and provide more comprehensive monitoring capability, thereby improving the coverage and accuracy of target identification. BRIEF DESCRIPTION OF DRAWINGS
[0049] In order to more clearly illustrate the schemes in the application, the drawings needed in the description of the embodiments of the application will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0050] Figure 1 is an exemplary system architecture diagram to which the application can be applied;
[0051] Figure 2 is a flowchart of one embodiment of the multi-target area dynamic vehicle and pedestrian identification method according to the application;
[0052] Figure 3is a structural schematic diagram of one embodiment of a multi-target regional dynamic vehicle pedestrian recognition system according to the present application;
[0053] Figure 4 is a structural schematic diagram of one embodiment of a computer device according to the present application. DETAILED DESCRIPTION
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs; the terminology used in the description herein is for describing particular embodiments only and is not intended to be limiting of the application; the description and claims herein and the above description of drawings herein utilize terms such as "comprising", "having" and "including" to convey construction that is inclusive of, but not limited to, such features unless otherwise indicated; the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0055] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrase "in an embodiment" in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily all directed to the same embodiment, or to a single alternative embodiment.
[0056] For those skilled in the art to better understand the technical solutions of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings below.
[0057] As shown in Figure 1 The system architecture 100 can include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a communication link medium between the terminal devices 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links, or optical fiber cables, etc.
[0058] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0059] The terminal devices 101, 102, and 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and the like.
[0060] The server 105 can be a server providing various services, such as a background server providing support for a page displayed on the terminal devices 101, 102, and 103.
[0061] It should be noted that the multi-target region dynamic vehicle and pedestrian identification method provided in the embodiments of the present application is generally executed by a server, and accordingly, a multi-target region dynamic vehicle and pedestrian identification system is generally provided in a server.
[0062] It should be understood that Figure 1 The number of terminal devices, networks, and servers in
[0063] With reference to Figure 2 , a flowchart of one embodiment of a multi-target region dynamic vehicle and pedestrian identification method according to the present application is shown. The multi-target region dynamic vehicle and pedestrian identification method includes the following steps:
[0064] In step S201, image data of a target region captured by a plurality of cameras is acquired.
[0065] In the present embodiment, an electronic device (for example, the server shown in Figure 1 The electronic device (for example, the server shown in
[0066] In this embodiment, image data is collected from multiple camera devices, and the collected image data contains objects in the target area (such as roads, parking lots, sidewalks, etc.) within the monitoring area. Each camera can be located at different positions or different angles to ensure coverage of different parts of the target area. To achieve the fusion of multi-camera data, time synchronization and spatial coordinate alignment are needed to ensure that image data from different cameras is consistent in time and space. For example, multiple cameras can work together to monitor a traffic area in real time, and simultaneously capture the same scene from different angles, which helps to achieve perspective complementation for target recognition and improve recognition accuracy.
[0067] Step S202, pre-processing the image data to extract features of the target area.
[0068] In this embodiment, the original image data obtained from multiple cameras is pre-processed to improve image quality and enhance the effect of target recognition. Specifically, the pre-processing operation includes:
[0069] Noise removal: using noise removal algorithms (such as median filtering, Gaussian filtering, etc.) to remove noise in the image to ensure the accuracy of subsequent processing.
[0070] Image enhancement: enhance image details through contrast enhancement, brightness adjustment, histogram equalization, etc. to make the target more obvious and facilitate subsequent target detection.
[0071] Edge detection: extract edge features in the image through edge detection algorithms (such as Canny operator, Sobel operator, etc.) to help separate the target and the background.
[0072] Through these pre-processing operations, the image quality is optimized, making the features of the target area more prominent and enhancing the effect of target recognition.
[0073] Step S203, target detection is performed on the image data to identify vehicle and pedestrian targets in the image data and determine the location and features of the targets.
[0074] In this embodiment, target detection is performed on the pre-processed image data to identify targets such as vehicles and pedestrians in the image. The target detection method that can be used includes a deep learning model based on convolutional neural network (CNN), such as:
[0075] YOLO (You Only Look Once): a real-time target detection algorithm that can detect multiple targets in an image simultaneously and return the position of each target's bounding box and its category.
[0076] Faster R-CNN: Efficient object detection through Region-CNN (R-CNN), suitable for accurate detection of multiple categories of objects.
[0077] RetinaNet: An object detection algorithm based on FPN (Feature Pyramid Network), capable of handling high-density small objects.
[0078] Each detected object is assigned a location identifier (such as the center coordinates or bounding box coordinates of the object), and further feature extraction is performed based on the appearance characteristics of the object (such as shape, color, texture). Feature extraction helps subsequent object classification, behavior analysis, and multi-camera data association. The output of this step is target data containing target location, category (vehicle or pedestrian), and features.
[0079] Step S204, preliminary positioning of each target, and prediction of the motion trajectory of the target through Kalman filtering algorithm or particle filtering algorithm.
[0080] In this embodiment, for each detected target, the position output by the target detection algorithm (such as the bounding box coordinates or center point coordinates) is used to achieve preliminary positioning of the target. Through the detection results, the system can determine the initial position of each target in the image. To achieve tracking and prediction of the target, the motion trajectory of the target in future frames needs to be predicted, which can be achieved using the following two filtering algorithms:
[0081] Kalman filtering algorithm: suitable for targets with stable motion, can provide optimal state estimation, in each frame, according to the previous position, velocity and acceleration of the target, predict the future position of the target. Through Kalman filtering, the difference between the measured value and the predicted value of the target can be used to optimize the trajectory of the target.
[0082] Particle filtering algorithm: suitable for targets with complex motion or non-linear, non-Gaussian noise. Particle filtering represents the state of the target through multiple particles and combines the measurement results for correction. Particle filtering can handle more complex motion trajectories and more chaotic dynamic environments.
[0083] This step improves the positioning accuracy of the target in the time dimension through continuous prediction of the target state, ensuring the stability of the target tracking.
[0084] Step S205, based on the joint data association method of the position, features and motion trajectory of the target, the targets recognized by multiple cameras are spatio-temporally associated.
[0085] In this embodiment, the spatio-temporal correlation is achieved by tracking and recognizing the target from multiple camera systems at multiple angles and perspectives to ensure that the same target is recognized by different cameras. The joint data correlation method includes the following key parts:
[0086] Temporal correlation: By comparing the timestamps of different cameras, it is ensured that the targets recognized by each camera are synchronized or close in time. If there is a deviation in the shooting time of the cameras, time synchronization processing is needed.
[0087] Spatial correlation: The positions of the target under different cameras are correlated. In order to overcome the difference in the perspective of different cameras, the world coordinate system can be used for coordinate transformation to convert the image coordinates of different cameras into a unified spatial coordinate system.
[0088] Target feature correlation: The appearance features (such as color, shape, texture, etc.) of the target are used to determine whether the targets recognized by different cameras are the same target. Targets with similar features are considered to be the same target.
[0089] Motion trajectory correlation: The motion trajectory of the target is used to further verify the spatio-temporal correlation. The motion trajectory of the target will change continuously in time and space, and if the motion trajectories of the targets recognized by two cameras are highly consistent, they can be considered to belong to the same target.
[0090] Data fusion: The positions, features, and motion trajectories of the target are combined to finally fuse the target recognition results of multiple cameras to ensure that the spatio-temporal data of each target remains consistent. For example, Kalman filtering can be used to weight and fuse the target data between different cameras to optimize the final position and trajectory of the target.
[0091] This step ensures that data obtained from different perspectives can be effectively merged to accurately identify and track targets, avoiding misidentification due to different camera perspectives or target occlusion.
[0092] The present application ensures that the system can efficiently identify and track targets in a wide area by defining the acquisition, preprocessing, target detection and tracking of multi-camera image data, and the spatio-temporal correlation method. Through the cooperation of multiple cameras, the system can overcome the blind area problem of a single camera perspective and provide more comprehensive monitoring capabilities, thereby improving the coverage and accuracy of target recognition.
[0093] In some optional implementations of the present embodiment, after the step of correlating the targets recognized by multiple cameras in space and time based on the joint data correlation method of the target's position, features, and motion trajectory, the method further includes:
[0094] The target trajectory is updated in real time through the data fusion model and trajectory optimization algorithm.
[0095] In this embodiment, based on the target detection data and spatio-temporal correlation results provided by multiple cameras, the motion trajectory of each target is updated in real time through a data fusion model (such as weighted average method, Kalman filter, particle filter, extended Kalman filter, etc.). The data fusion model combines spatio-temporal information from different cameras, can integrate measurement results from different sensors, and eliminate errors or omissions caused by incomplete information from a single perspective or a single camera. Through the weighted average algorithm, the state of the target is weighted according to the confidence of different cameras, and a more accurate target trajectory is obtained. The weight of each camera data can be dynamically adjusted according to factors such as the coverage range of the perspective, image quality, distance, etc. If the target's motion is linear and less affected by noise, Kalman filter is a suitable model. It recursively updates the difference between the predicted trajectory of the target and the actual measurement value, gradually corrects the state estimation of the target, and enhances the accuracy of the trajectory. For complex or nonlinear target motion, particle filter provides a flexible solution. It represents the target by the state of multiple particles and adjusts the weight of the particles according to the new observation value to finally obtain the optimal trajectory estimation. In the process of real-time tracking, the trajectory of the target may be discontinuous or erroneous due to occlusion, recognition error or noise, in order to solve this problem, a trajectory optimization algorithm is used for real-time correction:
[0096] Least squares method: use the least squares method to optimize the trajectory of the target, which is particularly suitable for smoothing the trajectory. This algorithm minimizes the sum of squares of errors between the target trajectory and the observed values, making the target trajectory more consistent with the actual motion pattern.
[0097] Continuity constraint: combine the motion characteristics of the target, such as speed, acceleration and motion direction, to optimize the continuity of the target trajectory, ensuring smooth transition of the target's motion trajectory in time and space.
[0098] Bayesian optimization: in the case of high uncertainty, Bayesian method can provide a probabilistic framework to optimize the target trajectory based on historical data and prior knowledge, and predict the future trajectory.
[0099] Real-time updating of target trajectory is achieved in the following way:
[0100] Dynamic tracking: the motion of the target is dynamic, so the trajectory needs to be updated in real time at each time. The data fusion model and trajectory optimization algorithm will process the target data of each frame in real time, update the current position, speed, acceleration and other information of the target.
[0101] Continuous prediction: through real-time updating of the target trajectory, the system can predict the future position of the target after processing each frame of image. This is crucial for target tracking and behavior prediction.
[0102] There is also a need to improve tracking accuracy and stability:
[0103] Error correction: Due to differences in image quality, viewing angle, and occlusion between different cameras, the initial target tracking may have errors. Through data fusion and trajectory optimization algorithms, the system can correct these errors to ensure the accuracy and stability of the target trajectory.
[0104] Robustness enhancement: When the target deviates due to external factors such as occlusion and light changes, the trajectory optimization algorithm can effectively correct the target trajectory and enhance the robustness of the system.
[0105] Improvement combined with historical data:
[0106] In the trajectory optimization process, the algorithm not only relies on the current detection results, but also refers to historical data for improvement. For example, the previous target position and motion pattern will affect the current trajectory prediction, so as to better adapt to the dynamic changes of the target.
[0107] Through the fusion of multiple camera data, the system can reduce error accumulation in complex dynamic environments, accurately track target positions, and update in real time. The ability to ensure that the system can still efficiently track in the case of rapid target motion changes, avoiding trajectory drift or loss.
[0108] In some optional implementations of the present embodiment, the above step of updating the target trajectory in real time through the data fusion model and the trajectory optimization algorithm includes:
[0109] Assign different weights to the target provided by each camera;
[0110] According to the weighted fusion algorithm, fuse the target of each camera according to its weight, calculate the weighted average position of the target, and generate the final target position;
[0111] Based on the target position, use the least squares method to fit the target trajectory;
[0112] Using the target trajectory, update the target's position, speed, and direction in real time.
[0113] In this embodiment, when multiple cameras simultaneously detect and track a target, the target information provided by each camera may have varying degrees of accuracy and reliability. Therefore, it is necessary to assign different weights to the target data provided by each camera to reflect the quality of each camera and its importance in the current scene. Weight assignment can be based on several factors, such as camera position and viewing angle: if a camera's viewing angle can more comprehensively cover the target, or if the camera is closer to the target, then the target data from that camera has a higher weight. Image quality: if a camera has high image quality (e.g., high resolution, proper exposure, no obstructions), then the data from that camera should be given a higher weight. Target salience: for detected targets, more salient target data (e.g., clear targets, no overlap) can be given a higher weight. The core objective of weight allocation is to reasonably adjust the influence of different camera data on the final target position and trajectory, ensuring that high-quality data dominates the fusion process and avoiding data with large noise or errors from affecting the results.
[0114] The purpose of weighted fusion algorithms is to calculate the weighted average position of the target based on the data provided by each camera and its weight. By combining data from different cameras, a more accurate and stable target position is obtained. Assuming multiple cameras detect the same target, each camera provides a position (e.g., the center coordinates of the bounding box). , ) and a weight The weighted average position of the target can then be calculated using the following formula:
[0115] ,
[0116] in, , , is the target location detected by each camera. This refers to the weight of the camera data. By calculating the weighted average position, the system can obtain a more accurate target position. The final target position is the weighted average of all camera data, representing the target position jointly confirmed by multiple cameras. This result will serve as the final position of the target in the global coordinate system.
[0117] The least squares method can fit the data to find a smooth trajectory and minimize the error between the target position and the fitted curve.
[0118] Trajectory Fitting: By performing least-squares fitting on the target's historical positions, a smoother and more continuous target trajectory can be obtained. Assume the target's position information at different time points is ( , ),(in is time), the least squares method fits the target trajectory by optimizing the following objective function:
[0119] ,
[0120] wherein, and is the function of trajectory fitting, representing the predicted position of the target at time The least squares method can minimize the difference between the historical position of the target and the fitted trajectory, thereby obtaining a smoother and more accurate target motion trajectory.
[0121] After fitting by the least squares method, the obtained target trajectory will remove short-term random fluctuations and provide a smooth and realistic motion pattern trajectory.
[0122] After the target trajectory is fitted by the least squares method, the trajectory information can be used to update the position, speed and direction of the target in real time. Based on the fitted trajectory, the accurate position of the target at the current time point is calculated. According to the motion trajectory of the target, the instantaneous speed of the target, i.e. the moving speed of the target at the current position, is calculated. The speed is usually calculated by the difference in position between two consecutive frames:
[0123] ,
[0124] wherein is the time difference between two frames. According to the motion trajectory of the target, the motion direction of the target, i.e. the angle of the target relative to the reference direction, can be calculated by the displacement vector between two consecutive frames.
[0125] Through the above updating method, the system can update the motion state of the target in each frame of image in real time, accurately grasp the position, speed and direction of the target, and help the prediction, behavior analysis and subsequent path planning of the target.
[0126] The present application optimizes the target position and trajectory through the weighted fusion algorithm and the least squares method, improves the accuracy of multi-camera data fusion; by giving different weights to different cameras, the system can integrate the advantages of each camera to ensure accurate positioning and tracking of the target under different conditions (such as camera angle difference, image quality difference, etc.), and the least squares method can effectively reduce the influence of noise and error by fitting the trajectory.
[0127] In some optional implementations of the embodiment, after the step of using the target trajectory to update the position, speed and direction of the target in real time, the method further comprises:
[0128] detecting abnormal data in the target tracking process;
[0129] The abnormal data is corrected using Kalman filtering or robust estimation algorithm.
[0130] In this embodiment, the definition of abnormal data is that during target tracking, some abnormal data may occur due to the following reasons:
[0131] Target occlusion: the target is partially or completely occluded by other objects, causing the camera to fail to accurately detect the target's position.
[0132] Sensor error: the camera or sensor has errors due to failure or external environmental influences (such as light changes, weather, reflections, etc.).
[0133] Fast motion change: the target makes a sharp turn or changes speed dramatically, which may cause position prediction deviation.
[0134] Multi-target interference: multiple targets gather in the same view, causing the detection system to fail to accurately distinguish different targets.
[0135] Abnormal data detection method: the purpose of detecting abnormal data is to identify data points that do not conform to the normal trajectory by analyzing the target's motion pattern or position change.
[0136] The following abnormal data detection methods can be used:
[0137] Trajectory-based detection: by analyzing the difference between the target's historical trajectory and the current predicted trajectory, if the difference between the current data point and the predicted trajectory exceeds the preset threshold, it can be considered as an abnormal value.
[0138] Speed and acceleration-based detection: if the target's speed or acceleration value in a frame suddenly changes dramatically, exceeding the reasonable range, it can also be considered as abnormal data.
[0139] Motion consistency detection: if the target's motion direction or trajectory changes between adjacent frames do not conform to expectations (for example, the target's turning angle changes too much), it can be considered as abnormal data.
[0140] Labeling of abnormal data: once abnormal data is detected, these data are distinguished from the normal trajectory by specific identification (such as setting a flag bit or segmenting the data), and are corrected later.
[0141] The purpose of modifying abnormal data is that the presence of abnormal data may affect the accuracy and stability of the target trajectory, especially in applications with high real-time requirements for target tracking. Therefore, it is necessary to modify abnormal data using effective algorithms to ensure the continuity and consistency of the target trajectory. Kalman filtering method can be used, which is an optimal linear estimation method. According to the historical state of the target and the current observation data, the prediction and observation information are combined to smooth and modify the abnormal data. If abnormal data is detected in the tracking of the target trajectory, Kalman filtering can modify the abnormal value by weighted average of the prediction and observation value of the abnormal data, so that the target trajectory is smoother. For example, Kalman filtering will predict the next position of the target according to the motion model of the target, and adjust the state estimation of the target by comparing the difference between the actual observation value and the predicted value. It can make optimal correction based on the previous target state and current observation value, and can effectively deal with occasional abnormal data and continuously adjust the target trajectory through recursion.
[0142] Robust estimation algorithm can also be used, which is specifically used to handle large noise or low-quality data. In the presence of abnormal data, robust estimation algorithm can modify by reducing the influence of abnormal value on overall data estimation. Similar to Kalman filtering, robust estimation algorithm considers the deviation of abnormal data and uses weighted method to reduce the influence of abnormal value. Common robust estimation algorithms include M-estimators, RANSAC, etc., which can correct the trajectory in the presence of uncertainty.
[0143] Compared with Kalman filtering, robust estimation method is more suitable for handling large abnormal data in data. It can effectively exclude the interference of false data and ensure the stability of trajectory correction.
[0144] After finding abnormal data, Kalman filtering or robust estimation algorithm will smooth and modify these abnormal data to update the position, speed and direction of the target, ensuring the continuity and consistency of the target trajectory.
[0145] The present application enhances the robustness of the system by detecting and modifying abnormal data in the target tracking process. The introduction of Kalman filtering or robust estimation algorithm enables the system to adaptively correct errors in complex environments, avoiding abnormal behaviors such as occlusion and rapid change in target tracking, thereby improving the stability and accuracy of the system.
[0146] In some optional implementations of the embodiment, the above steps of detecting the target in the image data, identifying the vehicle and pedestrian targets in the image data, and determining the position and characteristics of the target specifically include:
[0147] detect and classify targets, identify whether they are vehicles or pedestrians, and distinguish between different types of targets;
[0148] determine the location of each identified target;
[0149] extract features of each target using a convolutional neural network;
[0150] output the category, location, and features of the detected targets.
[0151] In this embodiment, target detection and classification is a key task in computer vision, which is used to detect and identify target objects from images. In the multi-target area dynamic vehicle and pedestrian recognition method, the target in the image data needs to be detected first, and then classified to identify whether it is a vehicle or a pedestrian, and possibly other target categories. Through target detection algorithms such as YOLO, Faster R-CNN, SSD, etc., the location of the target in the image is located, and these algorithms can find the target in the image and return the bounding box of each target, i.e. the rectangular area where the target is located. After detecting the target, it needs to be classified, and the image features of the target are extracted through the use of a convolutional neural network (CNN), and the classifier (such as a Softmax classifier) is used to determine whether the target belongs to a vehicle, a pedestrian, or other categories. This is a very important step in multi-target recognition, because vehicles and pedestrians have different motion characteristics, sizes, shapes, etc. First, distinguish the target type, once the target is detected, the classification process will ensure that each target is correctly classified into the corresponding type (such as vehicle, pedestrian, etc.). For example, vehicles usually have a larger size, a more regular shape, while pedestrians are usually smaller, and have a certain diversity in body parts and posture. Based on these differences, the classification model can distinguish between these different types of targets.
[0152] Position determination is the accurate calculation of the specific location of a target in an image after target detection. The position of a target is usually represented by the coordinates of the bounding box, which is usually represented by the coordinates of the top-left corner and the bottom-right corner of the rectangular box, or by the coordinates of the center point of the target. For each detected target, the target detection algorithm returns a rectangular box whose coordinates describe the spatial position of the target. In addition to the bounding box, the position of the target can also be represented by the coordinates of the center point of the target, which is of great significance in tracking and subsequent trajectory optimization. For each target, the coordinates of the center point ( , ) can be calculated from the upper and lower boundary values of the bounding box, which can more accurately represent the position of the target in the image, especially for subsequent tracking and behavior analysis of the target.
[0153] Convolutional Neural Network (CNN) is a deep learning model widely used in computer vision tasks, especially in object detection and image classification. In the process of object detection, CNN extracts features from images through multiple convolution operations, generating feature maps with deep semantic information, effectively capturing detailed features in the image. For each detected object, CNN processes the target region (part within the bounding box) to extract deep features that describe the object. These features can include texture, color, shape, edge information, etc. of the object. Through the extracted features, the object can be further classified, behavior analyzed, and even its attributes identified (such as license plate recognition, face recognition, etc.). These features can also play a role in object tracking, helping the system maintain consistency between multiple frames. CNN usually contains multiple convolutional layers and pooling layers, which gradually reduce the spatial dimensions of the image while extracting higher-order feature information. Finally, the features are converted into a feature vector of the object through a fully connected layer.
[0154] After completing object detection and feature extraction, the system will output the class, location, and features of each detected object. These three pieces of information are the basic outputs of a multi-object recognition system:
[0155] Class: The type of object obtained through the classification model (such as vehicles, pedestrians, cyclists, etc.).
[0156] Location: The location of the object usually includes its bounding box coordinates in the image or the center point coordinates of the object.
[0157] Features: The feature vector of the object extracted from the CNN, used to describe the appearance and morphology of the object. These features can be further used in subsequent tasks such as object tracking and behavior analysis.
[0158] Output format: The detection results are stored as a list or matrix containing target information, with each target's information including class, location coordinates, and feature vector.
[0159] This application extracts the features of the target through the convolutional neural network and classifies them, achieving efficient and accurate target recognition. This step can accurately distinguish between different types of targets such as vehicles and pedestrians, laying a solid foundation for subsequent tracking and recognition. The application of convolutional neural networks improves the robustness of object detection and adapts to the needs of target recognition in complex environments with different backgrounds and lighting conditions.
[0160] In some optional implementations of the present embodiment, the above steps of preliminarily positioning each target and predicting the motion trajectory of the target through Kalman filtering algorithm or particle filtering algorithm specifically include:
[0161] Initialize state variables for each detected object, including position and velocity information of the object;
[0162] Perform trajectory prediction based on the state variables;
[0163] Update the state variables of the object, and if there is a significant deviation between the actual position and the predicted position of the object in the current frame, perform trajectory correction.
[0164] In this embodiment, initializing state variables is the first step in object tracking. The motion state of each object can be described by a set of state variables, typically including position (such as the coordinates of the object in the image) and velocity (such as the velocity of the object in the axis and axis direction).
[0165] Object position: Position is usually the coordinates of the object in the image or the center point coordinates of the bounding box. Suppose the position of the object in the image is ( , ).
[0166] Object velocity: Object velocity refers to the moving speed of the object in the image coordinate system. Typically, this includes horizontal velocity ( ) and vertical velocity ( ).
[0167] State vector: In Kalman filtering or particle filtering, the state of the object is usually represented as a vector containing position and velocity information. For example, the state vector may be represented as:
[0168] ,
[0169] Here, and are the position of the object, and are the velocity of the object.
[0170] When detecting an object, the system will first assign an initial state variable to each object (for example, by initializing the detected object position and velocity in the current frame). If there is no velocity information, the initial velocity can be assumed to be zero, or initialized by the velocity of the previous frame.
[0171] Trajectory prediction is to estimate the possible position of the object in the next frame based on the current state variables (position and velocity). Kalman filtering and particle filtering algorithms can be used to predict the trajectory of the object. Kalman filtering is a recursive estimation algorithm that predicts the trajectory of the object based on a linear motion model. By knowing the position and velocity of the object, Kalman filtering can predict the expected position of the object in the next frame.
[0172] The prediction step is usually done according to the following equation:
[0173] ,
[0174] where, is the predicted state, is the state transition matrix, is the control matrix, is the control input (such as acceleration, etc.).
[0175] Particle filtering represents the motion state of the target by generating multiple "particles" and updating and resampling these particles based on the target motion model. Particle filtering is suitable for non-linear or non-Gaussian problems, and through multiple iterations, particle filtering can provide more accurate trajectory prediction.
[0176] In particle filtering, the prediction of the target trajectory is usually based on the known particle distribution (representing the possible states of the target) and is predicted by updating the weights and positions of the particles.
[0177] Predict the target position: Through the prediction formula or particle filtering algorithm, the system calculates the expected position of the target at the next time according to the state variables (position and velocity) of the target.
[0178] Update the state: The predicted position of the target is based on the state of the last frame, but due to the uncertainty of the environment and measurement errors, the actual observed position of the target may deviate from the predicted position. Therefore, it is necessary to update the state of the target.
[0179] State update of Kalman filter: Kalman filter updates the target state by combining the predicted value and the actual observed value. When there is an error between the actual observed position of the target and the predicted position , Kalman filter calculates a Kalman gain and corrects the state of the target through weighted average.
[0180] The update formula is usually:
[0181] ,
[0182] where, is the Kalman gain, is the actual observed value, is the predicted position.
[0183] Particle filter, on the other hand, corrects the target's state by adjusting the weight of particles and resampling. If the difference between the actual position and the predicted position of the target in the current frame is large, particle filter will increase the number of high-weight particles, thereby improving the estimation accuracy of the target. When the deviation between the predicted position and the actual position is too large, trajectory correction is needed. This can be achieved in the following two ways:
[0184] Error correction: directly adjust the state variables (position, velocity, etc.) of the target to better conform to the current observation.
[0185] Data association: re-associate the targets recognized by multiple cameras to ensure the consistency of the target's trajectory.
[0186] The present application provides accurate target positioning and tracking capability through Kalman filter or particle filter algorithm. The combination of preliminary positioning and trajectory prediction enables the system to accurately predict the position of the target in the case of discontinuous target motion trajectory, reduces the possibility of tracking breakage, and ensures continuity and stability.
[0187] In some optional implementations of the present embodiment, the joint data association method based on the position, features and motion trajectory of the target, which temporally and spatially associates the targets recognized by multiple cameras, specifically includes:
[0188] Sort the targets according to the timestamps;
[0189] Match the targets in different cameras within the same time period according to the timestamps for time alignment;
[0190] Convert the targets of different cameras to a unified coordinate system;
[0191] Spatially associate the targets according to their coordinate systems.
[0192] In the present embodiment, in a multi-camera system, the image data captured by the cameras usually has a timestamp to record the capture time of each image. Since multiple cameras may capture targets simultaneously or sequentially, it is necessary to sort the targets detected by different cameras according to the timestamps.
[0193] Purpose of sorting: the purpose of sorting is to ensure the time order of the targets, so that the subsequent time alignment and spatial association processes can be based on correct time information.
[0194] Implementation: sort the detection results from different cameras according to the timestamps, and assign a time order to each target. In this way, the time consistency of each target can be ensured, so that the subsequent steps can be processed based on the correct time order.
[0195] Time alignment refers to matching the targets detected in different cameras within the same time period through timestamps, ensuring the consistency of target data in different cameras in time.
[0196] Matching targets: In the detection results of different cameras, the targets within the same time period should be the same object (e.g., vehicles or pedestrians). Therefore, it is necessary to match the targets detected by different cameras at the same time point according to the timestamps.
[0197] Purpose of time alignment: The purpose of time alignment is to unify the detection data of targets in the time axis, so that in the subsequent spatial correlation step, the positions and motion trajectories of targets in different cameras can be accurately compared.
[0198] Implementation: Assuming that the timestamps of a certain moment in camera A and camera B are and If the two timestamps are very close or the same, the targets detected by camera A and B at this time point can be considered as detection data at the same time. At this time, subsequent analysis can be carried out according to these time-aligned data.
[0199] Coordinate transformation is a key step in multi-camera data fusion. Since the view angle, position and orientation of each camera may be different, the coordinate system of the targets captured by them may also be different. Therefore, it is necessary to convert the coordinate systems of different cameras into a unified coordinate system for accurate spatial correlation.
[0200] Coordinate system difference: Each camera has its own coordinate system, usually based on the image coordinate system (pixel coordinates) of the camera. These coordinate systems may not be consistent with the global coordinate system (such as the world coordinate system or the unified coordinate system). The differences between the coordinate systems of different cameras may include:
[0201] Image resolution and scale;
[0202] Rotation angle and position of the camera;
[0203] Field of view angle and projection method of the camera;
[0204] Coordinate transformation method: The method of converting the coordinate system of the camera to the unified coordinate system usually includes camera calibration and geometric transformation. For example, projection matrix and transformation matrix can be used to convert the position of the target in each camera to the position in the global coordinate system.
[0205] Implementation: By calibrating the internal parameters (focal length, distortion coefficients, etc.) and external parameters (position, orientation, etc.) of each camera, a projection matrix can be constructed. Then, by multiplying the pixel coordinates of the target by the corresponding projection matrix, the coordinate system of the camera can be converted to the global coordinate system.
[0206] Spatial correlation refers to matching and merging the targets recognized by different cameras in a unified coordinate system. Since the target may appear in the field of view of multiple cameras, the purpose of spatial correlation is to ensure that the target can be accurately tracked and identified in a unified coordinate system.
[0207] The goal of spatial correlation is to compare the spatial positions of the target with those of other targets in other cameras to determine whether they are the same object. Since the spatial positions of the target may be offset between different cameras (due to differences in viewing angles), it is necessary to match them through algorithms.
[0208] Spatial correlation method:
[0209] Distance threshold method: A distance threshold can be set, and only when the positional difference between two targets is less than a certain threshold, they are considered to be the same target. This method is simple and easy to implement, but may be affected by noise.
[0210] Feature matching method: In addition to position, other features of the target (such as color, shape, motion trajectory, etc.) can also be used as the basis for spatial correlation. Similarity measures (such as Euclidean distance, cosine similarity, etc.) can be used to measure the similarity of targets.
[0211] Implementation: In a unified coordinate system, spatial correlation is performed by calculating the difference in coordinates (or other features) of each target, ensuring consistency between multiple cameras. If multiple cameras recognize the same target, the system will track it as a unified target.
[0212] Through the spatio-temporal correlation method, the present application ensures that the target data from multiple cameras is effectively merged in a unified coordinate system. The sorting and matching of timestamps enable precise temporal alignment of targets detected by multiple cameras at the same time, solving the problem of time asynchronization. The unified conversion of the coordinate system and spatial correlation ensure that targets from different cameras can be accurately merged into a unified target, avoiding mis-matching caused by camera viewing angle differences.
[0213] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology and application system of using digital computer or machine controlled by digital computer to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain optimal results.
[0214] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The software technology of artificial intelligence mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0215] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through computer readable instructions, which can be stored in a computer readable storage medium. The program can include the processes of the above-mentioned embodiments when executed. Among them, the storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0216] It should be understood that although each step in the flowchart of the accompanying drawings is displayed in sequence according to the direction of the arrow, these steps are not necessarily executed in sequence according to the direction of the arrow. Unless otherwise stated herein, the execution of these steps is not strictly limited in sequence, and they can be executed in other orders. Moreover, at least part of the steps in the flowchart of the accompanying drawings can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution sequence is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.
[0217] Further referring to Figure 3 , as an implementation of the method shown in Figure 2 , the present application provides a multi-target area dynamic vehicle and pedestrian identification system. An embodiment of a multi-target area dynamic vehicle and pedestrian identification system, the system embodiment corresponds to the method embodiment shown in Figure 2 , the system can be applied to various electronic devices.
[0218] As Figure 3As shown, the multi-target area dynamic vehicle and pedestrian identification system described in the embodiment includes an acquisition module 301, a processing module 302, a detection module 303, a prediction module 304, and an association module 305. Among them:
[0219] The acquisition module 301 is configured to acquire image data of a target area collected by multiple cameras.
[0220] The processing module 302 is configured to pre-process the image data and extract features of the target area.
[0221] The detection module 303 is configured to perform target detection on the image data, identify vehicle and pedestrian targets in the image data, and determine the positions and features of the targets.
[0222] The prediction module 304 is configured to preliminarily locate each target and predict the motion trajectory of the target by using a Kalman filtering algorithm or a particle filtering algorithm.
[0223] The association module 305 is configured to perform spatio-temporal association of targets recognized by multiple cameras based on a joint data association method of the positions, features, and motion trajectories of the targets.
[0224] The multi-target area dynamic vehicle and pedestrian identification system provided by the application can ensure that the system can efficiently identify and track targets in a wide area by defining the acquisition, preprocessing, target detection and tracking of multi-camera image data, and the spatio-temporal association method. Through the cooperation of multiple cameras, the system can overcome the blind area problem of a single camera perspective and provide more comprehensive monitoring capability, thereby improving the coverage and accuracy of target identification.
[0225] In some optional implementation manners of the embodiment, the association module 305 is further configured to:
[0226] Update the target trajectory in real time by using a data fusion model and a trajectory optimization algorithm.
[0227] The multi-target area dynamic vehicle and pedestrian identification system provided by the application can reduce error accumulation in a complex dynamic environment by fusing multiple camera data, accurately track the target position, and update in real time. The ability ensures that the system can still efficiently track in the case of rapid target motion change, avoiding trajectory drift or loss.
[0228] In some optional implementation manners of the embodiment, the association module 305 is further configured to:
[0229] Different weights are assigned to the targets provided by each camera.
[0230] According to the weighted fusion algorithm, the target of each camera is fused according to its weight, the weighted average position of the target is calculated, and the final target position is generated.
[0231] Based on the target position, the target trajectory is fitted using the least square method;
[0232] Using the target trajectory, the position, speed and direction of the target are updated in real time.
[0233] The multi-target regional dynamic vehicle and pedestrian recognition system provided by the application optimizes the target position and trajectory through the weighted fusion algorithm and the least square method, improves the accuracy of multi-camera data fusion, and through assigning different weights to different cameras, the system can integrate the advantages of each camera, ensuring accurate positioning and tracking of the target under different conditions (such as camera angle difference, image quality difference, etc.), and the least square method effectively reduces the influence of noise and error by fitting the trajectory.
[0234] In some optional implementations of the embodiment, the association module 305 is further configured to:
[0235] Detect abnormal data in the target tracking process;
[0236] Correct the abnormal data using Kalman filtering or robust estimation algorithm.
[0237] The multi-target regional dynamic vehicle and pedestrian recognition system provided by the application detects and corrects abnormal data in the target tracking process, enhances the robustness of the system, and the introduction of Kalman filtering or robust estimation algorithm enables the system to adaptively correct errors in complex environments, avoiding abnormal behaviors in target tracking such as occlusion, rapid change, etc., thereby improving the stability and accuracy of the system.
[0238] In some optional implementations of the embodiment, the detection module 303 is further configured to:
[0239] Detect and classify the target, identify whether it is a vehicle or a pedestrian, and distinguish different types of targets;
[0240] Determine the position of each identified target;
[0241] Extract the features of each target using convolutional neural network;
[0242] Output the category, position and features of the detected target.
[0243] The multi-target region dynamic vehicle and pedestrian identification system provided by the application realizes efficient and accurate target identification by extracting the features of the target through a convolutional neural network, and this step can accurately distinguish different types of targets such as vehicles and pedestrians, thereby laying a solid foundation for subsequent tracking and identification, and the application of the convolutional neural network improves the robustness of target detection and adapts to the target identification requirements under different backgrounds and light conditions in complex environments.
[0244] In some optional implementation manners of the embodiment, the prediction module 304 is further configured to:
[0245] initializing a state variable for each detected target, the state variable including position and speed information of the target;
[0246] performing trajectory prediction according to the state variable;
[0247] updating the state variable of the target, and performing trajectory correction if there is a large deviation between the actual position and the predicted position of the target in the current frame.
[0248] The multi-target region dynamic vehicle and pedestrian identification system provided by the application performs target motion trajectory prediction through a Kalman filter or a particle filter algorithm, provides accurate target positioning and tracking capability, and the combination of preliminary positioning and trajectory prediction enables the system to accurately predict the position of the target in the case of discontinuous target motion trajectory, reduces the possibility of tracking breakage, and thus ensures continuity and stability.
[0249] In some optional implementation manners of the embodiment, the association module 305 is further configured to:
[0250] sorting the targets according to the timestamps;
[0251] matching the targets in different cameras in the same time period according to the timestamps to perform time sequence alignment;
[0252] converting the targets of different cameras to a unified coordinate system;
[0253] performing spatial association according to the coordinate system of the target.
[0254] The multi-target region dynamic vehicle and pedestrian identification system provided by the application ensures effective fusion of target data under multiple cameras in a unified coordinate system through a space-time association method, the sorting and matching of the timestamps enable the targets detected by multiple cameras at the same time to be accurately time sequence aligned, thereby solving the problem of time asynchronization, and the unified conversion and spatial association of the coordinate system ensure that the targets from different cameras can be accurately merged into a unified target, thereby avoiding the mismatch caused by the perspective difference of the cameras.
[0255] To solve the above technical problems, the embodiment of the present application further provides a computer device. For details, please refer to Figure 4 , Figure 4 The basic structure block diagram of the computer device of the embodiment is shown in the figure.
[0256] The computer device 4 comprises a memory 41, a processor 42 and a network interface 43 which are connected to each other through a system bus. It should be pointed out that only the computer device 4 with components 41-43 is shown in the figure, but it should be understood that all the shown components are not required to be implemented, and more or less components can be alternatively implemented. Among them, the computer device herein can be understood by those skilled in the art as a device capable of automatically performing numerical calculation and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0257] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server and other computing devices. The computer device can interact with the user through a keyboard, a mouse, a remote controller, a touchpad or a voice control device.
[0258] The memory 41 includes at least one type of readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (e.g., an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or a memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 4. Of course, the memory 41 can also include both the internal storage unit and the external storage device of the computer device 4. In this embodiment, the memory 41 is generally used to store an operating system and various application software installed on the computer device 4, such as computer readable instructions of a multi-target area dynamic vehicle and pedestrian identification method, etc. In addition, the memory 41 can also be used to temporarily store various data that have been output or will be output.
[0259] The processor 42 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run computer readable instructions or process data stored in the memory 41, such as computer readable instructions of the multi-target area dynamic vehicle and pedestrian identification method.
[0260] The network interface 43 can include a wireless network interface or a wired network interface, and is generally used to establish a communication connection between the computer device 4 and other electronic devices.
[0261] The computer device provided in the present application can ensure that the system can efficiently identify and track targets in a wide area by defining the acquisition, preprocessing, target detection and tracking of multi-camera image data, and the spatio-temporal correlation method. Through the cooperation of multiple cameras, the system can overcome the blind area problem of a single camera perspective and provide more comprehensive monitoring capabilities, thereby improving the coverage and accuracy of target identification.
[0262] The application also provides another implementation, namely providing a computer readable storage medium, the computer readable storage medium stores computer readable instructions, the computer readable instructions can be executed by at least one processor, so that the at least one processor executes the steps of a multi-target area dynamic vehicle and pedestrian identification method as described above.
[0263] The computer readable storage medium provided by the application ensures that the system can efficiently identify and track targets in a wide range by defining the acquisition, preprocessing, target detection and tracking of multi-camera image data, and the space-time correlation method; through the cooperation of multiple cameras, the system can overcome the blind area problem of single camera perspective, provide more comprehensive monitoring capability, thereby improving the coverage and accuracy of target identification.
[0264] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-mentioned embodiment method can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better implementation. Based on such understanding, the technical solutions of the application can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for making a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) execute the method described in each embodiment of the application.
[0265] Obviously, the above-described embodiments are only some of the embodiments of the application, not all the embodiments, and the preferred embodiments of the application are given in the drawings, but do not limit the patent scope of the application. The application can be implemented in many different forms, and on the contrary, the purpose of providing these embodiments is to make the disclosure of the application more thorough and comprehensive. Although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing specific embodiments, or make equivalent replacements to some technical features. Any equivalent structure made by using the content of the specification and drawings, directly or indirectly applied to other related technical fields, is also within the scope of the patent protection of the application.
Claims
1. A multi-target area dynamic vehicle pedestrian recognition method, characterized in that, The method comprises the following steps: obtaining image data of a target area collected by multiple cameras; preprocessing the image data to extract features of the target area; detecting targets in the image data to identify vehicle and pedestrian targets and determine the positions and features of the targets; initially locating each target and predicting the motion trajectory of the target by using a Kalman filtering algorithm or a particle filtering algorithm; spatiotemporally correlating the targets identified by the multiple cameras based on a joint data correlation method of the positions, features and motion trajectories of the targets; updating the target trajectory in real time by using a data fusion model and a trajectory optimization algorithm; wherein the step of updating the target trajectory in real time by using the data fusion model and the trajectory optimization algorithm comprises: assigning different weights to the targets provided by each camera; fusing the targets of each camera according to a weighted fusion algorithm, calculating the weighted average position of the targets and generating the final target position; fitting the target trajectory by using a least squares method based on the target position; updating the position, speed and direction of the target in real time by using the target trajectory; detecting abnormal data in the target tracking process; correcting the abnormal data by using a Kalman filtering or robust estimation algorithm.
2. The multi-target area dynamic vehicle pedestrian recognition method of claim 1, wherein, The step of detecting targets in the image data to identify vehicle and pedestrian targets and determine the positions and features of the targets specifically comprises: detecting and classifying the targets to identify whether they are vehicles or pedestrians and distinguish different types of targets; determining the position of each identified target; extracting the features of each target by using a convolutional neural network; outputting the category, position and features of the detected targets.
3. The multi-target area dynamic vehicle pedestrian recognition method of claim 1, wherein, The step of initially locating each target and predicting the motion trajectory of the target by using a Kalman filtering algorithm or a particle filtering algorithm specifically comprises: initializing state variables of each detected target, the state variables including the position and speed information of the target; predicting the trajectory according to the state variables; updating the state variables of the target, and correcting the trajectory if there is a large deviation between the actual position and the predicted position of the target in the current frame.
4. The multi-target area dynamic vehicle pedestrian recognition method of claim 1, wherein, The step of spatiotemporally correlating the targets identified by the multiple cameras based on a joint data correlation method of the positions, features and motion trajectories of the targets specifically comprises: sorting the targets according to timestamps; aligning the time sequence by matching the targets in different cameras within the same time period; converting the targets of different cameras to a unified coordinate system; spatially correlating the targets according to the coordinate system.
5. A multi-target area dynamic vehicle pedestrian recognition system, characterized by, The method comprises: an obtaining module configured to obtain image data of a target area collected by multiple cameras; a processing module configured to preprocess the image data to extract features of the target area; a detection module configured to detect targets in the image data to identify vehicle and pedestrian targets and determine the positions and features of the targets; a prediction module configured to initially locate each target and predict the motion trajectory of the target by using a Kalman filtering algorithm or a particle filtering algorithm; a spatiotemporal correlation module configured to spatiotemporally correlate the targets identified by the multiple cameras based on a joint data correlation method of the positions, features and motion trajectories of the targets; and a data fusion module configured to update the target trajectory in real time by using a data fusion model and a trajectory optimization algorithm. The association module is configured to perform spatio-temporal association on the targets recognized by the multiple cameras based on a joint data association method of the positions, features and motion trajectories of the targets.
6. A computer device, comprising: The computer readable storage medium has stored thereon computer readable instructions which, when executed by a processor, implement the steps of the multi-target area dynamic vehicle and pedestrian identification method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that, The computer readable storage medium has stored thereon computer readable instructions which, when executed by a processor, implement the steps of the multi-target area dynamic vehicle and pedestrian identification method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Cross-space-time correlation method and device for night target trajectory
CN117495913A
Dynamic target tracking method and system for vehicle-mounted mining intrinsic safety type video monitoring
CN119027860A