Multi-target area dynamic vehicle and pedestrian identification method and system

Through the multi-target area dynamic vehicle pedestrian recognition method, combined with the acquisition, preprocessing, object detection and space-time correlation technology of multi-camera data, the problems of mismatch and missed tracking in multi-camera target recognition are solved, and more efficient and accurate target recognition and tracking are achieved.

CN120071300AActive Publication Date: 2025-05-30SHENZHEN YUNYITONG DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510158035.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-30
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

Existing multi-camera target recognition and tracking technologies are prone to mismatch or miss tracking under high dynamic environments, complex backgrounds and target occlusion, resulting in missing targets or repeated recognition.

Method used

The multi-target area dynamic vehicle pedestrian recognition method is adopted, and the image data of multiple cameras is obtained, pre-processing and target detection is performed, vehicle and pedestrian targets are identified, and the target motion trajectory is predicted through Kalman filtering or particle filtering algorithms, and the target motion trajectory is associated with space-time and time, and the target trajectory is updated in real time.

Benefits of technology

It improves the coverage and accuracy of target recognition, reduces mismatch and missed tracking, and ensures that the system can efficiently identify and track targets in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071300A_ABST
    Figure CN120071300A_ABST
Patent Text Reader

Abstract

The embodiment of the invention belongs to the field of multi-camera systems, and relates to a multi-target area dynamic vehicle and pedestrian recognition method, which comprises the following steps: acquiring image data of a target area acquired by a plurality of cameras; preprocessing the image data, and extracting features of the target area; target detection is carried out on the image data, vehicle and pedestrian targets in the image data are identified, and the positions and features of the targets are determined; performing preliminary positioning on each target, and predicting the motion trail of the target through a Kalman filtering algorithm or a particle filtering algorithm; and carrying out space-time association on the targets identified by the plurality of cameras based on a joint data association method of the positions, the features and the motion trails of the targets. The invention also provides a multi-target area dynamic vehicle and pedestrian identification system. The coverage rate and accuracy of target identification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of multi-camera systems, and in particular, to a method and system for dynamically identifying vehicles and pedestrians in multiple target areas. Background Art

[0002] With the rapid development of fields such as intelligent transportation systems, public security monitoring, and driverless driving, multi-camera target recognition and tracking technology has been widely applied in various scenarios. By obtaining information about target objects (such as vehicles, pedestrians, etc.) through multiple cameras, more comprehensive and accurate target status and position data can be provided. However, existing multi-camera target recognition and tracking technology still faces several challenges, especially in high-dynamic environments, complex backgrounds, and target occlusion situations.

[0003] During the multi-camera target recognition and tracking process, due to issues such as camera viewing angles, position differences, and time synchronization, spatio-temporal correlation is a complex and error-prone link. Traditional methods often rely on simple spatial distances and time differences for correlation, lacking in-depth understanding of target motion trajectories and appearance features. Such methods are prone to false matching or missed tracking when faced with fast-moving targets, occlusion, or frequent cross-camera switching, resulting in target loss or repeated recognition. Summary of the Invention

[0004] The purpose of the embodiments of this application is to propose a method and system for dynamically identifying vehicles and pedestrians in multiple target areas to solve the technical problems of target repeated recognition and missed tracking.

[0005] To solve the above technical problems, the embodiments of this application provide a method for dynamically identifying vehicles and pedestrians in multiple target areas, adopting the following technical solutions: A method for dynamically identifying vehicles and pedestrians in multiple target areas includes the following steps: Obtain image data of the target area collected by multiple cameras; Preprocess the image data and extract the features of the target area; Perform target detection on the image data to identify vehicle and pedestrian targets in the image data, and determine the positions and features of the targets; Perform preliminary positioning on each target, and predict the motion trajectory of the target through the Kalman filter algorithm or the particle filter algorithm; Based on the joint data association method of the positions, features, and motion trajectories of the targets, perform spatio-temporal association on the targets identified by multiple cameras.

[0006] In a possible implementation manner, after the step of performing spatio-temporal association on the targets identified by multiple cameras based on the joint data association method of the positions, features, and motion trajectories of the targets, the method further includes: The target trajectory is updated in real time through a data fusion model and a trajectory optimization algorithm.

[0007] In a possible implementation, the steps of updating the target trajectory in real time through a data fusion model and a trajectory optimization algorithm include: Assign different weights to the targets provided by each camera; According to the weighted fusion algorithm, fuse the targets of each camera according to their weights, calculate the weighted average position of the targets, and generate the final target position; Based on the target position, use the least squares method to fit the target trajectory; Use the target trajectory to update the position, speed, and direction of the target in real time.

[0008] In a possible implementation, after the step of using the target trajectory to update the position, speed, and direction of the target in real time, it further includes: Detect abnormal data during the target tracking process; Use the Kalman filter or robust estimation algorithm to correct the abnormal data.

[0009] In a possible implementation, the steps of performing target detection on the image data, identifying vehicle and pedestrian targets in the image data, and determining the position and characteristics of the targets specifically include: Detect and classify the targets to identify whether they are vehicles or pedestrians and distinguish different types of targets; Determine the position of each identified target; Use a convolutional neural network to extract the characteristics of each target; Output the category, position, and characteristics of the detected targets.

[0010] In a possible implementation, the steps of performing preliminary positioning on each target and predicting the movement trajectory of the target through the Kalman filter algorithm or particle filter algorithm specifically include: Initialize the state variables for each detected target, and the state variables include the position and speed information of the target; Perform trajectory prediction according to the state variables; Update the state variables of the target. If there is a large deviation between the actual position and the predicted position of the target in the current frame, perform trajectory correction.

[0011] In a possible implementation, the steps of performing spatio-temporal association on the targets identified by multiple cameras through the joint data association method based on the position, characteristics, and movement trajectory of the target specifically include: Sort the targets according to the time stamp; Match the targets in different cameras within the same time period according to the time stamps for time series alignment; Convert the targets of different cameras to a unified coordinate system; Perform spatial association according to the coordinate system of the targets.

[0012] To solve the above technical problems, the embodiments of the present application further provide a multi-target area dynamic vehicle and pedestrian recognition system, which adopts the following technical solutions: A multi-target area dynamic vehicle and pedestrian recognition system, comprising: An acquisition module, configured to acquire image data of a target area collected by multiple cameras; A processing module, configured to preprocess the image data and extract the features of the target area; A detection module, configured to perform target detection on the image data, identify vehicle and pedestrian targets in the image data, and determine the positions and features of the targets; A prediction module, configured to perform preliminary positioning on each target and predict the movement trajectory of the target through a Kalman filtering algorithm or a particle filtering algorithm; An association module, configured to perform spatio-temporal association on the targets identified by multiple cameras based on a joint data association method of the positions, features, and movement trajectories of the targets.

[0013] To solve the above technical problems, the embodiments of the present application further provide a computer device, which adopts the following technical solutions: A computer device, comprising a memory and a processor, wherein computer-readable instructions are stored in the memory, and when the processor executes the computer-readable instructions, the steps of a multi-target area dynamic vehicle and pedestrian recognition method as described above are implemented.

[0014] To solve the above technical problems, the embodiments of the present application further provide a computer-readable storage medium, which adopts the following technical solutions: A computer-readable storage medium, wherein computer-readable instructions are stored on the computer-readable storage medium, and when the computer-readable instructions are executed by a processor, the steps of a multi-target area dynamic vehicle and pedestrian recognition method as described above are implemented.

[0015] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: A multi-target area dynamic vehicle and pedestrian recognition method disclosed in this application obtains image data of a target area collected by multiple cameras; preprocesses the image data to extract features of the target area; performs target detection on the image data to identify vehicle and pedestrian targets in the image data and determines the positions and features of the targets; performs preliminary positioning on each target and predicts the movement trajectory of the target through the Kalman filter algorithm or the particle filter algorithm; and performs spatio-temporal association on the targets recognized by multiple cameras based on a joint data association method of the positions, features, and movement trajectories of the targets. By defining the acquisition, preprocessing, target detection and tracking, and spatio-temporal association methods of multi-camera image data, this application ensures that the system can efficiently identify and track targets in a wide area; through the cooperation of multiple cameras, the system can overcome the blind spot problem of a single camera's perspective and provide more comprehensive monitoring capabilities, thereby improving the coverage and accuracy of target recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the solutions in this application, the following briefly introduces the drawings required for the description of the embodiments of this application. Obviously, the following drawings are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 is an exemplary system architecture diagram to which this application can be applied; Figure 2 is a flowchart of an embodiment of a multi-target area dynamic vehicle and pedestrian recognition method according to this application; Figure 3 is a schematic structural diagram of an embodiment of a multi-target area dynamic vehicle and pedestrian recognition system according to this application; Figure 4 is a schematic structural diagram of an embodiment of a computer device according to this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used in the description of the embodiments of this application herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the description and claims of this application and the above drawings are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the description and claims of this application or the above drawings are used to distinguish different objects and are not used to describe a specific order.

[0019] References to "embodiments" in this specification mean that specific features, structures, or characteristics described in connection with the embodiments can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and is not necessarily referring to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0020] To enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0021] As Figure 1 shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0022] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0023] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 (Moving Picture Experts Group Audio Layer III) players, MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, and desktop computers, etc.

[0024] The server 105 may be a server providing various services, such as a background server that provides support for the pages displayed on the terminal devices 101, 102, 103.

[0025] It should be noted that a multi-target area dynamic vehicle and pedestrian recognition method provided by the embodiments of the present application is generally executed by the server. Correspondingly, a multi-target area dynamic vehicle and pedestrian recognition system is generally set in the server.

[0026] It should be understood, Figure 1The numbers of the terminal devices, networks, and servers therein are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.

[0027] Continuing to refer to Figure 2 , a flowchart of an embodiment of a multi-target area dynamic vehicle and pedestrian recognition method according to the present application is shown. The multi-target area dynamic vehicle and pedestrian recognition method includes the following steps: Step S201, obtaining image data of a target area collected by multiple cameras.

[0028] In this embodiment, the electronic device (such as Figure 1 the server shown) on which a multi-target area dynamic vehicle and pedestrian recognition method runs can send or receive data through a wired connection method or a wireless connection method. It should be noted that the above wireless connection methods can include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.

[0029] In this embodiment, it involves collecting image data from multiple camera devices, and the collected image data contains objects in the target area (such as roads, parking lots, sidewalks, etc.) within the monitoring area. Each camera can be located at different positions or different perspectives to ensure coverage of different parts of the target area. To achieve the fusion of multi-camera data, time synchronization and spatial coordinate alignment are required to ensure the consistency of image data from different cameras in terms of time and space. For example, multiple cameras can work together to monitor a traffic area in real time, simultaneously shooting the same scene from different angles, which helps to achieve perspective supplementation during target recognition and improve the recognition accuracy.

[0030] Step S202, preprocessing the image data and extracting features of the target area.

[0031] In this embodiment, the raw image data obtained from multiple cameras is preprocessed to improve the image quality and enhance the effect of target recognition. Specifically, the preprocessing operations include: Denoising: Using denoising algorithms (such as median filtering, Gaussian filtering, etc.) to remove noise in the image and ensure the accuracy of subsequent processing.

[0032] Image enhancement: Enhancing image details through methods such as contrast enhancement, brightness adjustment, histogram equalization, etc., making the target more obvious and facilitating subsequent target detection.

[0033] Edge detection: Extract the edge features in the image through edge detection algorithms (such as Canny operator, Sobel operator, etc.) to help separate the target and the background.

[0034] Through these preprocessing operations, the image quality is optimized, making the features of the target area more prominent and enhancing the effect of target recognition.

[0035] Step S203: Perform target detection on the image data, identify the vehicle and pedestrian targets in the image data, and determine the positions and features of the targets.

[0036] In this embodiment, perform the target detection task on the preprocessed image data to identify targets such as vehicles and pedestrians in the image. The target detection methods that can be used include deep learning models based on convolutional neural networks (CNNs), such as: YOLO (You Only Look Once): A real-time target detection algorithm that can detect multiple targets in an image simultaneously and return the positions of the bounding boxes of each target and their categories.

[0037] Faster R-CNN: Achieve efficient target detection through Region Convolutional Neural Network (R-CNN), suitable for accurate detection of multi-class targets.

[0038] RetinaNet: A target detection algorithm based on FPN (Feature Pyramid Network) that can handle small objects with high density.

[0039] Each detected target will be assigned a position identifier (such as the center coordinates or bounding box coordinates of the target), and further feature extraction will be performed according to the appearance features of the target (such as shape, color, texture). Feature extraction helps with subsequent target classification, behavior analysis, and the association of multi-camera data. The output of this step is target data containing the target position, category (vehicle or pedestrian), and features.

[0040] Step S204: Perform preliminary localization on each target and predict the motion trajectory of the target through the Kalman filter algorithm or the particle filter algorithm.

[0041] In this embodiment, for each detected target, the preliminary localization of the target is achieved through the position (such as bounding box coordinates or center point coordinates) output by the target detection algorithm. Through the detection results, the system can determine the initial position of each target in the image. To achieve the tracking and prediction of the target, it is necessary to predict the motion trajectory of the target in future frames. The following two filtering algorithms can be used for prediction: Kalman Filter Algorithm: Suitable for targets with stable motion, it can provide the best state estimation. In each frame, based on the previous position, velocity, and acceleration of the target, it predicts the future position of the target. Through the Kalman filter, the trajectory of the target can be optimized according to the difference between the measured value and the predicted value of the target.

[0042] Particle Filter Algorithm: Suitable for situations where the target motion is complex or has non - linear and non - Gaussian noise. The particle filter represents the target state through multiple particles and corrects it in combination with the measurement results. The particle filter can handle more complex motion trajectories and relatively chaotic dynamic environments.

[0043] This step improves the positioning accuracy of the target in the time dimension through continuous prediction of the target state, ensuring the stability of target tracking.

[0044] Step S205, based on the joint data association method of the position, features, and motion trajectory of the target, spatially and temporally associate the targets identified by multiple cameras.

[0045] In this embodiment, spatial - temporal association is to track and identify the target from multiple angles and perspectives through a multi - camera system to ensure that the targets identified by different cameras are the same. The joint data association method includes the following key parts: Temporal Association: By comparing the timestamps of different cameras, ensure that the targets identified by each camera are synchronized or close in time. If there is a deviation in the shooting time of the cameras, time synchronization processing is required.

[0046] Spatial Association: Associate based on the positions of the target in different cameras. To overcome the perspective differences of different cameras, the world coordinate system can be used for coordinate transformation to convert the image coordinates of different cameras into a unified spatial coordinate system.

[0047] Target Feature Association: Determine whether the targets identified by different cameras are the same target through the appearance features of the target (such as color, shape, texture, etc.). Targets with similar features are considered to be the same target.

[0048] Motion Trajectory Association: Further verify the spatial - temporal association through the motion trajectory of the target. The motion trajectory of the target changes continuously in time and space. If the motion trajectories of the targets identified by two cameras are highly consistent, they can be considered to belong to the same target.

[0049] Data Fusion: Combine the position, features, and motion trajectory of the target, and finally fuse the target recognition results of multiple cameras to ensure that the spatial - temporal data of each target is consistent. For example, the Kalman filter can be used to perform weighted fusion on the target data between different cameras to optimize the final position and trajectory of the target.

[0050] This step ensures that the data obtained from different perspectives can be effectively merged, so as to accurately identify and track the target, and avoid misidentification caused by different camera perspectives or target occlusion.

[0051] This application ensures that the system can efficiently identify and track targets in a wide area by defining the acquisition, preprocessing, target detection and tracking of multi-camera image data, as well as spatio-temporal association methods; through the cooperation of multiple cameras, the system can overcome the blind area problem of a single camera perspective and provide more comprehensive monitoring capabilities, thereby improving the coverage rate and accuracy of target recognition.

[0052] In some alternative implementation manners of this embodiment, after the step of spatio-temporally associating the targets recognized by multiple cameras in the above-mentioned joint data association method based on the position, features and motion trajectory of the target, the following steps are further included: The target trajectory is updated in real time through a data fusion model and a trajectory optimization algorithm.

[0053] In this embodiment, based on the target detection data and spatio-temporal association results provided by multiple cameras, the motion trajectory of each target is updated in real time through a data fusion model (such as weighted average method, Kalman filter, particle filter, extended Kalman filter, etc.). The data fusion model combines spatio-temporal information from different cameras, can integrate the measurement results of different sensors, and eliminate errors or omissions caused by incomplete information from a single perspective or a single camera. Through the weighted average algorithm, the state of the target is weighted according to the confidence of different cameras to obtain a more accurate target trajectory. The weights of the data of each camera can be dynamically adjusted according to factors such as the coverage range of its perspective, image quality, distance, etc. If the motion of the target is linear and less affected by noise, the Kalman filter is a suitable model. It recursively updates according to the difference between the predicted trajectory of the target and the actual measurement value, and gradually corrects the state estimation of the target to enhance the accuracy of the trajectory. For the motion of complex or non-linear targets, the particle filter provides a flexible solution. It represents the target through the states of multiple particles and adjusts the weights of the particles according to new observations to finally obtain the optimal trajectory estimation. During the real-time tracking process, the trajectory of the target may be discontinuous or have errors due to occlusion, misidentification or noise. To solve this problem, a trajectory optimization algorithm is used for real-time correction: Least squares method: The least squares method is used to optimize the trajectory of the target, which is especially suitable for smoothing the trajectory. This algorithm makes the trajectory of the target more conform to the actual motion pattern by minimizing the sum of the squares of the errors between the target trajectory and the observations.

[0054] Continuity Constraint: Combining the motion characteristics of the target, such as speed, acceleration, and motion direction, etc., to optimize the continuity of the target trajectory and ensure the smooth transition of the target's motion trajectory in time and space.

[0055] Bayesian Optimization: In the case of high uncertainty, the Bayesian method can provide a probabilistic framework to optimize the target trajectory based on historical data and prior knowledge and infer future trajectories.

[0056] The real-time update of the target trajectory is carried out in the following ways: Dynamic Tracking: The motion of the target is dynamically changing, so the trajectory needs to be updated in real time at each moment. The data fusion model and the trajectory optimization algorithm will process the target data of each frame in real time to update information such as the current position, speed, and acceleration of the target.

[0057] Continuous Prediction: Through the real-time update of the target trajectory, the system can predict the future position of the target after processing each frame of image. This is crucial for target tracking and behavior prediction.

[0058] It is also necessary to improve the tracking accuracy and stability: Error Correction: Due to reasons such as the image quality, perspective difference, and occlusion of different cameras, there may be errors in the initial target tracking. Through the data fusion and trajectory optimization algorithms, the system can correct these errors to ensure the accuracy and stability of the target trajectory.

[0059] Robustness Enhancement: When the target deviates due to external factors (such as occlusion and light change), the trajectory optimization algorithm can effectively correct the target trajectory and enhance the robustness of the system.

[0060] Improvement by Combining Historical Data: In the process of trajectory optimization, the algorithm not only depends on the current detection results but also refers to historical data for improvement. For example, the previous target positions and motion patterns will affect the current trajectory prediction, so as to better adapt to the dynamic changes of the target.

[0061] Through the fusion of data from multiple cameras in this application, the system can reduce error accumulation in a complex dynamic environment, accurately track the target position, and the real-time update ability ensures that the system can still track efficiently when the target motion changes rapidly, avoiding trajectory drift or loss.

[0062] In some optional implementation manners of this embodiment, the steps of real-time updating the target trajectory through the data fusion model and the trajectory optimization algorithm include: Assign different weights to the targets provided by each camera; According to the weighted fusion algorithm, the targets of each camera are fused according to their weights, the weighted average position of the targets is calculated, and the final target position is generated; Based on the target position, the least squares method is used to fit the target trajectory; Using the target trajectory, the position, speed, and direction of the target are updated in real time.

[0063] In this embodiment, when multiple cameras detect and track a target simultaneously, the target information provided by each camera may have different degrees of accuracy and reliability. Therefore, different weights need to be assigned to the target data provided by each camera to reflect the quality of each camera and its importance in the current scenario. The assignment of weights can be based on multiple factors. For example, the position and perspective of the camera: if the perspective of a certain camera can cover the target more comprehensively, or the camera is closer to the target, then the weight of the target data of this camera is higher. Image quality: if the image quality of a certain camera is high (for example, high resolution, proper exposure, no occlusion, etc.), then the data of this camera should be assigned a higher weight. Salience of the target: for the detected target, the target data that is more salient (such as clear target, no overlap) can be given a higher weight. The core goal of weight assignment is to reasonably adjust the influence of data from different cameras on the final target position and trajectory, ensure that high-quality data dominates in the fusion process, and avoid data with large noise or errors from affecting the results.

[0064] The purpose of the weighted fusion algorithm is to calculate the weighted average position of the target based on the data provided by each camera and its weight. By combining the data from different cameras, a more accurate and stable target position is obtained. Suppose multiple cameras detect the same target, and each camera provides a position (such as the center coordinates of the bounding box , ), as well as a weight , then the weighted average position of the target can be calculated by the following formula: , where, , is the target position detected by each camera, is the weight of the data of this camera. By calculating the weighted average position, the system can obtain a more accurate target position. The final target position is the weighted average of the data of all cameras, representing the target position jointly confirmed by multiple cameras. This result will be used as the final position of the target in the global coordinate system.

[0065] The least squares method can fit the data to find a smooth trajectory and minimize the error between the target position and the fitted curve.

[0066] Trajectory Fitting: By performing least-squares fitting on the historical positions of the target, a smoother and more continuous target trajectory can be obtained. Assume that the position information of the target at different time points is ( , ), (where is time), and the least-squares method fits the target trajectory by optimizing the following objective function: , where, and are functions for trajectory fitting, representing the predicted position of the target at time . The least-squares method can minimize the difference between the target historical positions and the fitted trajectory, thereby obtaining a smoother and more accurate target motion trajectory.

[0067] After fitting by the least-squares method, the obtained target trajectory will remove short-term random fluctuations and provide a smooth trajectory that conforms to the actual motion pattern.

[0068] After the target trajectory is fitted by the least-squares method, the trajectory information can be used to update the position, speed, and direction of the target in real time. Based on the fitted trajectory, calculate the exact position of the target at the current time point. According to the target's motion trajectory, calculate the instantaneous speed of the target, that is, the moving rate of the target at the current position. Usually, the speed is calculated by the position difference between two consecutive frames of the target: , where is the time difference between two frames. According to the target's motion trajectory, calculate the motion direction of the target, that is, the angle of the target relative to the reference direction, and the direction angle can be calculated by the displacement vector between two consecutive frames of the target.

[0069] Through the above update method, the system can update the motion state of the target in real time in each frame of the image, maintain accurate control of the target's position, speed, and direction, and contribute to target prediction, behavior analysis, and subsequent path planning.

[0070] This application optimizes the target position and trajectory through a weighted fusion algorithm and the least-squares method, improving the accuracy of multi-camera data fusion; by assigning different weights to different cameras, the system can integrate the advantages of each camera, ensuring accurate positioning and tracking of the target under different conditions (such as differences in camera perspectives, image quality differences, etc.). The least-squares method effectively reduces the influence of noise and errors by fitting the trajectory.

[0071] In some alternative implementation manners of this embodiment, after the step of using the target trajectory to update the position, speed, and direction of the target in real time, the following is further included: Detect abnormal data during the target tracking process; Use Kalman filtering or robust estimation algorithms to correct the abnormal data.

[0072] In this embodiment, the definition of abnormal data is that during the target tracking process, some abnormal data may occur, which may be caused by the following reasons: Target occlusion: The target is partially or completely occluded by other objects, resulting in the camera being unable to accurately detect the position of the target.

[0073] Sensor error: The camera or sensor generates errors due to malfunctions or external environmental influences (such as light changes, weather, reflections, etc.).

[0074] Rapid movement change: The target makes a rapid turn or a drastic change in speed, which may lead to a deviation in position prediction.

[0075] Multi-target interference: Multiple targets gather in the same viewing angle, resulting in the detection system being unable to accurately distinguish different targets.

[0076] Method for detecting abnormal data: The purpose of detecting abnormal data is to identify data points that do not conform to the normal trajectory by analyzing the motion pattern or position change of the target.

[0077] The following abnormal detection methods can be adopted: Trajectory-based detection: By analyzing the difference between the historical trajectory of the target and the current predicted trajectory, if the difference between the current data point and the predicted trajectory exceeds a preset threshold, then this data can be considered an outlier.

[0078] Detection based on speed and acceleration: If the speed or acceleration value of the target suddenly changes drastically in one frame, exceeding a reasonable range, it can also be considered abnormal data.

[0079] Motion consistency detection: If the motion direction or trajectory change of the target between adjacent frames does not meet the expectation (for example, the turning angle change of the target is too large), then it can be regarded as abnormal data.

[0080] Marking of abnormal data: Once abnormal data is detected, these data are distinguished from the normal trajectory through specific identifiers (such as setting a flag bit or segmenting the data) for subsequent correction.

[0081] The purpose of correcting abnormal data is that the existence of abnormal data may affect the accuracy and stability of the target trajectory, especially in applications with high real-time requirements for target tracking. Therefore, it is necessary to use effective algorithms to correct abnormal data to ensure the continuity and consistency of the target trajectory. The Kalman filter method can be adopted. The Kalman filter is an optimal linear estimation method that can smooth and correct abnormal data based on the historical state of the target and the current observation data, combining prediction and observation information. If abnormal data is detected in the target trajectory tracking, the Kalman filter can correct the abnormal value by weighted averaging the prediction and observation values of the abnormal data, making the target trajectory smoother. For example, the Kalman filter predicts the next position of the target according to the motion model of the target and adjusts the state estimation of the target by comparing the difference between the actual observation value and the predicted value. It can make optimal corrections based on the previous target state and the current observation value, effectively handle occasionally occurring abnormal data, and continuously adjust the target trajectory recursively.

[0082] The robust estimation algorithm can also be used. The robust estimation algorithm is specifically used to handle situations with high noise or low data quality. When abnormal data appears, the robust estimation algorithm can correct it by reducing the impact of abnormal values on the overall data estimation. Similar to the Kalman filter, the robust estimation algorithm reduces the impact of abnormal values by considering the deviation degree of abnormal data and using a weighting method. Common robust estimation algorithms include M-estimators, RANSAC, etc., which can perform trajectory correction in the presence of uncertainty.

[0083] Compared with the Kalman filter, the robust estimation method is more suitable for handling situations with large abnormalities in the data. It can effectively exclude the interference of incorrect data and ensure the stability of trajectory correction.

[0084] After detecting abnormal data, the Kalman filter or the robust estimation algorithm will smooth and correct these abnormal data, update the position, speed, and direction of the target, and ensure the continuity and consistency of the target trajectory.

[0085] This application enhances the robustness of the system by detecting and correcting abnormal data in the target tracking process. The introduction of the Kalman filter or the robust estimation algorithm enables the system to adaptively correct errors in a complex environment, avoiding abnormal behaviors in target tracking, such as occlusion and rapid changes, thereby improving the stability and accuracy of the system.

[0086] In some optional implementation manners of this embodiment, the steps of performing target detection on the image data, identifying vehicle and pedestrian targets in the image data, and determining the positions and features of the targets specifically include: Detect and classify the targets to identify whether they are vehicles or pedestrians and distinguish different types of targets; Determine the location of each identified target; Extract the features of each target using a convolutional neural network; Output the category, location, and features of the detected targets.

[0087] In this embodiment, target detection and classification are key tasks in computer vision for detecting and identifying target objects from images. In the multi-target area dynamic vehicle and pedestrian recognition method, first, the targets in the image data need to be detected, and then classified to identify whether the target is a vehicle or a pedestrian and possible other target categories. The location of the target is determined by target detection algorithms (such as YOLO, Faster R-CNN, SSD, etc.) in the image. These algorithms can find the target in the image and return the bounding box of each target, that is, the rectangular area where the target is located. After the target is detected, it needs to be classified. The convolutional neural network (CNN) is used to extract the image features of the target, and a classifier (such as the Softmax classifier) is used to determine whether the target belongs to the vehicle, pedestrian, or other categories. This is a very important step in multi-target recognition because vehicles and pedestrians have different motion characteristics, sizes, shapes, etc. First, distinguish the target types. Once the target is detected, the classification process will ensure that each target is correctly classified into the corresponding type (such as vehicle, pedestrian, etc.). For example, vehicles usually have larger sizes and more regular shapes, while pedestrians are usually smaller, and there is a certain diversity in body parts and postures. Based on these differences, the classification model can distinguish these different types of targets.

[0088] Location determination is to accurately calculate the specific location of the target in the image after target detection. The target location is usually represented by the coordinates of the bounding box, usually represented by the upper-left corner coordinates and the lower-right corner coordinates of the rectangular box, or by the center point coordinates of the target. For each detected target, the target detection algorithm returns a rectangular box, and the coordinates of this box describe the spatial location of the target. In addition to the bounding box, the location of the target can usually also be represented by the center point coordinates of the target, which is of great significance in tracking and subsequent trajectory optimization. For each target, the coordinates of the center point ( , ) can be calculated from the upper, lower, left, and right boundary values of the bounding box, which can more accurately represent the location of the target in the image, especially for the subsequent tracking and behavior analysis of the target.

[0089] Convolutional Neural Network (CNN) is a deep learning model widely used in computer vision tasks, especially in object detection, image classification and other tasks. During the object detection process, CNN extracts features from images through multiple convolutional operations, generating feature maps with deep semantic information, thereby effectively capturing the detailed features in the images. For each detected object, CNN will process the object region (the part within the bounding box), extracting the deep features that describe the object. These features can include the texture, color, shape, edge information, etc. of the object. Through the extracted features, the object can be further classified, behavior analyzed, and even the attributes of the object can be recognized (such as license plate recognition, face recognition, etc.). These features can also play a role in object tracking, helping the system maintain consistency between multiple frames. CNN usually contains multiple convolutional layers and pooling layers. Through the processing of these layers, the spatial dimension of the image is gradually reduced while extracting higher and higher order feature information. Finally, the features will be converted into the feature vector of the object through the fully connected layer.

[0090] After completing object detection and feature extraction, the system will output the category, location, and features of each detected object. These three pieces of information are the basic outputs of the multi-object recognition system: Category: The type of object obtained through the classification model (such as vehicle, pedestrian, cyclist, etc.).

[0091] Location: The location of the object usually includes the bounding box coordinates of the object in the image, or the coordinates of the center point of the object.

[0092] Features: The feature vector of the object extracted from the CNN, used to describe information such as the appearance and form of the object. These features can be further used for subsequent object tracking, behavior analysis and other tasks.

[0093] Output format: The detection results will be stored as a list or matrix containing object information, and the information of each object includes the category, location coordinates, and feature vector.

[0094] This application extracts the features of the object through the convolutional neural network and classifies them, achieving efficient and accurate object recognition. This step can accurately distinguish different types of objects such as vehicles and pedestrians, laying a solid foundation for subsequent tracking and recognition. The application of the convolutional neural network improves the robustness of object detection and meets the object recognition requirements under different backgrounds and lighting conditions in complex environments.

[0095] In some optional implementation manners of this embodiment, the steps of initially positioning each object and predicting the motion trajectory of the object through the Kalman filter algorithm or the particle filter algorithm specifically include: Initialize the state variables for each detected object, and the state variables include the location and speed information of the object; Perform trajectory prediction based on the state variables; Update the state variables of the target. If there is a large deviation between the actual position and the predicted position of the target in the current frame, trajectory correction is performed.

[0096] In this embodiment, initializing the state variables is the first step in target tracking. The motion state of each target can be described by a set of state variables, usually including position (such as the coordinates of the target in the image) and velocity (such as the velocity of the target in the axis and axis directions).

[0097] Target position: The position is usually the coordinates of the target in the image or the center point coordinates of the bounding box. Assume the position of the target in the image is ([[]] , ).

[0098] Target velocity: The target velocity refers to the moving speed of the target in the image coordinate system. Usually, this includes the horizontal velocity ( ) and the vertical velocity ( ).

[0099] State vector: In Kalman filtering or particle filtering, the state of the target is usually represented as a vector containing information about the position and velocity. For example, the state vector might be represented as: , Here, and are the positions of the target, and are the velocities of the target.

[0100] During target detection, the system will first assign an initial state variable to each target (for example, initialize it through the position and velocity of the target detected in the current frame). If there is no velocity information, the initial velocity can be assumed to be zero, or initialized through the velocity of the previous frame.

[0101] Trajectory prediction is to estimate the possible position of the target in the next frame based on the current state variables (position and velocity). Both Kalman filtering and particle filtering algorithms can be used for target trajectory prediction. Kalman filtering is a recursive estimation algorithm that predicts the target's trajectory based on a linear motion model. Given the known position and velocity of the target, Kalman filtering can predict the expected position of the target in the next frame.

[0102] The prediction step is usually carried out according to the following equation: , where is the predicted state, is the state transition matrix, is the control matrix, is the control input (such as acceleration, etc.).

[0103] Particle filter represents the motion state of the target by generating multiple "particles", and updates and resamples these particles according to the target motion model. Particle filter is applicable to non - linear or non - Gaussian problems. Through multiple iterations, particle filter can provide relatively accurate trajectory prediction.

[0104] In particle filter, the prediction of the target trajectory is usually based on the known particle distribution (representing the possible states of the target), and the prediction is carried out by updating the weights and positions of the particles.

[0105] Predicting the target position: Through the prediction formula or particle filter algorithm, the system calculates the expected position of the target at the next moment according to the state variables (position and velocity) of the target.

[0106] State update: The predicted position of the target is based on the state of the previous frame. However, due to environmental uncertainty and measurement errors, the actually observed target position may deviate from the predicted position. Therefore, it is necessary to update the state of the target.

[0107] State update of Kalman filter: Kalman filter updates the target state by combining the predicted value and the actual observation value. When there is an error between the actually observed target position and the predicted position Kalman filter will calculate a Kalman gain and correct the target state by weighted average.

[0108] The update formula is usually: , where, is the Kalman gain, is the actual observation value, is the predicted position.

[0109] Particle filter corrects the target state by adjusting the weights of the particles and resampling. If the difference between the actual position and the predicted position of the target in the current frame is large, particle filter will increase the number of high - weight particles, thereby improving the estimation accuracy of the target. When the deviation between the predicted position and the actual position is too large, trajectory correction is required. This can be achieved in the following two ways: Error correction: Directly adjust the state variables (position, velocity, etc.) of the target to better conform to the current observation.

[0110] Data association: Re - associate the targets recognized by multiple cameras to ensure the consistency of the target trajectory.

[0111] This application predicts the target motion trajectory through the Kalman filter or particle filter algorithm, providing accurate target positioning and tracking capabilities. The combination of initial positioning and trajectory prediction enables the system to accurately predict the target's position in the case of discontinuous target motion trajectories, reducing the possibility of tracking breaks, thereby ensuring continuity and stability.

[0112] In some alternative implementation manners of this embodiment, the step of spatio-temporal association of the targets recognized by multiple cameras in the above-mentioned joint data association method based on the position, features, and motion trajectory of the target specifically includes: Sort the targets according to the time stamps; Match the targets within the same time period in different cameras according to the time stamps for time alignment; Convert the targets of different cameras to a unified coordinate system; Perform spatial association according to the coordinate system of the targets.

[0113] In this embodiment, in a multi-camera system, the image data collected by the cameras usually carries time stamps to record the capture time of each image. Since multiple cameras may capture the target simultaneously or sequentially, it is necessary to sort the targets detected by different cameras according to the time stamps.

[0114] Purpose of sorting: The purpose of sorting is to ensure the time order of the targets and ensure that the subsequent time alignment and spatial association processes can be based on the correct time information.

[0115] Implementation manner: Sort the detection results from different cameras according to the time stamps and assign a time order to each target. In this way, the time sequence consistency of each target can be ensured, enabling the subsequent steps to be processed based on the correct time order.

[0116] Time alignment refers to matching the targets detected in the same time period in different cameras through the time stamps, thereby ensuring the time consistency of the target data of different cameras.

[0117] Matching targets: Among the detection results of different cameras, the targets within the same time period should be the same object (for example, a vehicle or a pedestrian). Therefore, it is necessary to match the targets detected by different cameras at the same time point according to the time stamps.

[0118] Purpose of time alignment: The purpose of time alignment is to unify the detection data of the targets on the time axis, enabling the accurate comparison of the positions and motion trajectories of the targets under different cameras in the subsequent spatial association step.

[0119] Implementation manner: Assume that in camera A and camera B, the time stamps at a certain moment are respectively and If these two timestamps are very close or the same, the targets detected by cameras A and B at this time point can be regarded as detection data at the same moment. At this time, subsequent analysis can be carried out based on this temporally aligned data.

[0120] Coordinate transformation is a key step in multi-camera data fusion. Since the perspectives, positions, and orientations of each camera may be different, the coordinate systems of the targets they capture may also be different. Therefore, it is necessary to transform the coordinate systems of different cameras into a unified coordinate system for accurate spatial association.

[0121] Coordinate system differences: Each camera has its own coordinate system, usually based on the camera's image coordinate system (pixel coordinates). These coordinate systems are not necessarily consistent with the global coordinate system (such as the world coordinate system or the unified coordinate system). The differences between the coordinate systems of different cameras may include: Image resolution and scale; The rotation angle and position of the camera; The field of view angle and projection method of the camera; Coordinate transformation method: The method of transforming the coordinate system of the camera into a unified coordinate system usually includes camera calibration and geometric transformation. For example, projection matrices and transformation matrices can be used to transform the target positions in each camera into positions in the global coordinate system.

[0122] Implementation method: By calibrating the internal parameters (focal length, distortion coefficient, etc.) and external parameters (position, orientation, etc.) of each camera, a projection matrix can be constructed. Then, by multiplying the pixel coordinates of the target by the corresponding projection matrix, the coordinate system of the camera can be transformed into the global coordinate system.

[0123] Spatial association refers to matching and fusing the targets recognized by different cameras spatially under a unified coordinate system. Since the target may appear in the fields of view of multiple cameras, the purpose of spatial association is to ensure that the target can be accurately tracked and recognized under the unified coordinate system.

[0124] The goal of spatial association: Compare the spatial position of the target with the target positions in other cameras to determine whether they are the same object. Since the spatial position of the target may have an offset between different cameras (due to perspective differences), it is necessary to perform matching through algorithms.

[0125] Spatial association methods: Distance threshold method: A distance threshold can be set, and only when the position difference between two targets is less than a certain threshold is it considered that they are the same target. This method is simple and easy to implement, but may be affected by noise.

[0126] Feature matching method: In addition to position, other features of the target (such as color, shape, motion trajectory, etc.) can also be used as the basis for spatial association. Similarity metrics (such as Euclidean distance, cosine similarity, etc.) can be used to measure the similarity of the targets.

[0127] Implementation method: In a unified coordinate system, spatial association is performed by calculating the differences in the coordinates (or other features) of each target to ensure the consistency of the target among multiple cameras. If the same target is recognized by multiple cameras, the system will track it as a unified target.

[0128] Through the spatio-temporal association method, this application ensures the effective fusion of target data under multiple cameras in a unified coordinate system. The sorting and matching of timestamps enable the targets detected by multiple cameras at the same moment to be precisely aligned in time sequence, solving the problem of time asynchrony. The unified transformation of the coordinate system and spatial association ensure that the targets from different cameras can be accurately merged into a unified target, avoiding mis-matching caused by differences in camera perspectives.

[0129] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0130] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0131] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, Read-Only Memory (ROM), or a Random Access Memory (RAM), etc.

[0132] It should be understood that although the steps in the flowchart of the accompanying drawings are shown sequentially in the direction of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless otherwise clearly stated in this document, there is no strict order restriction for the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.

[0133] Further reference Figure 3 , as an implementation of the method shown above Figure 2 , an embodiment of a multi-target area dynamic vehicle and pedestrian recognition system is provided in this application. An embodiment of a multi-target area dynamic vehicle and pedestrian recognition system corresponds to the method embodiment shown in Figure 2 , and this system can be specifically applied to various electronic devices.

[0134] As Figure 3 shown, a multi-target area dynamic vehicle and pedestrian recognition system 300 described in this embodiment includes: an acquisition module 301, a processing module 302, a detection module 303, a prediction module 304, and an association module 305. Among them: The acquisition module 301 is used to acquire image data of a target area collected by multiple cameras; The processing module 302 is used to preprocess the image data and extract the features of the target area; The detection module 303 is used to perform target detection on the image data, identify vehicle and pedestrian targets in the image data, and determine the positions and features of the targets; The prediction module 304 is used to perform preliminary positioning on each target and predict the motion trajectory of the target through the Kalman filter algorithm or the particle filter algorithm; The association module 305 is used to perform spatio-temporal association on the targets recognized by multiple cameras based on the joint data association method of the positions, features, and motion trajectories of the targets.

[0135] The multi-target area dynamic vehicle and pedestrian recognition system provided in this application ensures that the system can efficiently identify and track targets in a wide area by defining the acquisition, preprocessing, target detection and tracking of multi-camera image data, and the spatio-temporal association method; through the cooperation of multiple cameras, the system can overcome the blind area problem of a single camera's perspective and provide more comprehensive monitoring capabilities, thereby improving the coverage rate and accuracy of target recognition.

[0136] In some alternative implementation manners of this embodiment, the association module 305 is further configured to: Update the target trajectory in real time through a data fusion model and a trajectory optimization algorithm.

[0137] A multi-target area dynamic vehicle and pedestrian recognition system provided by this application can reduce error accumulation in a complex dynamic environment, accurately track the target position through the fusion of data from multiple cameras, and the ability to update in real time ensures that the system can still efficiently track even when the target motion changes rapidly, avoiding trajectory drift or loss.

[0138] In some alternative implementation manners of this embodiment, the association module 305 is further configured to: Assign different weights to the targets provided by each camera; According to the weighted fusion algorithm, fuse the targets of each camera according to their weights, calculate the weighted average position of the targets, and generate the final target position; Based on the target position, use the least squares method to fit the target trajectory; Use the target trajectory to update the position, speed, and direction of the target in real time.

[0139] A multi-target area dynamic vehicle and pedestrian recognition system provided by this application optimizes the target position and trajectory through the weighted fusion algorithm and the least squares method, improving the accuracy of multi-camera data fusion; by assigning different weights to different cameras, the system can integrate the advantages of each camera to ensure accurate positioning and tracking of the target under different conditions (such as differences in camera viewing angles, image quality, etc.), and the least squares method effectively reduces the influence of noise and errors by fitting the trajectory.

[0140] In some alternative implementation manners of this embodiment, the association module 305 is further configured to: Detect abnormal data during the target tracking process; Use the Kalman filter or robust estimation algorithm to correct the abnormal data.

[0141] A multi-target area dynamic vehicle and pedestrian recognition system provided by this application enhances the robustness of the system by detecting and correcting abnormal data during the target tracking process. The introduction of the Kalman filter or robust estimation algorithm enables the system to adaptively correct errors in a complex environment, avoiding abnormal behaviors in target tracking, such as occlusion, rapid changes, etc., thereby improving the stability and accuracy of the system.

[0142] In some alternative implementation manners of this embodiment, the detection module 303 is further configured to: Detect and classify the targets to identify whether they are vehicles or pedestrians and distinguish different types of targets; Determine the positions of each identified target; Use a convolutional neural network to extract the features of each target; Output the categories, positions, and features of the detected targets.

[0143] A multi-target area dynamic vehicle and pedestrian recognition system provided by the present application extracts the features of targets through a convolutional neural network and classifies them, achieving efficient and accurate target recognition. This step can accurately distinguish different types of targets such as vehicles and pedestrians, laying a solid foundation for subsequent tracking and recognition. The application of the convolutional neural network improves the robustness of target detection and meets the target recognition requirements under different backgrounds and lighting conditions in complex environments.

[0144] In some optional implementation manners of this embodiment, the prediction module 304 is further configured to: Initialize state variables for each detected target, where the state variables include the position and speed information of the target; Perform trajectory prediction according to the state variables; Update the state variables of the target. If there is a large deviation between the actual position and the predicted position of the target in the current frame, trajectory correction is performed.

[0145] A multi-target area dynamic vehicle and pedestrian recognition system provided by the present application predicts the target motion trajectory through the Kalman filter or particle filter algorithm, providing accurate target positioning and tracking capabilities. The combination of initial positioning and trajectory prediction enables the system to accurately predict the position of the target in the case of discontinuous target motion trajectories, reducing the possibility of tracking breaks, thereby ensuring continuity and stability.

[0146] In some optional implementation manners of this embodiment, the association module 305 is further configured to: Sort the targets according to the timestamps; Match the targets in different cameras within the same time period according to the timestamps for temporal alignment; Convert the targets of different cameras to a unified coordinate system; Perform spatial association according to the coordinate systems of the targets.

[0147] A multi-target area dynamic vehicle and pedestrian recognition system provided by the present application ensures the effective fusion of target data under multiple cameras in a unified coordinate system through a spatio-temporal correlation method. The sorting and matching of timestamps enable the targets detected by multiple cameras at the same moment to be precisely aligned in time sequence, solving the problem of time asynchronization. The unified conversion of the coordinate system and spatial correlation ensure that targets from different cameras can be precisely merged into a unified target, avoiding false matching caused by differences in camera perspectives.

[0148] To solve the above technical problems, an embodiment of the present application also provides a computer device. For details, please refer to Figure 4 , Figure 4 which is the basic structural block diagram of the computer device in this embodiment.

[0149] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that communicate with each other through a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Among them, those skilled in the art of the present technology can understand that a computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0150] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server and other computing devices. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad or a voice control device, etc.

[0151] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disc, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc. equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and the external storage device of the computer device 4. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions of a multi-target area dynamic vehicle and pedestrian recognition method. In addition, the memory 41 may also be used to temporarily store various types of data that have been output or will be output.

[0152] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run the computer-readable instructions stored in the memory 41 or process data, such as running the computer-readable instructions of the multi-target area dynamic vehicle and pedestrian recognition method.

[0153] The network interface 43 may include a wireless network interface or a wired network interface, and the network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0154] The computer device provided in this application ensures that the system can efficiently identify and track targets in a wide area by defining the acquisition, preprocessing, target detection and tracking, and spatio-temporal association methods of multi-camera image data; through the cooperation of multiple cameras, the system can overcome the blind area problem of a single camera's perspective and provide more comprehensive monitoring capabilities, thereby improving the coverage and accuracy of target recognition.

[0155] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions, which can be executed by at least one processor to enable the at least one processor to execute the steps of a multi-target area dynamic vehicle and pedestrian recognition method as described above.

[0156] The computer-readable storage medium provided by the present application ensures that the system can efficiently identify and track targets in a wide area by defining the acquisition, preprocessing, target detection and tracking of multi-camera image data, as well as the spatio-temporal association method; through the cooperation of multiple cameras, the system can overcome the blind area problem of a single camera's perspective and provide more comprehensive monitoring capabilities, thereby improving the coverage and accuracy of target recognition.

[0157] Through the description of the above implementation manners, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0158] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show the preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific implementation manners, or perform equivalent replacements for some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is similarly within the scope of the patent protection of the present application.

Claims

1. A multi-target area dynamic vehicle and pedestrian recognition method, characterized in that: The steps include: Acquire image data of a target area collected by multiple cameras; Preprocessing the image data to extract features of the target area; Performing target detection on the image data, identifying vehicle and pedestrian targets in the image data, and determining the positions and features of the targets; Perform preliminary positioning of each target and predict the target's trajectory through the Kalman filter algorithm or particle filter algorithm; Based on the joint data association method of the position, characteristics and motion trajectory of the target, the targets recognized by multiple cameras are temporally and spatially associated.

2. The multi-target area dynamic vehicle and pedestrian recognition method according to claim 1 is characterized in that: The joint data association method based on the position, feature and motion trajectory of the target, after the step of performing spatiotemporal association of the targets identified by multiple cameras, further includes: The target trajectory is updated in real time through data fusion model and trajectory optimization algorithm.

3. The multi-target area dynamic vehicle and pedestrian recognition method according to claim 2 is characterized in that: The step of updating the target trajectory in real time through the data fusion model and the trajectory optimization algorithm includes: Assign different weights to the targets provided by each camera; According to the weighted fusion algorithm, the targets of each camera are fused according to their weights, the weighted average position of the target is calculated, and the final target position is generated; Based on the target position, fitting the target trajectory using the least squares method; Using the target trajectory, the position, speed, and direction of the target are updated in real time.

4. The multi-target area dynamic vehicle and pedestrian recognition method according to claim 3 is characterized in that: After the step of using the target trajectory to update the position, speed and direction of the target in real time, the method further includes: Detect abnormal data during target tracking; Use Kalman filtering or robust estimation algorithm to correct abnormal data.

5. The multi-target area dynamic vehicle and pedestrian recognition method according to claim 1, characterized in that: The step of performing target detection on the image data, identifying vehicle and pedestrian targets in the image data, and determining the positions and features of the targets specifically includes: Detect and classify targets, identify whether they are vehicles or pedestrians, and distinguish between different types of targets; Determine the location of each identified target; Use convolutional neural network to extract the features of each target; Output the category, location, and features of the detected object.

6. The multi-target area dynamic vehicle and pedestrian recognition method according to claim 1 is characterized in that: The step of preliminarily locating each target and predicting the target's motion trajectory by using a Kalman filter algorithm or a particle filter algorithm specifically includes: Initializing state variables for each detected target, the state variables including position and speed information of the target; Perform trajectory prediction based on the state variables; Update the state variables of the target. If there is a large deviation between the actual position of the target in the current frame and the predicted position, perform trajectory correction.

7. The multi-target area dynamic vehicle and pedestrian recognition method according to claim 1 is characterized in that: The joint data association method based on the position, features and motion trajectory of the target, the step of performing spatiotemporal association of the targets identified by multiple cameras, specifically includes: Sort the targets by timestamp; Match the targets in the same time period in different cameras according to the timestamps to perform time alignment; Convert targets from different cameras to a unified coordinate system; Spatial association is performed according to the target's coordinate system.

8. A multi-target area dynamic vehicle and pedestrian recognition system, characterized in that: include: An acquisition module, used to acquire image data of a target area collected by multiple cameras; A processing module, used for preprocessing the image data and extracting features of the target area; A detection module, used to perform target detection on the image data, identify vehicle and pedestrian targets in the image data, and determine the position and characteristics of the targets; The prediction module is used to perform preliminary positioning of each target and predict the target's trajectory through the Kalman filter algorithm or the particle filter algorithm; The association module is used to perform spatiotemporal association on the targets identified by multiple cameras based on a joint data association method of the position, features and motion trajectory of the targets.

9. A computer device, characterized in that: It comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the multi-target area dynamic vehicle and pedestrian recognition method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the multi-target area dynamic vehicle and pedestrian recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-source comprehensive processing bird protection device and method based on wind power plant

    CN114027288A

  • Expressway multi-target tracking method based on Leiyu fusion

    CN116863382A

  • Cross-space-time correlation method and device for night target trajectory

    CN117495913A

  • Multi-sound-source movement track prediction method for complex indoor sound field environment

    CN117590328A

  • Dynamic target tracking method and system for vehicle-mounted mining intrinsic safety type video monitoring

    CN119027860A

Cited By

  • Vehicle incidence relation reasoning method and system based on AI video analysis

    CN120783296A