Traffic flow track processing method and device, electronic equipment and storage medium

Through multiple perspective cameras, the local trajectories of traffic participants are obtained and matched, and the trajectories of traffic participants that meet real scenes are extracted in the autonomous driving technology, solving the problem of unreal trajectory in the existing technology, and promoting the high-quality construction of the scene library and the development of autonomous driving technology.

CN120070489APending Publication Date: 2025-05-30QINGKE LINGJING (ANHUI) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510126236.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, the trajectory of traffic participants is usually predicted by video data and does not conform to real scenarios, resulting in the inability to build a real and high-quality scenario library, hindering the development of autonomous driving technology.

Method used

A method for processing traffic flow trajectory is proposed. The local trajectory of traffic participants is obtained through multiple viewing cameras, and the trajectory matching process of the same traffic participant under multiple viewing cameras is performed. Finally, the local trajectory set of each traffic participant under multiple viewing cameras is fused to obtain the global trajectory.

Benefits of technology

The trajectory of traffic participants that meet the real scenes is extracted based on multiple perspective cameras, so that a real and high-quality scene library can be constructed in the future, and the development of autonomous driving technology is promoted.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070489A_ABST
    Figure CN120070489A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic flow track processing method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining a local track of each traffic participant under each single-view-angle camera based on a shot monitoring video for each single-view-angle camera in a plurality of view-angle cameras; performing local track matching processing of the same traffic participant under multiple view angle cameras on the obtained local tracks of all the traffic participants under all the single view angle cameras to obtain a local track set of each traffic participant under the multiple view angle cameras; and performing trajectory fusion processing on the local trajectory set of each traffic participant under the plurality of view angle cameras to obtain a global trajectory of each traffic participant under the plurality of view angle cameras. According to the method, extraction of traffic participant tracks conforming to a real scene can be realized based on multiple view angle cameras, so that a real and high-quality scene library can be constructed subsequently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of vehicle trajectory processing, and in particular to a method, apparatus, electronic device, and storage medium for processing traffic flow trajectories. Background Art

[0002] The verification of autonomous driving technology relies on a large number of simulation tests, and the simulation tests rely on the construction of a scenario library. The most important thing in the process of constructing the scenario library is the acquisition of traffic participant trajectories.

[0003] In the related art, traffic participant trajectories are often predicted from video data rather than extracted from captured videos. Therefore, the obtained trajectories do not conform to the real scenario, resulting in the inability to construct a real and high-quality scenario library subsequently, which hinders the development of autonomous driving technology. Summary of the Invention

[0004] Based on the above technical status quo, the present application proposes a method, apparatus, electronic device, and storage medium for processing traffic flow trajectories, which can extract the trajectories of traffic participants that conform to the real scenario according to cameras with different perspectives.

[0005] The first aspect of the present application provides a method for processing traffic flow trajectories, including:

[0006] For each single perspective camera among multiple perspective cameras, respectively obtain the local trajectory of each traffic participant under the single perspective camera based on the captured surveillance video;

[0007] Perform local trajectory matching processing of the same traffic participant under multiple perspective cameras on the obtained local trajectories of all traffic participants under all single perspective cameras to obtain the local trajectory set of each traffic participant under multiple perspective cameras;

[0008] Perform trajectory fusion processing on the local trajectory set of each traffic participant under multiple perspective cameras respectively to obtain the global trajectory of each traffic participant under multiple perspective cameras.

[0009] The second aspect of the present application provides a device for processing traffic flow trajectories, including:

[0010] A trajectory acquisition unit, configured to obtain the local trajectory of each traffic participant under each single perspective camera among multiple perspective cameras respectively based on the captured surveillance video;

[0011] A trajectory matching unit, configured to perform matching processing of the same traffic participant under multiple perspective cameras on the obtained local trajectories of all traffic participants under all single perspective cameras to obtain the local trajectory set of each traffic participant under multiple perspective cameras;

[0012] A trajectory fusion unit, configured to perform fusion processing on the local trajectory sets of each traffic participant under multiple perspective cameras respectively, to obtain the global trajectory of each traffic participant under multiple perspective cameras.

[0013] A third aspect of the present application provides an electronic device, including:

[0014] A memory and a processor;

[0015] The memory is connected to the processor and is used for storing programs;

[0016] The processor is configured to implement the above traffic flow trajectory processing method by running the programs in the memory.

[0017] A fourth aspect of the present application provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, the above traffic flow trajectory processing method is implemented.

[0018] For the traffic flow trajectory processing method provided by the present application, since the local trajectories of each traffic participant in each single perspective camera among multiple perspective cameras are respectively obtained, and then the local trajectories of the traffic participants under the obtained multiple perspective cameras are subjected to matching processing for the same traffic participant, to obtain the local trajectory sets of each traffic participant under multiple perspective cameras, and finally, fusion processing is respectively performed on each local trajectory set to obtain the global trajectory of each traffic participant under multiple perspective cameras. Therefore, the extraction of traffic participant trajectories conforming to the real scenario is realized based on multiple perspective cameras, enabling the subsequent construction of a real and high-quality scenario library and promoting the development of autonomous driving technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0020] Figure 1 It is a flowchart of a traffic flow trajectory processing method according to an embodiment of the present application;

[0021] Figure 2 It is a flowchart of a local trajectory acquisition according to an embodiment of the present application;

[0022] Figure 3 It is a flowchart of another local trajectory acquisition according to an embodiment of the present application;

[0023] Figure 4Schematic diagram of a multi - target tracking process according to an embodiment of the present application;

[0024] Figure 5 Schematic diagram of a process for obtaining a local trajectory set according to an embodiment of the present application;

[0025] Figure 6 Schematic diagram of a multi - target matching process according to an embodiment of the present application;

[0026] Figure 7 Schematic diagram of the structure of a traffic flow trajectory processing device according to an embodiment of the present application;

[0027] Figure 8 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0028] The technical solution of the embodiment of the present application is applicable to an application scenario of obtaining global trajectories under multiple perspective cameras. By adopting the technical solution of the embodiment of the present application, it is possible to extract the global trajectories of traffic participants under multiple perspective cameras, so as to enable subsequent construction of a real and high - quality scenario library.

[0029] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0030] Exemplary method

[0031] A traffic flow trajectory processing method according to an embodiment of the present application, as Figure 1 shown, the method includes:

[0032] Step 100: For each single perspective camera among multiple perspective cameras, respectively obtain the local trajectory of each traffic participant under the single perspective camera based on the captured surveillance video.

[0033] The multiple perspective cameras may refer to perspective cameras with multiple different perspectives.

[0034] The shooting of surveillance videos goes through a process. During this process, some traffic participants appear in the field of view of the perspective camera from the moment the surveillance video starts shooting and remain in the field of view of this perspective camera throughout the subsequent shooting process. Therefore, this part of the traffic participants exists in each frame of the surveillance video. Some traffic participants appear in the field of view of this perspective camera when the surveillance video starts shooting, but leave the field of view of this perspective camera during the subsequent shooting process. Therefore, this part of the traffic participants appear in the first part of the images of the surveillance video. Some traffic participants do not appear in the field of view of this perspective camera at the beginning of the surveillance video shooting and only enter the field of view of this perspective camera after a period of time and until the end of the surveillance video shooting. Therefore, this part of the traffic participants appear in the second part of the images of the surveillance video. Some traffic participants do not appear in the field of view of this perspective camera at the beginning of the surveillance video shooting, enter the field of view of this perspective camera after a period of time, and then leave the field of view of this perspective camera after another period of time. Therefore, this part of the traffic participants appear in the middle part of the images of the surveillance video. Traffic participants under a single perspective camera refer to all traffic participants appearing in the captured surveillance video, that is, as long as a traffic participant appears in one frame of the surveillance video, it is a traffic participant under a single perspective camera.

[0035] The process of obtaining the local trajectory of each traffic participant under the single perspective camera based on the captured surveillance video can be completed through the detection and tracking of the target.

[0036] Among them, the detection of the target can be carried out using target detection algorithms, such as YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), Faster R-CNN, etc. Among them, YOLO is a real-time object detection algorithm widely used in target detection tasks. Different from other traditional two-stage detectors (such as Faster R-CNN), YOLO adopts a single-stage end-to-end training method, which can achieve a very fast inference speed while ensuring a relatively high detection accuracy. SSD is also a single-stage target detection algorithm. SSD completes the classification and localization of the target simultaneously in one forward propagation, so it has a faster inference speed and is suitable for real-time applications. Faster R-CNN is a two-stage deep learning model widely used in target detection tasks. Different from single-stage detectors (such as YOLO, SSD), Faster R-CNN achieves more accurate target localization and classification by introducing a Region Proposal Network (RPN).

[0037] Tracking of targets can be carried out according to tracking algorithms. For example, SORT (Simple Online and Realtime Tracking), DeepSORT, FairMOT (Fair Multi Object Tracking), etc. Among them, SORT is a simple and efficient multi-object tracking algorithm mainly used for tracking multiple targets in real-time videos and is suitable for tracking tasks of objects such as pedestrians and vehicles. Compared with traditional multi-object tracking methods, the implementation of SORT is very concise and can achieve real-time performance while maintaining relatively high tracking accuracy. In each frame, SORT uses the Hungarian algorithm (also known as the Kuhn-Munkres algorithm) to match the detection results of the current frame with the trackers of the previous frame. The Hungarian algorithm matches by minimizing the distance (such as the IoU distance) between the detection bounding boxes and the trackers. SORT uses the Intersection over Union (IoU) as the matching metric. IoU measures the overlap degree of two bounding boxes, and the larger the value, the more overlap. SORT usually sets an IoU threshold, and only when the IoU is greater than this threshold is the detection bounding box considered to match the tracker. If the detection bounding box successfully matches a certain tracker, SORT uses the update step of the Kalman filter to adjust the state of the tracker. Specifically, if the detection bounding box does not match any tracker, SORT will create a new tracker for this detection bounding box and initialize the corresponding Kalman filter. If the detection bounding box does not match any tracker, SORT will continue to use the prediction step of the Kalman filter to estimate the future position of the target. If the detection bounding box is not matched for several consecutive frames, SORT will mark this tracker as "lost" and delete this tracker after a certain period of time. DeepSORT is an improved version of the SORT algorithm. Based on maintaining the simplicity and efficiency of SORT, DeepSORT introduces deep learning technology to enhance the ability to model the appearance features of targets, thus significantly improving the accuracy and robustness of multi-object tracking. FairMOT is a multi-object tracking algorithm that combines detection and ReID (Re-Identification) tasks. Different from traditional two-stage or multi-stage tracking methods, FairMOT simultaneously performs object detection and feature extraction through a single network, significantly improving the accuracy and speed of tracking.

[0038] Step 110: Perform local trajectory matching processing on the local trajectories of all traffic participants obtained under all single-view cameras for the same traffic participant under multiple-view cameras to obtain the local trajectory set of each traffic participant under multiple-view cameras.

[0039] The local trajectory of each traffic participant under a single perspective camera can be a series of position information with timestamps. Therefore, when performing local trajectory matching processing of the same traffic participant under multiple perspective cameras, preliminary screening can be carried out according to time and space information to exclude obviously impossible matching trajectories. For example, if the time interval between two local trajectories is too large or the spatial distance is too far, the matching possibility between them can be excluded. Then, based on the characteristics between the traffic participants to which the local trajectories belong, local trajectory matching processing is performed for different multiple perspective cameras to obtain the local trajectory sets of each traffic participant under multiple perspective cameras. Among them, when calculating the characteristics between traffic participants, methods such as Euclidean distance, cosine similarity, and Hamming distance can be used to measure the similarity between the feature vectors of different traffic participants.

[0040] The local trajectory sets of each traffic participant under multiple perspective cameras correspond to each traffic participant. For example, assume that the traffic participants include: traffic participant A, traffic participant B, and traffic participant C, and the multiple perspective cameras include: perspective camera 1, perspective camera 2, perspective camera 3, and perspective camera 4. Traffic participants A, B, and C each generate a local trajectory in the 4 perspective cameras respectively. Therefore, traffic participants A, B, and C each correspond to a local trajectory set in the 4 perspective cameras, and the trajectories in the corresponding local trajectory sets are the local trajectories under the 4 perspective cameras respectively.

[0041] Step 120: Perform trajectory fusion processing on the local trajectory sets of each traffic participant under multiple perspective cameras respectively to obtain the global trajectory of each traffic participant under multiple perspective cameras.

[0042] For the fusion of local trajectories in the local trajectory set, methods such as weighted average method and least squares method can be used. Among them, the weighted average method extracts trajectory points from multiple local trajectories in the local trajectory set according to the time dimension, performs weighted average on the trajectory points at the same time point, and obtains the weighted trajectory points corresponding to all time points, so as to obtain the fused global trajectory; the least squares method is to find an optimal global trajectory by combining the observed values of all trajectories, so that the difference (error) between all local trajectories and the global trajectory is minimized.

[0043] The traffic flow trajectory processing method provided by the embodiments of the present application obtains the local trajectories of each traffic participant in each single perspective camera among multiple perspective cameras, then performs matching processing on the local trajectories of traffic participants under multiple perspective cameras for the same traffic participant to obtain the local trajectory set of each traffic participant under multiple perspective cameras, and finally performs fusion processing on each local trajectory set respectively to obtain the global trajectory of each traffic participant under multiple perspective cameras. Therefore, the extraction of traffic participant trajectories that conform to the real scene is realized based on multiple perspective cameras, enabling the subsequent construction of a real and high-quality scene library and promoting the development of autonomous driving technology.

[0044] In an exemplary embodiment, as Figure 2 shown, obtaining the local trajectory of each traffic participant under the single perspective camera based on the captured surveillance video includes:

[0045] Step 200, obtain an image frame sequence corresponding to the captured surveillance video.

[0046] The acquisition of the image frame sequence can be implemented through a video processing library. The video processing library can include: OpenCV, FFmpeg, MoviePy, PyAV. Among them, OpenCV is powerful, supports multiple video formats, and provides rich image processing functions. It is one of the most commonly used choices. To obtain the image frame sequence using OpenCV, install OpenCV, adjust the frame extraction interval, image saving format and path according to needs, then open the video file, and OpenCV will read the video frame by frame and save the images using the adjusted extraction interval, image saving format and path.

[0047] Step 210, for each frame image in the obtained image frame sequence, respectively obtain the detection position, category and features of each traffic participant that appears in the image.

[0048] The acquisition of the detection positions and categories of each traffic participant in each frame of the obtained image sequence is usually carried out using an object detection model and an object tracking model. In the case of using an object detection model and an object tracking model to obtain detection positions and features, the acquisition process may include the following steps: Input each frame of the image into an object detection model (such as YOLO) for inference. Inside the object detection model, the convolutional neural network (CNN) is first used to divide the input image into multiple grids. Each grid will predict multiple bounding boxes, as well as the confidence and class probabilities of each bounding box. In the bounding box prediction stage, each bounding box is defined by four parameters: the coordinates (x, y) of the center point, the width w and height h of the bounding box. In addition, the object detection model will also output the confidence score of the bounding box, which indicates whether there is an object (i.e., a traffic participant) in the box and the degree of matching between the bounding box and the object. At the same time, the object detection model will output the probabilities belonging to different classes to help identify the class of the object in the bounding box. Next, by setting a confidence threshold, the bounding boxes with low confidence are filtered out. Specifically, the non-maximum suppression (NMS) method can be used to remove those bounding boxes with high overlap, ensuring that finally each object will be represented by only one bounding box (NMS will preferentially retain the bounding box with the highest confidence and remove other overlapping boxes). Finally, the object detection model outputs the bounding boxes, class labels, and confidence scores of each detected object. The output method is usually to draw these bounding boxes on the original image and label the class and confidence score of the object to visualize the results of multi-object detection. At this step, the object detection model has completed the detection of all objects in the image, obtaining the bounding boxes, classes, and confidence scores of each object. Next, this information is input into an object tracking model (such as DeepSORT) to obtain the detection positions, classes, and features of each object on the image.

[0049] The detection position of a traffic participant can refer to the position information of the bounding box corresponding to the traffic participant in the image.

[0050] According to different classification criteria, the categories of traffic participants can vary. Common classification methods include according to the type of vehicle, role, function, etc. of the participants. Traffic participants can be divided into: pedestrians, bicycles, motorcycles, and motor vehicles according to the type of vehicle.

[0051] The features of a traffic participant can refer to the appearance features of the traffic participant, such as the size, shape, and color represented by the traffic participant in the image.

[0052] Before obtaining the detection position, category, features, and movement direction of each traffic participant that appears in the image for each frame of the acquired image sequence, preprocessing of the image data of each acquired frame of the image can also be performed, and the preprocessing of the image data can be distortion correction processing.

[0053] Image distortion correction mainly involves the correction of radial distortion and tangential distortion, as follows:

[0054] 1. Radial distortion correction: Radial distortion correction mainly involves parameters such as k 1 , k 2 , k 3 etc. These parameters are used to correct the light distortion caused during the lens manufacturing process. Radial distortion can be divided into barrel distortion and pincushion distortion, where k 1 has a great effect and influence on the central region of the image with relatively small distortion, k 2 mainly acts on the edge region of the image with relatively large distortion, and k 3 is applicable to wide-angle lenses or fish-eye lenses. It can capture more complex distortion patterns and ensure more accurate correction results. The mathematical formula for radial distortion can be expressed as:

[0055] x distorted = x(1 + k 1 r 2 + k 2 r 4 + k 3 r 6 )

[0056] y distorted = y(1 + k 1 r 2 + k 2 r 4 + k 3 r 6 )

[0057] Among them, (x, y) are the distorted coordinates, (x distorted , y distorted ) are the coordinates before distortion, and r is the distance of this point from the center of the image.

[0058] 2. Tangential distortion correction: Different from radial distortion, tangential distortion is mainly due to the perspective transformation caused by the non-parallelism between the imaging plane and the lens plane. The correction of tangential distortion involves parameters such as p 1 , p 2 etc. These parameters are used to correct the image distortion caused by the manufacturing process. The mathematical formula for tangential distortion can be expressed as:

[0059] x corrected = x + [2p 1xy + p 2 (r 2 + 2x 2 )]

[0060] y corrected =y + [2p 1 (r 2 + 2y 2 ) + 2p 2 xy]

[0061] Wherein, (x, y) are the coordinates after distortion, and (x corrected , y corrected ) are the coordinates after tangential distortion correction.

[0062] Taking into account both radial distortion and tangential distortion, it can be achieved by directly adding the correction formulas of the two types of distortion.

[0063] Step 220: Obtain the local trajectory of each traffic participant under the single-view camera according to the detection positions, categories, and characteristics of all traffic participants appearing in all the acquired images.

[0064] As long as the traffic participants appearing in a frame of the surveillance video are traffic participants under the single-view camera, so the traffic participants appearing in each frame of the image are included in the scope of the traffic participants under the single-view camera.

[0065] To obtain the local trajectory of each traffic participant under the single-view camera, the traffic participants can be matched according to the three dimensions of the detection position, category, and characteristic of the traffic participants. In the application of the three dimensions, the matching can be first performed in the category dimension, and then the detection position and characteristic can be matched among the matching results after the category dimension matching, which can reduce the matching workload.

[0066] The traffic flow trajectory processing method provided by the embodiment of the present application matches traffic participants between different images under the same-view camera from the three dimensions of detection position, category, and characteristic, so as to improve the accuracy and speed of matching and accelerate the acquisition speed of local trajectories.

[0067] In an exemplary embodiment, as Figure 3 shown, obtaining the local trajectory of each traffic participant under the single-view camera based on the captured surveillance video includes:

[0068] Step 300: Obtain an image frame sequence corresponding to the captured surveillance video.

[0069] Step 310: For each frame of the obtained image frame sequence, respectively obtain the detection position, category, characteristic, and movement direction of each traffic participant appearing in the image.

[0070] The direction of a traffic participant can refer to the movement orientation of the traffic participant at a certain moment, usually referring to the direction of its travel or the direction it faces.

[0071] Step 320: Based on the detection positions, categories, features, and movement directions of all traffic participants that appear in all the acquired images, obtain the local trajectory of each traffic participant under the single-view camera.

[0072] In order to obtain the local trajectory of each traffic participant under the single-view camera, different from the way of matching traffic participants according to the detection position, category, feature, and three dimensions in the above embodiment, the present application also provides another alternative embodiment, that is, matching traffic participants according to four dimensions of the detection position, category, feature, and movement direction. In the application of the four matching dimensions, it is still possible to first perform matching in the category dimension, and then perform matching of the detection position, feature, and movement direction in the matching results obtained through the category dimension matching, which can also reduce the matching workload. Compared with the three-dimension matching provided in the above embodiment, after adding the movement direction dimension, it is possible to avoid the problem of final mis-matching caused by different movement directions but passing the other three-dimension matching, improve the accuracy of matching traffic participants between different images, and improve the accuracy of obtaining the local trajectory.

[0073] In an exemplary embodiment, the obtaining the local trajectory of each traffic participant under the single-view camera based on the detection positions, categories, and features of all traffic participants that appear in all the acquired images includes:

[0074] Take all the traffic participants that appear in the first frame image of the image frame sequence and their corresponding detection positions as the existing traffic participants and their corresponding existing trajectories respectively. Sequentially obtain each subsequent frame image, and whenever a frame image is obtained, take the obtained image as the current image and perform the following processing until there are no unprocessed images in the image frame sequence. Take the finally obtained existing traffic participants and their corresponding existing trajectories as the traffic participants and their corresponding local trajectories under the single-view camera respectively:

[0075] Determine the existing traffic participants to be matched with the current image among all the existing traffic participants. According to the detection positions, categories, and features of each traffic participant in the current image, as well as the categories, features, and existing trajectories of all the determined existing traffic participants, match each traffic participant in the current image with all the determined existing traffic participants; update the existing trajectories of the corresponding existing traffic participants according to the matching results, or update the range of the existing traffic participants and the existing trajectories of the corresponding existing traffic participants.

[0076] The determined existing traffic participants used for matching with the current image can be all of the existing traffic participants or some of them. When the determined existing traffic participants used for matching with the current image are some of all the existing traffic participants, it aims to screen the existing traffic participants used for matching with the current image. When screening the existing traffic participants used for matching with the current image, the existing traffic participants used for matching with the current image can be screened according to the appearance situation. For example, the existing traffic participants that appeared too early but have not appeared again later are not used for comparison, or the existing traffic participants with a short appearance duration are not used for comparison.

[0077] Taking the process of obtaining the next frame image of a certain intermediate frame as an example to illustrate the above process. Assume that when the certain intermediate frame image is processed, the existing traffic participants include: traffic participant A, traffic participant B, traffic participant C, traffic participant D, and traffic participant E. In the case of obtaining the next frame image of the intermediate frame (the obtained next frame image is the current image), first determine the existing traffic participants used for matching with the current image among the existing traffic participants A, B, C, D, and E. Assume they are existing traffic participants A, C, and D. Then, according to the detection positions, categories, and features of each traffic participant in the current image, as well as the categories, features, and existing trajectories of the existing traffic participants A, C, and D, match each traffic participant in the current image with the existing traffic participants A, C, and D; update the existing trajectories of the corresponding existing traffic participants A, C, and D according to the matching results, or update the range of the existing traffic participants and the existing trajectories of the corresponding existing traffic participants.

[0078] In an exemplary embodiment, obtaining the local trajectory of each traffic participant under the single perspective camera according to the detection positions, categories, features, and movement directions of all the traffic participants that appear in all the obtained images includes:

[0079] Taking all the traffic participants that appear in the first frame image of the image frame sequence and their corresponding detection positions as the existing traffic participants and the corresponding existing trajectories respectively, sequentially obtain each subsequent frame image, and whenever a frame image is obtained, take the obtained image as the current image and perform the following processing until there are no unprocessed images in the image frame sequence. Take the finally obtained existing traffic participants and the corresponding existing trajectories as the traffic participants and the corresponding local trajectories under the single perspective camera respectively:

[0080] Determine existing traffic participants for matching with the current image among all existing traffic participants. According to the detection positions, categories, features, and movement directions of each traffic participant in the current image, as well as the categories, features, and existing trajectories of all determined existing traffic participants, match each traffic participant in the current image with all the determined existing traffic participants; update the existing trajectories of the corresponding existing traffic participants according to the matching results, or update the scope of the existing traffic participants and the existing trajectories of the corresponding existing traffic participants.

[0081] Still taking the process of obtaining the next frame image of a certain intermediate frame as an example to illustrate the above process. Assume that when processing the certain intermediate frame image, the existing traffic participants include: traffic participant A, traffic participant B, traffic participant C, traffic participant D, and traffic participant E. When obtaining the next frame image of the certain intermediate frame image, first determine the existing traffic participants for matching with the current image among existing traffic participants A, B, C, D, and E. Assume they are existing traffic participants A, C, and D. Then, according to the detection positions, categories, features, and movement directions of each traffic participant in the current image, as well as the categories, features, and existing trajectories of existing traffic participants A, C, and D, match each traffic participant in the current image with existing traffic participants A, C, and D; update the existing trajectories of the corresponding existing traffic participants A, C, and D according to the matching results, or update the scope of the existing traffic participants and the existing trajectories of the corresponding existing traffic participants.

[0082] In an exemplary embodiment, the step of matching each traffic participant in the current image with all the determined existing traffic participants according to the detection positions and features of each traffic participant in the current image, as well as the features and existing trajectories of all the determined existing traffic participants, includes:

[0083] Predict the predicted positions where the determined existing traffic participants should be located if they appear in the current image according to the existing trajectories of each determined existing traffic participant.

[0084] Match each traffic participant in the current image with all the determined existing traffic participants together according to the feature matching requirements, category matching requirements, and position matching requirements; wherein, the feature matching requirements refer to the matching requirements that should be satisfied between the features of the traffic participant in the current image and the features of the determined existing traffic participant during matching; the category matching requirements refer to the matching requirements that should be satisfied between the categories of the traffic participant in the current image and the categories of the determined existing traffic participant during matching; the position matching requirements refer to the matching requirements that should be satisfied between the detection position of the traffic participant in the current image and the predicted position of the determined existing traffic participant during matching.

[0085] The category matching requirement can be that when matching, the category of the traffic participant in the current image is the same as the category of the determined existing traffic participant.

[0086] According to the above description, during the tracking of each traffic participant, it is not certain that each traffic participant will appear in every frame of the surveillance video. Therefore, it is not certain that each determined existing traffic participant will appear in the current image. However, before the traffic participants in the current image and the existing traffic participants are completely matched, it is also impossible to determine whether each existing traffic participant will appear in the current image. Therefore, the solution can be to first assume that each existing traffic participant will appear in the current image, predict the position where it appears, and then match the traffic participants in the current image and the existing traffic participants according to the characteristics, categories, and detection positions of the traffic participants in the current image and the characteristics, categories, and predicted positions of the existing traffic participants.

[0087] During the process of predicting the position where each determined existing traffic participant appears, the Kalman filter can be used to achieve this. The Kalman filter is based on the state of the previous frame to estimate the possible position (i.e., the predicted position) of the target (i.e., the existing traffic participant) in the current frame. For the existing traffic participants that did not appear in the previous frame, the possible position in the estimated previous frame can be used to continue estimating the possible position in the current frame.

[0088] In an exemplary embodiment, matching each traffic participant in the current image with all the determined existing traffic participants according to the detection position, characteristics, and movement direction of each traffic participant in the current image, and the characteristics and existing trajectories of all the determined existing traffic participants includes:

[0089] Predict the predicted position where the determined existing traffic participant should be if it appears in the current image according to the existing trajectory of each determined existing traffic participant;

[0090] Match each traffic participant in the current image with all the determined existing traffic participants together according to the feature matching requirement, category matching requirement, position matching requirement, and movement direction matching requirement; wherein, the feature matching requirement refers to the matching requirement that should be satisfied between the characteristics of the traffic participant in the current image and the characteristics of the determined existing traffic participant when matching; the category matching requirement refers to the matching requirement that should be satisfied between the category of the traffic participant in the current image and the category of the determined existing traffic participant when matching; the position matching requirement refers to the matching requirement that should be satisfied between the detection position of the traffic participant in the current image and the predicted position of the existing traffic participant when matching; the movement direction matching requirement refers to the matching requirement that should be satisfied between the movement direction of the traffic participant in the current image and the movement trajectory of the existing traffic participant when matching.

[0091] Different from the above embodiments, which match the traffic participants in the current image and the existing traffic participants according to the characteristics, categories, detection positions of the traffic participants in the current image and the characteristics, categories, predicted positions of the existing traffic participants, in this embodiment, the traffic participants in the current image and the existing traffic participants are matched according to the characteristics, categories, detection positions, movement directions of the traffic participants in the current image and the characteristics, categories, predicted positions of the existing traffic participants.

[0092] In an exemplary embodiment, updating the existing trajectory of the corresponding existing traffic participant according to the matching result, or updating the range of the existing traffic participant and the existing trajectory of the corresponding existing traffic participant includes:

[0093] When there are traffic participants with successful matches in the current image, for each traffic participant with a successful match in the current image, update the existing trajectory of the matched existing traffic participant according to the detection position of the traffic participant;

[0094] When there are traffic participants with unsuccessful matches in the current image, update the range of the existing traffic participants according to all the traffic participants with unsuccessful matches, and use the detection position of each traffic participant with an unsuccessful match as the existing trajectory of the corresponding traffic participant respectively.

[0095] Illustrate the above process by way of example. Assume that the existing traffic participants include: traffic participant A, traffic participant B, traffic participant C, traffic participant D, and traffic participant E. The existing traffic participants determined from the existing traffic participants A, B, C, D, and E for matching with the current image are: existing traffic participants A, C, and D. The traffic participants in the current image are initially denoted as existing traffic participant a, traffic participant b, traffic participant c, and traffic participant d. Each of the traffic participants a, b, c, and d in the current image is matched with the existing traffic participants A, C, and D. Assume that traffic participant a in the current image is successfully matched with existing traffic participant C, traffic participant c in the current image is successfully matched with existing traffic participant A, and traffic participant d in the current image is successfully matched with existing traffic participant D. Then, update the existing trajectory of existing traffic participant C using the detection position of traffic participant a in the current image, update the existing trajectory of existing traffic participant A using the detection position of traffic participant c in the current image, and update the existing trajectory of existing traffic participant D using the detection position of traffic participant d in the current image; and update the range of the existing traffic participants using traffic participant b in the current image, that is, regard traffic participant b in the current image as a new existing traffic participant, and update the range of the existing participants to: existing traffic participants A, C, D, and b, and use the detection positions of traffic participant b in the current image as the existing trajectories of the corresponding existing traffic participant b respectively.

[0096] In an exemplary embodiment, the method further includes:

[0097] When there are traffic participants with successful matches in the current image, determine whether there are target existing traffic participants among all the matched existing traffic participants. When there are target existing traffic participants, for each target existing traffic participant, update the existing trajectory of the target existing traffic participant according to the predicted position of the target existing traffic participant in the image where it does not appear; wherein, the target existing traffic participant refers to an existing traffic participant that has not appeared in at least the previous frame image of the current image.

[0098] The target existing traffic participant refers to an existing traffic participant that has not appeared in the previous frame image of the current image, or an existing traffic participant that has not appeared in several previous frame images of the current image. For these existing traffic participants, the predicted position in the image where they do not appear can be predicted based on their existing trajectories, and the existing trajectory can be updated again according to the predicted position. Therefore, the missing part in the middle of the two trajectories is complemented using the predicted position, thereby obtaining a complete trajectory.

[0099] The traffic flow trajectory processing method provided by the embodiments of the present application can not only match two interrupted trajectories caused by obstacles such as green belts, isolation belts, viaducts, and vehicles, but also perform trajectory completion processing on the occluded part through trajectory prediction, thereby improving the integrity of the matched trajectory.

[0100] In an exemplary embodiment, the feature matching requirements include:

[0101] The similarity between the features of traffic participants in the current image and the features of existing traffic participants is not less than a first preset similarity threshold.

[0102] The similarity between the features of traffic participants in the current image and the features of existing traffic participants can be calculated by means of color histogram, texture-based features, shape-based features, and deep learning features. Among them, the color histogram is a simple and effective method for representing image features, suitable for describing the overall color distribution of an image. By comparing the color histograms of two traffic participants, their appearance similarity can be evaluated; texture features are used to describe local patterns and structures in an image, suitable for capturing the surface texture information of traffic participants. Common texture feature extraction methods include Gray-Level Co-occurrence Matrix (GLCM), Local Binary Pattern (LBP), etc.; shape features are used to describe the contour or boundary of an object, suitable for capturing the geometric shape information of traffic participants. Common shape feature extraction methods include edge detection, Hu invariant moments, contour matching, etc.; deep learning features can directly learn high-level semantic features from an image through training a Convolutional Neural Network (CNN) and are used to calculate the similarity.

[0103] In an exemplary embodiment, the position matching requirements include:

[0104] The distance between the detection position of traffic participants in the current image and the predicted position of existing traffic participants is not greater than a first preset distance threshold.

[0105] The trajectory of an existing traffic participant reflects the historical movement of the existing traffic participant over a period of time. There must be a reasonable correlation between the historical movement of the existing traffic participant and the movement state at a new moment because the movement process is continuous rather than disjointed. In terms of position, there must also be a reasonably matching correlation between the historical trajectory of the existing traffic participant and the detected position of the traffic participant in the current image. Specifically, when the predicted position based on the historical trajectory of a certain existing traffic participant on the current image is denoted as point M, and the position of a traffic participant in the current image is point N, and the distance between point M and point N is very far, then from the perspective of the historical trajectory of this existing traffic participant, it is obvious that this existing traffic participant cannot change to point N within one frame of time. Therefore, the traffic participant corresponding to point N in the current image cannot match this existing traffic participant (i.e., the existing traffic participant corresponding to point M) in terms of the position matching requirement.

[0106] The setting of the first preset distance threshold can be determined according to the actual situation.

[0107] In an exemplary embodiment, the movement direction matching requirement includes:

[0108] The deviation between the movement direction of the traffic participant in the current image and the movement direction of the existing trajectory of the existing traffic participant is not greater than the first preset deviation threshold.

[0109] In terms of movement direction, there must also be a reasonable correlation between the historical trajectory of the existing traffic participant and the movement direction in the current image. Specifically, when the movement direction based on the historical trajectory of a certain existing traffic participant (which can be the movement direction of the existing traffic participant at the last position in the historical trajectory) is angle P (a reference direction is preset in advance, and the movement direction of the traffic participant is the angle relative to this reference direction), and the movement direction of a traffic participant in the current image is angle Q, and the deviation between angle P and angle Q is very large, then from the perspective of the historical trajectory of this existing traffic participant, it is obvious that this existing traffic participant cannot have such a large angle change within one frame of time. Therefore, the traffic participant corresponding to angle Q in the current image cannot match this existing traffic participant (i.e., the existing traffic participant corresponding to angle P) in terms of the position matching requirement.

[0110] The setting of the first preset deviation threshold can be determined according to the actual situation. The movement direction matching requirement can greatly improve the trajectory matching accuracy in a traffic flow environment with complex movement directions and avoid the matching of trajectories with the same category, similar features, close positions but large movement direction deviations.

[0111] In an exemplary embodiment, the existing traffic participants used for matching with the current image include:

[0112] All existing traffic participants that appear in the N frames of images successively forward from the current image; where N is an integer greater than or equal to 1.

[0113] For the existing traffic participants limited for matching with the current image, it can be achieved through the time when the existing traffic participant appears for the last time. Corresponding to the frame image, all existing traffic participants that appear in the N frames of images successively forward from the current image can be limited, while for other existing traffic participants in other frame images, they are not used as the existing traffic participants for matching with the current image.

[0114] N can be obtained according to the time when the last appearance of the existing traffic participant determined for filtering out the ones for comparison is. After determining the time, according to the acquisition interval between adjacent frame images, N can be obtained.

[0115] The above process is a process of detecting and tracking traffic participants in a video to extract the two-dimensional trajectories of traffic participants in the video, which can be called a multi-object detection and multi-object tracking process. The multi-object detection and multi-object tracking process can be summarized into the following steps:

[0116] 1. Multi-object detection

[0117] This process includes the following specific steps:

[0118] 1.1. Preprocess the input image

[0119] The image will be adjusted to the fixed size required by the object detection model while maintaining the aspect ratio of the image. This step also includes normalizing the pixel values of the image to between [0, 1], which can usually be completed by dividing the pixel values by 255. In addition, data augmentation techniques can also be used to increase the diversity of training data by means of rotation, cropping, translation, etc.

[0120] 1.2. Input the preprocessed image into the model (such as YOLO) for inference

[0121] Internally, the model first uses a convolutional neural network (CNN) to extract features from the image. The model will divide the input image into multiple grids, and each grid will predict multiple bounding boxes, as well as the confidence and class probabilities of each bounding box. In the bounding box prediction stage, each bounding box is defined by four parameters: the coordinates of the center point (x, y), the width w and height h of the bounding box. In addition, the model will also output the confidence score of the bounding box, which represents whether there is an object (i.e., a traffic participant) in the box and the matching degree between the bounding box and the object. At the same time, the model will output the probabilities belonging to different classes to help identify the class of the object in the bounding box.

[0122] 1.3. Process the output of the object detection model

[0123] First, set a confidence threshold to filter out the bounding boxes with low confidence.

[0124] Then, use the non-maximum suppression (NMS) method to remove the bounding boxes with high overlap, ensuring that each object is ultimately represented by only one bounding box.

[0125] Finally, the model outputs the detection results. The output includes the bounding boxes, class labels, and confidence scores of each detected object. These bounding boxes are usually drawn on the original image, and the class and confidence scores of the objects are annotated for visualizing the results of multi-object detection.

[0126] 2. Multi-object tracking

[0127] This process includes the following specific steps:

[0128] 2.1. Input the detection results of multi-object detection

[0129] At this stage, the model has completed the detection of all objects in the image and output the bounding boxes, class labels, and confidence scores of each object. DeepSORT takes these detection results as input to obtain the position and class information of each object.

[0130] 2.2. Initialize the tracker

[0131] DeepSORT uses the Kalman filter to predict the motion state of the object. Each detected object is assigned a unique track ID. The state initialization of the Kalman filter is set according to the detected bounding box output by the model, including information such as the position and velocity of the object.

[0132] 2.3. Extract object features

[0133] To enhance the tracking accuracy, DeepSORT adopts a deep neural network to extract the feature vector of each detected object and record the class ID of each object. This feature vector is used to describe the appearance of the object, help distinguish objects with similar appearances, and reduce the impact of object occlusion and appearance changes on tracking. Among them, the class ID is used to filter out invalid matches with different class IDs during subsequent feature matching.

[0134] 2.4. Predict the object state

[0135] The Kalman filter of DeepSORT predicts the position and motion state of the object in each frame of the image. The prediction process is based on the state of the previous frame to estimate the possible position of the object in the current frame, so that even if the object is not detected temporarily, the tracker can still maintain the estimation of the object state.

[0136] 2.5. Match detections with tracks

[0137] At this stage, DeepSORT uses the Hungarian algorithm to match the detection results with the existing tracks. The matching criteria include:

[0138] The predicted position of the Kalman filter: the distance between the detected object and the predicted position.

[0139] Appearance feature matching: Determine the similarity between the detected object and the tracked object by comparing the feature vectors of objects with the same class ID. This can effectively reduce matching errors caused by similar object appearances.

[0140] 2.6. Update object tracks

[0141] Successfully matched objects will update their tracks with the latest detection results. The state of the Kalman filter will be updated, including the position, speed, and appearance features of the object, to adapt to the new detection data.

[0142] 2.7. Process unmatched tracks

[0143] For those tracks that do not find a match in the current frame, DeepSORT will first perform IOU matching. For tracks that have been successfully matched, DeepSORT continues to track these tracks and maintains their states. If a track does not match a new detection result in multiple frames, the track will be marked as invalid, indicating that the object has left the field of view or disappeared.

[0144] 2.8. Output tracking results

[0145] DeepSORT will output the tracking results of each object, including the track ID, bounding box, and class information of the object. This information can be used for subsequent analysis or applications, providing continuous tracking and identification of objects in the video.

[0146] The process of multi-object tracking can be briefly summarized as the following process as shown Figure 4 First, detect the image, extract features of the objects in the image, then perform object matching based on the extracted features. If the matching fails, perform IOU matching. If the matching is successful, update the track and predict the motion state of the object in the next frame image, and continue to perform object matching for subsequent images. During the process of performing IOU matching, if the IOU matching fails, judge whether the number of frames in which the object does not appear is greater than the set maximum number of frames. If it is greater than the maximum number of frames, stop tracking the object. If it is not greater than the maximum number of frames, update the track and predict the motion state of the object in the next frame image, and continue to perform object matching for subsequent images.

[0147] In an exemplary embodiment, the matching process of local trajectories of all traffic participants under all single-view cameras for the same traffic participant under multiple view cameras is performed to obtain a set of local trajectories of each traffic participant under multiple view cameras, including:

[0148] Perform a coordinate system unification transformation process on the local trajectories of traffic participants under all single-view cameras;

[0149] Perform a trajectory matching process of the same traffic participant under different view cameras on the obtained local trajectories of traffic participants under all single-view cameras that have undergone the coordinate system transformation process to obtain a set of local trajectories of each traffic participant under multiple view cameras.

[0150] In the process of performing the coordinate system unification transformation process on the local trajectories of traffic participants under all single-view cameras, a coordinate system of one view camera can be first selected as the world coordinate system, and the coordinate systems of other view cameras are converted to the time coordinate system. During the coordinate system conversion process, the overlapping parts are respectively selected in the two views, the coordinate point sets are selected in the corresponding order, and it is ensured that the point pairs are the same point in reality during the selection process, and the coordinate transformation parameters are calculated using Singular Value Decomposition (SVD). The specific process can include the following steps:

[0151] 1. Calculate the centroids centroid of two sets of point sets A, of A and

[0152] where n is the number of all points in the point set, A i represents the coordinate of the i-th point in point set A, and B i represents the coordinate of the i-th point in point set B.

[0153] 2. Remove the centroids from the two sets of point sets A and B respectively to obtain the point sets after removing the centroids

[0154]

[0155] 3. Calculate the covariance matrix H:

[0156]

[0157] 4. Perform SVD decomposition on the calculated covariance matrix H:

[0158] H = USV T

[0159] 5. Obtain the rotation matrix R:

[0160] R = V T U T

[0161] 6. Calculate the scaling factor s:

[0162]

[0163] 7. Calculate the translation vector t:

[0164] t = centroid A - s·(R·centroid B )

[0165] 8. Perform coordinate transformation calculation:

[0166] A = s·(R·B) + t

[0167] In an exemplary embodiment, as Figure 5 shown, for all the local trajectories of traffic participants under each single - perspective camera obtained after coordinate system transformation processing, perform trajectory matching processing of the same traffic participant under different perspective cameras to obtain the local trajectory set of each traffic participant under multiple perspective cameras, including:

[0168] Step 400: According to the category matching requirement and the trajectory matching requirement, jointly perform matching processing on all the local trajectories of traffic participants under each single - perspective camera obtained after coordinate system transformation processing to obtain multiple pre - matched local trajectory sets.

[0169] Among them, the category matching requirement refers to the matching requirement that should be satisfied between traffic participants to which the mutually - matched local trajectories belong during matching; the trajectory matching requirement refers to the matching requirement that should be satisfied between the mutually - matched local trajectories during matching;

[0170] Step 410: For each pre - matched local trajectory set, respectively determine whether all the local trajectories in the pre - matched local trajectory set belong to the same traffic participant, and re - divide the pre - matched local trajectory sets where the local trajectories belong to multiple traffic participants;

[0171] Step 420: According to the local trajectory sets obtained after re - division and the pre - matched local trajectories where the local trajectories belong to the same traffic participant, obtain the local trajectory set of each traffic participant under multiple perspective cameras.

[0172] The category matching requirement can be that during matching, the categories of traffic participants to which the mutually - matched local trajectories belong are the same.

[0173] Since the local trajectories are extracted from the position information in multiple frames of images and the position information has a time attribute, the distance between local trajectories can be calculated as follows: Determine multiple moments for trajectory extraction, extract trajectory points from the local trajectories according to the multiple determined moments, calculate the distances between the trajectory points at the same time among the extracted ones, and for actual matching, for the convenience of matching, local trajectories can be matched pairwise.

[0174] In an exemplary embodiment, the trajectory matching requirements include:

[0175] The maximum distance between the local trajectories in a pre-matched local trajectory set does not exceed a second preset distance threshold, and the maximum direction deviation does not exceed a second preset deviation threshold.

[0176] The maximum distance between local trajectories may refer to the maximum distance between trajectory points. For example, assuming that the multiple moments for trajectory extraction are the t1 moment, the t2 moment, and the t3 moment respectively, then for each pair of local trajectories, calculate the distance between the two local trajectories at the t1 moment, calculate the distance between the two local trajectories at the t2 moment, calculate the distance between the two local trajectories at the t3 moment, and take the maximum calculated distance as the maximum distance between the two local trajectories.

[0177] In an exemplary embodiment, determining whether the local trajectories in the pre-matched local trajectory set all belong to the same traffic participant and re-partitioning the pre-matched local trajectory set where the local trajectories belong to multiple traffic participants includes:

[0178] Calculate the feature similarity between the traffic participants to which all pairs of trajectories in the pre-matched local trajectory set belong, determine whether all the feature similarities are greater than a second preset similarity threshold. When any one of the feature similarities is not greater than the second preset similarity threshold, determine that the local trajectories in the pre-matched local trajectory set belong to multiple traffic participants, and partition the pairs of trajectories with feature similarities greater than the second preset similarity threshold to the same local trajectory set.

[0179] For example, assume that there are 4 local trajectories in a pre-matched local trajectory set, namely local trajectories 1, 2, 3, and 4. Calculate the feature similarity between the traffic participants to which they belong respectively; and assume that the feature similarity between the traffic participant to which local trajectory 1 belongs and the traffic participant to which local trajectory 2 belongs, the feature similarity between the traffic participant to which local trajectory 1 belongs and the traffic participant to which local trajectory 4 belongs, the feature similarity between the traffic participant to which local trajectory 2 belongs and the traffic participant to which local trajectory 3 belongs, and the feature similarity between the traffic participant to which local trajectory 3 belongs and the traffic participant to which local trajectory 4 belongs are all not greater than the second preset similarity threshold, and the feature similarity between the traffic participant to which local trajectory 1 belongs and the traffic participant to which local trajectory 3 belongs, and the feature similarity between the traffic participant to which local trajectory 2 belongs and the traffic participant to which local trajectory 4 belongs are both greater than the second preset similarity threshold. Therefore, it is determined that the local trajectories in the pre-matched local trajectory set belong to multiple traffic participants, and local trajectory 1 and local trajectory 3 are divided into the same local trajectory set, and local trajectory 2 and local trajectory 4 are divided into another local trajectory set.

[0180] In an exemplary embodiment, after obtaining the global trajectory of each traffic participant under multiple perspective cameras, the method further includes:

[0181] Perform trajectory optimization processing on the global trajectories of each traffic participant under multiple perspective cameras respectively to obtain the processed global trajectories of each traffic participant under multiple perspective cameras; wherein, the trajectory optimization processing includes at least one of the following: moving average processing, Kalman filtering processing.

[0182] Moving average processing is a commonly used time series smoothing technique, widely used in fields such as data analysis, signal processing, and financial analysis. By calculating the average value within a certain window, the moving average can effectively reduce the noise in the data and reveal potential trends and patterns. According to different application scenarios and requirements, there are various forms of moving average, including simple moving average (SMA), weighted moving average (WMA), exponential moving average (EMA), etc.

[0183] Kalman filtering (KF) is a recursive, optimal estimation method based on linear dynamic systems, widely used in fields such as signal processing, control systems, navigation, and robotics. By combining the prediction model of the system and the observed data, it can effectively estimate the state of the system and reduce the influence of noise and uncertainty. The core idea of Kalman filtering is to recursively estimate the state of the system through two steps of prediction and update, while minimizing the variance of the estimation error.

[0184] The above process is the matching process of traffic participants in the video with cameras from different perspectives, which can be called the multi-camera perspective multi-object REID process. The multi-camera perspective multi-object REID process can be summarized into the following steps:

[0185] 1. Tracking result conversion

[0186] Convert the tracking results in each single perspective into 3D coordinates, and then convert them to the global coordinate system to unify the coordinate systems of multiple cameras.

[0187] 2. Create feature matching relationships

[0188] Match each pair of traffic participants from multiple camera perspectives and filter them through the following information:

[0189] 2.1 Delete the matching pairs with inconsistent class IDs.

[0190] 2.2 Delete the matching pairs where the position information of two traffic participants at the same moment is quite different.

[0191] If the deviation of the overlapping part of the trajectory of a traffic participant in one perspective from that in another perspective exceeds 3.5 meters, or the directions are opposite, or their class IDs are inconsistent, it is regarded as an incorrect matching relationship. Incorrect matching relationships do not perform feature matching and feature distance calculation.

[0192] 2.3 Delete the matching pairs with inconsistent driving directions according to the relative positions of the cameras in combination with their respective trajectories.

[0193] 3. Trajectory matching

[0194] Calculate the feature distances for the filtered matching relationships respectively, and retain the trajectory merge IDs with the shortest feature distance and less than a certain threshold.

[0195] When performing trajectory matching, the features of existing traffic participants in a camera perspective video can be converted from a list form to a dictionary for quick access to the feature vectors of each traffic participant. Then, for each existing traffic participant ID, calculate the cosine distance between the new feature vector and the feature vector of this target (i.e., the traffic participant) to obtain a similarity matrix. Next, extract the minimum distance value from this matrix, which represents the highest similarity between the new feature and the existing target feature.

[0196] Subsequently, record the minimum distance of each traffic participant ID and find the target traffic participant with the minimum distance, that is, the best-matched target. Finally, this method returns the ID of the target traffic participant that best matches the new feature vector, thus achieving accurate target association in the target tracking system.

[0197] After obtaining the global IDs of traffic participants from multiple cameras, the trajectories are optimized for a second time. The moving average and Kalman filter are used to smooth the trajectories respectively to reduce noise and jitter, and the trajectory information of the same global ID under multiple cameras is optimized to obtain the global trajectory information in smooth and continuous three-dimensional world coordinates. The final output is a simulation scenario file.

[0198] The process of multi-object matching can be as Figure 6 shown and briefly summarized as the following process: First, perform coordinate transformation, then create a matching relationship, and judge whether the timing constraint is satisfied. If not, discard the matching relationship. If satisfied, further judge whether the features can be matched. If not, discard the matching relationship. If matched, save the trajectory.

[0199] Exemplary device

[0200] The embodiment of the present application also provides a traffic flow trajectory processing device, as Figure 7 described, including:

[0201] A trajectory acquisition unit 500, configured to, for each single perspective camera among multiple perspective cameras, respectively obtain the local trajectory of each traffic participant under the single perspective camera based on the captured surveillance video;

[0202] A trajectory matching unit 510, configured to perform matching processing of the same traffic participant under multiple perspective cameras on the local trajectories of all traffic participants obtained under all single perspective cameras, to obtain a set of local trajectories of each traffic participant under multiple perspective cameras;

[0203] A trajectory fusion unit 520, configured to perform fusion processing on the set of local trajectories of each traffic participant under multiple perspective cameras respectively, to obtain the global trajectory of each traffic participant under multiple perspective cameras.

[0204] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0205] Obtain an image frame sequence corresponding to the captured surveillance video;

[0206] For each frame of the obtained image frame sequence, respectively obtain the detection position, category, and features of each traffic participant appearing in the image;

[0207] According to the detection positions, categories, and features of all traffic participants appearing in all the obtained images, obtain the local trajectory of each traffic participant under the single perspective camera.

[0208] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0209] Obtain a sequence of image frames corresponding to the captured surveillance video;

[0210] For each frame image in the obtained sequence of image frames, respectively obtain the detection positions, categories, features, and movement directions of each traffic participant that appears in the image;

[0211] Based on the detection positions, categories, features, and movement directions of all traffic participants that appear in all the obtained images, obtain the local trajectories of each traffic participant under the single perspective camera.

[0212] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0213] Take all traffic participants that appear in the first frame image of the sequence of image frames and their corresponding detection positions as existing traffic participants and corresponding existing trajectories respectively, sequentially obtain each subsequent frame image, and whenever a frame image is obtained, take the obtained image as the current image and perform the following processing until there are no unprocessed images in the sequence of image frames, and take the finally obtained existing traffic participants and corresponding existing trajectories as the traffic participants and corresponding local trajectories under the single perspective camera respectively:

[0214] Determine the existing traffic participants to be matched with the current image among all existing traffic participants, and perform matching for each traffic participant in the current image with all the determined existing traffic participants according to the detection positions, categories, and features of each traffic participant in the current image, and the categories, features, and existing trajectories of all the determined existing traffic participants; update the existing trajectories of the corresponding existing traffic participants according to the matching results, or update the scope of the existing traffic participants and the existing trajectories of the corresponding existing traffic participants.

[0215] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0216] Take all traffic participants that appear in the first frame image of the sequence of image frames and their corresponding detection positions as existing traffic participants and corresponding existing trajectories respectively, sequentially obtain each subsequent frame image, and whenever a frame image is obtained, take the obtained image as the current image and perform the following processing until there are no unprocessed images in the sequence of image frames, and take the finally obtained existing traffic participants and corresponding existing trajectories as the traffic participants and corresponding local trajectories under the single perspective camera respectively:

[0217] Identify existing traffic participants for matching with the current image among all existing traffic participants. Based on the detection positions, categories, features, and movement directions of each traffic participant in the current image, as well as the categories, features, and existing trajectories of all identified existing traffic participants, match each traffic participant in the current image with all the identified existing traffic participants; update the existing trajectories of the corresponding existing traffic participants according to the matching results, or update the scope of the existing traffic participants and the existing trajectories of the corresponding existing traffic participants.

[0218] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0219] Predict the predicted positions where the identified existing traffic participants should be located if they appear in the current image according to the existing trajectories of each identified existing traffic participant;

[0220] Match each traffic participant in the current image with all the identified existing traffic participants according to the feature matching requirement, category matching requirement, and position matching requirement together; wherein, the feature matching requirement refers to the matching requirement that should be satisfied between the features of the traffic participant in the current image and the features of the identified existing traffic participant during matching; the category matching requirement refers to the matching requirement that should be satisfied between the category of the traffic participant in the current image and the category of the identified existing traffic participant during matching; the position matching requirement refers to the matching requirement that should be satisfied between the detection position of the traffic participant in the current image and the predicted position of the identified existing traffic participant during matching.

[0221] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0222] Predict the predicted positions where the identified existing traffic participants should be located if they appear in the current image according to the existing trajectories of each identified existing traffic participant;

[0223] Match each traffic participant in the current image with all the identified existing traffic participants according to the feature matching requirement, category matching requirement, position matching requirement, and movement direction matching requirement together; wherein, the feature matching requirement refers to the matching requirement that should be satisfied between the features of the traffic participant in the current image and the features of the identified existing traffic participant during matching; the category matching requirement refers to the matching requirement that should be satisfied between the category of the traffic participant in the current image and the category of the identified existing traffic participant during matching; the position matching requirement refers to the matching requirement that should be satisfied between the detection position of the traffic participant in the current image and the predicted position of the identified existing traffic participant during matching; the movement direction matching requirement refers to the matching requirement that should be satisfied between the movement direction of the traffic participant in the current image and the movement trajectory of the identified existing traffic participant during matching.

[0224] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0225] When there are traffic participants with successful matches in the current image, for each traffic participant with a successful match in the current image, update the existing trajectory of the matched existing traffic participant according to the detected position of the traffic participant.

[0226] When there are traffic participants with unsuccessful matches in the current image, update the range of the existing traffic participants according to all the traffic participants with unsuccessful matches, and use the detected position of each traffic participant with an unsuccessful match as the existing trajectory of the corresponding traffic participant.

[0227] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0228] When there are traffic participants with successful matches in the current image, determine whether there are target existing traffic participants among all the matched existing traffic participants. When there are target existing traffic participants, for each target existing traffic participant, update the existing trajectory of the target existing traffic participant according to the predicted position of the target existing traffic participant in the image where it does not appear; wherein, the target existing traffic participant refers to an existing traffic participant that has not appeared in at least the previous frame image of the current image.

[0229] In an exemplary embodiment, the feature matching requirement includes:

[0230] The similarity between the features of the traffic participants in the current image and the features of the existing traffic participants is not less than a first preset similarity threshold.

[0231] In an exemplary embodiment, the position matching requirement includes:

[0232] The distance between the detected position of the traffic participant in the current image and the predicted position of the existing traffic participant is not greater than a first preset distance threshold.

[0233] In an exemplary embodiment, the motion direction matching requirement includes:

[0234] The deviation between the motion direction of the traffic participant in the current image and the motion direction of the existing trajectory of the existing traffic participant is not greater than a first preset deviation threshold.

[0235] In an exemplary embodiment, the existing traffic participants used for matching with the current image include:

[0236] All existing traffic participants that appear in N consecutive frames of images before the current image; wherein, N is an integer greater than or equal to 1.

[0237] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0238] Perform a coordinate system unification transformation process on the local trajectories of traffic participants under all single-view cameras;

[0239] Perform a trajectory matching process on the local trajectories of traffic participants under all single-view cameras that have undergone the coordinate system transformation process, to obtain a set of local trajectories of each traffic participant under multiple view cameras.

[0240] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0241] Perform a matching process on the local trajectories of traffic participants under all single-view cameras that have undergone the coordinate system transformation process according to the category matching requirement and the trajectory matching requirement together, to obtain multiple pre-matched local trajectory sets; wherein, the category matching requirement refers to the matching requirement that should be satisfied between traffic participants to which the mutually matching local trajectories belong during matching; the trajectory matching requirement refers to the matching requirement that should be satisfied between the mutually matching local trajectories during matching;

[0242] For each pre-matched local trajectory set, respectively determine whether the local trajectories in the pre-matched local trajectory set all belong to the same traffic participant, and re-partition the pre-matched local trajectory sets where the local trajectories belong to multiple traffic participants;

[0243] Based on the local trajectory sets obtained from the re-partitioning and the pre-matched local trajectories where the local trajectories belong to the same traffic participant, obtain a set of local trajectories of each traffic participant under multiple view cameras.

[0244] In an exemplary embodiment, the trajectory matching requirement includes:

[0245] The maximum distance between the local trajectories in a pre-matched local trajectory set is not greater than a second preset distance threshold, and the maximum direction deviation is not greater than a second preset deviation threshold.

[0246] In an exemplary embodiment, the trajectory matching unit 510 is further configured to:

[0247] Calculate the feature similarity between the traffic participants to which all pairwise trajectories in the pre-matched local trajectory set belong, determine whether all the feature similarities are greater than a second preset similarity threshold, when any one of the feature similarities is greater than the second preset similarity threshold, determine that the local trajectories in the pre-matched local trajectory set belong to multiple traffic participants, and partition the pairwise trajectories with feature similarities greater than the second preset similarity threshold to the same local trajectory set.

[0248] In an exemplary embodiment, the trajectory fusion unit 520 is further configured to:

[0249] Perform trajectory optimization processing on the global trajectories of each traffic participant under multiple perspective cameras respectively to obtain the processed global trajectories of each traffic participant under multiple perspective cameras; wherein, the trajectory optimization processing includes at least one of the following: moving average processing, Kalman filtering processing.

[0250] The traffic flow trajectory processing device provided in this embodiment belongs to the same inventive concept as the autonomous driving test scenario generation method provided in the above embodiments of the present application, and can execute the traffic flow trajectory processing method provided in any of the above embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference may be made to the specific processing content of the traffic flow trajectory processing method provided in the above embodiments of the present application, which will not be elaborated herein.

[0251] Exemplary electronic device

[0252] An embodiment of the present application further provides an electronic device, as Figure 8 shown, including: a memory 600 and a processor 610;

[0253] The memory 600 is connected to the processor 610 and is used to store programs;

[0254] The processor 610 is configured to implement the traffic flow trajectory processing method described in any of the above embodiments by running the programs in the memory 600.

[0255] Specifically, the above electronic device may further include: a bus, a communication interface 620, an input device 630, and an output device 640.

[0256] The processor 610, the memory 600, the communication interface 620, the input device 630, and the output device 640 are interconnected through the bus. Among them:

[0257] The bus may include a path for transmitting information between various components of the computer system.

[0258] The processor 610 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present invention. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0259] The processor 610 may include a main processor, and may also include a baseband chip, a modem, etc.

[0260] The memory 600 stores a program for implementing the technical solution of the present invention, and may also store an operating system and other key services. Specifically, the program may include program code, and the program code includes computer operation instructions. More specifically, the memory 600 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk memory, a flash memory, and so on.

[0261] The input device 630 may include devices for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor, etc.

[0262] The output device 640 may include devices for allowing information to be output to a user, such as a display screen, a printer, a speaker, etc.

[0263] The communication interface 620 may include devices of any transceiver type for communicating with other devices or communication networks, such as Ethernet, a radio access network (RAN), a wireless local area network (WLAN), etc.

[0264] The processor 610 executes the program stored in the memory 600 and calls other devices, and can be used to implement each step of any one of the traffic flow trajectory processing methods provided in the above embodiments of the present application.

[0265] Exemplary computer program product and storage medium

[0266] In addition to the above methods and devices, an embodiment of the present application may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the traffic flow trajectory processing method according to various embodiments of the present application described in any of the above embodiments of this specification.

[0267] The computer program product can be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code can be executed entirely on the user's computing device, partially on the user's device, executed as an independent software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0268] In addition, an embodiment of the present application also provides a storage medium, on which a computer program is stored. When the computer program is run by a processor, it implements the traffic flow trajectory processing method described in any of the above embodiments.

[0269] For the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0270] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0271] The steps in the methods of the embodiments of the present application can be adjusted, combined, and deleted according to actual needs. The technical features recorded in each embodiment can be replaced or combined.

[0272] The modules and sub-modules in the devices and terminals in the embodiments of the present application can be combined, divided, and deleted according to actual needs.

[0273] In several embodiments provided by the present application, it should be understood that the disclosed terminals, devices, and methods can be implemented in other ways. For example, the terminal embodiments described above are merely illustrative. For example, the division of modules or sub-modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple sub-modules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical, or other forms.

[0274] The modules or sub-modules described as separate components may or may not be physically separated. The components as modules or sub-modules may or may not be physical modules or sub-modules, that is, they can be located in one place, or can be distributed to multiple network modules or sub-modules. Some or all of the modules or sub-modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0275] In addition, each functional module or sub-module in various embodiments of the present application can be integrated in a processing module, or each module or sub-module can exist physically alone, or two or more modules or sub-modules can be integrated in one module. The above-mentioned integrated modules or sub-modules can be implemented in the form of hardware or in the form of software functional modules or sub-modules.

[0276] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0277] The steps of the methods or algorithms described in combination with the embodiments disclosed in this article can be directly implemented by hardware, software units executed by a processor, or a combination of the two. The software units can be placed in a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0278] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising said element.

[0279] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing traffic flow trajectories, characterized in that: include: For each single viewing angle camera among the multiple viewing angle cameras, a local trajectory of each traffic participant under the single viewing angle camera is obtained based on the captured surveillance video; For the local trajectories of all traffic participants obtained under all single-view cameras, the local trajectories of the same traffic participant under multiple-view cameras are matched to obtain a set of local trajectories of each traffic participant under multiple-view cameras; The local trajectory sets of each traffic participant under multiple view cameras are fused separately to obtain the global trajectory of each traffic participant under multiple view cameras.

2. The method according to claim 1, characterized in that The method of obtaining the local trajectory of each traffic participant under the single-view camera based on the captured surveillance video includes: Obtaining an image frame sequence corresponding to the captured surveillance video; For each frame of the image frame sequence obtained, respectively obtaining the detection position, category and features of each traffic participant appearing in the image; According to the respective detected positions, categories and features of all traffic participants appearing in all acquired images, the local trajectory of each traffic participant under the single-view camera is obtained.

3. The method according to claim 2, characterized in that The method of obtaining the local trajectory of each traffic participant under the single-view camera based on the captured surveillance video includes: Obtaining an image frame sequence corresponding to the captured surveillance video; For each frame of the image frame sequence obtained, respectively obtaining the detection position, category, feature and movement direction of each traffic participant appearing in the image; According to the respective detected positions, categories, features and movement directions of all traffic participants appearing in all acquired images, the local trajectory of each traffic participant under the single-view camera is obtained.

4. The method according to claim 2, characterized in that: The method of obtaining the local trajectory of each traffic participant under the single-view camera according to the respective detected positions, categories and features of all traffic participants appearing in all the obtained images includes: All traffic participants and corresponding detection positions appearing in the first frame of the image frame sequence are respectively regarded as existing traffic participants and corresponding existing tracks, and each subsequent frame of the image is obtained in sequence. Whenever a frame of the image is obtained, the obtained image is used as the current image for the following processing until there are no unprocessed images in the image frame sequence, and the finally obtained existing traffic participants and corresponding existing tracks are respectively regarded as traffic participants and corresponding local tracks under the single-view camera: Determine an existing traffic participant for matching with the current image from among all existing traffic participants, and match each traffic participant in the current image with all determined existing traffic participants based on the detected position, category and features of each traffic participant in the current image and the categories, features and existing trajectories of all determined existing traffic participants; and update the existing trajectories of the corresponding existing traffic participants based on the matching results, or update the range of the existing traffic participants and the existing trajectories of the corresponding existing traffic participants.

5. The method according to claim 3, characterized in that: The method of obtaining the local trajectory of each traffic participant under the single-view camera according to the respective detection positions, categories, features and movement directions of all traffic participants appearing in all the acquired images includes: All traffic participants and corresponding detection positions appearing in the first frame of the image frame sequence are respectively regarded as existing traffic participants and corresponding existing tracks, and each subsequent frame of the image is obtained in sequence. Whenever a frame of the image is obtained, the obtained image is used as the current image to perform the following processing until there are no unprocessed images in the image frame sequence, and the finally obtained existing traffic participants and corresponding existing tracks are respectively regarded as traffic participants and corresponding local tracks under the single-view camera: Determine an existing traffic participant for matching with the current image from among all existing traffic participants, and match each traffic participant in the current image with all determined existing traffic participants based on the detected position, category, feature and movement direction of each traffic participant in the current image and the category, feature and existing trajectory of all determined existing traffic participants; update the existing trajectory of the corresponding existing traffic participant based on the matching result, or update the range of the existing traffic participant and the existing trajectory of the corresponding existing traffic participant.

6. The method according to claim 4, characterized in that The matching of each traffic participant in the current image with all the determined existing traffic participants according to the detected position and features of each traffic participant in the current image and the features and existing trajectories of all the determined existing traffic participants includes: Predicting, based on the determined existing trajectory of each existing traffic participant, a predicted position where the determined existing traffic participant would be if it appeared on the current image; Each traffic participant in the current image is matched with all the existing traffic participants determined according to the feature matching requirements, category matching requirements and position matching requirements; wherein the feature matching requirements refer to the matching requirements that should be met between the features of the traffic participants in the current image and the features of the existing traffic participants determined during matching; the category matching requirements refer to the matching requirements that should be met between the analogy of the traffic participants in the current image and the category of the existing traffic participants determined during matching; the position matching requirements refer to the matching requirements that should be met between the detected positions of the traffic participants in the current image and the predicted positions of the existing traffic participants determined during matching.

7. The method according to claim 5, characterized in that The matching of each traffic participant in the current image with all the determined existing traffic participants according to the detected position, features and movement direction of each traffic participant in the current image and the features and existing trajectories of all the determined existing traffic participants includes: Predicting, based on the existing trajectory of each determined existing traffic participant, a predicted position where the determined existing traffic participant should be located if it appears on the current image; Each traffic participant in the current image is matched with all existing traffic participants determined according to feature matching requirements, category matching requirements, position matching requirements, and motion direction matching requirements; wherein the feature matching requirements refer to the matching requirements that should be met between the features of the traffic participants in the current image and the features of the existing traffic participants determined during matching; the category matching requirements refer to the matching requirements that should be met between the analogy of the traffic participants in the current image and the category of the existing traffic participants determined during matching; the position matching requirements refer to the matching requirements that should be met between the detected positions of the traffic participants in the current image and the predicted positions of the existing traffic participants during matching; the motion direction matching requirements refer to the matching requirements that should be met between the motion directions of the traffic participants in the current image and the motion trajectories of the existing traffic participants during matching.

8. A traffic flow trajectory processing device, characterized in that: include: A trajectory acquisition unit, configured to acquire, for each single viewing angle camera of the multiple viewing angle cameras, a local trajectory of each traffic participant under the single viewing angle camera based on the captured surveillance video; A trajectory matching unit is used to match the local trajectories of all traffic participants under all single-view cameras with the local trajectories of the same traffic participant under multiple-view cameras to obtain a set of local trajectories of each traffic participant under multiple-view cameras; The trajectory fusion unit is used to fuse the local trajectory sets of each traffic participant under multiple view cameras respectively to obtain the global trajectory of each traffic participant under multiple view cameras.

9. An electronic device, characterized in that: include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is used to implement the method for processing traffic flow trajectories as described in any one of claims 1 to 8 by running the program in the memory.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the method for processing the traffic flow trajectory according to any one of claims 1 to 8 is implemented.