Real-time multi-target tracking method based on multi-modal representation
By adopting the fusion of multimodal representation and lightweight network models in multi-objective tracking technology, combining intra-class and inter-class matching algorithms, the problem of insufficient real-time and accuracy in the existing technology is solved, and efficient multi-objective tracking is achieved.
Patent Information
- Application Number
- CN202510085608.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-10
AI Technical Summary
The existing multi-objective tracking technology is difficult to achieve real-time and high precision in practical application scenarios, especially in embedded systems with limited resources, the detection accuracy of lightweight detectors is insufficient, resulting in trajectory loss or mismatch problems.
A real-time multi-objective tracking method based on multimodal representation is adopted to adaptively integrate the features of depth and infrared data through a lightweight network model, combining intra-class matching and inter-class matching, and using IoU and area matching algorithms to update the target trajectory.
It effectively improves the accuracy of target detection and the real-time nature of multi-object tracking, reduces the problems of trajectory fragmentation and mismatch, and is suitable for a variety of practical application scenarios.
Smart Images

Figure CN120125613A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and in particular to a real-time multi-object tracking method based on multi-modal representation. Background Art
[0002] With the development of deep learning, multi-object tracking technology has received increasing attention in the field of computer vision and has been widely used in video surveillance, human-computer interaction, and virtual reality. Multi-object tracking aims to locate multiple target objects in a given video sequence, assign different identity IDs to different objects, and record the trajectories of each ID in the video. Currently, with the continuous development of neural network-based object detection technology, detection-based tracking algorithms have become the mainstream direction of multi-object tracking. Detection-based tracking algorithms first need to perform object detection on each video frame to obtain the detection results of each frame, and then perform data association based on the detection results to create the trajectories of each object in the video. Therefore, the accuracy of multi-object tracking is affected by the accuracy of the detection results. In practical application scenarios, inference is usually performed using an embedded system, etc. Compared with a server, the resources are limited, and real-time performance is a great challenge for existing multi-object tracking technologies.
[0003] In terms of object detection, with the rapid maturity of deep learning technology, object detection tends to use heavyweight detectors with a large amount of computation, making it difficult to achieve real-time performance in embedded systems commonly used in practical application scenarios, which disrupts social order. Although the detection performance of lightweight detectors is faster in inference speed, their detection accuracy is worse, easily leading to problems such as trajectory loss or incorrect matching. However, multiple modalities such as depth maps and infrared maps are easily obtained in practical application scenarios, and more abundant information can be obtained by fusing the features of multiple modalities. Currently, existing multi-modal fusion networks tend to use two-stream networks and perform multi-stage feature fusion, increasing the amount of computation.
[0004] In terms of data association, multi-object tracking algorithms are limited to intra-class matching and are difficult to handle the problem of false detection. At the same time, data association algorithms can be divided into two categories. One is feature-based matching, which increases the amount of computation; the other is motion-based matching, which uses IoU matching and is difficult to handle scenarios where the target changes drastically between adjacent frames.
[0005] In summary, how to design a multi-object tracking method based on multi-modal representation in a practical application scenario, while reducing the amount of computation and improving the accuracy of detection and data association is a challenging problem that needs to be solved. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies of the prior art and provide a real-time multi-object tracking method based on multi-modal representation.
[0007] To achieve the above object, the present invention adopts the following technical solutions.
[0008] In a first aspect, the present invention provides a real-time multi-object tracking method based on multi-modal representation, including:
[0009] Obtain the first modal data and the second modal data at the current moment and input them into the object detection model to obtain the object detection result at the current moment. The object detection model is a lightweight network model that can adaptively fuse multi-modal features. The object detection result includes the position, category, and confidence information of the object;
[0010] Perform intra-class matching and inter-class matching on the object trajectory at the previous moment and the object detection result at the current moment, and update the object trajectory at the previous moment according to the matching result to obtain the object trajectory at the current moment.
[0011] Further, the object detection model includes a feature extractor, a multi-modal fusion module, and an object detection module;
[0012] The input into the object detection model to obtain the object detection result at the current moment includes:
[0013] Extract features from the first modal data and the second modal data respectively through the feature extractor to obtain the first feature and the second feature;
[0014] Fuse the first feature and the second feature through the multi-modal fusion module to obtain the fused feature;
[0015] Perform object detection on the fused feature through the object detection module to obtain the object detection result at the current moment.
[0016] Further, the matching of the object trajectory at the previous moment and the object detection result at the current moment, and updating the object trajectory at the previous moment according to the matching result includes:
[0017] Predict the trajectory of the object trajectory at the previous moment at the current moment;
[0018] Perform intra-class matching on the predicted trajectory and the object detection result at the current moment according to the category information of the object to obtain the first matching result;
[0019] Determine whether there is an object trajectory at the previous moment that fails to match successfully in the predicted trajectory according to the first matching result;
[0020] If there is an object trajectory at the previous moment that fails to match successfully, perform inter-class matching on the predicted trajectory and the object detection results of other categories to obtain the second matching result;
[0021] Update the target trajectory at the previous moment according to the first matching result and the second matching result, and process the target trajectory at the previous moment that fails to match successfully and the target detection result at the current moment to obtain the final target trajectory at the current moment.
[0022] Further, both the first matching result and the second matching result are obtained through the IoU matching algorithm and / or the area matching algorithm.
[0023] Further, the updating the target trajectory at the previous moment according to the first matching result and the second matching result to obtain the final target trajectory at the current moment includes:
[0024] For the target trajectory at the previous moment that matches successfully and the target detection result at the current moment, use the target detection result at the current moment to update the matching target trajectory at the previous moment;
[0025] For the target trajectory at the previous moment that fails to match successfully, determine whether to retain the target trajectory at the previous moment that fails to match successfully according to the state of the target trajectory at the previous moment that fails to match successfully;
[0026] For the target detection result at the current moment that fails to match successfully, determine whether to initialize it as a new target trajectory according to the confidence level in the target detection result at the current moment that fails to match successfully.
[0027] Further, the target detection model is the NanoDet-Plus-m model and is accelerated for inference by TensorRT, the first modality data is depth data, and the second modality data is infrared data.
[0028] In a second aspect, the present invention also provides a real-time multi-object tracking device based on multi-modal representation, including:
[0029] A target detection module, configured to obtain the first modality data and the second modality data at the current moment and input them into the target detection model to obtain the target detection result at the current moment, where the target detection model is a lightweight network model that can adaptively fuse multi-modal features, and the target detection result includes the position, category, and confidence information of the target;
[0030] A trajectory matching module, configured to perform intra-class matching and inter-class matching on the target trajectory at the previous moment and the target detection result at the current moment, and update the target trajectory at the previous moment according to the matching result to obtain the target trajectory at the current moment
[0031] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the above-described method is implemented.
[0032] In a fourth aspect, the present invention further provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-described method.
[0033] In a fifth aspect, the present invention further provides a computer program product, including a computer program, characterized in that when the computer program is executed by a processor, the above-described method is implemented.
[0034] Advantages of the present invention: The real-time multi-object tracking method based on multi-modal representation provided by the present invention fuses the first feature and the second feature through a multi-modal fusion module to obtain a fused feature, which can effectively fuse features of multiple modalities with a small amount of computation and improve the accuracy of object detection. And a data matching based on hierarchical matching from easy to difficult is designed, breaking the thinking limitation of intra-class matching and proposing inter-class matching, which can alleviate the impact caused by misdetection of the object detection model; at the same time, area matching is also proposed on the basis of IoU matching, effectively dealing with the problem that the trajectory and the object cannot be associated due to the violent movement of the object. Therefore, the lightweight object detection model that can adaptively fuse multi-modal features combined with intra-class matching and inter-class matching can meet the requirements of real-time performance and accuracy at the same time and is well applied to various actual application scenarios. In addition, it should be emphasized that the combination of IoU matching and area matching can better handle the violent movement of the object in the previous and subsequent frames, alleviate the problem of object disappearance caused by the inability to match in IoU matching, and reduce trajectory fragmentation.
[0035] Additional aspects and advantages of the present invention will be given in part in the following description, which will become apparent from the following description or be understood through the practice of the present invention. Description of the Drawings
[0036] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0037] Figure 1 It is one of the flow diagrams of the real-time multi-object tracking method based on multi-modal representation provided by the embodiments of the present invention;
[0038] Figure 2It is the second flowchart of the real-time multi-object tracking method based on multi-modal representation provided by the embodiments of the present invention;
[0039] Figure 3 It is the third flowchart of the real-time multi-object tracking method based on multi-modal representation provided by the embodiments of the present invention. Detailed implementation manners
[0040] The following details the implementation manners of the present invention. Examples of the implementation manners are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The implementation manners described below by referring to the drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation to the present invention.
[0041] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an" and "the" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present invention means the presence of features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The phrase "and / or" used herein includes any unit and all combinations of one or more related listed items.
[0042] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the field to which the present invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as here.
[0043] For the convenience of understanding the embodiments of the present invention, the following will further explain with several specific embodiments as examples in conjunction with the drawings, and each embodiment does not constitute a limitation to the embodiments of the present invention.
[0044] Embodiment 1
[0045] See Figures 1 to 2 , a real-time multi-object tracking method based on multi-modal representation, includes the following steps:
[0046] S101. Obtain the first-modal data and the second-modal data at the current moment and input them into the target tracking model to obtain the target detection results (position, category, confidence) at the current moment. The target detection model is trained based on the first-modal training data and the second-modal training data.
[0047] Among them, the target detection model is a lightweight network model and is accelerated for inference through TensorRT. Specifically, it can be the NanoDet-Plus-m model. The first-modal data is depth data, and the second-modal data is infrared data. The target detection model includes a feature extractor, a multi-modal fusion module, and a target detection module.
[0048] Specifically, the input into the target detection model to obtain the target detection results at the current moment includes:
[0049] Use the feature extractor to perform feature extraction on the first-modal data and the second-modal data respectively to obtain the first feature and the second feature correspondingly. Schematically, use the first-layer convolutional neural network in the NanoDet-Plus-m model to perform feature extraction on the first-modal data and the second-modal data respectively.
[0050] Use the multi-modal fusion module to fuse the first feature and the second feature to obtain the fused feature. Specifically, the multi-modal fusion module uses two simple convolutional networks to fuse the first feature and the second feature together to obtain their initial fused feature, and then uses a simple convolutional network to adaptively combine their initial fused features, so that richer fused features can be obtained based on a small amount of computational effort.
[0051] Among them, the two simple convolutional networks used in the initial fusion process include: the first one refers to directly performing element-wise multiplication on the first feature and the second feature; the second one refers to passing the first feature and the second feature through a 2D convolutional network layer respectively, and then performing element-wise multiplication on the two obtained features. The simple convolutional network for adaptive fusion includes two parts: the first part includes two adaptively learned fusion factors, which respectively represent the influence proportions of the two initial fused features in the final fused feature, and perform a weighted summation operation on the two initial fused features based on the learned fusion factors; the second part passes the fused feature obtained in the first part through a 2D convolutional layer, a normalization layer, and a ReLU activation layer to obtain the final fused feature.
[0052] Use the target detection module to perform target detection on the fused feature to obtain the position, category, and confidence information of the target.
[0053] S102. Perform intra-class and inter-class matching between the target detection result at the previous moment and the target trajectory at the current moment, and update the target trajectory at the previous moment according to the matching result to obtain the target trajectory at the current moment.
[0054] Among them, intra-class matching refers to matching the target trajectory at the previous moment with the target detection result at the current moment of the same category, and inter-class matching refers to matching the target trajectory at the previous moment with the target detection results of other categories at the current moment.
[0055] This step specifically includes the following sub-steps:
[0056] S1021. Predict the trajectory of the target at the previous moment at the current moment. Specifically, it can be predicted through Kalman filter series algorithms.
[0057] S1022. Perform intra-class matching between the predicted trajectory and the target detection result at the current moment according to the category information of the target to obtain the first matching result. That is, match the predicted trajectory with the target detection results of the same category at the current moment to obtain the first matching result.
[0058] S1023. Determine whether there are trajectories that have not been successfully matched in the predicted trajectory according to the first matching result.
[0059] S1024. If there are trajectories that have not been successfully matched, perform inter-class matching between the predicted trajectory and the targets of other categories to obtain the second matching result. That is, associate the predicted trajectory with the detection frames of other categories that have not been matched to obtain the second matching result.
[0060] S1025. Update the target trajectory at the previous moment according to the first matching result and the second matching result to obtain the final target trajectory at the current moment, including:
[0061] For the target trajectory at the previous moment and the target detection result at the current moment that are matched, use the target detection result at the current moment to update the corresponding target trajectory at the previous moment.
[0062] For the target trajectory at the previous moment that has not been matched, judge whether to retain the target trajectory according to its state. For example, if the number of consecutive frames lost by the target trajectory at the previous moment does not reach the threshold, retain the trajectory, otherwise discard it.
[0063] For the target detection result at the current moment that has not been matched, if the confidence level is greater than the specified threshold, initialize it as a new trajectory, otherwise discard it.
[0064] Both the above-mentioned first matching result and the second matching result are obtained through the IoU matching algorithm and / or the area matching algorithm. Among them, the IoU matching algorithm matches the trajectory box and the detection box based on IoU, and if it exceeds a certain threshold, it is considered a match. The area matching algorithm determines whether there is a match based on the area inclusion relationship between the trajectory box and the detection box. For example, if the intersection area of the trajectory box and the detection box reaches a specified proportion of the area of a certain box and satisfies the specified IoU threshold relationship, it is considered to reach the threshold of area matching, and thus the two are considered to match.
[0065] The real-time multi-object tracking method based on multi-modal representation provided by the embodiments of the present invention fuses the first feature and the second feature through a multi-modal fusion module to obtain a fused feature. It can effectively fuse the features of multiple modalities with a small amount of calculation, improving the accuracy of object detection. And a data matching based on hierarchical matching from easy to difficult is designed, breaking the thinking limitation of intra-class matching, and inter-class matching is proposed, which can alleviate the impact caused by misdetection of the object tracking model. At the same time, area matching is proposed based on IoU matching, effectively dealing with the problem that the trajectory and the object cannot be associated due to the violent movement of the object. Therefore, the combination of the lightweight object tracking model and intra-class matching and inter-class matching can simultaneously meet the requirements of real-time performance and accuracy, and is well applied to various actual application scenarios. In addition, it should be emphasized that the combination of IoU matching and area matching can better handle the violent movement of the object in the previous and next few frames, alleviate the problem of object disappearance caused by the inability to match by IoU matching, and reduce trajectory fragmentation.
[0066] In some embodiments of the present invention, as Figure 3 shown, a real-time multi-object tracking method based on multi-modal representation includes the following steps:
[0067] S1. Based on the input multi-modal data, obtain multi-modal features.
[0068] Specifically, the video frames of two modalities (including the depth map video and the infrared map video) at time t are input into the feature extractor to extract the first feature and the second feature. Then, the first feature and the second feature are input into the multi-modal fusion module to obtain a more informative fused feature.
[0069] S2. Use the obtained fused feature to detect the objects in the current video frame.
[0070] Specifically, the fused feature is input into the object detection module to detect the objects at time t, including the category, location, and confidence of the objects.
[0071] S3. Match the objects detected in the current video frame with the trajectories at time t - 1, and update the object trajectories (data association).
[0072] Specifically, the target trajectory updated at time t-1 is associated with the targets detected at time t, including intra-class matching and inter-class matching, and both matching stages include IoU matching and area matching. The input for intra-class matching is the target trajectory updated at time t-1 and the targets detected at time t. The input for inter-class matching is the trajectories that were not matched in the intra-class matching stage and the detection boxes that were not matched, thereby improving the robustness to false detections of the detector and reducing the ID jumping situation to achieve data association. With the help of inter-class matching, the trajectories that were not matched in the intra-class matching stage can be matched with the detection results that were not matched in another category, which helps to utilize the results of possible misclassifications of the target detector and improve the robustness to detection errors of the target detector.
[0073] S4. The above steps S1 - S3 are repeatedly executed until the tracking results of the entire video frame are obtained.
[0074] In the stage from feature extraction to target detection, it is used to extract multi-modal features and detect the targets in the video frame. The parameters of these networks only need to be trained once and only need to be initialized once.
[0075] Embodiment 2
[0076] Based on Embodiment 1, this Embodiment 2 provides a real-time multi-object tracking device based on multi-modal representation. This real-time multi-object tracking device based on multi-modal representation corresponds to the above real-time multi-object tracking method based on multi-modal representation, and specifically includes:
[0077] A target detection module, which is used to obtain the first-modal data and the second-modal data at the current moment and input them into the target detection model to obtain the target detection results at the current moment. The target detection model is a lightweight network model that can adaptively fuse multi-modal features. The target detection results include the position, category, and confidence information of the targets;
[0078] A trajectory matching module, which is used to perform intra-class matching and inter-class matching between the target trajectory at the previous moment and the target detection results at the current moment, and update the target trajectory at the previous moment according to the matching results to obtain the target trajectory at the current moment.
[0079] For specific details, refer to the description in the part of the real-time multi-object tracking method based on multi-modal representation, which will not be elaborated here.
[0080] Embodiment 3
[0081] Embodiment 3 of the present invention provides an electronic device, including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor. The processor calls the program instructions to execute a real-time multi-object tracking method based on multi-modal representation. The method includes the following process steps:
[0082] Obtain the first-modal data and the second-modal data at the current moment and input them into the object detection model to obtain the object detection result at the current moment. The object detection model is a lightweight network model that can adaptively fuse multi-modal features. The object detection result includes the position, category, and confidence information of the object;
[0083] Perform intra-class matching and inter-class matching on the object trajectory at the previous moment and the object detection result at the current moment, and update the object trajectory at the previous moment according to the matching result to obtain the object trajectory at the current moment.
[0084] Embodiment 4
[0085] Embodiment 4 of the present invention provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements a real-time multi-object tracking method based on multi-modal representation. The method includes the following process steps:
[0086] Obtain the first-modal data and the second-modal data at the current moment and input them into the object detection model to obtain the object detection result at the current moment. The object detection model is a lightweight network model that can adaptively fuse multi-modal features. The object detection result includes the position, category, and confidence information of the object;
[0087] Perform intra-class matching and inter-class matching on the object trajectory at the previous moment and the object detection result at the current moment, and update the object trajectory at the previous moment according to the matching result to obtain the object trajectory at the current moment.
[0088] Embodiment 5
[0089] Embodiment 5 of the present invention provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements a real-time multi-object tracking method based on multi-modal representation. The method includes the following process steps:
[0090] Obtain the first-modal data and the second-modal data at the current moment and input them into the object detection model to obtain the object detection result at the current moment. The object detection model is a lightweight network model that can adaptively fuse multi-modal features. The object detection result includes the position, category, and confidence information of the object;
[0091] Perform intra-class matching and inter-class matching on the target trajectory at the previous moment and the target detection result at the current moment, and update the target trajectory at the previous moment according to the matching result to obtain the target trajectory at the current moment.
[0092] Those of ordinary skill in the art can understand that: The drawings are only schematic diagrams of an embodiment, and the modules or processes in the drawings are not necessarily essential for implementing the present invention.
[0093] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for method or system embodiments, since they are basically similar to method embodiments, they are described relatively simply. For related parts, refer to the partial description of the method embodiments. The method and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0094] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A real-time multi-target tracking method based on multimodal representation, characterized in that: include: Acquire the first modality data and the second modality data at the current moment and input them into the target detection model to obtain the target detection result at the current moment, wherein the target detection model is a lightweight network model that can adaptively fuse multimodal features, and the target detection result includes the location, category and confidence information of the target; The target trajectory at the previous moment is matched with the target detection result at the current moment by intra-class matching and inter-class matching, and the target trajectory at the previous moment is updated according to the matching result to obtain the target trajectory at the current moment.
2. The method according to claim 1, characterized in that: The target detection model includes a feature extractor, a multimodal fusion module and a target detection module; The input is fed into the target detection model to obtain the target detection result at the current moment, including: Performing feature extraction on the first modal data and the second modal data respectively by the feature extractor to obtain first features and second features accordingly; fusing the first feature and the second feature through the multimodal fusion module to obtain a fused feature; The target detection module performs target detection on the fused features to obtain the target detection result at the current moment.
3. The method according to claim 1, characterized in that The step of matching the target trajectory at the previous moment with the target detection result at the current moment, and updating the target trajectory at the previous moment according to the matching result, comprises: Predicting the trajectory of the target trajectory at the previous moment at the current moment; According to the category information of the target, the predicted trajectory is matched with the target detection result at the current moment to obtain a first matching result; Determining, according to the first matching result, whether there is a target trajectory at a previous moment that has not been successfully matched in the predicted trajectory; If there is a target trajectory at the previous moment that has not been successfully matched, the predicted trajectory is matched with the target detection results of other categories to obtain a second matching result; The target trajectory at the previous moment is updated according to the first matching result and the second matching result, and the target trajectory at the previous moment that was not successfully matched and the target detection result at the current moment are processed to obtain a final target trajectory at the current moment.
4. The method according to claim 3, characterized in that The first matching result and the second matching result are both obtained through an IoU matching algorithm and / or an area matching algorithm.
5. The method according to claim 3, characterized in that: The updating of the target trajectory at the previous moment according to the first matching result and the second matching result to obtain the final target trajectory at the current moment includes: For the successfully matched target trajectory at the previous moment and the target detection result at the current moment, using the target detection result at the current moment to update the matched target trajectory at the previous moment; For the target track at the last moment that was not successfully matched, determining whether to retain the target track at the last moment that was not successfully matched according to the state of the target track at the last moment that was not successfully matched; For the target detection result at the current moment that has not been successfully matched, it is determined whether to initialize it as a new target trajectory according to the confidence level in the target detection result at the current moment that has not been successfully matched.
6. The method according to any one of claims 1 to 5, characterized in that: The target detection model is a NanoDet-Plus-m model and the reasoning is accelerated by TensorRT. The first modality data is depth data, and the second modality data is infrared data.
7. A real-time multi-target tracking device based on multimodal representation, characterized in that: include: A target detection module is used to obtain the first modality data and the second modality data at the current moment and input them into the target detection model to obtain the target detection result at the current moment. The target detection model is a lightweight network model that can adaptively fuse multimodal features. The target detection result includes the location, category and confidence information of the target; The trajectory matching module is used to perform intra-class matching and inter-class matching on the target trajectory at the previous moment and the target detection result at the current moment, and update the target trajectory at the previous moment according to the matching result to obtain the target trajectory at the current moment.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.