Target tracking method and device based on multi-modal image, equipment, medium and product

By fusing image features of different imaging modes in the object detection model, the problem of low object detection and tracking accuracy in a single imaging mode is solved, and higher object detection and tracking accuracy and robustness are achieved.

CN119992267APending Publication Date: 2025-05-13BEIJING GUOYAN RONGXING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510011205.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the accuracy of the results of object detection and tracking is low, mainly due to the limitations of the information carried by images in a single imaging mode.

Method used

The multimodal image-based object tracking method is adopted, and images of different imaging modes (such as visible light images and infrared images) are acquired, and they are input into the trained object detection model to perform feature extraction and fusion to improve the accuracy of object detection and tracking.

Benefits of technology

Through the feature fusion of multimodal images, information from different modes can be made more fully utilized, the accuracy of object detection and tracking results can be improved, and the robustness of the system can be enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992267A_ABST
    Figure CN119992267A_ABST
Patent Text Reader

Abstract

The invention provides a target tracking method and device based on a multi-modal image, equipment, a medium and a product. The method comprises the following steps: acquiring a first image and a second image; respectively inputting the first image and the second image into a first feature extraction module and a second feature extraction module of a trained target detection model to obtain a first initial feature and a second initial feature; inputting the first initial feature and the second initial feature into a feature fusion module of a target detection model to obtain a fusion feature, and inputting the fusion feature into a detection module of the target detection model to obtain at least one target detection result; updating the tracking trajectory of the at least one target based on the target detection result; the feature fusion module comprises N convolution processing modules and N + 1 fusion modules, and each fusion module comprises a PAFPN module. According to the method, information from different modes can be fully utilized, and the accuracy of target detection and tracking results is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a target tracking method, device, equipment, medium and product based on multimodal images. Background Art

[0002] Target detection and tracking have important applications in intelligent monitoring, autonomous driving, military strikes, etc. However, the information carried by images obtained using a single imaging mode is limited. For example, a visible light camera can only capture the optical features of a scene in the visible light band, resulting in reduced accuracy of target detection and tracking results. Summary of the invention

[0003] The present invention provides a target tracking method, device, equipment, medium and product based on multimodal images, so as to solve the defect of low accuracy of target detection and tracking results in the prior art and improve the accuracy of target detection and tracking results.

[0004] The present invention provides a target tracking method based on multimodal images, comprising: Acquire a first image and a second image, wherein the first image and the second image are images respectively obtained by using different imaging modes for the same scene at the same time; Inputting the first image and the second image into a first feature extraction module and a second feature extraction module of a trained object detection model respectively to obtain a first initial feature and a second initial feature; Inputting the first initial feature and the second initial feature into a feature fusion module of the target detection model to obtain a fused feature, and inputting the fused feature into a detection module of the target detection model to obtain at least one target detection result, wherein the target detection result includes a target detection frame of the target and a confidence level of the target detection frame; Based on the target detection result, updating a tracking track of at least one of the targets, wherein the tracking track includes the target detection frame in the first image of the target at each time; Among them, the feature fusion module includes N convolution processing modules and N+1 fusion modules, the fusion module includes a PAFPN module, the input of the nth fusion module is the output of the n-1th convolution processing module and the input of the n-1th fusion module, and the input of the first fusion module is the first initial feature and the second initial feature.

[0005] According to a target tracking method based on multimodal images provided by the present invention, the N+1th fusion module further includes an SPPF module and a PPA module after the PAFPN module; through the PPA module, the following steps are performed on the input features input to the PPA module: Inputting the input features into the parallel local convolution branch, the global convolution branch and the serial convolution branch respectively, to obtain three output features outputted by the three branches respectively; Fusing the three output features to obtain a fused output feature; Processing the fused output features sequentially by a one-dimensional channel attention map and a two-dimensional spatial attention map to obtain processed output features; Rectified linear processing and batch normalization operations are performed on the processed output features to obtain the output of the PPA module.

[0006] According to a target tracking method based on multimodal images provided by the present invention, the first N fusion modules also include a convolution module and a C2f module after the PAFPN module, and the convolution processing module includes a convolution module and a C2f module.

[0007] According to a target tracking method based on multimodal images provided by the present invention, the fusion feature includes N feature images of different scales; the first initial feature and the second initial feature are input into the fusion module of the target detection model to obtain the fusion feature, including: Acquire each first intermediate feature image output by the 2nd to Nth fusion modules; Obtaining a second intermediate feature image output by the N+1th fusion module; The fusion feature is obtained based on each of the first intermediate feature images and the second intermediate feature images.

[0008] According to a target tracking method based on multimodal images provided by the present invention, updating the tracking trajectory of at least one target based on the target detection result includes: Based on the confidence in the target detection result at the target moment, classify the target detection frame in the target detection result at the target moment to obtain a first detection frame and a second detection frame, wherein the confidence of the first detection frame is higher than that of the second detection frame; Matching the first detection frame with the existing tracking tracks respectively, and when there is a tracking track matching the first detection frame, adding the first detection frame to the matching tracking track; Matching the second detection frame with the unmatched tracking track, and when there is a tracking track that matches the second detection frame, adding the second detection frame to the matched tracking track, and deleting the second detection frame that does not match the tracking track; The first detection frame that is not matched to the tracking trajectory is saved as the new tracking trajectory.

[0009] According to a target tracking method based on multimodal images provided by the present invention, matching the first detection frame with the existing tracking track respectively includes: Performing Kalman filtering on the tracking trajectory to obtain an estimation result of the target corresponding to the tracking trajectory at the target time; Matching the first detection frame with the estimation result.

[0010] The present invention also provides a target tracking device based on multimodal images, comprising: An image acquisition module, used to acquire a first image and a second image, wherein the first image and the second image are images obtained by using different imaging modes for the same scene at the same time; A feature extraction module, used to input the first image and the second image into a first feature extraction module and a second feature extraction module of a trained object detection model, respectively, to obtain a first initial feature and a second initial feature; A target detection module, used for inputting the first initial feature and the second initial feature into a feature fusion module of the target detection model to obtain a fused feature, and inputting the fused feature into a detection module of the target detection model to obtain at least one target detection result, wherein the target detection result includes a target detection frame of the target and a confidence level of the target detection frame; A tracking module, configured to update a tracking track of at least one of the targets based on the target detection result, wherein the tracking track includes the target detection frame in the first image of the target at each time; Among them, the feature fusion module includes N convolution processing modules and N+1 fusion modules, the fusion module includes a PAFPN module, the input of the nth fusion module is the output of the n-1th convolution processing module and the input of the n-1th fusion module, and the input of the first fusion module is the first initial feature and the second initial feature.

[0011] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, any of the above-mentioned target tracking methods based on multimodal images is implemented.

[0012] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the computer program implements any of the above-mentioned target tracking methods based on multimodal images.

[0013] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned target tracking methods based on multimodal images.

[0014] The target tracking method, device, equipment, medium and product based on multimodal images provided by the present invention obtain a first image and a second image obtained by using different imaging modes for the same scene at the same time, input the first image and the second image into a target detection model, and after extracting the first initial feature and the second initial feature respectively in the target detection model, the feature information of the features of the two modal images at adjacent levels is fused through a fusion module in the target detection model, thereby making fuller use of information from different modalities and improving the accuracy of target detection and tracking results. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0016] Figure 1 It is a flowchart of the target tracking method based on multimodal images provided by the present invention.

[0017] Figure 2 It is a structural schematic diagram of a target detection model in a target tracking method based on multimodal images provided by the present invention.

[0018] Figure 3 It is a structural schematic diagram of the PAFPN module in the target detection model in the target tracking method based on multimodal images provided by the present invention.

[0019] Figure 4 It is a structural schematic diagram of the PPA module in the target detection model in the target tracking method based on multimodal images provided by the present invention.

[0020] Figure 5 It is a schematic diagram of the target tracking process in the target tracking method based on multimodal images provided by the present invention.

[0021] Figure 6 It is a structural schematic diagram of a target tracking device based on multimodal images provided by the present invention.

[0022] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0024] Combine the following Figure 1-Figure 5 The target tracking method based on multimodal images provided by the present invention is described as follows: Figure 1 As shown, the method comprises the steps of: S110, acquiring a first image and a second image, where the first image and the second image are images obtained at the same time using different imaging modes for the same scene; S120, inputting the first image and the second image into a first feature extraction module and a second feature extraction module of a trained object detection model respectively to obtain a first initial feature and a second initial feature; S130, inputting the first initial feature and the second initial feature into a feature fusion module of a target detection model to obtain a fused feature, and inputting the fused feature into a detection module of the target detection model to obtain at least one target detection result, wherein the target detection result includes a target detection frame of the target and a confidence level of the target detection frame; S140: Based on the target detection result, update the tracking trajectory of at least one target, where the tracking trajectory includes the target detection box in the first image at each time.

[0025] Among them, the feature fusion module includes N convolution processing modules and N+1 fusion modules, the fusion module includes a PAFPN module, the input of the nth fusion module is the output of the n-1th convolution processing module and the input of the n-1th fusion module, and the input of the 1st fusion module is the first initial feature and the second initial feature.

[0026] In a possible implementation of the method provided by the present invention, the first image may be a visible light image, and the second image may be an infrared image. Visible light images intelligently capture the optical features of a scene in the visible light band, and infrared imaging devices have the characteristic of being less susceptible to changes in illumination. However, compared with visible light images, infrared images have only one color channel, and less feature information can be extracted. In addition, infrared images often have problems such as low resolution, blurred object edges, noise, and contrast. By comprehensively utilizing the advantages of visible light images and infrared images, target recognition that fuses infrared images and visible light images is beneficial to improving target detection and tracking performance. It can be understood that in the method provided by the present invention, the first image and the second image are not limited to visible light images and infrared images, but may also be images presented by other two imaging modes. The following is a specific description using the first image as a visible light image and the second image as an infrared image as an example.

[0027] In the method provided by the present invention, after the first image and the second image are acquired, they are input into a trained target detection model, and the target detection model is a model trained using multiple sets of training data, each set of training data includes a sample first image, a sample second image, and a target detection label. After training, the target detection model can output a target detection result based on the input first image and the second image. The target detection result includes a target detection frame and a confidence level of the target detection frame. The target detection frame is the area in the image that includes the target.

[0028] In the method provided by the present invention, the target detection model includes a first feature extraction module and a second feature extraction module, which are respectively used to perform initial feature extraction on the first image and the second image. Figure 4 As shown, in a possible implementation, the first feature extraction module and the second feature extraction module may include two Conv modules (convolution modules) and one C2f module, and both the Conv module and the C2f module are modules in the Yolov8 model (a model for target detection).

[0029] After performing initial feature extraction on the first image to obtain the first initial feature, and performing initial feature extraction on the second image to obtain the second initial feature, the first initial feature and the second initial feature are input into the feature fusion module. The feature fusion module includes a fusion module, and the fusion module includes a PAFPN (Path Aggregation Feature Pyramid Network) module, which is a module for fusing network features. PAFPN is improved based on the original FPN (Feature Pyramid Network) structure. It adds a downsampling module to downsample the high-level feature map so that it has the same size as the low-level feature map, so as to facilitate subsequent feature fusion; secondly, PAFPN introduces an additional 3×3 convolution in the feature fusion process, and further extracts and optimizes the fused feature map to eliminate the redundant information that may be introduced by downsampling fusion. Through 3×3 convolution, PAFPN can better integrate feature information at different levels and improve the feature representation capability. The structure of the PAFPN module is as shown in the figure. Figure 2 shown.

[0030] Furthermore, the first N fusion modules also include a convolution module and a C2f module after the PAFPN module, and the convolution processing module includes the convolution module and the C2f module.

[0031] The method provided by the present invention constructs two parallel feature extraction links, and introduces a PAFPN module after each C2f module to fuse the feature information of infrared image features and visible light RGB image features at adjacent levels. This mid-term fusion strategy can make more full use of information from different modalities, enhance the overall information of the multimodal target detection system, provide richer and more accurate data support for subsequent analysis steps, and improve the accuracy of target detection results.

[0032] Furthermore, in a possible implementation, the N+1th fusion module further includes an SPPF module and a PPA module after the PAFPN module; through the PPA module, the following steps are performed on the input features input to the PPA module: Input the input features to the parallel local convolution branch, global convolution branch and serial convolution branch respectively, and obtain three output features output by the three branches respectively; The three output features are fused to obtain fused output features; The fused output features are processed by the one-dimensional channel attention map and the two-dimensional spatial attention map in sequence to obtain the processed output features; Rectified linear processing and batch normalization operations are performed on the processed output features to obtain the output of the PPA module.

[0033] In this implementation of the method provided by the present invention, a PPA (Parallelized patch-aware attention) module is added after the SPPF (Spatial Pyramid Pooling Fast) module. The PPA module captures features of different scales and levels through parallel local branches, global branches, and serial convolution branches, so that the model can retain key local and global information, effectively reducing the important information that may be lost in the process of multiple downsampling of small infrared targets, thereby improving the accuracy and robustness of detection. The structure of the PPA module is as follows: Figure 3 shown.

[0034] For a given input feature tensor After the local branch, global branch and serial convolution branch, we can get , , Combining these three results, we get , that is: .

[0035] After feature extraction through multi-branch feature extraction, the attention mechanism is used for adaptive feature enhancement. The attention mechanism consists of a series of adaptive channel attention and spatial attention. The channel attention mechanism focuses on the importance of different feature channels. In this case Sequentially from the one-dimensional channel attention map and a 2D spatial attention map Carry out corresponding processing, this process can be summarized as follows: ; ; ; in, represents element-wise multiplication, and represents the characteristics after channel and space selection, Respectively represent the rectified linear unit (ReLU) and batch normalization (BN), It is the final output of the PPA module in the present invention.

[0036] The fusion feature includes N feature images of different scales; the first initial feature and the second initial feature are input into the fusion module of the target detection model to obtain the fusion feature, including: Obtaining each first intermediate feature image output by the 2nd to Nth fusion modules; Obtain the second intermediate feature image output by the N+1th fusion module; A fusion feature is obtained based on each of the first intermediate feature images and the second intermediate feature images.

[0037] In one possible implementation, Figure 4 As shown, N=3, based on each first intermediate feature image and the second intermediate feature image, a fusion feature is obtained, including: After upsampling the second intermediate feature image, the second intermediate feature image is spliced ​​with the second first intermediate feature image to obtain a first first spliced ​​feature image, the first spliced ​​feature image is input into the C2f module and then spliced ​​with the first convolution feature image to obtain a first second spliced ​​feature image, the first second spliced ​​feature image is input into the C2f module to obtain the second feature image in the fusion feature; After upsampling the first first spliced ​​feature image, it is spliced ​​with the first first intermediate feature image to obtain a second first spliced ​​feature image. After the second first spliced ​​feature image is input into the C2f module, the first feature image in the fusion feature is obtained. Inputting the second and first concatenated feature images into a convolution module to obtain a first convolution feature image; The second feature image in the fusion feature is input into the convolution module and then spliced ​​with the second intermediate feature image to obtain the second second spliced ​​feature image. The second second spliced ​​feature image is input into the C2f module to obtain the third feature image in the fusion feature.

[0038] After extracting the fusion features of the three scales, the feature maps of the three scales are passed through the detection module to obtain the target detection result. The target detection result includes the target detection box and the confidence of the target detection box. The detection module may include a Neck network layer and a Detect network layer.

[0039] After obtaining the target detection result, target tracking is performed based on the target detection result. The purpose of target tracking is to obtain a tracking trajectory of the target. The tracking trajectory of the target includes the target detection frame of the target in the first image at different times. Specifically, based on the target detection result, updating the tracking trajectory of at least one target includes: Based on the confidence in the target detection result at the target moment, classify the target detection frame in the target detection result at the target moment to obtain a first detection frame and a second detection frame, wherein the confidence of the first detection frame is higher than that of the second detection frame; Matching the first detection frame with the existing tracking tracks respectively, and when there is a tracking track matching the first detection frame, adding the first detection frame to the matching tracking track; Matching the second detection frame with the unmatched tracking track, and when there is a tracking track matching the second detection frame, adding the second detection frame to the matched tracking track, and deleting the second detection frame that is not matched to the tracking track; The first detection frame that is not matched to the tracking track is saved as a new tracking track.

[0040] When tracking begins, the target detection frame in the first image of the first picture can be used as the initial state of the tracking trajectory corresponding to each target. Then, based on the first image and the second image at each preset moment (the preset moment is the imaging moment of the first image), target detection is performed to obtain the target detection frame, and the tracking trajectory of the target is updated based on the newly generated target detection frame.

[0041] like Figure 5 As shown, for the newly generated target detection frame at the target moment, compared with the method of only performing data association on the high-confidence detection frame in other tracking methods, the method provided by the present invention retains almost all the target detection frames and performs identity matching. Specifically, first, all the target detection frames at the target moment are classified based on confidence, and are divided into two groups: high confidence (first detection frame) and low confidence (second detection frame). In other words, the target detection frame and the corresponding confidence score are obtained, and the target detection frame is classified. If the score is higher than the preset first threshold T_high, the target detection frame is classified into the high confidence group (first detection frame). When the score is lower than T_high and higher than the preset second threshold T_low, the target detection frame is classified into the low confidence group (second detection frame). For the detection frame below the second threshold T_low, it can be directly deleted.

[0042] First, the first detection frame with higher confidence is associated with the track, and then the second detection frame with lower confidence is associated with the unmatched tracking track to retain the second detection frame and filter the background.

[0043] Before matching and associating the target detection frame with the tracking trajectory, the state of the target at the target time is first predicted based on the existing tracking trajectory, and matching is performed based on the prediction result, that is, matching the first detection frame with the existing tracking trajectory respectively, including: Perform Kalman filtering on the tracking trajectory to obtain the estimated result of the target corresponding to the tracking trajectory at the target time; Match the first detection box with the estimated result.

[0044] Kalman filtering is performed on the target detection frame before the target moment in the tracking trajectory to estimate the target detection frame of the target in the tracking trajectory at the target moment, and an estimation result is obtained. According to the similarity between the first detection frame and the estimation result, the Hungarian algorithm is used for matching to obtain the matching result between the first detection frame and each existing tracking trajectory. The matching result includes two types: match and mismatch. The tracking trajectory corresponding to the estimation result with a matching degree with the first detection frame higher than a preset degree threshold and the highest matching degree can be used as the tracking trajectory matched with the first detection frame, and the first detection frame is added to the matched tracking trajectory, and the first detection frames that are not matched to the tracking trajectory and the tracking trajectory that are not matched to the detection frame are retained.

[0045] Then a second round of matching is performed. For the tracking trajectories that are not matched to the first detection frame in the first round, they are matched with the second detection frame. Similarly, when the second detection frame is matched with the tracking trajectory, it is also matched with the estimated result corresponding to the tracking trajectory. When there is a tracking trajectory that matches the second detection frame, the second detection frame is added to the matched tracking trajectory, and the second detection frames that are not matched to the tracking trajectory are deleted, because these second detection frames are determined to be not the background containing any objects.

[0046] The first detection boxes that are not matched to the tracking track are saved as new tracking tracks, because these detection boxes are likely to correspond to newly appeared targets.

[0047] The method provided by the present invention not only considers high-confidence detection frames but also reasonably considers low-confidence detection frames when tracking a target, and improves the accuracy and robustness of tracking through two rounds of matching processes.

[0048] The target tracking device based on multimodal images provided by the present invention is described below. The target tracking device based on multimodal images described below and the target tracking method based on multimodal images described above can be referred to each other. Figure 6 As shown, the target tracking device based on multimodal images provided by the present invention includes: An image acquisition module 610 is used to acquire a first image and a second image, where the first image and the second image are images obtained at the same time using different imaging modes for the same scene; A feature extraction module 620, used to input the first image and the second image into a first feature extraction module and a second feature extraction module of a trained object detection model, respectively, to obtain a first initial feature and a second initial feature; The target detection module 630 is used to input the first initial feature and the second initial feature into a feature fusion module of the target detection model to obtain a fused feature, and input the fused feature into a detection module of the target detection model to obtain at least one target detection result, wherein the target detection result includes a target detection frame of the target and a confidence level of the target detection frame; A tracking module 640, configured to update a tracking track of at least one target based on the target detection result, wherein the tracking track includes a target detection box of the target in the first image at each time point; Among them, the feature fusion module includes N convolution processing modules and N+1 fusion modules, the fusion module includes a PAFPN module, the input of the nth fusion module is the output of the n-1th convolution processing module and the input of the n-1th fusion module, and the input of the 1st fusion module is the first initial feature and the second initial feature.

[0049] Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7 As shown, the electronic device may include: a processor (processor) 710, a communication interface (Communications Interface) 720, a memory (memory) 730 and a communication bus 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other through the communication bus 740. The processor 710 can call the logic instructions in the memory 730 to execute a target tracking method based on a multimodal image, the method comprising: acquiring a first image and a second image, the first image and the second image being images obtained by using different imaging modes for the same scene at the same time; inputting the first image and the second image into a first feature extraction module and a second feature extraction module of a trained target detection model, respectively, to obtain a first initial feature and a second initial feature; inputting the first initial feature and the second initial feature into a feature fusion module of the target detection model to obtain a fusion feature, inputting the fusion feature into a detection module of the target detection model to obtain at least one target detection result, the target detection result including a target detection frame of the target and a confidence of the target detection frame; based on the target detection result, updating a tracking trajectory of at least one target, the tracking trajectory including a target detection frame of the target in the first image at each time; wherein the feature fusion module includes N convolution processing modules and N+1 fusion modules, the fusion module includes a PAFPN module, the input of the nth fusion module is the output of the n-1th convolution processing module and the input of the n-1th fusion module, and the input of the first fusion module is the first initial feature and the second initial feature.

[0050] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0051] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the target tracking method based on multimodal images provided by the above methods, and the method includes: acquiring a first image and a second image, wherein the first image and the second image are images obtained by using different imaging modes for the same scene at the same time; inputting the first image and the second image into a first feature extraction module and a second feature extraction module of a trained target detection model, respectively, to obtain a first initial feature and a second initial feature; inputting the first initial feature and the second initial feature into the target detection model In the feature fusion module of the target detection model, a fusion feature is obtained, and the fusion feature is input into the detection module of the target detection model to obtain at least one target detection result, wherein the target detection result includes a target detection frame of the target and a confidence level of the target detection frame; based on the target detection result, a tracking trajectory of at least one target is updated, wherein the tracking trajectory includes a target detection frame of the target in the first image at each moment; wherein the feature fusion module includes N convolution processing modules and N+1 fusion modules, the fusion module includes a PAFPN module, the input of the nth fusion module is the output of the n-1th convolution processing module and the input of the n-1th fusion module, and the input of the first fusion module is the first initial feature and the second initial feature.

[0052] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the target tracking method based on multimodal images provided by the above-mentioned methods, the method comprising: acquiring a first image and a second image, the first image and the second image being images obtained by adopting different imaging modes for the same scene at the same time; inputting the first image and the second image into a first feature extraction module and a second feature extraction module of a trained target detection model, respectively, to obtain a first initial feature and a second initial feature; inputting the first initial feature and the second initial feature into a feature fusion module of the target detection model, to obtain to fusion features, input the fusion features into the detection module of the target detection model to obtain at least one target detection result, the target detection result includes the target detection box of the target and the confidence of the target detection box; based on the target detection result, update the tracking trajectory of at least one target, the tracking trajectory includes the target detection box of the target in the first image at each moment; wherein the feature fusion module includes N convolution processing modules and N+1 fusion modules, the fusion module includes a PAFPN module, the input of the nth fusion module is the output of the n-1th convolution processing module and the input of the n-1th fusion module, and the input of the first fusion module is the first initial feature and the second initial feature.

[0053] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0054] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A target tracking method based on multimodal images, characterized in that: include: Acquire a first image and a second image, wherein the first image and the second image are images respectively obtained by using different imaging modes for the same scene at the same time; Inputting the first image and the second image into a first feature extraction module and a second feature extraction module of a trained object detection model respectively to obtain a first initial feature and a second initial feature; Inputting the first initial feature and the second initial feature into a feature fusion module of the target detection model to obtain a fused feature, and inputting the fused feature into a detection module of the target detection model to obtain at least one target detection result, wherein the target detection result includes a target detection frame of the target and a confidence level of the target detection frame; Based on the target detection result, updating a tracking track of at least one of the targets, wherein the tracking track includes the target detection frame in the first image of the target at each time; Among them, the feature fusion module includes N convolution processing modules and N+1 fusion modules, the fusion module includes a PAFPN module, the input of the nth fusion module is the output of the n-1th convolution processing module and the input of the n-1th fusion module, and the input of the first fusion module is the first initial feature and the second initial feature.

2. The target tracking method based on multimodal images according to claim 1, characterized in that: The N+1th fusion module further includes an SPPF module and a PPA module after the PAFPN module; through the PPA module, the following steps are performed on the input features input to the PPA module: Inputting the input features into the parallel local convolution branch, the global convolution branch and the serial convolution branch respectively, to obtain three output features outputted by the three branches respectively; Fusing the three output features to obtain a fused output feature; Processing the fused output features sequentially by a one-dimensional channel attention map and a two-dimensional spatial attention map to obtain processed output features; Rectified linear processing and batch normalization operations are performed on the processed output features to obtain the output of the PPA module.

3. The target tracking method based on multimodal images according to claim 1, characterized in that: The first N fusion modules also include a convolution module and a C2f module after the PAFPN module, and the convolution processing module includes a convolution module and a C2f module.

4. The target tracking method based on multimodal images according to claim 1, characterized in that: The fusion feature includes N feature images of different scales; the first initial feature and the second initial feature are input into the fusion module of the target detection model to obtain the fusion feature, including: Acquire each first intermediate feature image output by the 2nd to Nth fusion modules; Obtaining a second intermediate feature image output by the N+1th fusion module; The fusion feature is obtained based on each of the first intermediate feature images and the second intermediate feature images.

5. The target tracking method based on multimodal images according to claim 1, characterized in that: The updating of the tracking trajectory of at least one of the targets based on the target detection result includes: Based on the confidence in the target detection result at the target moment, classify the target detection frame in the target detection result at the target moment to obtain a first detection frame and a second detection frame, wherein the confidence of the first detection frame is higher than that of the second detection frame; Matching the first detection frame with the existing tracking tracks respectively, and when there is a tracking track matching the first detection frame, adding the first detection frame to the matching tracking track; Matching the second detection frame with the unmatched tracking track, and when there is a tracking track that matches the second detection frame, adding the second detection frame to the matched tracking track, and deleting the second detection frame that does not match the tracking track; The first detection frame that is not matched to the tracking trajectory is saved as the new tracking trajectory.

6. The target tracking method based on multimodal images according to claim 5, characterized in that: The matching the first detection frame with the existing tracking track respectively includes: Performing Kalman filtering on the tracking trajectory to obtain an estimation result of the target corresponding to the tracking trajectory at the target time; Matching the first detection frame with the estimation result.

7. A target tracking device based on multimodal images, characterized in that: The device comprises: An image acquisition module, used to acquire a first image and a second image, wherein the first image and the second image are images obtained by using different imaging modes for the same scene at the same time; A feature extraction module, used to input the first image and the second image into a first feature extraction module and a second feature extraction module of a trained object detection model, respectively, to obtain a first initial feature and a second initial feature; A target detection module, used for inputting the first initial feature and the second initial feature into a feature fusion module of the target detection model to obtain a fused feature, and inputting the fused feature into a detection module of the target detection model to obtain at least one target detection result, wherein the target detection result includes a target detection frame of the target and a confidence level of the target detection frame; A tracking module, configured to update a tracking track of at least one of the targets based on the target detection result, wherein the tracking track includes the target detection frame in the first image of the target at each time; Among them, the feature fusion module includes N convolution processing modules and N+1 fusion modules, the fusion module includes a PAFPN module, the input of the nth fusion module is the output of the n-1th convolution processing module and the input of the n-1th fusion module, and the input of the first fusion module is the first initial feature and the second initial feature.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the target tracking method based on multimodal images as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the target tracking method based on multimodal images as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the target tracking method based on multimodal images as described in any one of claims 1 to 6 is implemented.