Target detection method and system combined with multi-modal visual tracking
By combining multimodal visual tracking technology with visible light and infrared images, the problem of low accuracy and efficiency of single-technology detection is solved, and efficient target detection is achieved under complex backgrounds and high-definition image compression.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-31
- Publication Date
- 2026-03-31
AI Technical Summary
When using a single technology for target detection, existing technologies suffer from low detection accuracy and low detection efficiency, especially in complex background environments and high-definition image compression situations.
Multimodal visual tracking technology is employed, combining visible light and infrared images. Target video data is acquired through a multi-functional camera device, and the multimodal tracking results are fused. Matching and detection are then performed using a target detection database, thereby improving detection accuracy and efficiency.
By combining multimodal visual tracking technology, the accuracy and efficiency of target detection are improved, enabling effective target identification in complex backgrounds and under high-definition image compression conditions.
Smart Images

Figure CN116993772B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of target detection technology, and specifically to a target detection method and system that combines multimodal visual tracking. Background Technology
[0002] Currently, target detection algorithms based on visible light images have limited generalization capabilities. When the target background environment is complex, the algorithm's detection ability cannot meet the actual task requirements. Furthermore, due to platform computing power limitations, visible light images compress high-definition input images, causing the target's feature information to disappear after compression, making it difficult for the detection algorithm to effectively identify and detect the target. Infrared image-based target detection algorithms rely on infrared images to reflect the target's temperature and the scene's thermal radiation information. Due to the performance limitations of infrared sensors, infrared images have lower pixel resolution compared to visible light images. Currently, to reflect detailed target information, confirmation through visible light images is still necessary, leading to reduced detection efficiency.
[0003] In summary, existing technologies suffer from low accuracy and efficiency in target detection due to the use of a single technology. Summary of the Invention
[0004] This disclosure provides a target detection method and system that combines multimodal visual tracking to solve the technical problems in the prior art where the target detection accuracy and efficiency are low due to the use of a single technology.
[0005] According to a first aspect of this disclosure, a target detection method combining multimodal visual tracking is provided, comprising: obtaining a target detection database based on historical target detection data; connecting a multi-functional camera device to acquire target video data, and inputting the target video data into a dual-channel multimodal visual tracking system to obtain multimodal tracking results; fusing the multimodal tracking results to obtain a first fusion result; matching multiple targets in the first fusion result with the target detection database to determine a detection target; and detecting the detection target to obtain a target detection result.
[0006] According to a second aspect of this disclosure, a target detection system combining multimodal visual tracking is provided, comprising: a target detection database acquisition module, configured to acquire a target detection database based on historical target detection data; a multimodal tracking result acquisition module, configured to connect to a multi-functional camera device to acquire target video data and input the target video data into a dual-channel multimodal visual tracking system to obtain multimodal tracking results; a first fusion result acquisition module, configured to fuse the multimodal tracking results to obtain a first fusion result; a target detection acquisition module, configured to match multiple targets in the first fusion result with the target detection database to determine a target to be detected; and a target detection result acquisition module, configured to detect the target to obtain a target detection result.
[0007] According to a third aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method of the first aspect.
[0008] One or more technical solutions provided in this application have at least the following technical effects or advantages: According to this disclosure, a target detection database is obtained based on historical target detection data; target video data is acquired by connecting a multi-functional camera device, and the target video data is input into a multi-modal visual tracking dual-channel to obtain multi-modal tracking results; the multi-modal tracking results are fused to obtain a first fusion result; multiple targets in the first fusion result are matched with the target detection database to determine the detection target; the detection target is detected to obtain a target detection result. This method can combine multi-modal visual tracking technology for target detection, thereby improving detection accuracy and efficiency.
[0009] It should be understood that the description in this section is not intended to highlight key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0011] Figure 1 A schematic flowchart illustrating the target detection method combining multimodal visual tracking provided in this embodiment of the disclosure;
[0012] Figure 2 This is a schematic diagram of the process for obtaining multimodal tracking results using a target detection method combined with multimodal visual tracking, as described in an embodiment of this disclosure.
[0013] Figure 3 This is a schematic diagram of the process for obtaining the first fusion result using a target detection method combined with multimodal visual tracking, as described in an embodiment of this disclosure.
[0014] Figure 4 This is a schematic diagram of the structure of a target detection system combining multimodal visual tracking provided in an embodiment of the present disclosure;
[0015] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure.
[0016] Explanation of reference numerals in the attached figures: Target detection database acquisition module 11, multimodal tracking result acquisition module 12, first fusion result acquisition module 13, target detection acquisition module 14, target detection result acquisition module 15, electronic device 600, processor 601, memory 602, bus 603. Detailed Implementation
[0017] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0018] To address the technical problems of low accuracy and efficiency in target detection due to the use of a single technology in existing technologies, the inventors of this disclosure, through inventive effort, have obtained the target detection method and system that combines multimodal visual tracking:
[0019] Example 1
[0020] Figure 1 The target detection method combining multimodal visual tracking provided in this application includes:
[0021] Step S100: Obtain the target detection database based on historical target detection data;
[0022] Specifically, based on big data, historical object detection data is retrieved using historical object detection data as an index. This involves training an object detection database using this historical data until the loss data of the target database converges, saving the trained database, and then outputting the final object detection database.
[0023] Step S200: Connect the multi-functional camera device to acquire target video data, and input the target video data into the multi-modal visual tracking dual channel to obtain multi-modal tracking results respectively;
[0024] Specifically, a multi-functional camera device is connected to acquire target video data, which is the video data to be detected. For example, the target video data can be a video stream or a video file. Further, the target video data is input into a multi-modal visual tracking dual-channel system, where each channel can acquire different channel output results, which are then used as the multi-modal tracking results.
[0025] Step S300: Fuse the multimodal tracking results to obtain a first fusion result;
[0026] Specifically, multiple results from the multimodal tracking are fused along image edges to obtain a first fusion result. Further, the first fusion result is registered, specifically its grayscale values, to obtain a first registered image. Further, a fusion image model is constructed, comprising an input layer, a visible light bottom-level image acquisition layer, an infrared bottom-level image acquisition layer, a preliminary fusion layer, a reconstructed fusion layer, and an output layer. The visible light bottom-level image acquisition layer and the infrared bottom-level image acquisition layer are parallel layers. Further, the first registered image is input into the fusion image model, and a visible light bottom-level image is obtained through the visible light bottom-level image acquisition layer, and an infrared bottom-level image is obtained through the infrared bottom-level image acquisition layer. Further, the infrared bottom-level image and the visible light bottom-level image are combined to obtain a first combined result. Further, the first combined result undergoes line reconstruction to generate a redrawn line image, obtaining a reconstructed combined image. Further, the reconstructed combined image is enhanced, specifically its color, to obtain a first enhanced image. Furthermore, pseudo-color processing is applied to the first enhanced image, that is, color restoration and color detail enhancement are performed on the first enhanced image to obtain the first optimized fusion result. The first optimized fusion result is an optimization of the first fusion result.
[0027] Step S400: Match multiple targets in the first fusion result with the target database to determine the detection target;
[0028] Specifically, multiple targets in the first fusion result are sequentially matched with data in the target database to obtain the matching degree of multiple targets. Further, a preset matching degree is obtained by carefully comparing each target with the target database to obtain the matching degree of each target. If the matching degree of a target is high, the target with the higher matching degree is selected as the target to be detected. The method for determining the detection target based on the target's matching degree is as follows: the target's matching degree is compared with the preset matching degree; target matching degree values greater than or equal to the preset matching degree are extracted; the target corresponding to the target matching degree is obtained; and thus, the detection target is determined.
[0029] Step S500: Detect the target and obtain the target detection result.
[0030] Specifically, the identified target is tracked within the target video data to obtain its continuous motion trajectory. For example, during target tracking, a bounding box can be drawn around the target. Furthermore, by combining the target identification result and the target's motion trajectory, the target detection result is obtained.
[0031] In this embodiment, multimodal visual tracking technology can be combined for target detection, thereby improving detection accuracy and efficiency.
[0032] like Figure 2 As shown, step S200 in the method provided in this application embodiment includes:
[0033] S210: Extract image frames from the target video data to obtain multiple image frames, wherein the multiple image frames have time identifiers;
[0034] S220: Obtain a first image frame based on the plurality of image frames, and input the first image frame into the multimodal visual tracking dual channel to obtain multimodal tracking results respectively, wherein the multimodal visual tracking dual channel includes an infrared thermal imaging channel and a visible light imaging channel.
[0035] Specifically, the target video data is segmented to obtain multiple image frames. Each image frame is then time-stamped to obtain its time stamp. Further, a single image frame is randomly selected to obtain a first image frame, which is then input into a multimodal visual tracking dual-channel system. This system includes an infrared thermal imaging channel and a visible light imaging channel. Further, infrared thermal imaging results are obtained through the infrared thermal imaging channel, and visible light imaging results are obtained through the visible light imaging channel. This process ultimately yields the multimodal tracking result.
[0036] By inputting the first image frame into the multimodal visual tracking dual channels, multimodal visual tracking results can be obtained separately, enabling multimodal detection of the target detection video and thus improving detection accuracy.
[0037] like Figure 3 As shown, step S300 in the method provided in this application embodiment includes:
[0038] S310: Obtain the first imaging result based on the infrared thermal imaging channel;
[0039] S320: Obtain a second imaging result based on the visible light imaging channel;
[0040] S330: Obtain a first fusion result based on the first imaging result and the second imaging result.
[0041] Specifically, image processing is performed using the infrared thermal imaging channel in the multimodal visual tracking dual-channel system to obtain an infrared thermal imaging result, which serves as the first imaging result. Further, image processing is performed using the visible light thermal imaging channel in the multimodal visual tracking dual-channel system to obtain a visible light thermal imaging result, which serves as the second imaging result. Further, the first and second imaging results are then fused along the video edges to obtain a first fusion result.
[0042] By fusing visible light imaging results and infrared imaging results, multimodal visual tracking can be performed, which helps to improve detection accuracy.
[0043] Step S300 in the method provided in this application embodiment further includes:
[0044] S340: Extract the first fusion result and perform registration to obtain the first registered image;
[0045] S350: Construct a fused image model, wherein the fused image model includes an input layer, a visible light bottom image acquisition layer, an infrared bottom image acquisition layer, a preliminary fusion layer, a reconstructed fusion layer, and an output layer;
[0046] S360: Input the first registered image into the fused image model, obtain the first registered image transmitted to the visible light bottom image acquisition layer, extract the visible light pixel nodes of the first registered image, and obtain the visible light bottom image of the first registered image based on the visible light pixel nodes;
[0047] S370: Obtain the first registration image transmitted to the infrared bottom image acquisition layer, extract the infrared pixel nodes of the first registration image, and obtain the infrared bottom image of the first registration image based on the infrared pixel nodes;
[0048] S380: The visible light underlying image and the infrared underlying image of the first registration image are initially combined to obtain a first combination result.
[0049] Specifically, the first fusion result is extracted and registered to obtain the first registered image. Specifically, the grayscale nodes of the first fusion result are extracted, the grayscale of each grayscale node is obtained, and the first fusion result is registered according to the grayscale to obtain the first registered image.
[0050] Furthermore, a fused image model is constructed, which includes an input layer, a visible light bottom-level image acquisition layer, an infrared bottom-level image acquisition layer, a preliminary fusion layer, a reconstruction fusion layer, and an output layer. Furthermore, the visible light bottom-level image acquisition layer and the infrared bottom-level image acquisition layer are parallel layers.
[0051] Furthermore, the first registered image is input into the fusion image model to obtain the first registered image transmitted to the visible light bottom image acquisition layer. The visible light bottom image acquisition layer extracts pixel nodes from the first registered image to obtain visible light pixel nodes. Specifically, the visible light pixel nodes are combined by pixel elements representing lines, shapes, or edges to obtain the visible light bottom image of the first registered image.
[0052] Further, a first registration image transmitted to the infrared bottom image acquisition layer is acquired, and pixel nodes are extracted from the first registration image to obtain the infrared pixel nodes of the first registration image. Specifically, the infrared pixel nodes are combined by pixel elements representing lines, shapes, or edges to obtain the infrared bottom image of the first registration image.
[0053] Furthermore, the visible light under-layer image and the infrared under-layer image of the first registered image are initially combined. The combination method is to combine the infrared under-layer image and the visible light under-layer image according to the image edges to obtain the first combination result.
[0054] The first fusion result is input into the fusion image model, which can perform registration, low-level image acquisition, etc. on the first fusion image, which helps to improve the image accuracy of the first fusion image, and thus improve the target detection accuracy of the detected target in the first fusion image.
[0055] Step S300 in the method provided in this application embodiment further includes:
[0056] S390: Reconstruct the first combination result to obtain a reconstructed combined image;
[0057] S3100: Perform image enhancement on the reconstructed combined image to obtain a first enhanced image;
[0058] S3110: Perform pseudo-color processing on the first enhanced image to obtain the first optimized fusion result.
[0059] Specifically, the first combined result is reconstructed, whereby the image lines, edges, etc., of the first combined result are re-outlined to generate a basic image structure, thereby obtaining a reconstructed combined image. Further, the reconstructed combined image is enhanced, whereby the enhancement method is to enhance the reconstructed combined image according to color contrast to generate a basic color image structure, thereby obtaining a first enhanced image.
[0060] Furthermore, pseudo-color processing is applied to the first enhanced image. Pseudo-color processing refers to assigning color values to grayscale values according to certain criteria. In macroscopic terms, it converts a black-and-white image into a color image, or transforms a monochrome image into an image with a given color distribution. Since the human eye's ability to distinguish color is far greater than its ability to distinguish grayscale, converting a grayscale image into a color representation can improve the ability to distinguish image details. Therefore, the main purpose of pseudo-color processing is to improve the human eye's ability to distinguish image details, thereby achieving image enhancement. In this embodiment, pseudo-color processing of the first enhanced image can intelligently fuse the basic color structures in the first enhanced image to obtain more color details for further color enhancement. Specifically, by performing pseudo-color processing on the first enhanced image, a first optimized fusion result is obtained.
[0061] In particular, by reconstructing, color enhancing, and pseudo-color processing the first combined image, the image accuracy of the first combined image can be improved, thereby improving the accuracy of target detection.
[0062] Step S400 in the method provided in this application embodiment further includes:
[0063] S410: Obtain the preset matching coefficient;
[0064] S420: Traverse each target and match it with the target detection database to obtain multiple target matching coefficients;
[0065] S430: Extract the multiple target matching coefficients and compare them sequentially with the preset matching coefficients to obtain multiple comparison results;
[0066] S440: If the comparison result is that the target matching coefficient is greater than or equal to the preset matching coefficient, the target is determined to be the detection target.
[0067] Specifically, based on big data, a search is performed using a matching coefficient as a retrieval constraint to obtain a preset matching coefficient, which is used to represent the degree of matching of the matched objects. Further, multiple targets in the first fusion result are sequentially matched with the target database, and the degree of matching between the multiple targets and the data in the target database is calculated to obtain the matching coefficients of the multiple targets.
[0068] Furthermore, multiple target matching coefficients are extracted and compared sequentially with preset matching coefficients to obtain the comparison results of the target matching coefficients and the preset matching coefficients. If the comparison result shows that the target matching coefficient is greater than or equal to the preset matching coefficient, it is determined that the aforementioned target has a high degree of matching with the data in the target detection database, and thus the target is identified as the target to be detected.
[0069] By matching the target with a target detection database to determine the detection target, the influence of non-detection targets such as the background in the image can be reduced, thereby improving the efficiency and accuracy of target detection.
[0070] Step S500 in the method provided in this application embodiment further includes:
[0071] S510: The plurality of image frames are serialized according to the time identifier to obtain serialized image frames;
[0072] S520: Detect the target based on the serialized image frame, obtain the motion trajectory of the target, and generate the target detection result.
[0073] Specifically, multiple image frames are obtained from the target video data, and the time stamps of these image frames are extracted. The image frames are then serialized according to their time stamps. Specifically, the multiple image frames are arranged sequentially according to their chronological order to obtain an image frame sequence with a timeline.
[0074] Furthermore, the image frame sequence is merged to generate merged image frames with time stamps. The detected target is then tracked based on the merged image frames to obtain its motion trajectory. Finally, the target detection result is obtained by combining the target determination result and the target's motion trajectory.
[0075] Among them, obtaining the motion trajectory of the target can complete the target detection result.
[0076] Example 2
[0077] Based on the same inventive concept as the target detection method combined with multimodal visual tracking in the aforementioned embodiments, such as Figure 4 As shown, this application also provides a target detection system combining multimodal visual tracking, the system comprising:
[0078] The target detection database acquisition module 11 is used to acquire a target detection database based on historical target detection data.
[0079] The multimodal tracking result acquisition module 12 is used to connect to a multi-functional camera device to acquire target video data, and input the target video data into the multimodal visual tracking dual channel to obtain multimodal tracking results respectively;
[0080] The first fusion result acquisition module 13 is used to fuse the multimodal tracking results to obtain a first fusion result;
[0081] The target detection module 14 is used to determine the target detection by matching multiple targets in the first fusion result with the target detection database.
[0082] The target detection result acquisition module 15 is used to detect the target and obtain the target detection result.
[0083] Furthermore, the system also includes:
[0084] An image frame acquisition module is used to extract image frames from the target video data to obtain multiple image frames, wherein the multiple image frames have time identifiers;
[0085] A dual-channel acquisition module is used to obtain a first image frame based on the plurality of image frames, and input the first image frame into the multimodal visual tracking dual channels to obtain multimodal tracking results respectively. The multimodal visual tracking dual channels include an infrared thermal imaging channel and a visible light imaging channel.
[0086] Furthermore, the system also includes:
[0087] The first imaging result acquisition module is used to obtain a first imaging result based on the infrared thermal imaging channel;
[0088] The second imaging result acquisition module is used to obtain a second imaging result based on the visible light imaging channel;
[0089] The first fusion result acquisition module is used to obtain a first fusion result based on the first imaging result and the second imaging result.
[0090] Furthermore, the system also includes:
[0091] The first registration image acquisition module is used to extract the first fusion result for registration to obtain the first registration image;
[0092] A fused image model construction module is used to construct a fused image model, wherein the fused image model includes an input layer, a visible light bottom image acquisition layer, an infrared bottom image acquisition layer, a preliminary fusion layer, a reconstructed fusion layer, and an output layer;
[0093] The visible light underlying image acquisition module is used to input the first registered image into the fused image model, acquire the first registered image transmitted to the visible light underlying image acquisition layer, extract the visible light pixel nodes of the first registered image, and acquire the visible light underlying image of the first registered image based on the visible light pixel nodes.
[0094] An infrared underlying image acquisition module is used to acquire the first registration image transmitted to the infrared underlying image acquisition layer, extract the infrared pixel nodes of the first registration image, and acquire the infrared underlying image of the first registration image based on the infrared pixel nodes.
[0095] The first combination result obtaining module is used to initially combine the visible light underlying image and the infrared underlying image of the first registered image to obtain a first combination result.
[0096] Furthermore, the system also includes:
[0097] A reconstructed combined image acquisition module is used to reconstruct the first combination result to obtain a reconstructed combined image;
[0098] A first enhanced image acquisition module is used to enhance the reconstructed combined image to obtain a first enhanced image.
[0099] The first optimized fusion result acquisition module is used to perform pseudo-color processing on the first enhanced image to obtain the first optimized fusion result.
[0100] Furthermore, the system also includes:
[0101] A serialized image frame acquisition module is used to serialize the plurality of image frames according to the time identifier to obtain serialized image frames;
[0102] The target detection result acquisition module is used to detect the target based on the serialized image frames, obtain the motion trajectory of the target, and generate the target detection result.
[0103] Furthermore, the system also includes:
[0104] A preset matching coefficient acquisition module, wherein the preset matching coefficient acquisition module is used to obtain a preset matching coefficient;
[0105] A target matching coefficient acquisition module is used to traverse each target and match it with the target detection database to obtain multiple target matching coefficients.
[0106] The comparison result acquisition module is used to extract the multiple target matching coefficients and compare them sequentially with the preset matching coefficients to obtain multiple comparison results;
[0107] A target acquisition module is configured to determine the target as the detection target if the comparison result is that the target matching coefficient is greater than or equal to the preset matching coefficient.
[0108] The specific example of the target detection method combining multimodal visual tracking in Embodiment 1 described above is also applicable to the target detection system combining multimodal visual tracking in this embodiment. Through the foregoing detailed description of the target detection method combining multimodal visual tracking, those skilled in the art can clearly understand the target detection system combining multimodal visual tracking in this embodiment. Therefore, for the sake of brevity, it will not be described in detail here. As for the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant details can be found in the method section.
[0109] Example 3
[0110] Figure 5 This is a schematic diagram based on the third embodiment of the present disclosure, as shown below. Figure 5 As shown, the electronic device 600 in this disclosure may include a processor 601 and a memory 602.
[0111] Memory 602 is used to store programs. Memory 602 may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), double data rate synchronous dynamic random-access memory (DDR SDRAM), etc.; memory may also include non-volatile memory, such as flash memory. Memory 602 is used to store computer programs (such as application programs, functional modules, etc. that implement the above methods), computer instructions, etc. The computer programs, computer instructions, etc., can be partitioned and stored in one or more memories 602. Furthermore, the computer programs, computer instructions, data, etc., can be accessed by processor 601.
[0112] The aforementioned computer programs and instructions can be stored in one or more partitions of memory 602. Furthermore, the aforementioned computer programs and instructions can be invoked by processor 601.
[0113] The processor 601 is configured to execute the computer program stored in the memory 602 to implement the various steps in the methods described in the above embodiments.
[0114] For details, please refer to the relevant descriptions in the preceding method embodiments.
[0115] The processor 601 and the memory 602 can be independent structures or integrated structures. When the processor 601 and the memory 602 are independent structures, the memory 602 and the processor 601 can be coupled together via bus 603.
[0116] The electronic device in this embodiment can execute the technical solution in the above method. Its specific implementation process and technical principle are the same, and will not be repeated here.
[0117] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0118] According to embodiments of this disclosure, this disclosure also provides a computer program product comprising: a computer program stored in a readable storage medium, at least one processor of an electronic device being able to read the computer program from the readable storage medium, and the at least one processor executing the computer program causing the electronic device to perform the scheme provided in any of the above embodiments.
[0119] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this disclosure can be achieved, and this is not limited herein.
[0120] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A target detection method combined with multi-modal visual tracking, characterized in that, The method is applied to a target detection system combined with multi-modal visual tracking, the system is in communication connection with a multifunctional camera device, and the method comprises: obtaining a target detection database according to historical target detection data; connecting the multifunctional camera device to obtain target video data, inputting the target video data into a multi-modal visual tracking double channel to obtain multi-modal tracking results respectively; fusing the multi-modal tracking results to obtain a first fusion result; matching a plurality of targets in the first fusion result with the target detection database to determine a detection target; detecting the detection target to obtain a target detection result; the connecting the multifunctional camera device to obtain target video data, inputting the target video data into a multi-modal visual tracking double channel to obtain multi-modal tracking results respectively comprises: extracting image frames from the target video data to obtain a plurality of image frames, wherein the plurality of image frames have time identifiers; obtaining a first image frame from the plurality of image frames and inputting the first image frame into the multi-modal visual tracking double channel to obtain multi-modal tracking results respectively, wherein the multi-modal visual tracking double channel comprises an infrared thermal imaging channel and a visible light imaging channel; the fusing the multi-modal tracking results to obtain a first fusion result comprises: obtaining a first imaging result according to the infrared thermal imaging channel; obtaining a second imaging result according to the visible light imaging channel; obtaining a first fusion result according to the first imaging result and the second imaging result; after the fusing the multi-modal tracking results to obtain a first fusion result, further comprising: extracting the first fusion result for registration to obtain a first registered image; constructing a fusion image model, wherein the fusion image model comprises an input layer, a visible light bottom layer image acquisition layer, an infrared bottom layer image acquisition layer, a preliminary fusion layer, a reconstructed fusion layer, and an output layer; inputting the first registered image into the fusion image model to obtain a first registered image transmitted to the visible light bottom layer image acquisition layer, extracting visible light pixel nodes of the first registered image, and obtaining a visible light bottom layer image of the first registered image according to the visible light pixel nodes; obtaining the first registered image transmitted to the infrared bottom layer image acquisition layer, extracting infrared pixel nodes of the first registered image, and obtaining an infrared bottom layer image of the first registered image according to the infrared pixel nodes; preliminarily combining the visible light bottom layer image and the infrared bottom layer image of the first registered image to obtain a first combined result.
2. The method of claim 1, wherein, after the fusing the multi-modal tracking results to obtain a first fusion result, further comprising: reconstructing the first combined result to obtain a reconstructed combined image; performing image enhancement on the reconstructed combined image to obtain a first enhanced image; performing pseudo-color processing on the first enhanced image to obtain a first optimized fusion result.
3. The method of claim 1, wherein, the detecting the detection target to obtain a target detection result comprises: serializing the plurality of image frames according to the time identifiers to obtain serialized image frames; According to the serialized image frame, the detection target is detected, a motion trajectory of the detection target is obtained, and the target detection result is generated.
4. The method of claim 1, wherein, The matching of the multiple targets in the first fusion result with the target detection database is used to determine the detection target, including: A preset matching coefficient is obtained; Each target is matched with the target detection database to obtain multiple target matching coefficients; The multiple target matching coefficients are compared with the preset matching coefficient in sequence to obtain multiple comparison results; If the comparison result is that the target matching coefficient is greater than or equal to the preset matching coefficient, the target is determined as the detection target.
5. A target detection system incorporating multi-modal visual tracking, characterized by, A system for implementing the target detection method combined with multi-modal visual tracking according to any one of claims 1-4, the system comprising: A target detection database obtaining module, which is configured to obtain a target detection database according to historical target detection data; A multi-modal tracking result obtaining module, which is configured to connect a multi-functional camera device to obtain target video data, and input the target video data into a multi-modal visual tracking double channel to obtain a multi-modal tracking result; A first fusion result obtaining module, which is configured to fuse the multi-modal tracking result to obtain a first fusion result; A detection target obtaining module, which is configured to match multiple targets in the first fusion result with the target detection database to determine a detection target; A target detection result obtaining module, which is configured to detect the detection target to obtain a target detection result. 6.An electronic device, comprising: at least one processor; a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-4.
Citation Information
Patent Citations
Unmanned aerial vehicle target detection method, device, equipment and medium
CN113283411A