Target detection method and device based on dual-optical registration fusion and unmanned aerial vehicle system
By employing a dual-light registration and fusion target detection method in mountain search and rescue, and utilizing feature extraction and fusion techniques from dual-light images, the problem of low detection accuracy of traditional sensors in mountain search and rescue was solved, achieving high-precision target recognition and positioning.
Patent Information
- Application Number
- CN202411060633.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-05
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2044-08-05
AI Technical Summary
In mountain search and rescue operations, traditional single optical sensors are difficult to apply to complex lighting conditions and vegetation obstruction, resulting in low target detection accuracy. Existing dual-light equipment requires camera calibration during the registration process, which lacks universality and accuracy.
A target detection method based on dual-light registration and fusion is adopted. By extracting texture information from visible light images and temperature information from infrared images, the viewpoint deviation is calculated for image registration and feature fusion. A dual-channel target feature extraction network and a cross-modal attention module are used to improve detection accuracy.
The registration-fusion-detection process of dual-light images has been optimized, which improves the target detection accuracy, enables rapid identification and localization in complex environments, compensates for server processing latency, and enhances target localization accuracy.
Smart Images

Figure CN118864821B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing, and particularly relates to a target detection method, device and unmanned aerial vehicle system based on dual-light registration and fusion. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Due to poor weather conditions, complex terrain, and inconvenient transportation in mountainous areas, and because traditional search and rescue efforts rely primarily on large-scale manual searches, the search and rescue operations are lengthy, extremely difficult, and risky, posing significant challenges and potentially causing missed opportunities for optimal rescue. Therefore, rapid search, precise location, and timely response are crucial for improving the success rate of search and rescue. Furthermore, the dense vegetation in mountainous areas means that missing or trapped individuals are likely to be obscured, and disappearances often occur in the evening or late at night when lighting is poor. Drones equipped with single optical sensors are ill-suited for the complex search and rescue scenarios in mountainous regions.
[0004] Existing technologies, considering the greater applicability of dual-light devices in target search, propose an infrared-visible light fusion target detection technology for UAV applications using dual-light cameras. This technology first aligns and registers infrared and visible light image information, determining the target's position in the infrared image, and then detecting the target in the visible light image based on the position information. It fully utilizes the detailed information of the visible light image and the temperature information of the infrared image. However, in practical implementation, camera calibration is required during the registration stage, lacking universality. Furthermore, the application of traditional image processing algorithms to determine the target position in the infrared image has poor accuracy, thus affecting the target detection accuracy in the visible light image. Summary of the Invention
[0005] To address the technical problems mentioned above, this invention provides a target detection method, apparatus, and UAV system based on dual-light registration and fusion, which can optimize the registration-fusion-detection process based on dual-light images and improve target detection accuracy.
[0006] To achieve the above objectives, the present invention adopts the following technical solution:
[0007] The first aspect of the present invention provides a target detection method based on dual-light registration fusion.
[0008] A target detection method based on dual-light registration fusion includes:
[0009] Acquire visible light and infrared images of mountainous areas;
[0010] Texture information from visible light images of mountainous areas and temperature information from infrared images of mountainous areas are extracted to obtain visible light feature maps and infrared feature maps;
[0011] The viewing angle deviation between the infrared feature map and the visible light feature map is calculated, and then the mountain infrared image is spatially transformed into a mountain infrared image registered with the mountain visible light image based on the viewing angle deviation.
[0012] The registered infrared feature map is extracted from the registered mountain infrared image and then fused with the visible light feature map to obtain the fused feature;
[0013] Based on the relationship between the fused features and the semantic and location information of the target, the semantic and location information of the target corresponding to the current fused features is predicted.
[0014] As one implementation method, a dual-channel target feature extraction network is used to extract visible light feature maps and infrared feature maps.
[0015] In one implementation, each channel of the dual-channel target feature extraction network contains at least two layers, and the number of convolutional kernel channels in each layer increases progressively from top to bottom; the layers are processed by max pooling layers, so that the feature mapping size of the input of the next layer is reduced to a set ratio of the output of the previous layer.
[0016] As one implementation, a cross-modal attention module is used to extract visible light and infrared features separately from the dual-channel network, and then obtain the associated features of the two through convolution and dot product operations, which are used to calculate the viewing angle deviation between the infrared feature map and the visible light feature map.
[0017]
[0018] In the formula, It is an infrared feature map The index at a specific position in the text; Visible light characteristic map The index at a specific position in the text; Visible light feature map Features and locations at all locations Infrared feature map on The result of normalized and weighted summarization of features; the univariate function g is used to calculate the visible light feature map. In position The representation value of information; a binary function Calculate the visible light feature map using an embedded Gaussian function. Middle position Location information and feature mapping infrared feature map Middle position Relevance of information; Infrared feature maps exist Feature mapping points and visible light feature maps at the location exist Feature mapping points at; Visible light feature map any point above Infrared feature map Points on Correspondingly, the calculation yields .
[0019] A second aspect of the present invention provides a target detection device based on dual-light registration and fusion.
[0020] A target detection device based on dual-light registration and fusion includes:
[0021] Image acquisition module, which is used to acquire visible light images and infrared images of mountainous areas;
[0022] The feature extraction module is used to extract the texture information of the visible light image of the mountainous area and the temperature information of the infrared image of the mountainous area, respectively, to obtain the visible light feature map and the infrared feature map;
[0023] The image registration module is used to calculate the viewing angle deviation between the infrared feature map and the visible light feature map, and then, based on the viewing angle deviation, spatially transforms the mountain infrared image into a mountain infrared image registered with the mountain visible light image.
[0024] The feature fusion module is used to extract the registered infrared feature map from the registered mountain infrared image and then fuse it with the visible light feature map to obtain the fused feature.
[0025] The target prediction module is used to predict the semantic and location information of the target corresponding to the current fused feature based on the relationship between the fused features and the semantic and location information of the target.
[0026] As one implementation, in the feature extraction module, a dual-channel target feature extraction network is used to extract visible light feature maps and infrared feature maps.
[0027] In one implementation, in the feature extraction module, each channel of the dual-channel target feature extraction network contains at least two layers, and the number of convolutional kernel channels in each layer increases progressively from top to bottom; the layers are processed by max pooling layers, so that the feature mapping size of the input of the next layer is reduced to a set ratio of the output of the previous layer.
[0028] As one implementation, in the image registration module, a cross-modal attention module is used to extract visible light features and infrared features separately by the dual-channel network, and obtain the associated features of the two through convolution and dot multiplication operations, so as to calculate the viewing angle deviation between the infrared feature map and the visible light feature map.
[0029]
[0030] In the formula, It is an infrared feature map The index at a specific position in the text; Visible light characteristic map The index at a specific position in the text; Visible light feature map Features and locations at all locations Infrared feature map on The result of normalized and weighted summarization of features; the univariate function g is used to calculate the visible light feature map. In position The representation value of information; a binary function Calculate the visible light feature map using an embedded Gaussian function. Middle position Location information and feature mapping infrared feature map Middle position Relevance of information; Infrared feature maps exist Feature mapping points and visible light feature maps at the location exist Feature mapping points at; Visible light feature map any point above Infrared feature map Points on Correspondingly, the calculation yields .
[0031] A third aspect of the present invention provides an unmanned aerial vehicle (UAV) system.
[0032] An unmanned aerial vehicle (UAV) system includes an UAV terminal and a ground control center;
[0033] The UAV terminal and the ground control center communicate with each other through an ad-hoc network communication module.
[0034] The drone terminal includes a drone platform, an optoelectronic pod, and an AI (Artificial Intelligence) computing unit;
[0035] The optoelectronic pod is equipped with an image acquisition module, which is used to acquire visible light images and infrared images of the mountainous area and transmit them to the AI computing unit.
[0036] The AI computing unit includes a dual-light registration fusion detection module; the dual-light registration fusion detection module is configured to execute the steps in the target detection method based on dual-light registration fusion as described above, and to predict the semantic and location information of the target based on the image transmitted by the image acquisition module and transmit it to the UAV platform.
[0037] In one implementation, the UAV platform is equipped with a flight control unit and a high-precision positioning module; the flight control unit is used to receive and execute flight tasks uploaded by the ground control center; the high-precision positioning module is used for positioning and measuring the UAV's position coordinates; the high-precision positioning module is communicatively connected to an AI computing unit for calculating the target position.
[0038] Compared with the prior art, the beneficial effects of the present invention are:
[0039] (1) This invention utilizes the viewpoint deviation between the feature maps of visible light images and infrared images of mountainous areas, and performs spatial transformation to register the infrared images and visible light images of mountainous areas. It then uses the corresponding feature map of the registered infrared images of mountainous areas to fuse with the feature map of the visible light images of mountainous areas. Finally, it uses the relationship between the fused features and the target to predict the semantic and location information of the target corresponding to the current fused features. This optimizes the registration-fusion-detection method based on dual-light images and improves the target detection accuracy in visible light images.
[0040] (2) The present invention is equipped with an AI computing unit on the UAV terminal. The dual-light registration and fusion detection module in the AI computing unit is used to process and analyze the visible light and infrared dual-modal data collected by the photoelectric pod, identify and track the set target in the field of view, make up for the time delay caused by the back-end processing of the server, and improve the target positioning accuracy.
[0041] (3) The present invention has a high-precision positioning module installed inside the drone platform on the drone terminal. This can be combined with the AI computing unit to calculate the position coordinates of the target by the position and attitude of the drone and the pod, thus realizing the rapid identification and positioning of the target.
[0042] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0043] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0044] Figure 1 This is a flowchart of the target detection method based on dual-light registration and fusion according to an embodiment of the present invention;
[0045] Figure 2 This is a schematic diagram of the target detection principle based on dual-light registration and fusion according to an embodiment of the present invention;
[0046] Figure 3 This is a schematic diagram of the cross-modal attention module according to an embodiment of the present invention;
[0047] Figure 4 This is a schematic diagram illustrating the working principle of the spatial transformation layer in an embodiment of the present invention.
[0048] Figure 5 This is a schematic diagram of the target detection device based on dual-light registration and fusion according to an embodiment of the present invention;
[0049] Figure 6 This is a schematic diagram of the unmanned aerial vehicle system structure according to an embodiment of the present invention;
[0050] Figure 7 This is a schematic diagram of the structure of the drone terminal according to an embodiment of the present invention;
[0051] Figure 8 This is a schematic diagram of the missing persons identification and location calculation structure according to an embodiment of the present invention. Detailed Implementation
[0052] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0053] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0054] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0055] Figure 1 This is a flowchart of a target detection method based on dual-light registration and fusion according to an embodiment of the present invention. Figure 1This invention provides a target detection method based on dual-light registration fusion, comprising:
[0056] S101: Acquire visible light and infrared images of mountainous areas;
[0057] S102: Extract texture information from visible light images of mountainous areas and temperature information from infrared images of mountainous areas to obtain visible light feature maps and infrared feature maps;
[0058] S103: Calculate the viewing angle deviation between the infrared feature map and the visible light feature map, and then spatially transform the mountain infrared image into a mountain infrared image registered with the mountain visible light image based on the viewing angle deviation.
[0059] S104: Extract the registered infrared feature map from the registered mountain infrared image, and then fuse it with the visible light feature map to obtain the fused feature;
[0060] S105: Based on the relationship between the fused features and the semantic and location information of the target, predict the semantic and location information of the target corresponding to the current fused features.
[0061] In step S102, a dual-channel target feature extraction network is used to extract visible light feature maps and infrared feature maps.
[0062] In the dual-channel target feature extraction network, each channel contains at least two layers (e.g., three layers), and the number of convolutional kernel channels in each layer increases progressively from top to bottom. The layers are processed by max pooling layers (e.g., 2 × 2 max pooling layers) so that the feature mapping size of the input of the next layer is reduced to a set ratio (e.g., 1 / 2) of the output of the previous layer.
[0063] Specifically, each layer contains a double convolution process, consisting of two 3 × 3 convolution kernels and a non-linear activation unit (ReLU).
[0064] As one implementation, a cross-modal attention module is used to extract visible light and infrared features separately from the dual-channel network, and then obtain the associated features of the two through convolution and dot product operations, which are used to calculate the viewing angle deviation between the infrared feature map and the visible light feature map.
[0065]
[0066] In the formula, It is an infrared feature map The index at a specific position in the text; Visible light characteristic map The index at a specific position in the text; Visible light feature map Features and locations at all locations Infrared feature map on The result of normalized and weighted summarization of features; the univariate function g is used to calculate the visible light feature map. In position The representation value of information; a binary function Calculate the visible light feature map using an embedded Gaussian function. Middle position Location information and feature mapping infrared feature map Middle position Relevance of information; Infrared feature maps exist Feature mapping points and visible light feature maps at the location exist Feature mapping points at; Visible light feature map any point above Infrared feature map Points on Correspondingly, the calculation yields .
[0067] In step S103: the viewing angle deviation between the infrared feature map and the visible light feature map is characterized by a multi-parameter offset matrix.
[0068] In the cross-modal feature registration stage, a cross-modal attention module is first introduced into each feature extraction channel. Then, after two layers of convolution and pooling, a fully connected operation is performed to regress and predict an 8-parameter offset matrix. The infrared image is then input into the STN layer as a spatial transformation parameter, and the registered infrared image can be obtained after the infrared image is input into the STN.
[0069] The Spatial Transformation Layer (STN) mainly consists of a grid generator and a sampler. It reflects the spatial mapping relationship between image pixels, which can be expressed by the following formula:
[0070]
[0071] in, offset matrix The homography matrix obtained by the direct linear transformation algorithm has nine parameters that can be used to represent rotation bias, translation bias, and tilt bias, respectively, to perform point-by-point registration and correction on the input image. These are the pixel coordinates of key points on the original infrared image. These are the coordinates of the key point after the infrared image and the visible light image have been registered through an affine transformation operation.
[0072] In the cross-modal feature fusion and target recognition stage, the feature extraction process of visible light images is omitted, which effectively reduces the redundancy of the model and improves the computation speed of the model. At the same time, a multimodal attention fusion mechanism is introduced, which fuses the downsampled multi-scale feature maps of the dual-light images separately, enhances the information exchange between the visible light and infrared channels and the enhancement of multi-scale features, and ensures the preservation of information in the fused image. This effectively improves the accuracy of identifying missing persons under vegetation cover and low light conditions, and reduces the occurrence of missed or false detections.
[0073] In this embodiment, the cross-modal feature fusion and target recognition stage is based on a feature pyramid network for multi-scale target detection and a multi-modal attention fusion mechanism. The feature pyramid network can be roughly divided into three parts: downsampling, upsampling, and intermediate connections. The multi-modal attention fusion mechanism is applied in the downsampling stage to fuse visible light and infrared feature maps, ultimately predicting the semantic and positional information of the target through regression. The multi-modal attention fusion mechanism is a dual-input residual network that takes the visible light and infrared feature maps from the same layer in the downsampling of the feature pyramid network as input. After six convolutional layers, the features of both are fused. The fused features are then horizontally connected to the corresponding downsampling layer in the feature pyramid network to identify the target's position and semantics. The downsampled feature map undergoes a 1×1 convolutional operation to change the number of channels, making it the same as the number of channels in the upsampling feature map of the corresponding layer, and then pixel-wise addition is performed between them.
[0074] Figure 5 This is a schematic diagram of the target detection device based on dual-light registration and fusion according to an embodiment of the present invention. Figure 5 As shown, this embodiment of the invention provides a target detection device based on dual-light registration and fusion, comprising:
[0075] Image acquisition module 501 is used to acquire visible light images and infrared images of mountainous areas;
[0076] The feature extraction module 502 is used to extract the texture information of the visible light image of the mountain area and the temperature information of the infrared image of the mountain area, respectively, to obtain the visible light feature map and the infrared feature map;
[0077] The image registration module 503 is used to calculate the viewing angle deviation between the infrared feature map and the visible light feature map, and then, based on the viewing angle deviation, spatially transforms the mountain infrared image into a mountain infrared image registered with the mountain visible light image.
[0078] The feature fusion module 504 is used to extract the registered infrared feature map from the registered mountain infrared image and then fuse it with the visible light feature map to obtain the fused feature.
[0079] The target prediction module 505 is used to predict the semantic and location information of the target corresponding to the current fused feature based on the relationship between the fused features and the semantic and location information of the target.
[0080] Specifically, in the feature extraction module 502, a dual-channel target feature extraction network is used to extract visible light feature maps and infrared feature maps.
[0081] In the specific implementation process, in the feature extraction module 502, each channel of the dual-channel target feature extraction network contains at least two layers, and the number of convolutional kernel channels in each layer increases progressively from top to bottom; the layers are processed by max pooling layers so that the feature mapping size of the input of the next layer is reduced to a set ratio of the output of the previous layer.
[0082] In the image registration module 503, the cross-modal attention module is used to extract the visible light features and infrared features separately by the dual-channel network, and obtain the associated features of the two through convolution and dot multiplication operations, so as to calculate the viewing angle deviation between the infrared feature map and the visible light feature map.
[0083]
[0084] In the formula, It is an infrared feature map The index at a specific position in the text; Visible light characteristic map The index at a specific position in the text; Visible light feature map Features and locations at all locations Infrared feature map on The result of normalized and weighted summarization of features; the univariate function g is used to calculate the visible light feature map. In position The representation value of information; a binary function Calculate the visible light feature map using an embedded Gaussian function. Middle position Location information and feature mapping infrared feature map Middle position Relevance of information; Infrared feature maps exist Feature mapping points and visible light feature maps at the location exist Feature mapping points at; Visible light feature map any point above Infrared feature map Points on Correspondingly, the calculation yields .
[0085] It should be noted that each module in the target detection device based on dual-light registration and fusion in this embodiment corresponds one-to-one with each step in the target detection method based on dual-light registration and fusion described above, and their specific implementation processes are the same, so they will not be described in detail here.
[0086] according to Figure 6 It also provides an unmanned aerial vehicle (UAV) system, including the UAV terminal and a ground control center;
[0087] The UAV terminal and the ground control center communicate with each other through the self-organizing network communication module 2;
[0088] The drone terminal includes a drone platform 1, an optoelectronic pod 3, and an AI computing unit 4, such as Figure 7 As shown;
[0089] The photoelectric pod 3 is equipped with an image acquisition module, which is used to acquire visible light images and infrared images of the mountainous area and transmit them to the AI computing unit.
[0090] The AI computing unit 4 includes a dual-light registration fusion detection module; the dual-light registration fusion detection module is configured to perform the above-described... Figure 1 The steps in the target detection method based on dual-light registration fusion are used to predict the semantic and location information of the target based on the image transmitted by the image acquisition module and transmit it to the UAV platform.
[0091] Specifically, the optoelectronic pod is mounted on the bottom of the drone platform, and the AI computing unit is fixed on the top of the drone platform. Both the optoelectronic pod and the AI computing unit are powered by the drone platform's power supply.
[0092] In this embodiment, the UAV platform is equipped with a flight control unit and a high-precision positioning module; the flight control unit is used to receive and execute flight tasks uploaded by the ground control center; the high-precision positioning module is used for positioning and measuring the UAV's position coordinates; the high-precision positioning module is communicatively connected to the AI computing unit for calculating the target position.
[0093] In some embodiments, in addition to an image acquisition module (including a high-definition visible light camera and an infrared thermal imager) within the electro-optical pod, the pod also houses a laser rangefinder, a gyroscope sensor, and a pod drive module, all electrically connected to the AI computing unit. The high-definition visible light camera and infrared thermal imager are used to capture visible light and infrared images from the UAV's current perspective, and are closely arranged. The laser rangefinder is used to obtain the straight-line distance between the pod and the target. The gyroscope sensor is used to sense the pod's attitude information. The pod drive module drives the electro-optical pod's pitch and translation. The ground control center acquires the raw video data, ranging information, and pod attitude through the self-organizing network communication module on the ground end, and drives the electro-optical pod's movement by sending commands.
[0094] In addition to the dual-light registration and fusion detection module mentioned above, the AI computing unit also includes a pod control module and a position calculation module.
[0095] The pod control module is used to: when the target recognition module detects a missing person, send instructions to the pod drive module inside the optoelectronic pod to align the center of the optoelectronic pod's field of view with the target through pitch and translation, thereby achieving the purpose of tracking. The position calculation module calculates the position coordinates of the identified target by processing the UAV attitude and position coordinates sent by the high-precision positioning module, the straight-line distance to the target obtained by the laser rangefinder, and the pod attitude information sensed by the gyroscope sensor.
[0096] In some specific embodiments, the AI computing unit may adopt a quad-core processor architecture, a neural network processing unit with 8 TOPS computing power, and integrate numerous high-speed communication interfaces. It can simultaneously access multiple sensor signals and perform deep learning and image processing algorithm operations. The lightweight and integration level is improved, the interface is rich and the scalability is strong, which indirectly improves the search endurance of the drone.
[0097] The ground control center includes a search and rescue command platform and a ground-based ad hoc network communication module. The ground-based ad hoc network communication module communicates with the UAVs. The search and rescue command platform's functional modules include a UAV interaction module, a pod interaction module, and an AI function module. The UAV interaction module is used to plan flight routes and set flight parameters for the UAVs, and monitor the UAV's flight data in real time. The pod interaction module controls the attitude of the electro-optical pod, retrieves raw video images acquired by visible light cameras and infrared thermal imagers, and monitors the pod's attitude information in real time. The AI function module controls the activation and deactivation of AI detection functions and receives target identification results and location coordinates. Detailed implementation method:
[0099] After receiving the alarm and understanding the approximate location of the missing person, the search and rescue command platform at the ground control center plans the search route of the drone and sets important parameters such as the drone's flight altitude and speed before starting the search mission. The drone will then cruise along the predetermined route.
[0100] Once the search mission begins, the search and rescue command platform is activated. The electro-optical pod mounted on the drone automatically adjusts its pitch angle to a set angle, such as 45°, while the visible light camera and infrared thermal imager tilt to capture video images of the ground. On one hand, the electro-optical pod is electrically connected to the onboard self-organizing network communication module, allowing the raw video images to be transmitted to the search and rescue command platform for real-time viewing. On the other hand, the electro-optical pod is electrically connected to the AI computing unit, and the dual-modal images of visible light and infrared light within the same field of view are used by the AI computing unit's dual-light registration and fusion detection module to detect missing persons in real-time. Once a missing person is detected, the identification result is transmitted to the search and rescue command platform via the self-organizing network communication module. Simultaneously, the pod control module controls the pitch and translation angles of the electro-optical pod to keep the target centered in the field of view, achieving real-time target tracking. Finally, the position calculation module calculates the specific coordinates of the identified target by combining the drone's azimuth, attitude, and coordinates, as well as the azimuth, attitude, and distance information from the laser rangefinder. This coordinate information is simultaneously transmitted to the search and rescue command platform, notifying search and rescue personnel to proceed to the coordinates for rescue operations.
[0101] In one or more embodiments, a computer device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the target detection method based on dual-light registration fusion as described above.
[0102] Specifically, according to embodiments of this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including functions for executing... Figure 1 The program code for the method shown. In such an embodiment, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by the central processing unit, it performs the various functions defined in the apparatus of this application.
[0103] in, Figure 1 The computer program instructions corresponding to the method shown may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in the process. Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0104] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0105] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A target detection method based on dual-light registration and fusion, characterized in that, include: Acquire visible light and infrared images of mountainous areas; Texture information from visible light images of mountainous areas and temperature information from infrared images of mountainous areas are extracted to obtain visible light feature maps and infrared feature maps; By utilizing a cross-modal attention module, the visible light and infrared features extracted separately by the dual-channel network are obtained by convolution and dot product operations to obtain the associated features between the two, which are then used to calculate the viewing angle deviation between the infrared feature map and the visible light feature map. In the formula, It is an infrared feature map The index at a specific position in the text; Visible light characteristic map The index at a specific position in the text; Visible light feature map Features and locations at all locations Infrared feature map on The result of normalized and weighted summarization of features; the univariate function g is used to calculate the visible light feature map. In position The representation value of information; a binary function Calculate the visible light feature map using an embedded Gaussian function. Middle position Location information and feature mapping infrared feature map Middle position Relevance of information; Infrared feature maps exist Feature mapping points and visible light feature maps at the location exist Feature mapping points at; Visible light feature map any point above Infrared feature map Points on Correspondingly, the calculation yields ; The viewing angle deviation between the infrared feature map and the visible light feature map is calculated, and then the mountain infrared image is spatially transformed into a mountain infrared image registered with the mountain visible light image based on the viewing angle deviation. The registered infrared feature map is extracted from the registered mountain infrared image and then fused with the visible light feature map to obtain the fused feature; Based on the relationship between the fused features and the semantic and location information of the target, the semantic and location information of the target corresponding to the current fused features is predicted; The drone is equipped with an AI computing unit, which uses a dual-light registration and fusion detection module to process and analyze the visible light and infrared dual-modal data collected by the optoelectronic pod, and to identify and track the set target in the field of view. The drone platform is equipped with a high-precision positioning module, which, combined with an AI computing unit, calculates the target's position coordinates by determining the position and attitude of the drone and the pod. The multimodal attention fusion mechanism is a dual-input residual network that takes the visible light and infrared feature maps of the same layer in the downsampled feature pyramid network as inputs. After 6 convolutional layers, the features of the two are fused. The fused features are laterally connected with the corresponding downsampled layer of the feature pyramid network to identify the location and semantics of the target. The downsampled feature map is convolved with a 1×1 convolution kernel to change the number of channels of the feature map so that it is the same as the number of channels of the upsampled feature map of the corresponding layer, and then the pixels are added to it.
2. The target detection method based on dual-light registration and fusion as described in claim 1, characterized in that, A dual-channel target feature extraction network is used to extract visible light feature maps and infrared feature maps.
3. The target detection method based on dual-light registration and fusion as described in claim 2, characterized in that, Each channel in the dual-channel target feature extraction network contains at least two layers, and the number of convolutional kernel channels in each layer increases progressively from top to bottom. The layers are processed by max pooling layers, which reduces the feature mapping size of the input of the next layer to a set ratio of the output of the previous layer.
4. A target detection device based on dual-light registration and fusion, characterized in that, include: Image acquisition module, which is used to acquire visible light images and infrared images of mountainous areas; The feature extraction module is used to extract the texture information of the visible light image of the mountainous area and the temperature information of the infrared image of the mountainous area, respectively, to obtain the visible light feature map and the infrared feature map; By utilizing a cross-modal attention module, the visible light and infrared features extracted separately by the dual-channel network are obtained by convolution and dot product operations to obtain the associated features between the two, which are then used to calculate the viewing angle deviation between the infrared feature map and the visible light feature map. In the formula, It is an infrared feature map The index at a specific position in the text; Visible light characteristic map The index at a specific position in the text; Visible light feature map Features and locations at all locations Infrared feature map on The result of normalized and weighted summarization of features; the univariate function g is used to calculate the visible light feature map. In position The representation value of information; a binary function Calculate the visible light feature map using an embedded Gaussian function. Middle position Location information and feature mapping infrared feature map Middle position Relevance of information; Infrared feature maps exist Feature mapping points and visible light feature maps at the location exist Feature mapping points at; Visible light feature map any point above Infrared feature map Points on Correspondingly, the calculation yields ; The image registration module is used to calculate the viewing angle deviation between the infrared feature map and the visible light feature map, and then, based on the viewing angle deviation, spatially transforms the mountain infrared image into a mountain infrared image registered with the mountain visible light image. The feature fusion module is used to extract the registered infrared feature map from the registered mountain infrared image and then fuse it with the visible light feature map to obtain the fused feature. The target prediction module is used to predict the semantic and location information of the target corresponding to the current fused feature based on the relationship between the fused features and the semantic and location information of the target. An AI computing unit is mounted on the drone. The dual-light registration and fusion detection module in the AI computing unit processes and analyzes the visible light and infrared dual-modal data collected by the optoelectronic pod to identify and track the set target in the field of view. The drone platform is equipped with a high-precision positioning module, which, combined with an AI computing unit, calculates the target's position coordinates by determining the position and attitude of the drone and the pod. The multimodal attention fusion mechanism is a dual-input residual network that takes the visible light and infrared feature maps of the same layer in the downsampled feature pyramid network as inputs. After 6 convolutional layers, the features of the two are fused. The fused features are laterally connected with the corresponding downsampled layer of the feature pyramid network to identify the location and semantics of the target. The downsampled feature map is convolved with a 1×1 convolution kernel to change the number of channels of the feature map so that it is the same as the number of channels of the upsampled feature map of the corresponding layer, and then the pixels are added to it.
5. The target detection device based on dual-light registration and fusion as described in claim 4, characterized in that, In the feature extraction module, a dual-channel target feature extraction network is used to extract visible light feature maps and infrared feature maps.
6. The target detection device based on dual-light registration and fusion as described in claim 4, characterized in that, In the feature extraction module, each channel of the dual-channel target feature extraction network contains at least two layers, and the number of convolutional kernel channels in each layer increases progressively from top to bottom; the layers are processed by max pooling layers so that the feature mapping size of the input of the next layer is reduced to a set ratio of the output of the previous layer.
7. An unmanned aerial vehicle (UAV) system, characterized in that, Including the drone terminal and the ground control center; The UAV terminal and the ground control center communicate with each other through an ad-hoc network communication module. The drone terminal includes a drone platform, an optoelectronic pod, and an AI computing unit; The optoelectronic pod is equipped with an image acquisition module, which is used to acquire visible light images and infrared images of the mountainous area and transmit them to the AI computing unit. The AI computing unit includes a dual-light registration fusion detection module; the dual-light registration fusion detection module is configured to perform the steps in the target detection method based on dual-light registration fusion as described in any one of claims 1-3, for predicting the semantic and location information of the target based on the image transmitted by the image acquisition module and transmitting it to the UAV platform.
8. The unmanned aerial vehicle system as described in claim 7, characterized in that, The UAV platform is equipped with a flight control unit and a high-precision positioning module; the flight control unit is used to receive and execute flight tasks uploaded by the ground control center; the high-precision positioning module is used for positioning and measuring the UAV's position coordinates.
Citation Information
Patent Citations
Infrared and visible light image fusion method based on spatial correlation attention
CN116704274A
Infrared and visible image collaborative target detection method and device, equipment and medium
CN117576371A
Unmanned aerial vehicle mountainous area rescue method and system, medium and electronic equipment
CN118155105A