Mooring unmanned aerial vehicle multi-mode target detection method

Through the multimodal object detection method of tethered drone and combined with multiple data sources and edge detection algorithms, the problem of identification and positioning of disaster-affected personnel in disaster rescue is solved, and efficient and accurate target detection and rescue support is achieved.

CN119929202APending Publication Date: 2025-05-06SHAOXING SHUHONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411980373.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and locate disaster-affected personnel in disaster rescue, especially in the messy ruins, resulting in inefficient identification and repeated identification problems.

Method used

The multimodal object detection method of tethered drone is adopted. Through the image acquisition module, thermal imaging acquisition module, ultrasonic imaging acquisition module and acoustic wave acquisition module, a variety of data sources are collected and combined, including regional ground images, thermal imaging images, ultrasonic images and sound source positioning information, and combined with edge detection algorithms and coordinate systems to establish, human shapes are filtered and the sound source position is judged to output the results of suspected disaster-affected personnel.

Benefits of technology

Multimodal object detection is realized, the accuracy and efficiency of identifying and positioning affected personnel is improved, the blindness and inefficiency of manual monitoring is avoided, the rescue efficiency is greatly improved, and the system's adaptability in complex environments is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119929202A_ABST
    Figure CN119929202A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal target detection method for a mooring unmanned aerial vehicle, belongs to the technical field of unmanned aerial vehicle recognition, and solves the problem that disaster victims cannot be found in disordered ruins by a relatively accurate method for simply performing image recognition through visual assistance, and disaster relief personnel can be repeatedly recognized even if the disaster victims are recognized. And finally, the method can only be used for outputting pictures so that disaster relief personnel can cooperate to carry out manual monitoring, and the recognition efficiency is poor. Comprising the following steps: acquiring an area ground picture through an image acquisition module in the mooring unmanned aerial vehicle; according to the invention, multi-modal target detection is realized, whether suspected disaster victims exist in the area is judged by combining direct shooting, thermal imaging, ultrasonic image and sound source positioning, disaster relief personnel are helped to carry out rescue more efficiently, and blind searching is avoided. Meanwhile, it is not needed to completely rely on manual monitoring of images of the mooring unmanned aerial vehicle, the rescue efficiency is greatly improved, and rescue workers are prevented from being repeatedly recognized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned aerial vehicle identification, and in particular to a multi-modal target detection method for a tethered unmanned aerial vehicle. Background Art

[0002] A tethered drone is a drone that is connected to the ground via a cable and can fly stably in the air for a long time. They obtain uninterrupted energy by connecting to the ground power supply system through an optoelectronic composite cable, thus breaking through the limitations of the endurance of traditional drones. This type of drone can carry various optoelectronic and communication application payloads such as pods, radars, cameras, radio stations, base stations, antennas, etc., and achieve functions such as 24-hour hovering, large payload, and one-button automatic take-off and landing. Tethered drones are widely used in many fields such as power line inspection, agricultural plant protection, urban management, and aerial monitoring. Their main advantage is that they can provide continuous power supply and stable data transmission, allowing drones to perform tasks in specific areas for a long time, such as fixed-point monitoring and emergency communications. In addition, tethered drones also have the advantages of small size, light weight, and simple operation, which makes them perform well in scenarios that require long-term, high-stability tasks.

[0003] Tethered drones play an important role in scenarios that require long-term, uninterrupted monitoring. For example, in disaster relief and emergency response, they can be equipped with high-definition cameras and thermal imaging equipment to monitor changes in disaster areas for a long time and help search for trapped people. In military reconnaissance and border monitoring, tethered drones can conduct real-time intelligence collection and image monitoring, and provide continuous aerial monitoring support. In addition, in communication relay and temporary network construction, they serve as air communication relay stations, providing important communication support for maritime information networks, especially in areas without communication network coverage, to achieve signal relay and conversion. These application scenarios require drones to continuously identify targets to ensure the real-time and accuracy of information. Tethered drones are ideal for these tasks due to their long endurance and hovering capabilities.

[0004] In the actual disaster relief process, a tethered drone is often set up in the center of an area to identify and detect the stranded people in the area, thereby ensuring that the stranded people can be discovered and rescued in time to avoid harm to their personal safety and health.

[0005] However, simply using visual assistance to perform image recognition does not provide a more accurate way to find disaster victims in the cluttered ruins. Even if they are identified, it may be a duplicate identification of the rescue workers, which ultimately results in it being able to only be used to output images so that the rescue workers can cooperate in manual monitoring, and the recognition efficiency is poor.

[0006] Therefore, a multimodal target detection method for tethered UAV is proposed to solve or alleviate the above problems. Summary of the invention

[0007] The purpose of the present invention is to solve the shortcomings of the prior art and propose a multi-modal target detection method for a tethered drone.

[0008] In order to achieve the above object, the present invention adopts the following technical solutions:

[0009] A multimodal target detection method for a tethered UAV comprises the following steps:

[0010] The image acquisition module in the tethered drone collects the regional ground images, and the thermal imaging acquisition module in the tethered drone collects the regional thermal imaging images and transmits them to the ground system simultaneously;

[0011] After receiving the regional ground image and the regional thermal imaging image, the ground system collects the regional hidden image of the ground in the area through the ultrasonic imaging acquisition module, and simultaneously locates the sound source through the sound wave acquisition module;

[0012] The regional hidden picture and the regional thermal imaging picture are combined to obtain the regional life picture;

[0013] Determine a humanoid shape with a contour maintained in the regional life picture by using an edge detection algorithm, determine a human shape with a contour maintained in the regional ground picture by using an edge detection algorithm, and filter the human shape after comparing the humanoid shape with the human shape in the regional life picture;

[0014] A coordinate system is established for the regional life picture, and the coordinate area of ​​the remaining human-like shapes in the regional life picture is determined. It is compared whether the sound source is located in the coordinate area of ​​the remaining human-like shapes. If so, the output result is a suspected disaster victim, if not, no result is output.

[0015] Preferably, the tethered drone includes an image acquisition module, a thermal imaging acquisition module, a drone processor, and a drone memory coupled to the drone processor, the image acquisition module is used to acquire regional ground images and transmit them to the drone processor, the thermal imaging acquisition module is used to acquire regional thermal imaging images and transmit them to the drone processor, the drone processor is used to receive regional ground images and regional thermal imaging images and process and transmit them, and the drone memory is used to temporarily store data in the drone processor.

[0016] Preferably, the ground system includes an acoustic wave acquisition module, an ultrasonic imaging acquisition module, a ground processor, and a ground storage device. The acoustic wave acquisition module is used to locate the sound source and transmit it to the ground processor. The ultrasonic imaging acquisition module is used to collect regional hidden images and transmit them to the ground processor. The ground processor is used to receive and process the sound source, regional hidden images, and data from the drone processor. The ground storage device is used to temporarily store data in the ground processor.

[0017] Preferably, the ground system and the tethered drone are coupled via a cable.

[0018] Preferably, the regional hidden picture is combined with the regional thermal imaging picture to obtain the regional life picture, comprising the following steps:

[0019] Preprocessing the regional hidden image and the regional thermal imaging image, wherein the preprocessing includes filtering, denoising, and contrast enhancement;

[0020] Matching the regional hidden image with the regional thermal imaging image to determine a number of common reference points;

[0021] Perform linear weighted fusion on the regional hidden image and the regional thermal imaging image to obtain a fused image;

[0022] The fused image is sharpened and pseudo-colored to generate a regional life picture.

[0023] Preferably, the steps of determining the human-like shape with contours maintained in the regional life picture by an edge detection algorithm, determining the human shape with contours maintained in the regional ground picture by an edge detection algorithm, and filtering the human shape after comparing the human-like shape and the human shape in the regional life picture include the following steps:

[0024] The Canny edge detection algorithm is used to detect the edges of the regional life images and the regional ground images, and the edge detection threshold is set;

[0025] Using a contour detection algorithm to extract contours in the regional life picture and the regional ground picture, and performing shape matching on the contours in the regional life picture and the regional ground picture, thereby determining humanoid shapes and human shapes;

[0026] Calculate the Hausdorff distance between humanoid shapes and human shapes and set a similarity threshold, and output whether the result is a human shape. If so, filter the human shape in the regional life picture through masking operation. If not, do not perform masking operation.

[0027] Preferably, the method of establishing a coordinate system for the regional life picture, determining the coordinate area of ​​the remaining human-like shapes in the regional life picture, and comparing whether the sound source is located in the coordinate area of ​​the remaining human-like shapes, if so, outputting the result as a suspected disaster victim, if not, not outputting the result, includes the following steps:

[0028] Establish a coordinate system for the regional life picture after filtering the human shape;

[0029] A humanoid coordinate region is established in the coordinate system according to the outline of the remaining humanoid shapes in the regional life picture;

[0030] Establish the coordinates of the sound source according to the positions of the intersection coordinate points of the sound sources;

[0031] Determine whether the sound source coordinates are within the human-like coordinate area. If the judgment result is yes, the output result is the suspected disaster victim. If not, no result is output.

[0032] The present invention has the following beneficial effects:

[0033] The present invention can perform multi-modal target detection actions, and judge whether there are suspected disaster victims in the current area based on the images obtained by direct shooting, thermal imaging, and ultrasound, combined with the sound source, so that disaster relief personnel can assist in disaster relief according to the results output by the method to avoid random searches. At the same time, it also ensures that there is no need to rely entirely on manual monitoring of the output images of tethered drones to find disaster victims, which greatly improves the efficiency of the search and avoids repeated identification and output of disaster relief personnel in the image. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.

[0035] Figure 1 It is a flowchart of the present invention;

[0036] Figure 2 The structural frame of the present invention Figure 1 ;

[0037] Figure 3 The structural frame of the present invention Figure 2 . DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.

[0039] Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0040] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, further definition and explanation thereof is not required in subsequent drawings.

[0041] In the description of the present invention, it should be understood that the terms "center", "up", "down", "left", "right", "vertical", "horizontal", "inside", "outside" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, or are the orientations or positional relationships in which the product of the invention is conventionally placed when in use, or are the orientations or positional relationships conventionally understood by those skilled in the art. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be understood as a limitation on the present invention.

[0042] Furthermore, the terms “first”, “second”, “third”, etc. are merely used for distinguishing descriptions and are not to be understood as indicating or implying relative importance.

[0043] In the description of the present invention, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set", "install", "connect", and "connect" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0044] A multimodal target detection method for tethered UAV, such as Figure 1 As shown, the following steps are included:

[0045] The image acquisition module in the tethered drone collects the regional ground images, and the thermal imaging acquisition module in the tethered drone collects the regional thermal imaging images and transmits them to the ground system simultaneously;

[0046] After receiving the regional ground image and the regional thermal imaging image, the ground system collects the regional hidden image of the ground in the area through the ultrasonic imaging acquisition module, and simultaneously locates the sound source through the sound wave acquisition module;

[0047] The regional hidden picture and the regional thermal imaging picture are combined to obtain the regional life picture;

[0048] Determine a humanoid shape with a contour maintained in the regional life picture by using an edge detection algorithm, determine a human shape with a contour maintained in the regional ground picture by using an edge detection algorithm, and filter the human shape after comparing the humanoid shape with the human shape in the regional life picture;

[0049] A coordinate system is established for the regional life picture, and the coordinate area of ​​the remaining human-like shapes in the regional life picture is determined. It is compared whether the sound source is located in the coordinate area of ​​the remaining human-like shapes. If so, the output result is a suspected disaster victim, if not, no result is output.

[0050] The present invention uses multimodal target detection technology, combined with multiple data sources, including traditional shooting pictures, thermal imaging images, ultrasonic images and sound source positioning information, to conduct a comprehensive analysis of the current area and accurately determine whether there are suspected disaster victims. Through this method, disaster relief personnel can efficiently locate the disaster-stricken area based on the results output by the system, avoiding the blindness and inefficiency in the traditional manual search process, thereby greatly improving the rescue efficiency. At the same time, the present invention does not need to rely entirely on the real-time images taken by the tethered drone for manual monitoring, which not only reduces the burden of manual monitoring, but also ensures that the disaster victims can be found quickly, avoiding the problem of repeated identification of rescue personnel in the picture. In addition, combined with sound source positioning technology, the adaptability of the system in complex environments is further enhanced, ensuring that the rescue work can be carried out stably under various conditions, especially in post-disaster environments with difficult communication or limited line of sight, it can still provide accurate target positioning to ensure the smooth development of disaster relief work. Therefore, the invention not only improves the rescue efficiency and avoids waste of resources, but also enhances the intelligence and automation of the disaster relief process, significantly improving the overall rescue effect.

[0051] Preferably, if Figure 2 and Figure 3As shown, the tethered drone includes an image acquisition module, a thermal imaging acquisition module, a drone processor, and a drone memory coupled to the drone processor. The image acquisition module is used to acquire regional ground images and transmit them to the drone processor. The thermal imaging acquisition module is used to acquire regional thermal imaging images and transmit them to the drone processor. The drone processor is used to receive regional ground images and regional thermal imaging images and process and transmit them. The drone memory is used to temporarily store data in the drone processor.

[0052] The tethered drone realizes data collection, processing and storage. Through the collaborative work of the drone processor and memory, it realizes real-time processing and temporary storage of collected data, providing support for subsequent data analysis.

[0053] Preferably, if Figure 2 and Figure 3 As shown, the ground system includes an acoustic wave acquisition module, an ultrasonic imaging acquisition module, a ground processor, and a ground storage. The acoustic wave acquisition module is used to locate the sound source and transmit it to the ground processor. The ultrasonic imaging acquisition module is used to acquire regional hidden images and transmit them to the ground processor. The ground processor is used to receive and process the sound source, regional hidden images, and data from the UAV processor. The ground storage is used to temporarily store the data in the ground processor.

[0054] The ground system is capable of receiving data transmitted by the drone and combining it with sound and ultrasonic data for comprehensive processing. It uses the powerful computing power of the ground processor to deeply fuse and analyze the data collected by the drone and the data collected on the ground to improve the accuracy of target detection.

[0055] Preferably, if Figure 2 and Figure 3 As shown, the ground system and the tethered drone are coupled via cables, which ensures the stability and real-time performance of data transmission while also maintaining power supply.

[0056] Preferably, the regional hidden picture and the regional thermal imaging picture are combined to obtain the regional life picture, including the following steps:

[0057] Preprocessing the regional hidden images and regional thermal imaging images, including filtering, denoising, and contrast enhancement;

[0058] Matching the regional hidden image with the regional thermal imaging image to determine a number of common reference points;

[0059] Perform linear weighted fusion on the regional hidden image and the regional thermal imaging image to obtain a fused image;

[0060] The fused image is sharpened and pseudo-colored to generate a regional life picture.

[0061] Through image preprocessing and fusion technology, the quality of life pictures is improved and the accuracy of target detection is enhanced. By filtering and denoising the hidden pictures and thermal imaging pictures, contrast enhancement, linear weighted fusion and other steps, a clearer regional life picture is generated.

[0062] Preferably, determining a humanoid shape with a contour maintained in the regional life picture by an edge detection algorithm, determining a human shape with a contour maintained in the regional ground picture by an edge detection algorithm, and filtering the human shape after comparing the humanoid shape and the human shape in the regional life picture includes the following steps:

[0063] The Canny edge detection algorithm is used to detect the edges of the regional life images and the regional ground images, and the edge detection threshold is set;

[0064] Using a contour detection algorithm to extract contours in the regional life picture and the regional ground picture, and performing shape matching on the contours in the regional life picture and the regional ground picture, thereby determining humanoid shapes and human shapes;

[0065] Calculate the Hausdorff distance between humanoid shapes and human shapes and set a similarity threshold, and output whether the result is a human shape. If so, filter the human shape in the regional life picture through masking operation. If not, do not perform masking operation.

[0066] Through edge detection and shape matching technology, human shapes are accurately identified and filtered, and the specificity of detection is improved. The Canny edge detection algorithm and contour detection algorithm are used to extract contours, calculate the Hausdorff distance and set the similarity threshold to distinguish between pseudo-human shapes and human shapes.

[0067] Preferably, a coordinate system is established for the regional life picture, and the coordinate area of ​​the remaining human-like shapes in the regional life picture is determined, and the sound source is compared to see whether it is located in the coordinate area of ​​the remaining human-like shapes. If so, the output result is a suspected disaster victim, and if not, no result is output, including the following steps:

[0068] Establish a coordinate system for the regional life picture after filtering the human shape;

[0069] A humanoid coordinate region is established in the coordinate system according to the outline of the remaining humanoid shapes in the regional life picture;

[0070] Establish the coordinates of the sound source according to the positions of the intersection coordinate points of the sound sources;

[0071] Determine whether the sound source coordinates are within the human-like coordinate area. If the judgment result is yes, the output result is the suspected disaster victim. If not, no result is output.

[0072] By establishing a coordinate system and locating the sound source, the accuracy of identifying suspected disaster victims is improved. A coordinate system is established in the filtered regional life picture, the coordinate area of ​​the human shape is determined, and compared with the sound source coordinates to determine whether it is a suspected disaster victim.

[0073] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A multimodal target detection method for a tethered UAV, characterized in that: The following steps are included: The image acquisition module in the tethered drone collects the regional ground images, and the thermal imaging acquisition module in the tethered drone collects the regional thermal imaging images and transmits them to the ground system simultaneously; After receiving the regional ground image and the regional thermal imaging image, the ground system collects the regional hidden image of the ground in the area through the ultrasonic imaging acquisition module, and simultaneously locates the sound source through the sound wave acquisition module; The regional hidden picture and the regional thermal imaging picture are combined to obtain the regional life picture; Determine a humanoid shape with a contour maintained in the regional life picture by using an edge detection algorithm, determine a human shape with a contour maintained in the regional ground picture by using an edge detection algorithm, and filter the human shape after comparing the humanoid shape with the human shape in the regional life picture; A coordinate system is established for the regional life picture, and the coordinate area of ​​the remaining human-like shapes in the regional life picture is determined. It is compared whether the sound source is located in the coordinate area of ​​the remaining human-like shapes. If so, the output result is a suspected disaster victim, if not, no result is output.

2. A tethered UAV multimodal target detection method according to claim 1, characterized in that: The tethered drone includes an image acquisition module, a thermal imaging acquisition module, a drone processor, and a drone memory coupled to the drone processor. The image acquisition module is used to acquire regional ground images and transmit them to the drone processor. The thermal imaging acquisition module is used to acquire regional thermal imaging images and transmit them to the drone processor. The drone processor is used to receive, process and transmit regional ground images and regional thermal imaging images. The drone memory is used to temporarily store data in the drone processor.

3. A tethered UAV multimodal target detection method according to claim 2, characterized in that: The ground system includes an acoustic wave acquisition module, an ultrasonic imaging acquisition module, a ground processor, and a ground storage. The acoustic wave acquisition module is used to locate the sound source and transmit it to the ground processor. The ultrasonic imaging acquisition module is used to collect regional hidden images and transmit them to the ground processor. The ground processor is used to receive and process the sound source, regional hidden images, and data from the drone processor. The ground storage is used to temporarily store data in the ground processor.

4. A tethered UAV multimodal target detection method according to claim 3, characterized in that: The ground system and the tethered drone are coupled via a cable.

5. The multimodal target detection method for a tethered UAV according to claim 1, characterized in that: The regional hidden picture and the regional thermal imaging picture are combined to obtain the regional life picture, including the following steps: Preprocessing the regional hidden image and the regional thermal imaging image, wherein the preprocessing includes filtering, denoising, and contrast enhancement; Matching the regional hidden image with the regional thermal imaging image to determine a number of common reference points; Perform linear weighted fusion on the regional hidden image and the regional thermal imaging image to obtain a fused image; The fused image is sharpened and pseudo-colored to generate a regional life picture.

6. A tethered UAV multimodal target detection method according to claim 5, characterized in that: The method of determining a humanoid shape with a contour maintained in the regional life picture by an edge detection algorithm, determining a human shape with a contour maintained in the regional ground picture by an edge detection algorithm, and filtering the human shape after comparing the humanoid shape and the human shape in the regional life picture includes the following steps: The Canny edge detection algorithm is used to detect the edges of the regional life images and the regional ground images, and the edge detection threshold is set; Using a contour detection algorithm to extract contours in the regional life picture and the regional ground picture, and performing shape matching on the contours in the regional life picture and the regional ground picture, thereby determining humanoid shapes and human shapes; Calculate the Hausdorff distance between humanoid shapes and human shapes and set a similarity threshold, and output whether the result is a human shape. If so, filter the human shape in the regional life picture through masking operation. If not, do not perform masking operation.

7. A tethered UAV multimodal target detection method according to claim 6, characterized in that: The method of establishing a coordinate system of the regional life picture and determining the coordinate area of ​​the remaining human-like shapes in the regional life picture, comparing whether the sound source is located in the coordinate area of ​​the remaining human-like shapes, and if so, outputting the result as a suspected disaster victim, and if not, not outputting the result, includes the following steps: Establish a coordinate system for the regional life picture after filtering the human shape; A humanoid coordinate region is established in the coordinate system according to the outline of the remaining humanoid shapes in the regional life picture; Establish the coordinates of the sound source according to the positions of the intersection coordinate points of the sound sources; Determine whether the sound source coordinates are within the human-like coordinate area. If the judgment result is yes, the output result is the suspected disaster victim. If not, no result is output.