Method and device for accurately identifying persons in distress in multiple scenes
By determining the search area and flight altitude, collecting and processing video images to achieve detailed restoration and multi-scale feature extraction, the problem of relying on artificial experience and environmental factors in traditional drone search and rescue is solved, and the efficiency and accuracy of water rescue is improved.
Patent Information
- Application Number
- CN202510529255.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-04-25
AI Technical Summary
Traditional drones lack a systematic search system in water search and rescue, rely on manual experience and the quality of video data is greatly affected by environmental factors, resulting in the inability to obtain comprehensive information.
By acquiring drone flight data and distress information, determining the search area and optimal flight altitude, collecting and processing video images for details restoration and multi-scale feature extraction, the YOLOv5 target detection network and image restoration network are used to improve video quality and recognition accuracy.
It improves search efficiency, obtains high-quality video data, realizes accurate identification and positioning of people in distress, and enhances the accuracy of water rescue.
Smart Images

Figure CN120544071A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of water rescue technology, and in particular to a method and device for accurately identifying persons in distress in multiple scenarios. Background Art
[0002] As maritime transport demand continues to expand, ship traffic is becoming increasingly dense and complex, making maritime accidents more likely. In complex weather conditions such as fog, rain, and low illumination, or in dense traffic conditions, people in distress are easily affected by wind, waves, and currents, causing them to drift, making the search for them extremely difficult. Drones, due to their high flexibility, low cost, and strong controllability, are widely used in maritime emergency rescue missions.
[0003] However, traditional drones lack a systematic search system and the ability to autonomously detect people in distress on the water. Their operations rely on manual experience, resulting in a haphazard and unsystematic approach that reduces the efficiency of search and rescue personnel. Furthermore, the video data collected by traditional drone imaging equipment is significantly affected by environmental factors, resulting in low quality and preventing rescue personnel from obtaining comprehensive information from the video data.
[0004] Therefore, it is urgent to propose a method and device for accurately identifying people in distress in multiple scenarios to solve the technical problem in the existing technology that the video data of traditional drones is greatly affected by environmental factors, resulting in low video data quality and the inability of search and rescue personnel to obtain comprehensive information from the video data. Summary of the Invention
[0005] In view of this, it is necessary to provide a method and device for accurately identifying people in distress in multiple scenarios to solve the technical problem in the existing technology that the video data of traditional drones is greatly affected by environmental factors, resulting in low video data quality and the inability of search and rescue personnel to obtain comprehensive information from the video data.
[0006] In order to solve the above problems, in a first aspect, the present invention provides a method for accurately identifying people in distress in multiple scenarios, comprising: Obtain UAV flight data and mission requirements, as well as distress information of people in distress; determining a search area based on the distress information and the mission requirements, and determining an optimal flight altitude based on the search area and the flight data; When the UAV reaches the optimal flight altitude of the search area, capturing a current video image from the current viewing angle of the UAV, and performing detail restoration and degradation recovery on the current video image to obtain a clear video image; Multi-scale feature extraction is performed on the clear video image to obtain an identification result of the person in distress.
[0007] In one possible implementation, the distress information includes a probability density distribution of the initial position of the person in distress; and determining the search area according to the distress information and the task requirement includes: Determining a scannable area of the UAV on the water according to a parallel line scanning search method and the mission requirements; Establishing a person-in-distress drift model based on the influence of multiple factors on the person in distress and the probability density distribution; The drift model of the person in distress and the scannable area are simulated and optimized to obtain a search area.
[0008] In a possible implementation, simulating and optimizing the drift model of the person in distress and the scannable area to obtain a search area includes: Estimate the wind speed at a preset sea level according to a preset wind pressure model and determine the disturbance coefficient; According to the disturbance coefficient, the wind-induced drift velocity is obtained; Estimate the flow velocity at the preset water depth to obtain the flow-induced drift velocity; Obtaining a drift velocity according to the wind-induced drift velocity and the flow-induced drift velocity; The drift velocity is input into the person-in-distress drift model to perform random particle simulation optimization on the scannable area to obtain a search area.
[0009] In a possible implementation, determining the optimal flight altitude according to the search area and the flight data includes: Determining decision variables for the flight data in the search area according to a preset probability relationship model; the decision variables include coverage and route spacing; Constructing an objective function that maximizes the probability of discovery based on the coverage range and the route interval; Optimizing the maximum discovery probability of the objective function, the coverage range, and the route interval according to preset constraints and a preset parallel selection genetic algorithm to determine a target sweep width; According to the target sweep width, the optimal search flight altitude is obtained.
[0010] In a possible implementation, performing detail restoration and degradation recovery on the current video image to obtain a clear video image includes: determining current weather based on the drone; Obtaining a historical dataset from a severe weather image dataset according to the current weather; The details of the current video image are restored and degradation is restored according to the historical data set to obtain a clear video image.
[0011] In one possible implementation, the generation process of the severe weather image dataset includes: Generate fog images based on the atmospheric scattering model to obtain a fog image dataset; Generate rainy day images based on the rain map model to obtain a rainy day image dataset; A severe weather image dataset is obtained according to the fog image dataset and the rainy day image dataset.
[0012] In a possible implementation, performing detail restoration and degradation recovery on the current video image according to the historical data set to obtain a clear video image includes: Comparing the historical data set and the current video image according to a comparative degradation encoder to obtain a potential degradation representation; The current video image and the potential degradation representation are input into a degradation guided restoration network for degradation restoration to obtain a clear video image.
[0013] In a possible implementation, performing multi-scale feature extraction on the clear video image to obtain the identification result of the person in distress includes: Construct a multi-scale target detection model; the multi-scale target detection model includes a backbone feature extraction network, a BiFormer attention module, a CBAM attention module and a decoupling head module; Extracting features from the clear video image according to the backbone feature extraction network to obtain initial features; Performing adaptive feature fusion on the initial features according to the BiFormer attention module to obtain fused features; Performing weighted processing on the fused features according to the CBAM attention module to obtain refined features; The refined features are identified according to the decoupling head module to obtain an identification result.
[0014] In a possible implementation, the comparing the historical data set and the current video image according to the comparative degradation encoder to obtain a potential degradation representation includes: Comparing the historical data set and the current video image according to the contrast degradation encoder to determine positive samples and negative samples; A potential degradation representation is obtained according to maximizing the consistency between positive samples in the positive samples and minimizing the consistency between negative samples in the negative samples.
[0015] In a second aspect, the present invention further provides a multi-scenario accurate identification device for persons in distress, comprising: The information acquisition module is used to obtain the flight data and mission requirements of the UAV, as well as the distress information of people in distress; an area determination module, configured to determine a search area based on the distress information and the mission requirements, and to determine an optimal flight altitude based on the search area and the flight data; an image processing module, configured to capture a current video image from a current viewing angle of the drone when the drone reaches the optimal flight altitude of the search area, and perform detail restoration and degradation recovery on the current video image to obtain a clear video image; The result recognition module is used to perform multi-scale feature extraction on the clear video image to obtain the recognition result of the person in distress.
[0016] The beneficial effects of the present invention are as follows: the search area and the optimal flight altitude can be determined based on the distress information of the person in distress and the flight data and mission requirements of the drone, so that when the drone reaches the optimal flight altitude of the search area, the drone can collect the most appropriate video image data at the current viewing angle, thereby improving the search efficiency; the current video image can also be restored in detail and degraded to obtain a clear video image, thereby improving the quality of the video data; the clear video image can also be subjected to multi-scale feature extraction to obtain the identification result of the person in distress, thereby obtaining comprehensive information about the person in distress from the video data. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A schematic flow chart of an embodiment of the multi-scenario method for accurately identifying persons in distress provided by the present invention; Figure 2 For the present invention Figure 1 A schematic flow chart of an embodiment of step S102; Figure 3 A schematic diagram of the structure of an embodiment of the UAV search path provided by the present invention; Figure 4 A schematic structural diagram of an embodiment of the drift model for persons in distress provided by the present invention; Figure 5 A schematic structural diagram of an embodiment of the final search area provided by the present invention; Figure 6 A schematic diagram of a flow chart of an embodiment of the present invention for providing a clear video image; Figure 7 A schematic structural diagram of an embodiment of the auxiliary rescue system for people in distress on water provided by the present invention; Figure 8 This is a schematic structural diagram of an embodiment of the multi-scenario person-in-distress accurate identification device provided by the present invention. DETAILED DESCRIPTION
[0018] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, and are not used to limit the scope of the present invention.
[0019] YOLOv5 (You Only Look Once version 5) is a popular object detection algorithm in the fields of artificial intelligence and computer vision, renowned for its speed and accuracy. YOLOv5 is the latest version of the YOLO series (including YOLO, YOLOv2, YOLOv3, and YOLOv4). It inherits the strengths of previous versions and incorporates several improvements. YOLOv5 adopts a single-stage object detection approach, directly classifying and localizing objects in an image, rather than using the traditional two-stage approach (such as Faster R-CNN: region proposal generation and classification and localization). This approach reduces computational effort and improves detection speed.
[0020] The All-in-One Image Restoration (AiOIR) network is a unified framework designed to address multiple image degradation issues. By integrating advanced deep learning techniques, it can handle multiple degradation types, such as noise, blur, and weather effects, in a single network, providing a more convenient and versatile solution.
[0021] like Figure 1 As shown, a specific embodiment of the present invention discloses a method for accurately identifying people in distress in multiple scenarios, including: S101. Obtain the flight data and mission requirements of the UAV, as well as the distress information of the person in distress.
[0022] The multi-scenario precise identification method for persons in distress provided in the embodiments of the present application can be applied to a multi-scenario precise identification system for persons in distress, wherein the multi-scenario precise identification system for persons in distress can be a software system based on a terminal device, and the terminal device can be a server, a tablet computer, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a mobile phone, or other terminal device. The embodiments of the present application do not impose any restrictions on the specific type of the terminal device.
[0023] The embodiments of the present invention can be applied to situations where the search area is large and the position of the person in distress is uncertain, and a coverage search of the area is required. The search area can be divided according to the task requirements and a clear scanning path can be determined, so that the drone can fly continuously on parallel lines, reducing the complexity of operation and improving work efficiency, so that the flight data of the drone can be obtained. When the person in distress encounters danger at sea, he can send a distress message, and the drone can receive the distress information of the person in distress. The distress information may include information such as the quality and initial position of the person in distress. Due to its strong maneuverability, wide coverage, and high detection efficiency, the drone can effectively solve the problem of limited vision of the human eye and fixed port and shipping monitoring equipment in the search and rescue scenario of people in distress on the water, and the inability to achieve full coverage supervision of the water area. Among them, the task requirements can be the tasks set by the rescue personnel for the drone according to the actual situation.
[0024] S102. Determine a search area based on the distress information and mission requirements, and determine an optimal flight altitude based on the search area and flight data.
[0025] Among them, drones can search at sea using the parallel line sweeping method. Parallel line sweeping is the most commonly used and simplest visual search method in maritime search and rescue. It is mainly used when the search area is large and the location of the person in distress is uncertain. The search area can be determined based on the distress information and mission requirements. The optimal flight altitude can also be calculated based on the search area and the drone's flight data.
[0026] S103. When the UAV reaches the optimal flight altitude of the search area, the current video image under the current viewing angle of the UAV is collected, and the details of the current video image and degradation restoration are performed to obtain a clear video image.
[0027] Among them, after determining the search area and the optimal flight altitude, the drone can be controlled to fly to the optimal flight altitude of the search area, and then the current video image can be collected through the current perspective of the drone. The current video image can then be restored in detail and degraded to obtain a clear video image, providing high-quality visual information for the drone during water rescue, assisting rescuers to more accurately identify and locate people in distress.
[0028] S104: Perform multi-scale feature extraction on the clear video image to obtain identification results of the person in distress.
[0029] After obtaining a clear video image, multi-scale feature extraction can be performed on the clear video image based on the YOLOv5 object detection network to obtain more accurate identification results for people in distress. The YOLOv5 object detection network overcomes the shortcomings of traditional detection methods in multi-scale object detection, ensuring accurate detection under different shooting conditions and object scales, and improving the target recognition capabilities of drones in water rescue scenarios.
[0030] Compared with the existing technology, the present embodiment provides a method for determining a search area and an optimal flight altitude based on the distress information of the person in distress and the flight data and mission requirements of the drone. Therefore, when the drone reaches the optimal flight altitude of the search area, the drone can collect the most appropriate video image data at the current perspective, thereby improving the search efficiency. The present embodiment also allows for detail restoration and degradation recovery of the current video image, thereby obtaining a clear video image and improving the quality of the video data. The present embodiment also allows for multi-scale feature extraction of the clear video image to obtain identification results of the person in distress, thereby obtaining comprehensive information about the person in distress from the video data.
[0031] In some embodiments of the present invention, the distress information includes the probability density distribution of the initial position of the person in distress; Figure 2 As shown, step S102 includes: S201. Determine a scannable area of the UAV on the water based on a parallel line scanning search method and mission requirements.
[0032] Among them, the parallel line scanning search method can determine the area to be searched according to the mission requirements and divide it into scannable areas. According to the set scanning path, the UAV flies continuously on the parallel line. Parallel line search reduces the number of turns, reduces the complexity of operation and improves work efficiency. Figure 3 As shown in the figure, when the search area is a rectangle, the search starting point is usually a vertex of the rectangle. When performing the search mission, the UAV moves from the search starting point to 1 / 2 of the route distance. S The search route is parallel to the long side of the rectangle, thus avoiding continuous turns and saving time. According to the set scanning path, the drone flies continuously on the parallel line until the search end point.
[0033] S202. Based on the impact of multiple factors on the persons in distress and the probability density distribution, a drift model of the persons in distress is established.
[0034] Among them, the force analysis of people in distress, people in distress are affected by wind, waves and currents on the sea surface and drift. In order to simplify the force analysis of people in distress at sea, the Coriolis force is ignored. It is possible to determine multiple factors that affect people in distress, comprehensively consider the force analysis of people in distress under multiple external factors, and establish a drift model for people in distress. Multiple factors can include wind-induced drift and current-induced drift. The drift model for people in distress is as follows: Figure 4As shown in the figure, wind drift (LEEWAY), or wind pressure, is the directionally related movement of the part of a person in distress exposed above the water surface caused by the surface wind (10 meters above sea level). The magnitude of wind pressure is the speed of the person in distress relative to the sea, and the direction of wind pressure is expressed by the wind pressure angle. Therefore, the amount of wind drift of a person in distress can be expressed by wind pressure. Figure 4 The person in distress is affected by the surface current at the starting point of drift, which generates different wind pressure angles according to the direction of wind pressure. The drift position is calculated based on different wind pressures and wind pressure angles. The drift motion equation of the person in distress drift model is shown in formula (1): (1) Where, For the quality of people in distress; is the drift motion speed; is the wind force; is the force of ocean current; is the force of the waves; The mass of the person in distress and the action time can be obtained based on the distress information of the person in distress, and the wind force, current force, and wave force can be obtained by analyzing the actual situation and work experience.
[0035] S203: Simulate and optimize the drift model of the person in distress and the scannable area to obtain a search area.
[0036] After the drift model of the person in distress is established, the drift model of the person in distress can be simulated by random particles, so as to simulate and optimize the scannable area and obtain the optimized final search area.
[0037] In some embodiments of the present invention, step S203 includes: The wind speed at a preset sea level is estimated based on a preset wind pressure model to determine the disturbance coefficient.
[0038] Among them, the preset wind pressure model can be a drift model for the person in distress, and the drift motion speed can be calculated by analyzing the impact of wind-induced drift and flow-induced drift on the person in distress. The specific process is: in the offshore search mission, ocean currents are an important factor in considering the location of the person in distress. The part below the waterline of the person in distress will be affected by the surface current, and the flow rate is generally about 0.3-1 meters. Considering that the target of the embodiment of the present invention is a person in distress in a vertical posture, the flow rate at 0.5 meters is selected as the estimated input for the prediction of the drift motion of the person in distress. In order to quantify the uncertainty caused by the errors in the wind-induced drift and flow-induced drift of the person in distress to the experiment, random disturbances need to be set. Assuming that the wind speed disturbance obeys the normal distribution, the wind speed at a height of 10m above the sea surface is expressed as shown in formula (2): (2) Where, is the estimated wind speed; is the predicted wind speed; is the wind speed disturbance, and , for For the 0-Model wind pressure model, after adding the disturbance coefficient, it is shown as formula (3): (3) Where, is the wind pressure coefficient after adding disturbance, i.e. disturbance coefficient; is a random perturbation, and , is the wind pressure coefficient The standard deviation of .
[0039] According to the disturbance coefficient, the wind-induced drift velocity is obtained.
[0040] Among them, the disturbance coefficient is added to the drift velocity, and the wind-induced drift velocity after the disturbance is added is As shown in formula (4): (4) The flow velocity at the preset water depth is estimated to obtain the flow-induced drift velocity.
[0041] The preset water depth and flow rate can be set is the flow velocity at a water depth of 0.5 m, which will be obtained through the ocean current field forecast at the place where the drowning event occurs. After adding random disturbances, it can be expressed as shown in formula (5): (5) Where, is the flow-induced drift velocity; is a random perturbation, and , is the random disturbance of flow velocity The standard deviation of .
[0042] The drift velocity is obtained according to the wind-induced drift velocity and the flow-induced drift velocity.
[0043] Among them, according to the description of wind drift and flow drift Figure 4 From the drift model of the person in distress, it can be seen that the drift speed of the person in distress at sea can be expressed as the vector superposition of the wind-induced drift speed and the ocean current speed, as shown in formula (6): (6) Where, is the drift velocity, that is, the drift motion velocity in formula (1) .
[0044] The drift velocity is input into the distress person drift model to perform random particle simulation optimization on the scannable area to obtain the search area.
[0045] A random particle simulation experiment determined the scannable area, where the Proof of Concept (POC) for the presence of the person in distress was maximized, reaching 100%. However, given the tight timelines, demanding workload, and limited available search resources, an area that is too large may not effectively cover the search, resulting in a lower search success rate. Conversely, if the area is too small, the probability of the person in distress being within the search area decreases, resulting in a too-small POC. While resources can fully cover the area, the random nature of the search task can easily lead to search failure. Therefore, to improve the success rate and timeliness of search operations, it is necessary to rationally optimize the search area for persons in distress.
[0046] In order to optimize and reduce the area, the scannable area where the particles are finally located during the simulation time period is gridded through the drift model of people in distress, divided into 225 square sub-areas of 2000×2000m, and the number of particles in each sub-area is counted. According to the particle position distribution diagram of the initial search area, it is found that the distribution of particles follows the rule that there are more particles in the middle and fewer particles in the surrounding areas. The largest number of particles is likely to be distributed in the center of the rectangular area, and the particles distributed around it may drift out of the search area as the search progresses, resulting in a problem of increased search time but reduced efficiency. Therefore, the initial search area should be reduced while ensuring a certain POC, the search time should be reduced, and the search efficiency should be improved. The embodiment of the present invention will diffuse outward from the sub-area with the largest number of particles in the center until the inclusion probability POC is above 90%, and the rectangular area is determined to be the final search area. As shown in the figure, Figure 5 As shown, Figure 5 middle x The axis represents nautical miles in the east-west direction. y The axis represents the number of nautical miles in the north-south direction. The rectangular area is the final search area after regional optimization. Compared with the initial search area, the probability of inclusion (POC) reaches 95%, which is greater than the 90% requirement. At the same time, the search area is a 9×9 sub-area grid, which is 64% smaller than the initial 15×15 search area. This greatly shortens the search operation time, improves the search efficiency, and solves the problem of determining the search area.
[0047] In some embodiments of the present invention, step S102 includes: Determine decision variables for flight data in the search area based on a preset probability relationship model; the decision variables include coverage and route spacing; According to the coverage area and route interval, an objective function that maximizes the probability of discovery is constructed.
[0048] Among them, the UAV relies on visible light / infrared payloads to scan the two-dimensional plane (sea surface) in the air. The scanning width of the UAV is related to the flight altitude, which can be specifically expressed as ,in , is the flight altitude; Theoretically, the solution of increasing the scanning width by significantly increasing the flight altitude is not practical, because whether the people in distress on the water can be detected by the drone on the water surface is a probabilistic event. The target detection rate of the detection payload depends on the flight altitude of the drone. The higher the altitude, the greater the probability of detection. It is necessary to combine the characteristic parameters of the drone's own payload and its preset probability relationship model. It can be expressed as shown in formula (7): (7) Where, is the target detection rate; is the preset target detection rate; is the failure alarm rate; 、 is the characteristic parameter of the load. According to formula (7), when the flight altitude h At the best detection height When the target detection rate is within the preset target detection rate, 100%; when h When this value is exceeded, the target detection rate is linearly related to the flight altitude; when h Greater than maximum height When the target detection rate reaches the lowest, it is impossible to complete the search and detection task. At the same time, the constraints of the drone's target detection rate and the drone's own performance must be considered, so that the flight altitude can be optimized based on the constraints and the drone's own performance. h , thereby improving the search success rate.
[0049] Scanning width with drone W and route spacing S As the decision variable, the objective function of maximizing the probability of discovery in the UAV coverage trajectory planning is established to determine the optimal search flight altitude of the UAV. The objective function is shown in formula (8): (8) Where, It represents the objective function of maximizing the probability of discovery, that is, when the search area is determined, the probability of including people in distress on the water is certain, and the search success rate of the drone is improved by maximizing the probability of discovery. is the probability of discovery in the search area, and its calculation expression is , and coverage C By scan width W and route spacing S Joint decision, expressed as , is the flight altitude; It is half of the UAV payload field of view.
[0050] According to the preset constraints and the preset parallel selection genetic algorithm, the objective function of maximizing the probability of discovery, coverage and route interval is optimized to determine the target scanning width; The optimal search flight altitude is obtained based on the target scanning width.
[0051] Among them, according to the requirements of the UAV search mission, the target detection rate, sweep width, and route interval of the UAV payload are used as constraints to conditionally restrict the UAV search track planning. The preset constraints can be expressed as follows: (8) (9) (10) Where, is the minimum target detection rate for search, This is the minimum warning flight altitude for drones.
[0052] Under the preset constraints, the parallel selection genetic algorithm is used to select the scanning width. and route spacing To optimize, specifically, to maximize the probability of finding the objective function Bring it into the parallel selection genetic algorithm for genetic processing, so as to adjust the scan width and route spacing Iterative optimization is performed to obtain the optimized target sweep width and target route interval, and then the optimal flight altitude is obtained through the target sweep width. h , thereby improving the target detection rate through the optimal flight altitude , thereby improving the search success rate.
[0053] In some embodiments of the present invention, the process of generating a severe weather image dataset includes: Generate fog images based on the atmospheric scattering model to obtain a fog image dataset; Generate rainy day images based on the rain map model to obtain a rainy day image dataset; According to the fog image dataset and the rainy image dataset, a severe weather image dataset is obtained.
[0054] Given the limited availability of comparative image data from drone-perspective images of complex scenes, which severely impacts the performance of the image restoration network, the dataset can be expanded by artificially synthesizing images from drone perspectives under adverse weather conditions to enhance image restoration capabilities in complex scenes. To further enhance the richness of the dataset, the present invention employs two methods for synthesizing images under adverse weather conditions: one based on traditional methods and the other using a GAN network.
[0055] The fog image is generated using the atmospheric scattering model, which is shown in formula (11): (11) Where, is the coordinate value in the image, It is an artificially synthesized foggy image. It is the original clear image. represents the transmittance, Indicates the atmospheric light value. The transmittance at any point It decays exponentially with the increase of propagation distance, which can be described as shown in formula (12): (12) Where, represents the scattering coefficient caused by atmospheric scattering particles and absorbed light, Indicates the distance from the scene to the camera.
[0056] Then the fog image dataset can be obtained through formula (11) and formula (12).
[0057] The generation of rainy day images adopts the rain map model, which is shown in formula (13): (13) Where, is the pixel value in the image. The superposition reduces the clarity of the scene By setting the density, length, and angle of the rain lines, you can simulate any possible degree of rain in real scenes. To simulate light rain, the linear superposition method in the model is changed to weighted superposition.
[0058] The generation of low-light images uses the Retinex model, which is shown in formula (14):
[0059] Where, is the pixel value in the image, They represent the image observed by the human eye, the reflectivity and reflection component of the object to the light, and the light and incident component irradiated on the object. According to the Retinex theory, when the light is very weak, the incident light irradiated on the surface of the object is Very dark, and the reflection component It is an inherent property of the object and will not be affected by light. Therefore, from the above formula, we can see that when the amount of light is weak, the brightness of the final image will be darker.
[0060] The rainy day image dataset can be obtained by using formula (13) and formula (14), and then the foggy day image dataset and the rainy day image dataset can be merged to obtain the severe weather image dataset.
[0061] Furthermore, embodiments of the present invention utilize generative adversarial networks to construct a deep style transfer model with powerful semantic information representation capabilities. This model simultaneously extracts global, attentional, and edge features from multi-scene video images. Based on the extracted feature maps, it then uses a multi-scale generator to perform style transfer, simulating visible light video images affected by various weather conditions. To address the potential loss of detail during style transfer, a multi-scale style generator based on multi-scale codecs and multi-receptive field residual blocks is constructed to preserve structural and texture information in video images.
[0062] In some embodiments of the present invention, step S103 includes: Determine the current weather based on the drone; Obtain historical datasets from severe weather image datasets based on current weather; The details of the current video image are restored and degradation is restored based on the historical data set to obtain a clear video image.
[0063] The current weather can be determined based on the weather at the time the drone captured the video image. This can then be used to search the severe weather image dataset to obtain image data for the corresponding weather conditions, thereby generating a historical dataset for the current weather. This historical dataset can then be used to restore details and degrade the current video image, resulting in a clear video image.
[0064] In some embodiments of the present invention, Figure 6 As shown in the figure, the details of the current video image are restored and the degradation is restored based on the historical data set to obtain a clear video image, including: The historical dataset and the current video image are compared according to the contrastive degradation encoder to obtain the potential degradation representation.
[0065] The general image enhancement model is based on an integrated image restoration network (AirNet), a network model that can enhance images with various degradation types and degrees. Existing image restoration methods often only process images with specific degradation types and degrees and require prior knowledge of the damage. In real-world applications, the type and degree of degradation are constantly changing, and models struggle to automatically distinguish the type of damage, making existing methods difficult to handle. However, AirNet is unaffected by prior knowledge of degradation type and degree. It uses only observed degraded images for inference and restoration of various degraded images. The model consists of two modules: a contrastive-based degradation encoder (CBDE) and a degradation-guided restoration network (DGRN). Its integration and versatility across multiple scenarios allow AirNet to infer degraded images from historical datasets, restoring details of the current video image to produce a restored image. This significantly improves the success rate of UAV missions in complex scenarios.
[0066] In some embodiments of the present invention, comparing a historical data set and a current video image according to a comparative degradation encoder to obtain a potential degradation representation includes: Compare the historical data set and the current video image according to the contrast degradation encoder to determine the positive and negative samples; The latent degradation representation is obtained by maximizing the consistency between positive samples and minimizing the consistency between negative samples in positive samples.
[0067] CBDE, consisting of multiple convolutional layers (Conv), extracts a latent degradation representation from the degraded image of the input restored image. AirNet randomly crops a given degraded image. Since the degradation of the same image is consistent, this is considered a positive sample. Image patches from other images in the historical dataset are considered negative samples. Using the obtained contrast image patches, CBDE obtains a latent degradation representation by maximizing the consistency between positive samples while minimizing the consistency between negative samples.
[0068] The current video image and potential degradation representation are input into the degradation-guided restoration network for degradation restoration to obtain a clear video image.
[0069] The Degradation-Guided Restoration Network (DGRN) restores input images with unknown degradation types and degrees into clear video images using the latent degradation representation learned by CBDE. The DGRN consists of five Degradation-Guided Groups (DGGs), each of which is further composed of five Degradation-Guided Modules (DGMs).
[0070] The DGM module primarily consists of a deformable convolution layer (DCN) and a spatial feature transform layer (SFN). To adapt to different degradation types, the network combines the features output by the previous DGM with the underlying degradation representation and feeds them into the DGM's convolutional layer to learn offsets and masks. The receptive field can be dynamically adjusted based on the modulated offsets and masks. To narrow the distribution gap and achieve stronger resilience to multiple degradations, the SFN learns a mapping function that outputs the modulation parameters of a given underlying degradation representation. The SFN performs an affine transformation by scaling and shifting the features output by the previous DGM using the modulation parameters, and then outputs a new feature. This process is repeated across multiple DGMs, and the DGRN ultimately outputs a clear video image.
[0071] In some embodiments of the present invention, multi-scale feature extraction is performed on a clear video image to obtain a person-in-distress identification result, including: S601. Build a multi-scale target detection model; the multi-scale target detection model includes a backbone feature extraction network, a BiFormer attention module, a CBAM attention module, and a decoupling head module.
[0072] To address the issues of missed and false detections of small targets, the present invention proposes a multi-scale target detection model based on YOLOv5. This model constructs a dataset of over 10,000 images of people in distress on water, including real people in distress, simulated people in distress on water, and drowning dummies. By training a multi-scale target detection network based on YOLOv5 using this dataset, the team achieves efficient and accurate detection of targets in distress on water, effectively improving the search and rescue efficiency of aerial visual perception systems.
[0073] S602: Extract features from the clear video image according to the backbone feature extraction network to obtain initial features.
[0074] The multi-scale object detection module uses the backbone feature extraction network (Backbone) in the YOLOv5 network to extract features from clear video images and obtain initial features. The neck network (Neck) then enhances feature extraction through a feature pyramid. Finally, the prediction network (Head) predicts the object corresponding to the feature points and obtains the recognition result. The addition of the spatial pyramid pooling layer (SPPF) solves the convolutional network's need to fix the input image scale, reducing image information loss. The neck network incorporates the BiFormer attention module and the CBAM attention module, which are built based on the dynamic sparse attention module. The BiFormer attention module dynamically reduces the network's computational overhead while preserving fine-grained object features. Furthermore, the newly added CBAM attention module in the neck network combines spatial and channel-wise attention mechanisms to perform feature weighting on image data of people in distress on the water to obtain recognition results, helping to improve the detection accuracy of multi-scale objects in distress on the water. This effectively improves the efficiency and accuracy of multi-scale overboard object detection.
[0075] S603: Adaptively fuse the initial features according to the BiFormer attention module to obtain fused features.
[0076] In the traditional YOLO model, feature maps are typically processed through convolutional layers and then predicted through fully connected layers. However, this approach loses detailed information in the image. The Bi-level Routing Attention Vision Transformer (BiFormer) introduces a dual-level routing attention mechanism while employing sparse sampling to extract feature information of objects in the image. BiFormer consists of two attention layers. The first attention layer adaptively fuses features from the input feature map to extract richer semantic information. The second attention layer weights the fused features to highlight important object regions. This dual-layer attention mechanism can significantly improve the accuracy and robustness of object detection. However, while the traditional two-layer architecture design offers performance improvements, it also comes with issues such as high memory usage and computational cost. BiFormer reduces network parameters and computational complexity by collecting key-value pairs in relevant windows and leveraging sparsity to skip computation of the least relevant regions.
[0077] S604: Perform weighted processing on the fused features according to the CBAM attention module to obtain refined features.
[0078] The Convolutional Block Attention Module (CBAM) combines the spatial and channel attention mechanisms. The channel attention module extracts color features from the convolutional layer outputs. The weighted results are then compared with the spatial attention module, ultimately yielding refined features.
[0079] S605: Identify the refined features according to the decoupling head module to obtain a recognition result.
[0080] Among them, the multi-scale target detection model, that is, the improved YOLOv5 target detection network structure, can include a head structure (that is, a decoupling head module), and the head structure can include four detection heads. The four detection heads can respectively recognize the four refined features output by the CBAM attention module to obtain recognition results. The specific recognition process can be set according to actual conditions. The embodiment of the present invention is not limited here. The module can improve the network's detection accuracy for water distress targets with specific color features.
[0081] In some embodiments of the present invention, after performing multi-scale feature extraction on a clear video image to obtain a result of identifying the person in distress, the method further includes: Based on the recognition results, the drone is controlled to reach the location of the person in distress, and when the drone is controlled to descend to a preset distance, the person in distress is rescued.
[0082] In a specific embodiment of the present invention, wind is an unavoidable factor in water rescue environments. Conventional maritime drones struggle to maintain a stable flight during the release of a lifebuoy. To address this issue, the present invention improves the drone's lifebuoy release mechanism, mitigating the impact of wind-induced swaying on the drone's stability. After the maritime drone locates the person in distress based on identification results, it controls the drone to reach the person's location. The drone then descends to a preset distance (e.g., approximately 1.2 meters above the water surface) to release the inflatable lifebuoy, allowing the person to maintain their strength and thus perform a rescue.
[0083] The auxiliary rescue system for people in distress on water mainly consists of two parts: search and rescue platform and drone platform. Figure 7As shown in the figure, when water rescuers receive a water rescue mission (i.e., mission requirements) from a search and rescue platform, a search and rescue drone (UAV) takes off urgently from the base and arrives at the waters where the person in distress is located. The UAV's intelligent search path planning module plans the drone's search path. During flight, the drone's onboard imaging device captures real-time video stream data of the waters and transmits it back to the search and rescue platform. The search and rescue platform uses a visual perception enhancement module to improve the quality of the video stream in real time. The platform then uses a multi-scale object detection module to search for people in distress within the drone's field of view. Upon discovering a person in distress, the rescue platform is notified and rescue objects, such as an inflatable lifebuoy, are dropped to prolong the person's survival and probability, buying valuable time for rescuers to take further action. Rescuers then receive the drone's information and dispatch a rescue vessel to rescue the person in distress. This system improves search efficiency and rescuers' on-scene perception, enabling more accurate searches for people in distress and increasing the success rate of search and rescue missions.
[0084] In order to better implement the multi-scenario accurate identification method for persons in distress in the embodiment of the present invention, based on the multi-scenario accurate identification method for persons in distress, the embodiment of the present invention also provides a multi-scenario accurate identification device for persons in distress, such as Figure 8 As shown, the multi-scenario person-in-distress accurate identification device 800 includes: The information acquisition module 801 is used to obtain the flight data and mission requirements of the UAV, as well as the distress information of the person in distress; an area determination module 802 for determining a search area based on distress information and mission requirements, and determining an optimal flight altitude based on the search area and flight data; The image processing module 803 is used to collect the current video image from the current perspective of the drone when the drone reaches the optimal flight altitude in the search area, and perform detail restoration and degradation recovery on the current video image to obtain a clear video image; The result recognition module 804 is used to perform multi-scale feature extraction on the clear video image to obtain the recognition result of the person in distress.
[0085] The multi-scenario precise identification device 800 provided in the above embodiment can implement the technical solution described in the above-mentioned multi-scenario precise identification method embodiment. The specific implementation principles of the above-mentioned modules or units can refer to the corresponding contents in the above-mentioned multi-scenario precise identification method embodiment, which will not be repeated here.
[0086] The above is a detailed introduction to the multi-scenario method and device for accurately identifying people in distress provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A multi-scenario accurate identification method for people in distress, characterized by: include: Obtain UAV flight data and mission requirements, as well as distress information of people in distress; determining a search area based on the distress information and the mission requirements, and determining an optimal flight altitude based on the search area and the flight data; When the UAV reaches the optimal flight altitude of the search area, capturing a current video image from the current viewing angle of the UAV, and performing detail restoration and degradation recovery on the current video image to obtain a clear video image; Multi-scale feature extraction is performed on the clear video image to obtain an identification result of the person in distress.
2. The multi-scenario accurate identification method for persons in distress according to claim 1 is characterized in that: The distress information includes a probability density distribution of the initial position of the person in distress; and determining a search area according to the distress information and the task requirement includes: Determining a scannable area of the UAV on the water according to a parallel line scanning search method and the mission requirements; Establishing a person-in-distress drift model based on the influence of multiple factors on the person in distress and the probability density distribution; The drift model of the person in distress and the scannable area are simulated and optimized to obtain a search area.
3. The multi-scenario accurate identification method for persons in distress according to claim 2 is characterized in that: The simulation optimization of the drift model of the person in distress and the scannable area to obtain a search area includes: Estimate the wind speed at a preset sea level according to a preset wind pressure model and determine the disturbance coefficient; According to the disturbance coefficient, the wind-induced drift velocity is obtained; Estimate the flow velocity at the preset water depth to obtain the flow-induced drift velocity; Obtaining a drift velocity according to the wind-induced drift velocity and the flow-induced drift velocity; The drift velocity is input into the person-in-distress drift model to perform random particle simulation optimization on the scannable area to obtain a search area.
4. The multi-scenario accurate identification method for persons in distress according to claim 1 is characterized in that: The determining of the optimal flight altitude according to the search area and the flight data includes: Determining decision variables for the flight data in the search area according to a preset probability relationship model; the decision variables include coverage and route spacing; Constructing an objective function that maximizes the probability of discovery based on the coverage range and the route interval; Optimizing the maximum discovery probability of the objective function, the coverage range, and the route interval according to preset constraints and a preset parallel selection genetic algorithm to determine a target sweep width; According to the target sweep width, the optimal search flight altitude is obtained.
5. The multi-scenario accurate identification method for persons in distress according to claim 1 is characterized in that: The performing detail restoration and degradation recovery on the current video image to obtain a clear video image includes: determining current weather based on the drone; Obtaining a historical dataset from a severe weather image dataset according to the current weather; The details of the current video image are restored and degradation is restored according to the historical data set to obtain a clear video image.
6. The multi-scenario accurate identification method for persons in distress according to claim 5 is characterized in that: The generation process of the severe weather image dataset includes: Generate fog images based on the atmospheric scattering model to obtain a fog image dataset; Generate rainy day images based on the rain map model to obtain a rainy day image dataset; A severe weather image dataset is obtained according to the fog image dataset and the rainy day image dataset.
7. The multi-scenario accurate identification method for persons in distress according to claim 5, characterized in that: The performing detail restoration and degradation recovery on the current video image according to the historical data set to obtain a clear video image includes: Comparing the historical data set and the current video image according to a comparative degradation encoder to obtain a potential degradation representation; The current video image and the potential degradation representation are input into a degradation guided restoration network for degradation restoration to obtain a clear video image.
8. The multi-scenario accurate identification method for persons in distress according to claim 1 is characterized in that: The performing multi-scale feature extraction on the clear video image to obtain the identification result of the person in distress includes: Construct a multi-scale target detection model; the multi-scale target detection model includes a backbone feature extraction network, a BiFormer attention module, a CBAM attention module and a decoupling head module; Extracting features from the clear video image according to the backbone feature extraction network to obtain initial features; Performing adaptive feature fusion on the initial features according to the BiFormer attention module to obtain fused features; Performing weighted processing on the fused features according to the CBAM attention module to obtain refined features; The refined features are identified according to the decoupling head module to obtain an identification result.
9. The multi-scenario accurate identification method for persons in distress according to claim 7, characterized in that: The comparing the historical data set and the current video image according to the comparative degradation encoder to obtain a potential degradation representation includes: Comparing the historical data set and the current video image according to the contrast degradation encoder to determine positive samples and negative samples; A potential degradation representation is obtained according to maximizing the consistency between positive samples in the positive samples and minimizing the consistency between negative samples in the negative samples.
10. A multi-scenario accurate identification device for people in distress, characterized by: include: The information acquisition module is used to obtain the flight data and mission requirements of the UAV, as well as the distress information of people in distress; an area determination module, configured to determine a search area based on the distress information and the mission requirements, and to determine an optimal flight altitude based on the search area and the flight data; an image processing module, configured to capture a current video image from a current viewing angle of the drone when the drone reaches the optimal flight altitude of the search area, and perform detail restoration and degradation recovery on the current video image to obtain a clear video image; The result recognition module is used to perform multi-scale feature extraction on the clear video image to obtain the recognition result of the person in distress.
Citation Information
Patent Citations
Method for autonomously identifying people falling into water based on multi-rotor unmanned aerial vehicle
CN110321775A
Method and system for detecting and tracking marine persons in distress based on unmanned aerial vehicle
CN111709308A
Intelligent auxiliary search and rescue system for people in distress on water based on visual perception and calculation
CN112016373A
Unmanned aerial vehicle monitoring method and system for basin-wide flood scene
WO2020221284A1
KR20240083962A