Water area safety patrol method and system based on multi-modal feature analysis
By using drones to collect multimodal data and perform multimodal feature analysis, the problems of monitoring blind spots and misjudgment in reservoir water safety patrols in the existing technology are solved, and a higher patrol coverage rate and abnormal discovery accuracy are achieved.
Patent Information
- Application Number
- CN202510653383.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-21
AI Technical Summary
The prior art has blind spots in monitoring, lag in target identification or misjudgment in safety patrols of reservoir waters, especially in complex environments, which are difficult to achieve continuous dynamic monitoring of large-scale waters.
The water safety patrol method based on multimodal feature analysis is adopted to collect visible light images, infrared images and lidar data through drones, and combine target profile extraction, thermal anomaly screening and three-dimensional dynamic feature analysis to improve patrol coverage and abnormal discovery accuracy.
It effectively improves patrol coverage and abnormal discovery accuracy in complex water environments, reduces monitoring blind spots and misjudgment of target recognition, and improves the accuracy and response targeting of abnormal target recognition.
Smart Images

Figure CN120198808A_ABST
Abstract
Description
Background Art
[0002] Currently, for the safety inspection of reservoir waters, the related technologies mainly rely on deploying fixed monitoring devices around the reservoir and combining them with manual inspections for on-site inspections. In practical applications, usually, the back-end monitoring personnel observe the water conditions by remotely viewing the monitoring images, and in key areas such as around the flood discharge channels, manual inspectors are arranged to conduct on-site inspections and abnormal situation investigations.
[0003] However, due to the vast reservoir environment and the frequent fluctuations of the water surface conditions with the environment, the fixed monitoring devices are limited by the installation location, monitoring angle, and coverage area, resulting in monitoring blind spots and it is difficult to cover all key areas in a timely manner. At the same time, the back-end personnel's method of monitoring through images is limited by factors such as image clarity and environmental visibility changes, and it is easy to have situations of lagging target recognition or misjudgment. Although manual inspections can, to a certain extent, supplement the deficiencies of fixed monitoring, they are limited by human resources, inspection frequency, and emergency response speed, and it is difficult to achieve continuous dynamic monitoring of a large area of water.
[0004] Especially during flood discharge periods or under complex weather conditions, the traditional method relying on fixed monitoring and manual inspections has the risks of delayed discovery of abnormal targets, untimely response, and missed detection of potential water safety hazards. Therefore, how to improve the inspection coverage rate and the accuracy of abnormal situation discovery in complex water environments has become an urgent problem to be solved in the current water area safety inspection technology. Summary of the Invention
[0005] The purpose of the embodiments of the present disclosure is to provide a water area safety inspection method based on multi-modal feature analysis, a water area safety inspection system based on multi-modal feature analysis, an electronic device, and a computer-readable storage medium, which can improve the inspection coverage rate and the accuracy of abnormal situation discovery in complex water environments based on multi-source data collection and synchronous processing, combined with target contour extraction, thermal anomaly screening, and three-dimensional dynamic feature analysis.
[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.
[0007] According to the first aspect of the embodiments of the present disclosure, a water area safety inspection method based on multimodal feature analysis is provided, including: controlling a drone to collect visible light images, infrared images, and lidar data corresponding to a target water area based on a preset data collection strategy; performing time synchronization and spatial alignment on the visible light images, infrared images, and lidar data to obtain visible light aligned data, infrared aligned data, and lidar aligned data; in response to the clarity of the visible light image being higher than a quality reference value, performing edge detection and region segmentation processing on the visible light aligned data to generate a set of candidate regions including target contours; performing heat source feature extraction and thermal anomaly screening processing on the infrared aligned data corresponding to the set of candidate regions, and screening out thermal anomaly regions from the set of candidate regions; performing three-dimensional spatial feature extraction and motion state feature extraction on the lidar aligned data corresponding to the thermal anomaly regions, and performing anomaly recognition on the thermal anomaly regions based on the three-dimensional spatial features and the motion state features.
[0008] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the water area safety inspection method based on multimodal feature analysis further includes: screening out thermally normal regions from the set of candidate regions, and matching the target contours corresponding to the thermally normal regions with a reference contour; in response to the matching result being a human body, determining the three-dimensional spatial features corresponding to the target contour according to the lidar aligned data corresponding to the target contour; performing anomaly recognition on the thermally normal regions based on the three-dimensional spatial features.
[0009] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the water area safety inspection method based on multimodal feature analysis further includes: determining the clarity corresponding to the visible light image; in response to the clarity being lower than the quality reference value, respectively performing feature extraction on the visible light aligned data, infrared aligned data, and lidar aligned data to obtain first image features, second image features, and third image features; performing feature fusion on the first image features, second image features, and third image features to obtain multimodal fusion features; performing anomaly recognition based on the multimodal fusion features.
[0010] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the determining the clarity corresponding to the visible light image includes: performing texture analysis on the visible light image to obtain texture information of the visible light image; performing grayscale histogram statistical processing on the visible light image to obtain contrast information of the visible light image; determining the clarity corresponding to the visible light image based on the texture information and the contrast information.
[0011] In some exemplary embodiments of the present disclosure, based on the foregoing solution, determining the clarity corresponding to the visible light image includes: dividing the visible light image into regions to obtain a plurality of local image regions; determining the local gradient mean corresponding to each local image region, performing weighted fusion on the local gradient means, and determining the clarity corresponding to the visible light image based on the fusion result.
[0012] In some exemplary embodiments of the present disclosure, based on the foregoing solution, performing feature fusion on the first image feature, the second image feature, and the third image feature to obtain a multi-modal fusion feature includes: determining the fusion weights corresponding to the first image feature, the second image feature, and the third image feature according to the environmental information corresponding to the visible light image; performing weighted fusion processing on the first image feature, the second image feature, and the third image feature based on the fusion weights to obtain the multi-modal fusion feature.
[0013] In some exemplary embodiments of the present disclosure, based on the foregoing solution, before controlling the drone based on a preset data collection strategy, it further includes: determining the data collection strategy corresponding to the drone according to the current environmental information of the target water area.
[0014] In some exemplary embodiments of the present disclosure, based on the foregoing solution, determining the data collection strategy corresponding to the drone according to the current environmental information of the target water area includes: obtaining the current light intensity information, visibility information, and wind speed information of the target water area; determining the data collection strategy corresponding to the drone based on the light intensity information, visibility information, and wind speed information; wherein the data collection strategy includes the collection frequency and collection range corresponding to each type of data.
[0015] In some exemplary embodiments of the present disclosure, based on the foregoing solution, performing heat source feature extraction and thermal anomaly screening processing on the infrared alignment data corresponding to the candidate region set, and screening out thermal anomaly regions from the candidate region set includes: extracting multi-scale heat source features of the infrared alignment data corresponding to the candidate region set at multiple scales; determining the temperature distribution features corresponding to each candidate region based on the multi-scale heat source features; extracting the heat source morphology features corresponding to each candidate region based on the temperature distribution features; screening out the thermal anomaly regions having irregular edge characteristics or concentrated heat source characteristics from the candidate region set based on the heat source morphology features.
[0016] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the extraction of three-dimensional spatial features and motion state features from the lidar alignment data corresponding to the thermal anomaly region includes: based on the lidar alignment data, extracting three-dimensional structural features and surface curvature features corresponding to the thermal anomaly region, where the three-dimensional structural features include one or more of size information, volume information, bounding box compactness, and spatial information; based on the lidar alignment data collected at different times, extracting a set of multi-time trajectory points corresponding to the target contour in the thermal anomaly region; fitting a motion trajectory curve based on the set of multi-time trajectory points corresponding to the target contour, and extracting the motion state features based on the motion trajectory curve.
[0017] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the abnormal identification of the thermal anomaly region based on the three-dimensional spatial features and the motion state features includes: inputting the three-dimensional structural features, the surface curvature features, and the motion state features into an abnormal identification model to obtain an abnormal identification result; in response to the abnormal identification result being a preset target category, predicting the motion trajectory of the target contour according to the set of multi-time trajectory points.
[0018] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the above-mentioned water area safety inspection method based on multi-modal feature analysis further includes: extracting a visible light image corresponding to the abnormal region according to the abnormal identification result; visually displaying the visible light image, the abnormal identification result, and the position information of the abnormal region; generating a corresponding early warning response instruction according to the abnormal type corresponding to the abnormal identification result.
[0019] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the above-mentioned water area safety inspection method based on multi-modal feature analysis further includes: determining the spatial position information corresponding to the abnormal target based on the abnormal identification result; generating a trajectory adjustment instruction based on the spatial position information and the current position of the unmanned aerial vehicle; adjusting the flight path of the unmanned aerial vehicle based on the trajectory adjustment instruction to make the unmanned aerial vehicle approach the abnormal target; updating the trajectory adjustment instruction based on the real-time spatial relationship between the abnormal target and the unmanned aerial vehicle to optimize the observation angle and cruising radius of the unmanned aerial vehicle.
[0020] According to a second aspect of the embodiments of the present disclosure, there is provided a water area safety inspection system based on multimodal feature analysis, including: a data acquisition module, configured to control a drone to acquire visible light images, infrared images, and lidar data corresponding to a target water area based on a preset data acquisition strategy; a data alignment module, configured to perform time synchronization and spatial alignment on the visible light images, infrared images, and lidar data to obtain visible light aligned data, infrared aligned data, and lidar aligned data; a contour extraction module, configured to, in response to the clarity of the visible light image being higher than a quality reference value, perform edge detection and region segmentation processing on the visible light aligned data to generate a set of candidate regions including target contours; a thermal anomaly screening module, configured to perform heat source feature extraction and thermal anomaly screening processing on the infrared aligned data corresponding to the set of candidate regions, and screen out thermal anomaly regions from the set of candidate regions; an anomaly recognition module, configured to perform three-dimensional spatial feature extraction and motion state feature extraction on the lidar aligned data corresponding to the thermal anomaly regions, and perform anomaly recognition on the thermal anomaly regions based on the three-dimensional spatial features and the motion state features.
[0021] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method for water area safety inspection based on multimodal feature analysis as in the first aspect is implemented.
[0022] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for water area safety inspection based on multimodal feature analysis as in the first aspect is implemented.
[0023] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: In the method for water area safety inspection based on multimodal feature analysis in the exemplary embodiments of the present disclosure, first, visible light images, infrared images, and lidar data corresponding to a target water area are acquired by a drone, improving the inspection coverage rate, and through time synchronization and spatial alignment processing of data from different sources, on the basis of ensuring that various types of perception data have a unified spatio-temporal reference system, the consistency and accuracy of subsequent feature analysis are improved.
[0024] On the one hand, by judging the clarity of visible light images, when the clarity is higher than the quality reference value, edge detection and region segmentation processing are performed on the visible light alignment data, which can extract clear target contours when the image quality is good, reduce the risk of false detection and missed detection caused by environmental light changes, and enhance the reliability of the candidate region set. On the other hand, heat source feature extraction and thermal anomaly screening processing are performed based on the infrared alignment data corresponding to the candidate region set, which can effectively utilize the prominent response characteristics of infrared sensing to thermal anomalies under low light or complex meteorological conditions, thereby compensating for the recognition limitations of a single visible light image when visibility decreases, and further ensuring the accurate screening of thermal anomaly regions under different environmental conditions. On yet another hand, three-dimensional spatial feature extraction and motion state feature extraction are performed on the lidar alignment data corresponding to the thermal anomaly region, which can further determine the dynamic behavior of the abnormal target through three-dimensional structural features and trajectory change features, effectively identify different abnormal categories such as floating objects and drowning persons, and improve the accuracy of abnormal target recognition in complex water environments.
[0025] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0027] Figure 1 Schematically shows a flowchart of a water area safety inspection method based on multi-modal feature analysis according to some embodiments of the present disclosure.
[0028] Figure 2 Schematically shows a flowchart of abnormal recognition when the clarity of a visible light image is insufficient according to some embodiments of the present disclosure.
[0029] Figure 3 Schematically shows a flowchart of screening thermal anomaly regions according to some embodiments of the present disclosure.
[0030] Figure 4 Schematically shows a system architecture diagram of a water area safety inspection system based on multi-modal feature analysis according to some embodiments of the present disclosure.
[0031] Figure 5 Schematically shows a monitoring interface diagram of a water area safety inspection system based on multi-modal feature analysis according to some embodiments of the present disclosure.
[0032] Figure 6 A block diagram of a water area safety inspection system based on multi-modal feature analysis according to some embodiments of the present disclosure is schematically shown.
[0033] Figure 7 A schematic structural diagram of a computer system of an electronic device according to some embodiments of the present disclosure is schematically shown.
[0034] Figure 8 A schematic diagram of a computer-readable storage medium according to some embodiments of the present disclosure is schematically shown.
[0035] In the drawings, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed implementation manners
[0036] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0037] The terms used in this specification are for the purpose of describing particular embodiments only and are not intended to limit this specification. The singular forms "a", "the", and "said" used in this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. Example embodiments will now be described more fully with reference to the drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art.
[0038] In addition, the drawings are only schematic diagrams and are not necessarily drawn to scale. The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0039] In the present exemplary embodiment, first, a water area safety inspection method based on multi-modal feature analysis is provided. Figure 1A flowchart schematically shows a water area safety inspection method based on multimodal feature analysis according to some embodiments of the present disclosure. Refer to Figure 1 As shown, the water area safety inspection method based on multimodal feature analysis may include the following steps: Step S110, controlling a drone to collect visible light images, infrared images, and lidar data corresponding to a target water area based on a preset data collection strategy; Step S120, performing time synchronization and spatial alignment on the visible light images, infrared images, and lidar data to obtain visible light aligned data, infrared aligned data, and lidar aligned data; Step S130, in response to the clarity of the visible light image being higher than a quality reference value, performing edge detection and region segmentation processing on the visible light aligned data to generate a set of candidate regions containing target contours; Step S140, performing heat source feature extraction and heat anomaly screening processing on the infrared aligned data corresponding to the set of candidate regions, and screening out heat anomaly regions from the set of candidate regions; Step S150, performing three-dimensional spatial feature extraction and motion state feature extraction on the lidar aligned data corresponding to the heat anomaly region, and performing anomaly recognition on the heat anomaly region based on the three-dimensional spatial features and motion state features.
[0040] According to the water area safety inspection method based on multimodal feature analysis in this exemplary embodiment, visible light images, infrared images, and lidar data corresponding to a target water area are collected by a drone, improving the inspection coverage rate. On the one hand, by judging the clarity of the visible light image, when the clarity is higher than the quality reference value, edge detection and region segmentation processing are performed on the visible light aligned data, which can extract clear target contours when the image quality is good, reduce the risk of false detection and missed detection caused by environmental light changes, and enhance the reliability of the set of candidate regions. On the other hand, heat source feature extraction and heat anomaly screening processing are performed based on the infrared aligned data corresponding to the set of candidate regions, which can effectively utilize the prominent response characteristics of infrared sensing to heat source anomalies under low light or complex meteorological conditions, thereby compensating for the recognition limitations of a single visible light image when visibility decreases, and further ensuring the accurate screening of heat anomaly regions under different environmental conditions. On the further hand, three-dimensional spatial feature extraction and motion state feature extraction are performed on the lidar aligned data corresponding to the heat anomaly region, which can further distinguish the dynamic behavior of abnormal targets through three-dimensional structural features and trajectory change features, effectively identify different abnormal categories such as floating objects and drowning persons, and improve the accuracy of identifying abnormal targets in complex water area environments.
[0041] Next, the water area safety inspection method based on multimodal feature analysis in this exemplary embodiment will be further described.
[0042] Step S110: Control the drone to collect visible light images, infrared images, and lidar data corresponding to the target water area based on a preset data collection strategy.
[0043] Among them, the data collection strategy can represent a data collection plan preset for different environmental conditions, and can include startup methods of collection devices, operating frequencies, collection time intervals, collection modes, collection ranges, and other suitable device configuration methods. The target water area can represent the spatial area where the inspection task is to be performed, and can include water body areas such as reservoirs, flood discharge areas, lakes, and rivers with specific geographical boundaries and safety monitoring requirements. Visible light images can represent two-dimensional image data in the visible light band obtained by the imaging device carried on the drone, and are used to reflect the visual feature information of the target water area surface under natural light conditions. Infrared images can represent thermal imaging image data formed based on the thermal radiation intensity obtained by the infrared imaging device, usually reflecting the temperature distribution characteristics of different objects or regions within the target area, and are used to identify heat source targets. Lidar data can represent three-dimensional point cloud data of the target water area collected by the lidar sensor, and the spatial position information of the target surface is obtained based on the time difference between the emitted laser pulse and the received echo, and is used to reconstruct the spatial structure characteristics of the target area.
[0044] By using drones for water area safety inspections, high-efficiency monitoring of large-scale water area regions can be achieved. Compared with fixed monitoring devices, drones have the ability to dynamically adjust flight routes and can obtain data from multiple angles and viewpoints within the task range, thus significantly improving the coverage ability of water area inspections. Further, by carrying multiple types of sensors to simultaneously obtain visible light images, infrared images, and lidar data, multi-source perception information with complementary characteristics can be obtained under complex environmental conditions, providing better data support for subsequent data fusion, target recognition, and anomaly detection.
[0045] Step S120: Perform time synchronization and spatial alignment on the visible light images, infrared images, and lidar data to obtain visible light aligned data, infrared aligned data, and lidar aligned data.
[0046] Among them, time synchronization can represent the process of uniformly adjusting the acquisition timings of visible light images, infrared images, and lidar data based on the timestamp information of various types of data, which is used to ensure the correspondence of various types of data in the time dimension. Spatial alignment can represent the process of converting and registering the spatial coordinates of the data collected by multiple types of sensors, so as to achieve the position consistency of multi-source data in the spatial dimension. Visible light aligned data can represent the image data obtained by extracting from the visible light image under the unified spatio-temporal reference after time synchronization and spatial alignment processing. Infrared aligned data can represent the infrared image data obtained by time synchronization and spatial alignment based on the infrared image, which retains the temperature distribution information on the premise of sharing the same spatio-temporal coordinate system with other data types. Lidar aligned data can represent the point cloud data generated after time synchronization and spatial coordinate conversion of the lidar data, which has the same timing and spatial position reference as the image data.
[0047] Due to the differences in the acquisition frequencies, viewing angle parameters, and imaging mechanisms of different modality sensors, if joint processing is directly performed, it is easy to cause inconsistent target feature expressions due to time offset or spatial misalignment. Through time synchronization processing, it is possible to ensure that various types of data are corresponding and matched under the same time reference, and avoid the problem of target misalignment caused by time offset. Through spatial alignment processing, it is possible to accurately map the data of each modality into a unified spatial framework, and ensure the correspondence and comparability of the features of each modality in terms of position during the subsequent analysis process.
[0048] Step S130, in response to the clarity of the visible light image being higher than the quality reference value, perform edge detection and region segmentation processing on the visible light aligned data to generate a set of candidate regions containing the target contour.
[0049] Among them, clarity can be an indicator used to characterize the ability of the local structural details distribution of an image, and can be used to reflect the image quality. The quality reference value can be a clarity threshold used to determine whether an image meets the requirements of subsequent processing, and is used to judge whether the current image has the conditions for edge detection and region segmentation processing. Edge detection can be an image processing process for identifying and extracting positions with significant gray or color changes in an image, and is used to highlight the boundary information between objects and the background in the image. Region segmentation processing can be a processing method for dividing image regions in an image based on pixel feature similarity, edge information, or semantic information, and is used to divide the image into multiple sub-regions with structural consistency to extract the spatial range of potential targets. The target contour can be a closed or approximately closed boundary structure extracted through edge detection and region segmentation in an image, and is used to describe the geometric shape characteristics of possible abnormal targets on the image plane. The target contour can include human contours, ship contours, floating object contours, obstacle contours, and other suitable monitoring target contours within the water area. The candidate region set can be a set composed of multiple image sub-regions containing target contours.
[0050] Among them, subsequent edge detection and region segmentation processing procedures are only carried out when the clarity of the visible light image is higher than the quality reference value, avoiding the risk of unstable extraction results in the case of blurred or severely interfered images. By performing edge detection and region segmentation processing on visible light alignment data with satisfactory clarity, image regions with structural continuity and prominent boundaries can be effectively extracted, and then a candidate region set containing various target contours such as human contours, ship contours, floating object contours, and obstacle contours can be generated.
[0051] Step S140: Perform heat source feature extraction and thermal anomaly screening processing on the infrared alignment data corresponding to the candidate region set, and screen out thermal anomaly regions from the candidate region set.
[0052] Among them, heat source feature extraction can be a process of extracting image features used to characterize the temperature characteristics of the target region based on the temperature distribution information of each candidate region in the infrared alignment data. Thermal anomaly screening can be a processing operation for identifying and screening regions in the candidate region set whose temperature characteristics significantly deviate from the background distribution or exceed the preset temperature threshold based on the heat source feature extraction results. The thermal anomaly region can be a target region in the candidate region set determined to have significant thermal radiation characteristics through the thermal anomaly screening process.
[0053] Performing heat source feature extraction and thermal anomaly screening processing based on the infrared alignment data corresponding to the candidate region set can utilize the characteristic that infrared images are sensitive to temperature distribution changes to identify regions with significant thermal radiation characteristics without relying on the clarity of visible light. Since drowning persons are usually accompanied by obvious heat source signals, small boats and their power devices, some floating objects or obstacles may also show local temperature difference responses in thermal imaging. Therefore, by extracting the heat source characteristics of each candidate region in the infrared alignment data and combining the background temperature distribution for anomaly screening, regions with insufficient thermal response or similar to the water body background can be effectively removed, and thermal anomaly regions can be further screened out. The above processing not only improves the positioning ability of targets such as drowning persons, but also provides a significant preliminary screening basis for subsequent multi-modal recognition based on spatial structure and motion characteristics.
[0054] Step S150: Extract three-dimensional spatial features and motion state features from the lidar alignment data corresponding to the thermal anomaly region, and perform anomaly recognition on the thermal anomaly region based on the three-dimensional spatial features and motion state features.
[0055] Among them, the three-dimensional spatial features can represent a set of spatial structure parameters extracted based on the lidar alignment data, which can be used to describe the geometric shape and volume distribution of the target in the three-dimensional coordinate system. The motion state features can represent a set of target dynamic behavior parameters extracted based on the lidar alignment data at multiple time points, and can include the spatial displacement change, velocity vector, acceleration trend, trajectory fitting curve, and trajectory stability of the target in consecutive time frames, etc., which are used to describe the motion characteristics of the target in the water area. Anomaly recognition can represent a comprehensive judgment process based on the three-dimensional spatial features and motion state features, which is used to identify whether the thermal anomaly region shows spatial shape and dynamic behavior characteristics consistent with known target types (such as drowning persons, floating objects, boats, obstacles, etc.), and distinguish and mark target regions with abnormal contour features or abnormal motion patterns from the thermal anomaly regions.
[0056] Performing three-dimensional spatial feature extraction and motion state feature extraction on the lidar alignment data corresponding to the thermal anomaly region can further utilize the advantages of lidar in structural modeling and dynamic tracking on the basis of the initial screening of heat sources, and complement the target shape and behavior information. The three-dimensional spatial features can reflect the external dimensions, structural volume, and curvature complexity of the thermal anomaly region in space, and are used to distinguish regular-shaped targets (such as ships and obstacles) from irregular or partially submerged targets (such as drowning persons). The motion state features can extract the trajectory continuity, speed change, and stability trend of the target based on multiple frames of lidar data, and are used to identify motion behavior patterns such as floating objects drifting with the water, small boats advancing smoothly, and drowning persons struggling and shaking. Using the above two types of features for anomaly recognition processing can effectively distinguish different types of monitoring objects in the thermal anomaly region, strengthen the discrimination ability for drowning persons, and improve the recognition accuracy and response pertinence in complex dynamic water environments.
[0057] Next, the content in steps S110 to S150 will be described in detail.
[0058] In some embodiments, the time synchronization and spatial alignment of the visible light image, infrared image, and lidar data in step S120 can be achieved through the following technical process to obtain visible light alignment data, infrared alignment data, and lidar alignment data, specifically including: obtaining the timestamp information corresponding to the visible light image, infrared image, and lidar data to establish the acquisition time sequence of various types of data. Based on the timestamp information, perform time synchronization processing on the visible light image, infrared image, and lidar data to generate intermediate synchronization data with a unified time reference. Obtain the calibration parameters for registration, and perform coordinate transformation processing on the intermediate synchronization data based on the calibration parameters. Based on the results of the coordinate transformation processing, generate visible light alignment data, infrared alignment data, and lidar alignment data corresponding to the unified spatial reference coordinate system respectively.
[0059] Among them, obtaining the calibration parameters for registration and performing coordinate transformation processing on the intermediate synchronization data based on the calibration parameters can include: obtaining the relative pose information of various sensors in the UAV installation structure as the basic calibration parameters for constructing the coordinate system relationship between sensors, where the calibration parameters include the rotation matrix and translation vector of each sensor relative to the unified spatial reference coordinate system. Substitute the original coordinate values corresponding to various types of perception data in the intermediate synchronization data into the coordinate transformation model composed of the rotation matrix and translation vector to perform three-dimensional coordinate transformation operations to obtain the transformed data set in the unified coordinate system.
[0060] Based on the results of coordinate transformation processing, visible light alignment data, infrared alignment data, and lidar alignment data corresponding to the unified spatial reference coordinate system can be generated respectively, which may include: classifying the transformation data set according to the sensor source, respectively extracting the transformed image pixel coordinates and point cloud coordinate values, and constructing visible light alignment data, infrared alignment data, and lidar alignment data consistent with the unified spatial reference coordinate system. Among them, the visible light alignment data and the infrared alignment data are respectively corrected for coordinate mapping by means of image reprojection, while the lidar alignment data retains the original point cloud spatial accuracy to support subsequent multi-modal spatial consistency analysis and feature fusion processing.
[0061] In some embodiments, the edge detection and region segmentation processing of the visible light alignment data in step S130 can be implemented through the following technical process to generate a set of candidate regions containing the target contour, which specifically includes: performing image preprocessing operations on the visible light alignment data to enhance the edge contrast of the image and suppress noise interference; based on the preprocessed visible light alignment data, performing edge detection processing and extracting the edge pixel distribution in the image to generate an edge image; performing region segmentation processing on the edge image to divide the edge image into multiple image sub-regions with boundary structures; screening out the regions with closed boundary features and complete contours from the image sub-regions to construct a set of candidate regions containing the target contour.
[0062] In specific implementation, when performing image preprocessing operations on the visible light alignment data, the visible light alignment data can be subjected to image grayscale conversion, histogram equalization, and edge enhancement filtering processing to enhance the brightness contrast between the boundary region and the background region in the image, and reduce the noise interference introduced by factors such as environmental reflection and water surface disturbance through median filtering or Gaussian filtering to improve the response clarity of the edge region. When generating the edge image, Sobel operator, Canny operator, or other edge detection algorithms based on gradient change can be used to calculate the response to the pixel gradient change in the image, extract the pixel set at the strong edge positions in the image, and convert the set into an edge image with a binary edge structure. Then, using the connected component analysis algorithm or other suitable region segmentation algorithms, the adjacent edge pixels in the edge image are aggregated, image sub-regions composed of multiple continuous pixels are extracted, and the boundary range and geometric shape information of each image sub-region are recorded. Finally, the shape structure of the image sub-regions is analyzed to determine whether the region edges form a closed curve structure, and regions with complete geometric contours are screened out in combination with region area, aspect ratio, and boundary coherence indicators, and such regions are constructed into a set of candidate regions.
[0063] In some embodiments, it can be through Figure 2The steps S210 to S240 shown above implement the heat source feature extraction and thermal anomaly screening process for the infrared alignment data corresponding to the candidate region set in step S140 above, and screen out the thermal anomaly regions from the candidate region set, specifically including: Step S210, extract the multi-scale heat source features of the infrared alignment data corresponding to the candidate region set at multiple scales.
[0064] Among them, the multi-scale heat source features can represent a set of temperature response feature sets obtained by extracting features from the infrared alignment data corresponding to the candidate region set at different spatial scales. Based on the infrared alignment data corresponding to the candidate region set, a scale-variable window operator is used to perform heat source feature extraction operations on each candidate region respectively, and a multi-scale heat source feature vector is constructed: Among them, represents the multi-scale heat source feature vector corresponding to the candidate region ; represents the heat intensity response calculated under the th scale window, represents the spatial standard deviation of the th scale window, represents the total number of scales.
[0065] Step S220, based on the multi-scale heat source features, determine the temperature distribution features corresponding to each candidate region.
[0066] Among them, the temperature distribution features can represent a set of features used to describe the overall distribution structure of the temperature values of each pixel point in the candidate region in the infrared alignment data. Based on the multi-scale heat source feature vector, the temperature distribution features at each scale within the candidate region are calculated respectively: Among them, represents the temperature distribution feature of the candidate region ; represents the temperature value of the pixel point at position within the candidate region at the th scale, represents the average temperature of the candidate region at the th scale, represents the number of pixel points included in the candidate region , represents the total number of scales.
[0067] Step S230, based on the temperature distribution features, extract the heat source morphology features corresponding to each candidate region.
[0068] Among them, the heat source morphological features can represent a set of image features extracted from an infrared image for describing the boundary structure of the heat source and the shape of the thermal energy distribution. Based on the temperature distribution features, the heat source morphological features are extracted for each candidate region respectively: Among them, represents the heat source morphological features of the candidate region . represents the regional edge irregularity parameter, which is defined as the variance of the regional edge curvature change. represents the heat source concentration parameter, which is defined as the ratio of the number of pixels above the temperature threshold in the region to the total number of pixels in the region. represents the heat source centroid offset parameter, which is defined as the distance from the heat source centroid position to the geometric center of the region, and the weight coefficient is the fusion weight parameter, satisfying .
[0069] Step S240, based on the heat source morphological features, screen out the thermal anomaly regions with irregular edge characteristics or concentrated heat source characteristics from the candidate region set.
[0070] Among them, the irregular edge characteristics can represent the manifestation form that the heat source region has an irregular boundary contour in the infrared image. This boundary can be a broken, serrated, uneven diffusion or a closed or semi-closed structure without obvious geometric rules, and can be used to indicate the thermal contour characteristics of abnormal heat sources such as humans, floating objects and other non-structural targets. According to the heat source morphological features , set the thermal anomaly screening threshold , and select the regions from the candidate region set that meet the following conditions: Mark the candidate regions that meet the above conditions as thermal anomaly regions, and output the thermal anomaly regions for subsequent abnormal target recognition.
[0071] In some embodiments, the above-mentioned three-dimensional space feature extraction and motion state feature extraction of the lidar alignment data corresponding to the thermal anomaly region in step S150 can be implemented through the following technical process, specifically including: First, based on the lidar alignment data, extract the three-dimensional structure features and surface curvature features corresponding to the thermal anomaly region. Among them, the three-dimensional structure features include one or more of size information, volume information, bounding box compactness, and spatial information.
[0072] Among them, the bounding box compactness can represent the spatial utilization efficiency of the outer bounding box of the three-dimensional point cloud corresponding to the thermal anomaly region in the lidar-aligned data, which is calculated by the ratio of the actual point cloud volume of the thermal anomaly region to the volume of its minimum circumscribed cube. The spatial information can represent the three-dimensional position, distribution range, and spatial density and other structural characteristics of the point cloud corresponding to the thermal anomaly region in the lidar-aligned data. The surface curvature feature can represent the curvature change of the fitted surface constructed based on the local point cloud in the lidar-aligned data at each point, which can be characterized by the mean curvature, Gaussian curvature, or local curvature variance, etc.
[0073] Then, based on the lidar-aligned data collected at different times, a set of multi-time trajectory points corresponding to the target contour in the thermal anomaly region is extracted.
[0074] Specifically, at multiple sampling times, the centroid coordinate points of the target contour in the thermal anomaly region are respectively extracted from the lidar-aligned data and a set of multi-time trajectory points of the target contour in the time series is constructed where represents the three-dimensional spatial centroid position of the target contour at time , , , respectively represent the coordinate components of the target in the X, Y, and Z axis directions at the corresponding times, represents the total number of times of trajectory sampling. Among them, the target contour can be a human contour, a ship contour, a floating object contour, etc., and the specific category can be adaptively set according to the application scenario.
[0075] Finally, based on the set of multi-time trajectory points corresponding to the target contour, a motion trajectory curve is fitted, and motion state features are extracted based on the motion trajectory curve.
[0076] Specifically, the set of multi-time trajectory points is processed by the least squares fitting method for time parameterization to construct a three-dimensional smooth trajectory function to obtain the motion trajectory curve: where represents the three-dimensional position vector of the fitted target trajectory point at time , is the polynomial coefficient vector of the three-dimensional trajectory fitting, represents the order of the fitted curve, , , respectively can represent the position coordinate functions of the target in the X-axis, Y-axis, and Z-axis directions at time , may represent a time variable used to describe the change of the target trajectory, and may represent the term index in polynomial trajectory fitting.
[0077] After obtaining the motion trajectory fitting curve, other appropriate motion state features such as the average speed feature, the acceleration fluctuation feature, and the average trajectory curvature can be calculated based on the motion trajectory fitting curve.
[0078] In some embodiments, the above-mentioned abnormal recognition of the thermal abnormal area based on the three-dimensional space feature and the motion state feature in step S150 can be implemented through the following technical process, specifically including: inputting the three-dimensional structure feature, the surface curvature feature, and the motion state feature into the abnormal recognition model to obtain an abnormal recognition result; in response to the abnormal recognition result being a preset target category, predicting the motion trajectory of the target contour according to the multi-moment trajectory point set.
[0079] Among them, the abnormal recognition model may represent a calculation model for classifying the target corresponding to the area based on the three-dimensional space feature and the motion state feature of the thermal abnormal area. Exemplarily, the abnormal recognition model can be constructed based on other appropriate classification models such as a multi-layer perceptron neural network, a support vector machine classifier, and a random forest model. The abnormal recognition result may represent the category prediction output of the abnormal recognition model for the input thermal abnormal area, and the result may include the category label of the target and its corresponding probability score. Among them, the category label may include a person falling into the water, a floating object, a boat, or an obstacle, etc. The target category may represent the recognition target type with a response priority or special warning significance predefined in the water area safety inspection task, specifically a person falling into the water or a floating object, etc.
[0080] Preferably, an abnormal recognition model can be constructed based on a multi-layer perceptron neural network, and the motion trajectory of the target contour can be predicted by a long short-term memory network model. Specifically, in the case where the abnormal recognition result indicates that the target of the thermal abnormal area belongs to the preset target category, the multi-moment trajectory point set corresponding to the target contour is encoded into a trajectory input sequence in chronological order. The trajectory input sequence is input into the trained long short-term memory network model, and the predicted values of the trajectory positions of the target at multiple future moments are obtained through progressive time series modeling. The long short-term memory network can record the hidden state information and historical memory information within multiple time steps, and is used to learn the continuous evolution pattern of the target contour in the time dimension, and then output the three-dimensional motion trajectory sequence of the target contour within the prediction time interval.
[0081] When the recognition result indicates that the target is a preset target category, such as a person falling into the water, further predicting the movement trajectory of the target contour based on the multi-moment trajectory point set can effectively restore the movement trend of the target in the future period and dynamically reflect its drift path, thereby providing a more accurate decision-making basis for water rescue response and improving the rescue success rate of the drowning person.
[0082] In some embodiments, the data acquisition strategy corresponding to the UAV can be determined according to the current environmental information of the target water area through the following technical steps. Among them, by determining the data acquisition strategy corresponding to the UAV according to the current environmental information of the target water area, the dynamic adjustment of the sensor working mode and acquisition frequency can be realized, thereby improving the effectiveness and adaptability of data acquisition. For example, in the daytime scene with sufficient light, high-frequency visible light image acquisition is preferentially adopted, and in the low visibility or night scene, it is dynamically switched to infrared imaging and lidar encrypted scanning, effectively avoiding the problem of insufficient performance of a single sensor in a specific environment.
[0083] In some embodiments, determining the data acquisition strategy corresponding to the UAV according to the current environmental information of the target water area specifically includes the following technical processes: obtaining the current light intensity information, visibility information, and wind speed information of the target water area; determining the data acquisition strategy corresponding to the UAV based on the light intensity information, visibility information, and wind speed information; wherein, the data acquisition strategy includes the acquisition frequency and acquisition range corresponding to each type of data.
[0084] Among them, the acquisition frequency can represent the number of times the image sensor or lidar carried by the UAV performs data acquisition per unit time. The acquisition frequency can represent the spatial area or field of view boundary covered by the image sensor or lidar carried by the UAV.
[0085] Exemplarily, the light intensity information can represent the level of light radiation energy received per unit area of the target area, preferably characterized in lux (Lux); the visibility information can represent the maximum horizontal distance at which the UAV sensor can clearly identify the target under the current meteorological conditions; the wind speed information can represent the real-time wind speed magnitude at the spatial position where the UAV is located. Based on the above three types of environmental information, an environmental state vector is constructed, and combined with the preset data acquisition strategy mapping rules, the acquisition frequency and acquisition range corresponding to visible light images, infrared images, and lidar data are respectively determined. Preferably, when the light intensity is high, the visibility is good, and the wind speed is low, high-frequency acquisition of visible light images (such as 30 frames per second) and wide-angle coverage (such as a 120° horizontal viewing angle) are configured; when the light intensity decreases or the visibility drops to the set threshold, the acquisition frequency of infrared images is increased, and the acquisition range of visible light images is reduced to reduce the risk of defocus; when the wind speed increases beyond the flight stability threshold, the acquisition range of lidar is appropriately reduced and its sampling density is increased to enhance the structural recognition ability of distant targets.
[0086] In some embodiments, for the thermally normal regions in the candidate region set, anomaly recognition can be performed through the following technical steps: screen out the thermally normal regions from the candidate region set, and match the target contour corresponding to the thermally normal region with the reference contour; in response to the matching result being a human body, determine the three-dimensional spatial features corresponding to the target contour according to the lidar alignment data corresponding to the target contour; and perform anomaly recognition on the thermally normal region based on the three-dimensional spatial features.
[0087] Among them, the reference contour can represent a set of standardized target shape model contours pre-constructed and stored in the system, and is used for spatial shape matching and similarity analysis with the actually recognized target contour. The reference contour can be generated according to the two-dimensional projection contour features of the typical human body structure in different postures and perspectives. Preferably, it includes the boundary shapes of common drowning human body postures such as standing sideways, floating horizontally, and curling and floating, and is represented in the form of a contour point set or an edge function, and is used for calculating the contour coincidence degree or shape template matching with the target contour extracted from the thermally normal region. The three-dimensional spatial features can represent a set of three-dimensional structure information of the thermally normal region target constructed based on the lidar alignment data.
[0088] In specific implementation, the regions other than the thermally abnormal regions in the candidate region set can be used as the thermally normal regions, and then the contour point set of the target contour of the thermally normal region in the visible light alignment data is obtained to construct a target contour set , where represents the -th edge point coordinate of the target contour in the image coordinate system, represents the number of contour points. The target contour is subjected to spatial shape matching with the preset reference contour set , and by constructing a normalized contour distance loss function: Among them, represents the minimum matching loss between the target contour and the reference contour, represents the two-dimensional Euclidean distance, represents the two-dimensional spatial coordinate difference between the target contour point and the reference contour point in the thermally normal region.
[0089] Further, to enhance the structural stability of the contour shape matching, a contour edge direction gradient similarity index is introduced: Among them, represents the edge direction similarity score, represents the number of contour points, represents the target contour at the point The normal direction at is represented as being matched with the reference contour point of the normal direction. The higher the similarity, the closer it approaches 1. When both is less than the maximum matching distance error threshold and is greater than the minimum requirement threshold of the edge direction gradient similarity, it is determined that the target contour matches the reference contour, and the matching result is of the human body category.
[0090] Finally, in response to the matching result being a human body, the spatial point cloud data of the corresponding contour area is extracted from the lidar alignment data, and based on the spatial projection and boundary mapping method, its three-dimensional spatial feature set is constructed. The three-dimensional spatial features include, but are not limited to, structural feature parameters such as size information, volume information, boundary compactness, and point cloud density. The three-dimensional spatial features can be input into the anomaly recognition and determination model to perform a secondary determination on whether there is a situation of a person falling into the water.
[0091] In some embodiments, referring to Figure 3 as shown, when the clarity of the visible light image does not meet the preset conditions, anomaly recognition can be performed through the following technical steps, specifically including: Step S310, determining the clarity of the visible light image.
[0092] Among them, the clarity of the visible light image can be represented by other suitable image information such as texture information, contrast information, and gradient mean.
[0093] Step S320, in response to the clarity being lower than the quality reference value, feature extraction is respectively performed on the visible light alignment data, the infrared alignment data, and the lidar alignment data to obtain the first image feature, the second image feature, and the third image feature.
[0094] Among them, the first image feature can represent the image structure information and texture information extracted based on the visible light alignment data. The second image feature can represent the thermal radiation response information extracted based on the infrared alignment data. The third image feature can represent the spatial structure feature extracted based on the lidar alignment data.
[0095] Step S330, performing feature fusion on the first image feature, the second image feature, and the third image feature to obtain a multi-modal fusion feature. Among them, feature fusion can adopt other suitable methods such as weighted fusion and feature splicing.
[0096] Step S340, performing anomaly recognition based on the multi-modal fusion feature.
[0097] Preferably, the anomaly recognition model can be constructed based on a multi-layer perceptron neural network, which includes an input layer, at least two hidden layers, and an output layer. Each hidden layer uses a non-linear activation function to perform layer-by-layer transformation on the fused feature vector. The input layer receives the fused vector composed of three types of modal features, extracts high-order discriminant features through fully connected operations and activation mapping in the hidden layer, and finally generates an anomaly recognition result in the output layer. The output result can be in the form of classification, used to determine whether the detected target is a drowning person, a floating object, a ship, or a non-anomalous background area, and output the corresponding confidence score.
[0098] In this embodiment, when the clarity of the visible light image is lower than the preset quality reference value, anomaly recognition based on the fused features can make full use of the infrared thermal radiation information and lidar spatial structure information for compensation analysis when the quality of the visible light image is insufficient to support effective recognition, thereby improving the recognition stability and accuracy of anomalous targets in complex environments such as low light, strong reflection, or blurred images.
[0099] In some embodiments, the clarity corresponding to the visible light image can be determined through the following technical process: performing texture analysis on the visible light image to obtain the texture information of the visible light image; performing gray histogram statistical processing on the visible light image to obtain the contrast information of the visible light image; and determining the clarity corresponding to the visible light image based on the texture information and the contrast information.
[0100] Specifically, first, perform texture analysis on the visible light image based on the gray change situation within the local neighborhood of the image to construct a texture complexity matrix , where represents the local texture intensity of the pixel coordinate point in the image. Preferably, the texture intensity can be calculated using the mode distribution frequency or gradient response amplitude after local binary pattern encoding. Perform normalized mean statistics on the texture intensity within the entire image area to obtain the global texture index of the image: where, represents the image texture complexity score, and represent the width and height of the image respectively, represent the pixel coordinates in the horizontal and vertical directions of the image respectively.
[0101] Subsequently, perform contrast evaluation based on the gray histogram of the visible light image to construct an image gray intensity histogram , where represents the gray value, represents the pixel frequency corresponding to the gray value. Calculate the standard deviation of this histogram: Among them, represents the gray-scale contrast score of the image, represents the average gray value of the image, represents the total number of pixels in the image.
[0102] Finally, the image texture complexity score and the gray-scale contrast score are weighted and fused to obtain the image quality score corresponding to the visible light image, that is, the clarity corresponding to the visible light image.
[0103] In some embodiments, the clarity corresponding to the visible light image can also be determined through the following technical process: the visible light image is divided into regions to obtain multiple local image regions; the local gradient mean corresponding to each local image region is determined, the local gradient means are weighted and fused, and the clarity corresponding to the visible light image is determined based on the fusion result.
[0104] Among them, the local gradient mean can represent the average of the gradient intensities of all pixel points in a local image region of the image, which is used to characterize the severity of the brightness change in the region and further reflect the clarity of the edge details. Preferably, the local gradient intensity can be calculated by performing a first-order gradient operator, such as the Sobel operator, on each pixel point. In a specific implementation, the visible light image can be divided into several image sub-regions of the same size. For the pixel points in each sub-region, the gray-scale gradients in the horizontal and vertical directions are calculated respectively using the Sobel operator, and the gradient amplitude of each pixel point is obtained accordingly; the gradient amplitudes of all pixel points in the sub-region are averaged to obtain the local gradient mean of the region; then, the local gradient means are weighted and fused according to the weights set by the regional position or brightness distribution to obtain a fusion gradient index, and finally, this fusion index is used as the clarity score of the image to measure whether the current visible light image meets the abnormal recognition requirements.
[0105] In some embodiments, feature fusion is performed on the first image feature, the second image feature, and the third image feature to obtain a multi-modal fusion feature, which specifically includes: determining the fusion weights corresponding to the first image feature, the second image feature, and the third image feature according to the environmental information corresponding to the visible light image; performing a weighted fusion process on the first image feature, the second image feature, and the third image feature based on the fusion weights to obtain a multi-modal fusion feature.
[0106] Among them, by dynamically generating fusion weights based on the environmental information corresponding to the visible light image, the system can adaptively select the optimal feature source, reasonably allocate the weights of visible light, infrared, and lidar features under different lighting, visibility, or imaging conditions, thereby improving the adaptability of multi-modal feature fusion and the accuracy of anomaly recognition. In the specific implementation process, the fusion weights corresponding to the first image feature, the second image feature, and the third image feature can be determined according to the lighting intensity information, visibility information, and wind speed information corresponding to the acquisition of the visible light image. For example, in an environmental condition with high lighting intensity, good visibility, and low wind speed, the fusion weight of the first image feature is preferentially increased, and the fusion weights of the second image feature and the third image feature are appropriately reduced; when the lighting intensity is low or the visibility decreases, the fusion weight of the second image feature is correspondingly increased to enhance the discrimination ability for heat source targets; when the wind speed increases, the fusion weight of the third image feature is enhanced.
[0107] In some embodiments, after obtaining the anomaly recognition result, an early warning can also be performed based on the following technical steps, which specifically include: extracting the visible light image corresponding to the anomaly area according to the anomaly recognition result; visually displaying the visible light image, the anomaly recognition result, and the position information of the anomaly area; and generating a corresponding early warning response instruction according to the anomaly type corresponding to the anomaly recognition result.
[0108] In the specific implementation, when the anomaly recognition is completed, the visible light alignment data corresponding to the anomaly area in time and space is obtained, and the image segment covering the anomaly area is extracted based on the visible light alignment data to construct the visible light image corresponding to the anomaly area. Based on the above visible light image, the target category information and the corresponding confidence information included in the anomaly recognition result are marked, and combined with the three-dimensional position information of the anomaly area, a visual image interface with recognition results and spatial annotations is constructed to highlight the anomaly area and category annotation in the image. Preferably, the distribution of trajectory points of the target at historical and predicted times can also be superimposed on the interface to assist in judging its movement direction and trend. Further, according to the anomaly type characterized by the anomaly recognition result, the response rule corresponding to this type is called, and an early warning response instruction is generated. The early warning response instruction can include response level, response method, and rescue strategy information. Among them, the response level is used to distinguish the severity of the event, the response method is used to specify the output form of the warning content (such as interface pop-up window, sound and light prompt, or remote notification), and the linkage strategy is used to control subsequent response actions such as drone tracking, rescue equipment activation, or command terminal synchronization.
[0109] In some embodiments, after obtaining the anomaly recognition result, the following technical steps can be further taken for precise inspection of the drone, which specifically include: based on the anomaly recognition result, determining the spatial position information corresponding to the anomaly target; based on the spatial position information and the current position of the drone, generating a flight path adjustment instruction; based on the flight path adjustment instruction, adjusting the flight path of the drone to make the drone approach the anomaly target; based on the real-time spatial relationship between the anomaly target and the drone, updating the flight path adjustment instruction to optimize the observation angle and cruising radius of the drone.
[0110] When the recognition result output by the anomaly recognition model indicates that the target belongs to a preset target category, such as a person falling into the water or an abnormal vessel, the three-dimensional spatial position information corresponding to the anomaly target is extracted from the lidar alignment data, and this position is used as the target navigation reference point. Further, the flight attitude parameters and spatial position information of the drone at the current moment are obtained, the spatial vector relationship between the anomaly target and the drone is constructed, and an initial flight path adjustment instruction is generated based on the principle of the shortest flight distance and the safety approach constraint. Execute the flight path adjustment instruction to dynamically guide the drone to deflect the current flight path and correct the attitude, so that its flight direction gradually deviates towards the anomaly target space area. During the flight, continuously update the spatial position of the anomaly target and the current pose data of the drone, calculate the relative azimuth angle and distance information between the two in real time, and optimize the adjustment instruction based on strategies such as multi-angle coverage priority, constant radius approach, or low-energy stable hover, so that the drone can track and approach the anomaly target and achieve the best perspective coverage while maintaining observation stability, thereby improving the inspection response ability and image acquisition quality in a high-dynamic target environment.
[0111] Furthermore, Figure 4 Schematically shows a system architecture diagram of a water area safety inspection system based on multi-modal feature analysis according to some embodiments of the present disclosure, which is applied to the safety inspection and early warning response of a target reservoir.
[0112] Among them, the inspection system is composed of a public platform, an application layer, a service layer, a data access layer, a persistence layer, and a security management module. The public platform module integrates function entrances such as a console, inspection trajectory, real-time image, data analysis, early warning setting, information recording, and remote monitoring. Users can control the drone through the console to collect visible light images, infrared images, and lidar data corresponding to the target water area based on a preset data collection strategy. The relevant image data is displayed through the real-time image module for the data analysis module to perform edge detection and region segmentation processing to generate a set of candidate regions containing the target contour. The system application layer includes sub-modules such as inspection and monitoring, anomaly early warning, data analysis, image recognition, and device management. The specific functions interact with the service layer through WS protocol, HTTP protocol, and FTP protocol to complete image processing, data analysis, anomaly detection, data collection, and early warning notification services.
[0113] In the service layer, the data acquisition service is used to collect visible light images, infrared images, and lidar data corresponding to the target water area. The image processing service is used to perform preprocessing operations on the visible light images and extract the first image features. The anomaly detection service is used to extract heat source features and perform heat anomaly screening processing based on the infrared alignment data, and screen out the heat anomaly areas from the candidate area set. The data analysis service extracts three-dimensional spatial features and motion state features from the lidar alignment data, and performs anomaly recognition on the heat anomaly areas based on the three-dimensional spatial features and motion state features. The early warning analysis service is used to generate corresponding early warning response instructions according to the anomaly types corresponding to the anomaly recognition results. The data access layer supports data writing, deletion, update, and query, and realizes the dynamic management of the inspection data. The data is written into the reservoir information database, inspection record database, early warning information database, equipment status database, and monitoring database in the persistence layer through the access interface, supporting the traceability analysis of historical inspection results and early warning responses. The security management module provides a unified authentication, unified authorization, unified audit, and unified permission management mechanism to ensure the data security and operation compliance during the system operation. This system architecture completely covers the entire process from perception data acquisition, multi-modal fusion analysis to anomaly recognition and response. Through the collaborative work of different functional layers, it realizes high-precision, low-latency, and multi-angle intelligent security inspection in complex water area environments.
[0114] Figure 5 Schematically shows a schematic diagram of the monitoring interface of the water area security inspection system based on multi-modal feature analysis according to some embodiments of the present disclosure. It is designed for the real-time inspection and anomaly early warning scenarios in the reservoir area. This interface mainly includes four functional modules: aerial intelligent inspection, image recognition analysis, reservoir real-time monitoring, and early warning information release. Among them: The aerial intelligent inspection module is used to display the real-time progress and data processing status of the drone inspection task. Among them, the data shows that the current coverage rate of the drone for the target reservoir area is 25%, the number last month was 63, and the number this month is 114, reflecting the current inspection frequency and time distribution; the early warning information release data shows that the current number of pushed early warning information is 170, and it was 141 last month, reflecting the anomaly discovery ability during the inspection and recognition process; the image recognition analysis index shows that the system has processed 886 images this month, and the recognition efficiency has increased by 422% compared with last month. The early warning information release module provides the real-time early warning status of the current reservoir area. The figure shows that the current early warning level is a first-level early warning, and the completion degree is shown in the form of a circular diagram as 60%. This result is the response level generated by the system based on the three-dimensional spatial features and motion state features of the targets in the anomaly area for anomaly recognition and combined with the preset conditions.
[0115] The image recognition and analysis module further classifies warning events in tabular form, including three columns: warning level, triggering conditions, and related matters. The warning levels are divided into level 1, level 2, and level 3. The corresponding triggering conditions include personnel falling into the water, abnormal vessels, and abnormal floating objects respectively. Each type of abnormal situation corresponds to specific related matters. For example, personnel falling into the water is related to the location of the personnel, abnormal vessels are related to the location of the vessels, and abnormal floating objects are related to the category of the floating objects. These can be automatically generated by the image segmentation and feature recognition module and presented in a structured manner on the interface. The reservoir real-time monitoring module is a time series graph, which is used to show the change trend of warning events at different levels in the past year. In the graph, the occurrence numbers of level 1 warning, level 2 warning, and level 3 warning are respectively plotted in the form of a line graph for each month, reflecting the system's ability to continuously perceive and respond to the safety situation of the target water area. The data of this graph can be sourced from the system's extraction of heat source features and screening of thermal anomalies from the candidate area set, and combined with the target contour movement trajectory extracted from the lidar alignment data to complete anomaly recognition and level classification.
[0116] It should be noted that although the steps of the method in the present disclosure are described in a specific order in the accompanying drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.
[0117] In addition, in this exemplary embodiment, a water area safety inspection system based on multi-modal feature analysis is also provided. Referring to Figure 6 As shown, the water area safety inspection system 600 based on multi-modal feature analysis includes: a data acquisition module 610, a data alignment module 620, a contour extraction module 630, a thermal anomaly screening module 640, and an anomaly recognition module 650. Among them: The data acquisition module 610 can be used to control the drone to collect visible light images, infrared images, and lidar data corresponding to the target water area based on a preset data acquisition strategy; The data alignment module 620 can be used to perform time synchronization and spatial alignment on the visible light image, infrared image, and lidar data to obtain visible light alignment data, infrared alignment data, and lidar alignment data; The contour extraction module 630 can be used to perform edge detection and region segmentation processing on the visible light alignment data in response to the clarity of the visible light image being higher than the quality reference value, and generate a set of candidate regions containing the target contour; The thermal anomaly screening module 640 can be used to perform heat source feature extraction and thermal anomaly screening processing on the infrared alignment data corresponding to the set of candidate regions, and screen out thermal anomaly regions from the set of candidate regions; The anomaly recognition module 650 can be used to extract three-dimensional spatial features and motion state features from the lidar alignment data corresponding to the thermal anomaly region, and perform anomaly recognition on the thermal anomaly region based on the three-dimensional spatial features and motion state features.
[0118] The specific details of each module of the above water area safety inspection system based on multi-modal feature analysis have been described in detail in the corresponding water area safety inspection method based on multi-modal feature analysis, so they will not be elaborated here.
[0119] It should be noted that although several modules or units of the water area safety inspection system based on multi-modal feature analysis are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0120] In addition, in the exemplary embodiments of the present disclosure, an electronic device capable of implementing the above water area safety inspection method based on multi-modal feature analysis is also provided.
[0121] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, method, or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0122] Next, refer to Figure 7 to describe the electronic device 700 according to the embodiments of the present disclosure. Figure 7 The illustrated electronic device 700 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0123] As Figure 7 shown, the electronic device 700 is presented in the form of a general-purpose computing device. The components of the electronic device 700 may include, but are not limited to: at least one of the above processing units 710, at least one of the above storage units 720, a bus 730 connecting different system components (including the storage unit 720 and the processing unit 710), and a display unit 740.
[0124] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 710, so that the processing unit 710 executes the steps according to various exemplary embodiments of the present disclosure described in the above "exemplary method" section of this specification.
[0125] The storage unit 720 may include a readable medium in the form of a volatile storage unit, such as a random access memory (RAM) 721 and / or a cache storage unit 722, and may further include a read-only memory (ROM) 723.
[0126] The storage unit 720 may also include a program / utilities 724 having a set (at least one) of program modules 725. Such program modules 725 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.
[0127] The bus 730 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.
[0128] The electronic device 700 may also communicate with one or more external devices 770 (such as a keyboard, a pointing device, a Bluetooth device, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 700, and / or may communicate with any device that enables the electronic device 700 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be through an input / output (I / O) interface 750. Moreover, the electronic device 700 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 760. As shown in the figure, the network adapter 760 communicates with other modules of the electronic device 700 through the bus 730. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0129] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0130] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above methods in this specification is stored. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0131] Reference Figure 8 As shown, a program product 800 for implementing the above water area safety inspection method based on multi-modal feature analysis according to an embodiment of the present disclosure is described. It can be a portable compact disc read-only memory and includes program code, and can run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited to this. In this document, a readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0132] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disc read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0133] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0134] The program code contained on the readable medium can be transmitted by any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0135] Program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, can be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0136] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.
[0137] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, server, touch terminal, or network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0138] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0139] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A water area safety inspection method based on multimodal feature analysis, characterized in that: include: Control the UAV to collect visible light images, infrared images and lidar data corresponding to the target water area based on the preset data collection strategy; Performing time synchronization and spatial alignment on the visible light image, infrared image and lidar data to obtain visible light alignment data, infrared alignment data and lidar alignment data; In response to the clarity of the visible light image being higher than a quality reference value, performing edge detection and region segmentation processing on the visible light alignment data to generate a set of candidate regions containing a target contour; Performing heat source feature extraction and thermal anomaly screening processing on the infrared alignment data corresponding to the candidate area set, and screening out thermal anomaly areas from the candidate area set; Three-dimensional spatial feature extraction and motion state feature extraction are performed on the laser radar alignment data corresponding to the thermal anomaly area, and anomaly identification is performed on the thermal anomaly area based on the three-dimensional spatial feature and the motion state feature.
2. The water area safety inspection method based on multimodal feature analysis according to claim 1 is characterized in that: Also includes: Screening out a thermally normal region from the candidate region set, and matching a target contour corresponding to the thermally normal region with a reference contour; In response to the matching result being a human body, determining a three-dimensional spatial feature corresponding to the target contour according to the laser radar alignment data corresponding to the target contour; The abnormality of the normal thermal area is identified based on the three-dimensional spatial features.
3. The water area safety inspection method based on multimodal feature analysis according to claim 1 is characterized in that: Also includes: Determining the clarity corresponding to the visible light image; In response to the clarity being lower than the quality reference value, feature extraction is performed on the visible light alignment data, the infrared alignment data, and the lidar alignment data to obtain a first image feature, a second image feature, and a third image feature; Performing feature fusion on the first image feature, the second image feature, and the third image feature to obtain a multimodal fusion feature; Anomaly identification is performed based on the multimodal fusion features.
4. The water area safety inspection method based on multimodal feature analysis according to claim 3 is characterized in that: The determining the clarity corresponding to the visible light image includes: Performing texture analysis on the visible light image to obtain texture information of the visible light image; Performing grayscale histogram statistical processing on the visible light image to obtain contrast information of the visible light image; Based on the texture information and the contrast information, the clarity corresponding to the visible light image is determined.
5. The water area safety inspection method based on multimodal feature analysis according to claim 3 is characterized in that: The determining the clarity corresponding to the visible light image includes: Dividing the visible light image into regions to obtain a plurality of local image regions; A local gradient mean corresponding to each of the local image regions is determined, the local gradient mean values are weightedly fused, and the clarity corresponding to the visible light image is determined based on the fusion result.
6. The water area safety inspection method based on multimodal feature analysis according to claim 3 is characterized in that: The step of fusing the first image feature, the second image feature, and the third image feature to obtain a multimodal fusion feature includes: Determining, according to the environmental information corresponding to the visible light image, fusion weights corresponding to the first image feature, the second image feature, and the third image feature; The first image feature, the second image feature and the third image feature are weightedly fused based on the fusion weight to obtain the multimodal fusion feature.
7. The water area safety inspection method based on multimodal feature analysis according to claim 1 is characterized in that: Before controlling the drone based on a preset data collection strategy, the method further includes: According to the current environmental information of the target water area, a data collection strategy corresponding to the UAV is determined.
8. The water area safety inspection method based on multimodal feature analysis according to claim 7 is characterized in that: The step of determining the data collection strategy corresponding to the UAV according to the current environmental information of the target water area includes: Obtaining current light intensity information, visibility information and wind speed information of the target water area; Determining the data collection strategy corresponding to the drone based on the light intensity information, visibility information and wind speed information; The data collection strategy includes the collection frequency and collection range corresponding to each type of data.
9. The water area safety inspection method based on multimodal feature analysis according to claim 1 is characterized in that: The step of performing heat source feature extraction and thermal anomaly screening processing on the infrared alignment data corresponding to the candidate area set, and screening out thermal anomaly areas from the candidate area set, includes: Extracting multi-scale heat source features of the infrared alignment data corresponding to the candidate region set at multiple scales; Based on the multi-scale heat source characteristics, determining the temperature distribution characteristics corresponding to each candidate area; Based on the temperature distribution characteristics, extracting the heat source morphological characteristics corresponding to each of the candidate areas; Based on the heat source morphological characteristics, the thermal anomaly area having irregular edge characteristics or concentrated heat source characteristics is screened out from the candidate area set.
10. The water area safety inspection method based on multimodal feature analysis according to claim 1 is characterized in that: The extracting three-dimensional spatial features and motion state features of the laser radar alignment data corresponding to the thermal anomaly area includes: Based on the laser radar alignment data, extracting three-dimensional structural features and surface curvature features corresponding to the thermal anomaly area, wherein the three-dimensional structural features include one or more of size information, volume information, bounding box compactness, and spatial information; Based on the laser radar alignment data collected at different times, extracting a set of multi-time trajectory points corresponding to the target contour in the thermal anomaly area; A motion trajectory curve is fitted based on a set of multi-time trajectory points corresponding to the target contour, and the motion state feature is extracted based on the motion trajectory curve.
11. The water area safety inspection method based on multimodal feature analysis according to claim 10 is characterized in that: The performing abnormality identification on the thermal abnormality area based on the three-dimensional space feature and the motion state feature includes: Inputting the three-dimensional structural features, the surface curvature features, and the motion state features into an abnormality recognition model to obtain an abnormality recognition result; In response to the abnormal recognition result being a preset target category, a motion trajectory of the target contour is predicted according to the multi-time trajectory point set.
12. The water area safety inspection method based on multimodal feature analysis according to claim 1 is characterized in that: Also includes: Extracting a visible light image corresponding to the abnormal area according to the abnormality recognition result; Visually displaying the visible light image, the abnormality recognition result, and the location information of the abnormal area; According to the abnormality type corresponding to the abnormality identification result, a corresponding early warning response instruction is generated.
13. The water area safety inspection method based on multimodal feature analysis according to claim 1 is characterized in that: Also includes: Based on the abnormality identification result, determining the spatial position information corresponding to the abnormal target; Generate a track adjustment instruction based on the spatial position information and the current position of the drone; Based on the track adjustment instruction, adjust the flight path of the UAV so that the UAV approaches the abnormal target; Based on the real-time spatial relationship between the abnormal target and the UAV, the track adjustment instruction is updated to optimize the observation angle and cruising radius of the UAV.
14. A water area safety inspection system based on multimodal feature analysis, characterized in that: include: The data acquisition module is used to control the UAV to collect visible light images, infrared images and lidar data corresponding to the target water area based on a preset data acquisition strategy; A data alignment module is used to perform time synchronization and spatial alignment on the visible light image, infrared image and lidar data to obtain visible light alignment data, infrared alignment data and lidar alignment data; A contour extraction module, configured to perform edge detection and region segmentation processing on the visible light alignment data in response to the clarity of the visible light image being higher than a quality reference value, and generate a candidate region set containing a target contour; A thermal anomaly screening module, used for performing heat source feature extraction and thermal anomaly screening processing on the infrared alignment data corresponding to the candidate area set, and screening out thermal anomaly areas from the candidate area set; The anomaly recognition module is used to extract three-dimensional spatial features and motion state features from the laser radar alignment data corresponding to the thermal anomaly area, and to recognize anomalies in the thermal anomaly area based on the three-dimensional spatial features and the motion state features.
Citation Information
Patent Citations
Multi-dimensional decision fusion unmanned aerial vehicle cluster collaborative decision-making method
CN118707968A
Intelligent security and protection monitoring method and system based on image recognition
CN119495054A
Method and device for identifying unmanned aerial vehicle based on motion features of high-point camera holder
CN119625633A
System and Method for Extremely Efficient Image and Pattern Recognition and Artificial Intelligence Platform
US20200184278A1
Multi-sensor data fusion method and device
WO2020135810A1
Cited By
Parcel boundary identification method based on fusion of laser point cloud and visible light
CN120451799A
Land boundary recognition method based on fusion of laser point cloud and visible light
CN120451799B
Rainfall recharge control method and device
CN120495996A
Unmanned aerial vehicle detection method and system based on infrared image analysis
CN120766171A
Property automation patrol method, device and equipment and storage medium
CN120786031A