Water area safety inspection method and system based on multi-modal feature analysis
Through the UAV collecting multimodal data and performing feature analysis, the problems of monitoring blind spots and target identification lag in reservoir water safety patrols are solved, and efficient and accurate identification and monitoring of abnormal targets are achieved.
Patent Information
- Application Number
- CN202510653383.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-05-21
AI Technical Summary
In the existing reservoir water safety patrol technology, fixed monitoring equipment has monitoring blind spots and target recognition lag, and manual patrols are difficult to achieve large-scale continuous dynamic monitoring, resulting in delays in abnormal discovery and missed safety hazards.
Using a multimodal feature analysis method, the UAV collects visible light images, infrared images and lidar data, performs time synchronization and spatial alignment, and combines target profile extraction, thermal anomaly screening and three-dimensional dynamic feature analysis to improve patrol coverage and abnormal discovery accuracy.
It realizes efficient and accurate abnormal target recognition of complex water environments, reduces the risk of missed detection caused by environmental light changes, and improves the ability to identify abnormal categories such as floating objects and people falling into the water.
Smart Images

Figure CN120198808B_ABST
Abstract
Description
Background Art
[0002] At present, for the safety inspection of reservoir waters, the relevant technologies mainly rely on deploying fixed monitoring devices around the reservoir and combining them with manual inspections for on-site inspections. In practical applications, usually, the background monitoring personnel observe the water conditions by remotely viewing the monitoring images, and in key areas such as around the flood discharge channels, manual inspectors are arranged to conduct on-site inspections and abnormal situation investigations.
[0003] However, due to the vast reservoir environment and the frequent fluctuations of the water surface conditions with the environment, the fixed monitoring devices are limited by the installation location, monitoring angle, and coverage area, resulting in monitoring blind spots and it is difficult to cover all key areas in a timely manner. At the same time, the background personnel's method of monitoring through images is limited by factors such as image clarity and environmental visibility changes, and it is prone to situations of lagging target recognition or misjudgment. Although manual inspections can, to a certain extent, supplement the deficiencies of fixed monitoring, due to human resources, inspection frequency, and emergency response speed limitations, it is difficult to achieve continuous dynamic monitoring of a large area of waters.
[0004] Especially during flood discharge periods or under complex weather conditions, the traditional method relying on fixed monitoring and manual inspections has the risks of delayed discovery of abnormal targets, untimely response, and missed detection of potential water safety hazards. Therefore, how to improve the inspection coverage rate and the accuracy of abnormal situation discovery in complex water environments has become an urgent problem to be solved in the current water area safety inspection technology. Summary of the Invention
[0005] The purpose of the embodiments of the present disclosure is to provide a water area safety inspection method based on multi-modal feature analysis, a water area safety inspection system based on multi-modal feature analysis, an electronic device, and a computer-readable storage medium, which can, based on multi-source data collection and synchronous processing, combined with target contour extraction, thermal anomaly screening, and three-dimensional dynamic feature analysis, improve the inspection coverage rate and the accuracy of abnormal situation discovery in complex water environments.
[0006] Other features and advantages of the present disclosure will become apparent through the following detailed description, or be learned in part through the practice of the present disclosure.
[0007] According to the first aspect of the embodiments of the present disclosure, a water area safety inspection method based on multi-modal feature analysis is provided, including: controlling a drone to collect visible light images, infrared images, and lidar data corresponding to a target water area based on a preset data collection strategy; performing time synchronization and spatial alignment on the visible light images, infrared images, and lidar data to obtain visible light aligned data, infrared aligned data, and lidar aligned data; in response to the clarity of the visible light image being higher than a quality reference value, performing edge detection and region segmentation processing on the visible light aligned data to generate a set of candidate regions including target contours; performing heat source feature extraction and heat anomaly screening processing on the infrared aligned data corresponding to the set of candidate regions, and screening out heat anomaly regions from the set of candidate regions; performing three-dimensional spatial feature extraction and motion state feature extraction on the lidar aligned data corresponding to the heat anomaly regions, and performing anomaly recognition on the heat anomaly regions based on the three-dimensional spatial features and the motion state features.
[0008] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the water area safety inspection method based on multi-modal feature analysis further includes: screening out heat-normal regions from the set of candidate regions, and matching the target contours corresponding to the heat-normal regions with a reference contour; in response to the matching result being a human body, determining the three-dimensional spatial features corresponding to the target contour according to the lidar aligned data corresponding to the target contour; performing anomaly recognition on the heat-normal regions based on the three-dimensional spatial features.
[0009] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the water area safety inspection method based on multi-modal feature analysis further includes: determining the clarity corresponding to the visible light image; in response to the clarity being lower than the quality reference value, respectively performing feature extraction on the visible light aligned data, infrared aligned data, and lidar aligned data to obtain a first image feature, a second image feature, and a third image feature; performing feature fusion on the first image feature, the second image feature, and the third image feature to obtain a multi-modal fusion feature; performing anomaly recognition based on the multi-modal fusion feature.
[0010] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the determining the clarity corresponding to the visible light image includes: performing texture analysis on the visible light image to obtain texture information of the visible light image; performing grayscale histogram statistical processing on the visible light image to obtain contrast information of the visible light image; determining the clarity corresponding to the visible light image based on the texture information and the contrast information.
[0011] In some exemplary embodiments of the present disclosure, based on the foregoing solution, determining the sharpness corresponding to the visible light image includes: dividing the visible light image into regions to obtain a plurality of local image regions; determining the local gradient mean corresponding to each local image region, performing weighted fusion on the local gradient means, and determining the sharpness corresponding to the visible light image based on the fusion result.
[0012] In some exemplary embodiments of the present disclosure, based on the foregoing solution, performing feature fusion on the first image feature, the second image feature, and the third image feature to obtain a multi-modal fusion feature includes: determining the fusion weights corresponding to the first image feature, the second image feature, and the third image feature according to the environmental information corresponding to the visible light image; performing weighted fusion processing on the first image feature, the second image feature, and the third image feature based on the fusion weights to obtain the multi-modal fusion feature.
[0013] In some exemplary embodiments of the present disclosure, based on the foregoing solution, before controlling the drone based on a preset data acquisition strategy, it further includes: determining the data acquisition strategy corresponding to the drone according to the current environmental information of the target water area.
[0014] In some exemplary embodiments of the present disclosure, based on the foregoing solution, determining the data acquisition strategy corresponding to the drone according to the current environmental information of the target water area includes: obtaining the current light intensity information, visibility information, and wind speed information of the target water area; determining the data acquisition strategy corresponding to the drone based on the light intensity information, visibility information, and wind speed information; wherein the data acquisition strategy includes the acquisition frequency and acquisition range corresponding to each type of data.
[0015] In some exemplary embodiments of the present disclosure, based on the foregoing solution, performing heat source feature extraction and thermal anomaly screening processing on the infrared alignment data corresponding to the candidate region set, and screening out thermal anomaly regions from the candidate region set includes: extracting multi-scale heat source features of the infrared alignment data corresponding to the candidate region set at multiple scales; determining the temperature distribution features corresponding to each candidate region based on the multi-scale heat source features; extracting the heat source morphology features corresponding to each candidate region based on the temperature distribution features; screening out the thermal anomaly regions having irregular edge characteristics or concentrated heat source characteristics from the candidate region set based on the heat source morphology features.
[0016] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the extraction of three-dimensional spatial features and motion state features from the lidar alignment data corresponding to the thermal anomaly region includes: based on the lidar alignment data, extracting three-dimensional structural features and surface curvature features corresponding to the thermal anomaly region, wherein the three-dimensional structural features include one or more of size information, volume information, bounding box compactness, and spatial information; based on the lidar alignment data collected at different times, extracting a set of multi-time trajectory points corresponding to the target contour in the thermal anomaly region; fitting a motion trajectory curve based on the set of multi-time trajectory points corresponding to the target contour, and extracting the motion state features based on the motion trajectory curve.
[0017] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the abnormal recognition of the thermal anomaly region based on the three-dimensional spatial features and the motion state features includes: inputting the three-dimensional structural features, the surface curvature features, and the motion state features into an abnormal recognition model to obtain an abnormal recognition result; in response to the abnormal recognition result being a preset target category, predicting the motion trajectory of the target contour according to the set of multi-time trajectory points.
[0018] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the above-mentioned water area safety inspection method based on multi-modal feature analysis further includes: extracting a visible light image corresponding to the abnormal region according to the abnormal recognition result; visually displaying the visible light image, the abnormal recognition result, and the position information of the abnormal region; generating a corresponding early warning response instruction according to the abnormal type corresponding to the abnormal recognition result.
[0019] In some exemplary embodiments of the present disclosure, based on the foregoing solution, the above-mentioned water area safety inspection method based on multi-modal feature analysis further includes: determining the spatial position information corresponding to the abnormal target based on the abnormal recognition result; generating a flight path adjustment instruction based on the spatial position information and the current position of the unmanned aerial vehicle; adjusting the flight path of the unmanned aerial vehicle based on the flight path adjustment instruction to make the unmanned aerial vehicle approach the abnormal target; updating the flight path adjustment instruction based on the real-time spatial relationship between the abnormal target and the unmanned aerial vehicle to optimize the observation angle and cruising radius of the unmanned aerial vehicle.
[0020] According to a second aspect of the embodiments of the present disclosure, there is provided a water area safety inspection system based on multimodal feature analysis, including: a data acquisition module, configured to control a drone to acquire visible light images, infrared images, and lidar data corresponding to a target water area based on a preset data acquisition strategy; a data alignment module, configured to perform time synchronization and spatial alignment on the visible light images, infrared images, and lidar data to obtain visible light aligned data, infrared aligned data, and lidar aligned data; a contour extraction module, configured to, in response to the clarity of the visible light image being higher than a quality reference value, perform edge detection and region segmentation processing on the visible light aligned data to generate a set of candidate regions including target contours; a thermal anomaly screening module, configured to perform heat source feature extraction and thermal anomaly screening processing on the infrared aligned data corresponding to the set of candidate regions, and screen out thermal anomaly regions from the set of candidate regions; an anomaly recognition module, configured to perform three-dimensional spatial feature extraction and motion state feature extraction on the lidar aligned data corresponding to the thermal anomaly regions, and perform anomaly recognition on the thermal anomaly regions based on the three-dimensional spatial features and the motion state features.
[0021] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a processor; and a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method for water area safety inspection based on multimodal feature analysis as in the first aspect is implemented.
[0022] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for water area safety inspection based on multimodal feature analysis as in the first aspect is implemented.
[0023] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects:
[0024] In the method for water area safety inspection based on multimodal feature analysis in the exemplary embodiments of the present disclosure, first, visible light images, infrared images, and lidar data corresponding to a target water area are acquired by a drone, improving the inspection coverage rate, and through time synchronization and spatial alignment processing of data from different sources, on the basis of ensuring that various types of perception data have a unified spatio-temporal reference system, the consistency and accuracy of subsequent feature analysis are improved.
[0025] On the one hand, by judging the clarity of visible light images, when the clarity is higher than the quality reference value, edge detection and region segmentation processing are performed on the visible light alignment data, which can extract clear target contours when the image quality is good, reduce the risk of false detection and missed detection caused by environmental light changes, and enhance the reliability of the candidate region set. On the other hand, heat source feature extraction and thermal anomaly screening processing are performed based on the infrared alignment data corresponding to the candidate region set, which can effectively utilize the prominent response characteristics of infrared sensing to heat source anomalies under low light or complex meteorological conditions, thereby compensating for the recognition limitations of a single visible light image when visibility decreases, and further ensuring the accurate screening of thermal anomaly regions under different environmental conditions. On the further hand, three-dimensional spatial feature extraction and motion state feature extraction are performed on the lidar alignment data corresponding to the thermal anomaly region, which can further determine the dynamic behavior of the abnormal target through three-dimensional structure features and trajectory change features, effectively identify different abnormal categories such as floating objects and drowning persons, and improve the accuracy of abnormal target recognition in complex water environments.
[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0028] Figure 1 Schematically shows a flowchart of a water area safety inspection method based on multi-modal feature analysis according to some embodiments of the present disclosure.
[0029] Figure 2 Schematically shows a flowchart of abnormal recognition when the clarity of a visible light image is insufficient according to some embodiments of the present disclosure.
[0030] Figure 3 Schematically shows a flowchart of screening thermal anomaly regions according to some embodiments of the present disclosure.
[0031] Figure 4 Schematically shows a system architecture diagram of a water area safety inspection system based on multi-modal feature analysis according to some embodiments of the present disclosure.
[0032] Figure 5 Schematically shows a monitoring interface diagram of a water area safety inspection system based on multi-modal feature analysis according to some embodiments of the present disclosure.
[0033] Figure 6 A block diagram schematically shows a water area safety inspection system based on multimodal feature analysis according to some embodiments of the present disclosure.
[0034] Figure 7 A schematic diagram shows the structure of a computer system of an electronic device according to some embodiments of the present disclosure.
[0035] Figure 8 A schematic diagram shows a computer-readable storage medium according to some embodiments of the present disclosure.
[0036] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. Detailed implementation manners
[0037] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this specification. On the contrary, they are merely examples of devices and methods consistent with some aspects of this specification as detailed in the appended claims.
[0038] The terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit this specification. The singular forms "a", "the", and "said" used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and includes any or all possible combinations of one or more of the associated listed items. Now, the exemplary embodiments will be described more fully with reference to the drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; on the contrary, these embodiments are provided so that this disclosure will be more complete and comprehensive, and the concept of the exemplary embodiments will be fully conveyed to those skilled in the art.
[0039] In addition, the drawings are only schematic diagrams and are not necessarily drawn to scale. The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0040] In the present exemplary embodiment, first, a water area safety inspection method based on multimodal feature analysis is provided. Figure 1A flowchart schematically shows a water area safety inspection method based on multimodal feature analysis according to some embodiments of the present disclosure. Refer to Figure 1 As shown, the water area safety inspection method based on multimodal feature analysis may include the following steps:
[0041] Step S110, controlling a drone to collect visible light images, infrared images, and lidar data corresponding to a target water area based on a preset data collection strategy;
[0042] Step S120, performing time synchronization and spatial alignment on the visible light images, infrared images, and lidar data to obtain visible light aligned data, infrared aligned data, and lidar aligned data;
[0043] Step S130, in response to the clarity of the visible light image being higher than a quality reference value, performing edge detection and region segmentation processing on the visible light aligned data to generate a set of candidate regions containing target contours;
[0044] Step S140, performing heat source feature extraction and thermal anomaly screening processing on the infrared aligned data corresponding to the set of candidate regions, and screening out thermal anomaly regions from the set of candidate regions;
[0045] Step S150, performing three-dimensional spatial feature extraction and motion state feature extraction on the lidar aligned data corresponding to the thermal anomaly regions, and performing anomaly recognition on the thermal anomaly regions based on the three-dimensional spatial features and motion state features.
[0046] According to the water area safety inspection method based on multimodal feature analysis in this exemplary embodiment, by using a drone to collect visible light images, infrared images, and lidar data corresponding to a target water area, the inspection coverage rate is improved. On the one hand, by judging the clarity of the visible light image, when the clarity is higher than the quality reference value, edge detection and region segmentation processing are performed on the visible light aligned data, which can extract clear target contours when the image quality is good, reduce the risk of false detection and missed detection caused by environmental light changes, and enhance the reliability of the set of candidate regions. On the other hand, performing heat source feature extraction and thermal anomaly screening processing based on the infrared aligned data corresponding to the set of candidate regions can effectively utilize the prominent response characteristics of infrared sensing to heat source anomalies under low light or complex meteorological conditions, thereby compensating for the recognition limitations of a single visible light image when visibility decreases, and further ensuring the accurate screening of thermal anomaly regions under different environmental conditions. On the other hand, performing three-dimensional spatial feature extraction and motion state feature extraction on the lidar aligned data corresponding to the thermal anomaly regions can further distinguish the dynamic behaviors of abnormal targets through three-dimensional structural features and trajectory change features, effectively identify different abnormal categories such as floating objects and drowning persons, and improve the accuracy of identifying abnormal targets in complex water area environments.
[0047] Next, the water area safety inspection method based on multimodal feature analysis in this exemplary embodiment will be further described.
[0048] Step S110, control the drone to collect visible light images, infrared images, and lidar data corresponding to the target water area based on a preset data collection strategy.
[0049] Among them, the data collection strategy can represent a data collection plan preset for different environmental conditions, which can include startup methods of collection devices, working frequencies, collection time intervals, collection modes, and other appropriate device configuration methods such as collection ranges. The target water area can represent the spatial area where the inspection task is to be performed, which can include water body areas such as reservoirs, flood discharge areas, lakes, and rivers with specific geographical boundaries and safety monitoring requirements. Visible light images can represent two-dimensional image data in the visible light band obtained by an imaging device mounted on the drone, used to reflect the visual feature information of the surface of the target water area under natural lighting conditions. Infrared images can represent thermal imaging image data formed based on the intensity of thermal radiation obtained by an infrared imaging device, usually reflecting the temperature distribution characteristics of different objects or areas within the target area, used to identify heat source targets. Lidar data can represent three-dimensional point cloud data of the target water area collected by a lidar sensor, and the spatial position information of the target surface is obtained based on the time difference between the emitted laser pulse and the received echo, used to reconstruct the spatial structure characteristics of the target area.
[0050] By using a drone for water area safety inspection, high-efficiency monitoring of a large range of water area regions can be achieved. Compared with fixed monitoring devices, the drone has the ability to dynamically adjust the flight path and can obtain data from multiple angles and viewpoints within the task range, thus significantly improving the coverage ability of water area inspection. Further, by carrying multiple types of sensors to simultaneously obtain visible light images, infrared images, and lidar data, multi-source perception information with complementary features can be obtained under complex environmental conditions, providing better data support for subsequent data fusion, target recognition, and anomaly detection.
[0051] Step S120, perform time synchronization and spatial alignment on the visible light images, infrared images, and lidar data to obtain visible light aligned data, infrared aligned data, and lidar aligned data.
[0052] Among them, time synchronization can represent the process of uniformly adjusting the acquisition timings of visible light images, infrared images, and lidar data based on timestamp information of various types of data, which is used to ensure the corresponding relationship of various types of data in the time dimension. Spatial alignment can represent the process of converting and registering the spatial coordinates of data collected by multiple types of sensors, so as to achieve the position consistency of multi-source data in the spatial dimension. Visible light aligned data can represent the image data obtained by extracting from visible light images under a unified spatio-temporal reference after time synchronization and spatial alignment processing. Infrared aligned data can represent the infrared image data obtained by time synchronization and spatial alignment based on infrared images, which retains the temperature distribution information on the premise of sharing the same spatio-temporal coordinate system with other data types. Lidar aligned data can represent the point cloud data generated after time synchronization and spatial coordinate conversion of lidar data, which has the same timing and spatial position reference as the image data.
[0053] Due to the differences in the acquisition frequencies, viewing angle parameters, and imaging mechanisms of different modality sensors, if joint processing is directly performed, it is easy to cause inconsistent target feature expressions due to time offset or spatial misalignment. Through time synchronization processing, it is possible to ensure that various types of data are correspondingly matched under the same time reference, avoiding the problem of target misalignment caused by time offset. Through spatial alignment processing, it is possible to accurately map the data of each modality to a unified spatial framework, ensuring the correspondence and comparability of the features of each modality in terms of position during subsequent analysis.
[0054] Step S130, in response to the clarity of the visible light image being higher than the quality reference value, perform edge detection and region segmentation processing on the visible light aligned data to generate a set of candidate regions containing the target contour.
[0055] Among them, clarity can be an index used to characterize the ability of the local structural details distribution of an image, and can be used to reflect the image quality. The quality reference value can be a clarity threshold used to determine whether an image meets the requirements of subsequent processing, and is used to judge whether the current image has the conditions for edge detection and region segmentation processing. Edge detection can be an image processing process for identifying and extracting positions with significant gray or color changes in an image, and is used to highlight the boundary information between objects and the background in the image. Region segmentation processing can be a processing method for dividing image regions in an image based on pixel feature similarity, edge information or semantic information, and is used to divide the image into multiple sub-regions with structural consistency to extract the spatial range of potential targets. The target contour can be a closed or approximately closed boundary structure extracted through edge detection and region segmentation in the image, and is used to describe the geometric shape characteristics of possible abnormal targets on the image plane. The target contour can include human contours, ship contours, floating object contours, obstacle contours, and other suitable monitoring target contours in the water area. The candidate region set can be a set composed of multiple image sub-regions containing target contours.
[0056] Among them, when the clarity of the visible light image is higher than the quality reference value, the subsequent edge detection and region segmentation processing procedures are carried out, avoiding the risk of unstable extraction results in the case of blurred or severely interfered images. By performing edge detection and region segmentation processing on the visible light alignment data whose clarity meets the requirements, image regions with structural continuity and prominent boundaries can be effectively extracted, and then a candidate region set containing various target contours such as human contours, ship contours, floating object contours, and obstacle contours can be generated.
[0057] Step S140: Perform heat source feature extraction and thermal anomaly screening processing on the infrared alignment data corresponding to the candidate region set, and screen out thermal anomaly regions from the candidate region set.
[0058] Among them, heat source feature extraction can be a process of extracting image features used to characterize the temperature characteristics of the target region based on the temperature distribution information of each candidate region in the infrared alignment data. Thermal anomaly screening can be a processing operation for identifying and screening regions in the candidate region set whose temperature characteristics significantly deviate from the background distribution or exceed the preset temperature threshold based on the heat source feature extraction results. The thermal anomaly region can be a target region in the candidate region set determined to have significant thermal radiation characteristics through the thermal anomaly screening process.
[0059] Performing heat source feature extraction and thermal anomaly screening on the infrared alignment data corresponding to the candidate region set can utilize the characteristic that infrared images are sensitive to temperature distribution changes to identify regions with significant thermal radiation characteristics without relying on the clarity of visible light. Since drowning victims are usually accompanied by obvious heat source signals, small boats and their power devices, some floating objects or obstacles may also show local temperature difference responses in thermal imaging. Therefore, by extracting the heat source characteristics of each candidate region in the infrared alignment data and combining the background temperature distribution for anomaly screening, regions with insufficient thermal responses or similar to the water body background can be effectively removed, and thermal anomaly regions can be further screened out. The above processing not only improves the positioning ability of targets such as drowning victims, but also provides a significant preliminary screening basis for subsequent multi-modal recognition based on spatial structure and motion characteristics.
[0060] Step S150: Extract three-dimensional spatial features and motion state features from the lidar alignment data corresponding to the thermal anomaly region, and perform anomaly recognition on the thermal anomaly region based on the three-dimensional spatial features and motion state features.
[0061] Among them, the three-dimensional spatial features can represent a set of spatial structure parameters extracted from the lidar alignment data, which can be used to describe the geometric shape and volume distribution of the target in a three-dimensional coordinate system. The motion state features can represent a set of target dynamic behavior parameters extracted from the lidar alignment data at multiple time points, which can include the spatial displacement change, velocity vector, acceleration trend, trajectory fitting curve, and trajectory stability of the target in consecutive time frames, etc., and are used to describe the motion characteristics of the target in the water area. Anomaly recognition can represent a comprehensive judgment process based on the three-dimensional spatial features and motion state features, which is used to identify whether the thermal anomaly region shows spatial form and dynamic behavior characteristics consistent with known target types (such as drowning victims, floating objects, boats, obstacles, etc.), and distinguish and mark target regions with abnormal contour features or abnormal motion patterns from the thermal anomaly region.
[0062] Performing three-dimensional spatial feature extraction and motion state feature extraction on the lidar alignment data corresponding to the thermal anomaly region can further utilize the advantages of lidar in structural modeling and dynamic tracking on the basis of initial heat source screening, and complement the target shape and behavior information. The three-dimensional spatial features can reflect the external dimensions, structural volume, and curvature complexity of the thermal anomaly region in space, and are used to distinguish regular-shaped targets (such as ships, obstacles) from irregular or partially submerged targets (such as drowning persons). The motion state features can extract the trajectory continuity, speed change, and stability trend of the target based on multiple frames of lidar data, and are used to identify motion behavior patterns such as floating objects drifting with the water, small boats advancing smoothly, and drowning persons struggling and jittering. Using the above two types of features for anomaly recognition processing can effectively distinguish different types of monitoring objects in the thermal anomaly region, strengthen the discrimination ability of drowning persons, and improve the recognition accuracy and response pertinence in complex dynamic water environments.
[0063] Next, the content in steps S110 to S150 will be described in detail.
[0064] In some embodiments, the time synchronization and spatial alignment of the visible light image, infrared image, and lidar data in step S120 can be achieved through the following technical process to obtain visible light alignment data, infrared alignment data, and lidar alignment data, specifically including: obtaining the timestamp information corresponding to the visible light image, infrared image, and lidar data to establish the acquisition timing sequence of various types of data. Based on the timestamp information, perform time synchronization processing on the visible light image, infrared image, and lidar data to generate intermediate synchronization data with a unified time reference. Obtain the calibration parameters for registration, and perform coordinate transformation processing on the intermediate synchronization data based on the calibration parameters. Based on the results of the coordinate transformation processing, generate visible light alignment data, infrared alignment data, and lidar alignment data corresponding to the unified spatial reference coordinate system respectively.
[0065] Among them, obtaining the calibration parameters for registration and performing coordinate transformation processing on the intermediate synchronization data based on the calibration parameters can include: obtaining the relative pose information of various sensors in the UAV installation structure as the basic calibration parameters for constructing the coordinate system relationship between sensors, where the calibration parameters include the rotation matrix and translation vector of each sensor relative to the unified spatial reference coordinate system. Substitute the original coordinate values corresponding to various types of perception data in the intermediate synchronization data into the coordinate transformation model composed of the rotation matrix and translation vector to perform three-dimensional coordinate transformation operations to obtain a transformation data set in the unified coordinate system.
[0066] Based on the results of coordinate transformation processing, visible light alignment data, infrared alignment data, and lidar alignment data corresponding to the unified spatial reference coordinate system can be generated respectively, which may include: classifying the transformation data set according to the sensor source, respectively extracting the transformed image pixel coordinates and point cloud coordinate values, and constructing visible light alignment data, infrared alignment data, and lidar alignment data consistent with the unified spatial reference coordinate system. Among them, the visible light alignment data and the infrared alignment data are respectively corrected for coordinate mapping by means of image reprojection, while the lidar alignment data retains the original point cloud spatial accuracy to support subsequent multi-modal spatial consistency analysis and feature fusion processing.
[0067] In some embodiments, the edge detection and region segmentation processing of the visible light alignment data in step S130 can be implemented through the following technical process to generate a set of candidate regions containing the target contour, specifically including: performing image preprocessing operations on the visible light alignment data to enhance the image edge contrast and suppress noise interference; based on the preprocessed visible light alignment data, performing edge detection processing and extracting the edge pixel distribution in the image to generate an edge image; performing region segmentation processing on the edge image to divide the edge image into multiple image sub-regions with boundary structures; screening out regions with closed boundary features and complete contours from the image sub-regions to construct a set of candidate regions containing the target contour.
[0068] In specific implementation, when performing image preprocessing operations on the visible light alignment data, the visible light alignment data can be subjected to image grayscale conversion, histogram equalization, and edge enhancement filtering processing to enhance the brightness contrast between the boundary region and the background region in the image, and reduce the noise interference introduced by factors such as environmental reflection and water surface disturbance through median filtering or Gaussian filtering to improve the response clarity of the edge region. When generating the edge image, Sobel operator, Canny operator, or other edge detection algorithms based on gradient change can be used to calculate the response to the pixel gradient change in the image, extract the pixel set at the strong edge position in the image, and convert the set into an edge image with a binary edge structure. Then, using the connected component analysis algorithm or other appropriate region segmentation algorithms, aggregate the adjacent edge pixels in the edge image, extract the image sub-regions composed of multiple continuous pixels, and record the boundary range and geometric shape information of each image sub-region. Finally, perform shape structure analysis on the image sub-regions, judge whether the region edge forms a closed curve structure, and combine region area, aspect ratio, and boundary coherence indicators to screen out regions with complete geometric contours, and construct such regions into a set of candidate regions.
[0069] In some embodiments, it can be achieved through Figure 2The steps S210 to S240 shown above implement the heat source feature extraction and thermal anomaly screening process on the infrared alignment data corresponding to the candidate region set in step S140 above, and screen out the thermal anomaly regions from the candidate region set, specifically including:
[0070] Step S210: Extract multi-scale heat source features of the infrared alignment data corresponding to the candidate region set at multiple scales.
[0071] Among them, the multi-scale heat source features can represent a set of temperature response feature sets obtained by performing feature extraction on the infrared alignment data corresponding to the candidate region set at different spatial scales. Based on the infrared alignment data corresponding to the candidate region set, a scale-variable window operator is used to perform heat source feature extraction operations on each candidate region respectively, and a multi-scale heat source feature vector is constructed:
[0072]
[0073] Among them, represents the multi-scale heat source feature vector corresponding to the candidate region , represents the heat intensity response calculated under the th scale window, represents the spatial standard deviation of the th scale window, represents the total number of scales.
[0074] Step S220: Determine the temperature distribution characteristics corresponding to each candidate region based on the multi-scale heat source features.
[0075] Among them, the temperature distribution characteristics can represent a set of characteristics used to describe the overall distribution structure of the temperature values of each pixel point in the candidate region in the infrared alignment data. Based on the multi-scale heat source feature vector, the temperature distribution characteristics at each scale within the candidate region are calculated respectively:
[0076]
[0077] Among them, represents the temperature distribution characteristics of the candidate region , represents the temperature value of the pixel point at position within the candidate region at the th scale, represents the temperature mean of the candidate region at the th scale, represents the number of pixel points included in the candidate region , represents the total number of scales.
[0078] Step S230: Based on the temperature distribution characteristics, extract the heat source morphological characteristics corresponding to each candidate region.
[0079] Among them, the heat source morphological characteristics can represent the set of image characteristics obtained by extraction in the infrared image and used to describe the heat source boundary structure and the shape of the heat energy distribution. Based on the temperature distribution characteristics, the heat source morphological characteristics are extracted for each candidate region respectively:
[0080]
[0081] Among them, represents the heat source morphological characteristics of the candidate region , represents the non - regularity parameter of the region edge, defined as the variance of the curvature change of the region edge, represents the heat source concentration parameter, defined as the ratio of the number of pixels above the temperature threshold in the region to the total number of pixels in the region, represents the heat source centroid offset parameter, defined as the distance from the heat source centroid position to the geometric center of the region, and the weight coefficient is the fusion weight parameter, satisfying .
[0082] Step S240: Based on the heat source morphological characteristics, screen out the thermal anomaly regions with irregular edge characteristics or concentrated heat source characteristics from the candidate region set.
[0083] Among them, the irregular edge characteristic can represent the manifestation form that the heat source region has an irregular boundary contour in the infrared image. This boundary can be a broken, serrated, unevenly diffused or closed or semi - closed structure without obvious geometric rules, and can be used to indicate the thermal contour characteristics of abnormal heat sources such as human bodies, floating objects and other non - structural targets. According to the heat source morphological characteristics , set the thermal anomaly screening threshold , and select the regions from the candidate region set that meet the following conditions:
[0084]
[0085] Mark the candidate regions that meet the above conditions as thermal anomaly regions, and output the thermal anomaly regions for subsequent abnormal target recognition.
[0086] In some embodiments, the above - mentioned extraction of three - dimensional space characteristics and motion state characteristics from the lidar alignment data corresponding to the thermal anomaly region in step S150 can be achieved through the following technical process, specifically including:
[0087] First, based on the lidar-aligned data, extract the three-dimensional structural features and surface curvature features corresponding to the thermal anomaly region. Among them, the three-dimensional structural features include one or more of size information, volume information, bounding box compactness, and spatial information.
[0088] Among them, the bounding box compactness can represent the spatial utilization efficiency of the outer bounding box of the three-dimensional point cloud corresponding to the thermal anomaly region in the lidar-aligned data, which is calculated by the ratio of the actual point cloud volume of the thermal anomaly region to the volume of its minimum circumscribed cube. The spatial information can represent the three-dimensional position, distribution range, and spatial density and other structural characteristics of the point cloud corresponding to the thermal anomaly region in the lidar-aligned data. The surface curvature feature can represent the curvature change of the fitting surface constructed based on the local point cloud in the lidar-aligned data at each point, which can be characterized by mean curvature, Gaussian curvature, or local curvature variance, etc.
[0089] Then, based on the lidar-aligned data collected at different times, extract the multi-time trajectory point set corresponding to the target contour in the thermal anomaly region.
[0090] Specifically, at multiple sampling times, respectively extract the centroid coordinate points of the target contour in the thermal anomaly region from the lidar-aligned data and construct a multi-time trajectory point set of the target contour in the time series , where represents the three-dimensional spatial centroid position of the target contour at time , , , respectively represent the coordinate components of the target in the X, Y, and Z axis directions at the corresponding times, represents the total number of times of trajectory sampling. Among them, the target contour can be a human contour, a ship contour, a floating object contour, etc., and the specific category can be adaptively set according to the application scenario.
[0091] Finally, based on the multi-time trajectory point set corresponding to the target contour, fit the motion trajectory curve and extract the motion state features based on the motion trajectory curve.
[0092] Specifically, perform time parameterization processing on the multi-time trajectory point set using the least squares fitting method to construct a three-dimensional smooth trajectory function to obtain the motion trajectory curve:
[0093]
[0094] where represents the three-dimensional position vector of the fitted target trajectory point at time , is the polynomial coefficient vector of the three-dimensional trajectory fitting, Indicates the order of the fitting curve, , , respectively can represent the position coordinate functions of the target in the X-axis, Y-axis, and Z-axis directions at time moment, Can represent the time variable used to describe the change of the target trajectory, Can represent the term index in polynomial trajectory fitting.
[0095] After obtaining the motion trajectory fitting curve, other appropriate motion state characteristics such as the velocity mean feature, acceleration fluctuation feature, and trajectory curvature average value can be calculated based on the motion trajectory fitting curve.
[0096] In some embodiments, the above-mentioned step S150 of abnormal recognition of the thermal abnormal area based on three-dimensional space features and motion state features can be implemented through the following technical process, specifically including: inputting the three-dimensional structure feature, surface curvature feature, and motion state feature into the abnormal recognition model to obtain the abnormal recognition result; in response to the abnormal recognition result being a preset target category, predicting the motion trajectory of the target contour according to the multi-moment trajectory point set.
[0097] Among them, the abnormal recognition model can represent a calculation model for classifying the corresponding target in the area based on the three-dimensional space features and motion state features of the thermal abnormal area. Exemplarily, the abnormal recognition model can be constructed based on other appropriate classification models such as a multi-layer perceptron neural network, a support vector machine classifier, and a random forest model. The abnormal recognition result can represent the class prediction output made by the abnormal recognition model for the input thermal abnormal area, and its result can include the class label of the target and its corresponding probability score. Among them, the class label can include person falling into water, floating object, boat, or obstacle, etc. The target category can represent the recognition target type with response priority or special warning significance predefined in the water area safety inspection task, specifically person falling into water or floating object, etc.
[0098] Preferably, an anomaly recognition model can be constructed based on a multi-layer perceptron neural network, and the motion trajectory of the target contour can be predicted through a long short-term memory network model. Specifically, in the case where the anomaly recognition result indicates that the target to which the thermal anomaly region belongs is a preset target category, the multi-moment trajectory point sets corresponding to the target contour are encoded into a trajectory input sequence in chronological order. The trajectory input sequence is input into the trained long short-term memory network model, and the predicted values of the trajectory positions of the target at multiple future moments are obtained through progressive time series modeling. The long short-term memory network can record the hidden state information and historical memory information within multiple time steps for learning the continuous evolution pattern of the target contour in the time dimension, and then output the three-dimensional motion trajectory sequence of the target contour within the predicted time interval.
[0099] When the recognition result indicates that the target is a preset target category, such as a person falling into the water, further predicting the motion trajectory of the target contour based on the multi-moment trajectory point sets can effectively restore the motion trend of the target in the future period, dynamically reflect its drift path, thereby providing a more accurate decision-making basis for water rescue response and improving the rescue success rate of the person falling into the water.
[0100] In some embodiments, the data acquisition strategy corresponding to the unmanned aerial vehicle can be determined according to the current environmental information of the target water area through the following technical steps. Among them, by determining the data acquisition strategy corresponding to the unmanned aerial vehicle according to the current environmental information of the target water area, the dynamic adjustment of the sensor working mode and acquisition frequency can be realized, thereby improving the effectiveness and adaptability of data acquisition. For example, high-frequency visible light image acquisition is preferentially adopted in the daytime scene with sufficient light, and infrared imaging and lidar encrypted scanning are dynamically switched in the low visibility or night scene, effectively avoiding the problem of insufficient performance of a single sensor in a specific environment.
[0101] In some embodiments, determining the data acquisition strategy corresponding to the unmanned aerial vehicle according to the current environmental information of the target water area specifically includes the following technical processes: obtaining the current light intensity information, visibility information and wind speed information of the target water area; determining the data acquisition strategy corresponding to the unmanned aerial vehicle based on the light intensity information, visibility information and wind speed information; among them, the data acquisition strategy includes the acquisition frequency and acquisition range corresponding to each type of data.
[0102] Among them, the acquisition frequency can represent the number of times the image sensor or lidar carried by the unmanned aerial vehicle performs data acquisition per unit time. The acquisition frequency can represent the spatial area or field of view boundary covered by the image sensor or lidar carried by the unmanned aerial vehicle.
[0103] Exemplarily, the light intensity information can represent the level of light radiation energy received per unit area of the target area, preferably characterized in lux (Lux); the visibility information can represent the maximum horizontal distance at which the UAV sensor can clearly identify the target under the current meteorological conditions; the wind speed information can represent the real-time wind speed magnitude at the spatial position where the UAV is located. An environmental state vector is constructed based on the above three types of environmental information, and combined with the preset data acquisition strategy mapping rules, the acquisition frequencies and acquisition ranges corresponding to the visible light image, infrared image, and lidar data are respectively determined. Preferably, when the light intensity is high, the visibility is good, and the wind speed is low, high-frequency acquisition (such as 30 frames per second) and wide-angle coverage (such as 120° horizontal viewing angle) of the visible light image are configured; when the light intensity decreases or the visibility drops to a set threshold, the acquisition frequency of the infrared image is increased, and the acquisition range of the visible light image is reduced to reduce the risk of defocus; when the wind speed increases beyond the flight stability threshold, the acquisition range of the lidar is appropriately reduced and its sampling density is increased to enhance the structural recognition ability of distant targets.
[0104] In some embodiments, for the thermally normal regions in the candidate region set, anomaly recognition can be performed through the following technical steps: screen out the thermally normal regions from the candidate region set, and match the target contour corresponding to the thermally normal region with the reference contour; in response to the matching result being a human body, determine the three-dimensional spatial features corresponding to the target contour according to the lidar alignment data corresponding to the target contour; perform anomaly recognition on the thermally normal region based on the three-dimensional spatial features.
[0105] Among them, the reference contour can represent a set of standardized target shape model contours pre-constructed and stored in the system, used for spatial shape matching and similarity analysis with the actually recognized target contour. The reference contour can be generated according to the two-dimensional projection contour features of the typical human body structure in different postures and perspectives. Preferably, it includes the boundary shapes of common drowning human body postures such as standing sideways, floating horizontally, and curling and floating, and is represented in the form of a contour point set or an edge function, used for calculating the contour coincidence degree or shape template matching with the target contour extracted from the thermally normal region. The three-dimensional spatial features can represent a set of three-dimensional structure information of the thermally normal region target constructed based on the lidar alignment data.
[0106] In specific implementation, the regions in the candidate region set other than the thermally abnormal regions can be used as thermally normal regions, and then the contour point set of the target contour in the visible light alignment data of the thermally normal regions is obtained to construct a target contour set , where, represents the coordinate of the th edge point of the target contour in the image coordinate system, represents the number of contour points. The target contour is compared with the preset reference contour set Perform spatial shape matching by constructing a normalized contour distance loss function:
[0107]
[0108] Among them, represents the minimum matching loss between the target contour and the reference contour, represents the two-dimensional Euclidean distance, represents the target contour point in the thermal normal area and the reference contour point the two-dimensional spatial coordinate difference between them.
[0109] Furthermore, to enhance the structural stability of contour shape matching, a contour edge direction gradient similarity index is introduced:
[0110]
[0111] Among them, represents the edge direction similarity score, represents the number of contour points, represents the normal direction of the target contour at point , represents the normal direction of the reference contour point matched with , the higher the similarity, approaches 1. When both is less than the maximum matching distance error threshold and is greater than the minimum required threshold of the edge direction gradient similarity, it is determined that the target contour matches the reference contour, and the matching result is the human body category.
[0112] Finally, in response to the matching result being a human body, extract the spatial point cloud data of the corresponding contour area from the lidar alignment data, and construct its three-dimensional spatial feature set based on the spatial projection and boundary mapping method. The three-dimensional spatial features include, but are not limited to, structural feature parameters such as size information, volume information, boundary compactness, and point cloud density. The three-dimensional spatial features can be input into the anomaly recognition and determination model to perform a secondary determination on whether there is a situation of a person falling into the water.
[0113] In some embodiments, as shown in Figure 3 , when the clarity of the visible light image does not meet the preset conditions, anomaly recognition can be performed through the following technical steps, specifically including:
[0114] Step S310, determine the clarity of the visible light image.
[0115] Among them, the clarity of the visible light image can be represented by other suitable image information such as texture information, contrast information, and gradient mean.
[0116] Step S320, in response to the clarity being lower than the quality reference value, perform feature extraction on the visible light alignment data, the infrared alignment data, and the lidar alignment data respectively to obtain a first image feature, a second image feature, and a third image feature.
[0117] Among them, the first image feature can represent the image structure information and texture information extracted based on the visible light alignment data. The second image feature can represent the thermal radiation response information extracted based on the infrared alignment data. The third image feature can represent the spatial structure feature extracted based on the lidar alignment data.
[0118] Step S330, perform feature fusion on the first image feature, the second image feature, and the third image feature to obtain a multi-modal fusion feature. Among them, the feature fusion can adopt other appropriate methods such as weighted fusion and feature splicing.
[0119] Step S340, perform anomaly recognition based on the multi-modal fusion feature.
[0120] Preferably, the anomaly recognition model can be constructed based on a multi-layer perceptron neural network. The multi-layer perceptron includes an input layer, at least two hidden layers, and an output layer. Each hidden layer uses a non-linear activation function to perform layer-by-layer transformation on the fusion feature vector. The input layer receives the fusion vector composed of three types of modal features, and extracts high-order discriminant features through full connection operations and activation mappings in the hidden layer. Finally, the anomaly recognition result is generated in the output layer. The output result can be in the form of classification, used to determine whether the detected target is a drowning person, a floating object, a ship, or a non-anomalous background area, and output the corresponding confidence score.
[0121] In this embodiment, when the clarity of the visible light image is lower than the preset quality reference value, anomaly recognition is performed based on the fusion feature, which can make full use of the infrared thermal radiation information and the lidar spatial structure information for compensation analysis when the quality of the visible light image is insufficient to support effective recognition, thereby improving the recognition stability and accuracy of abnormal targets in complex environments such as low light, strong reflection, or blurred images.
[0122] In some embodiments, the clarity corresponding to the visible light image can be determined through the following technical process: perform texture analysis on the visible light image to obtain the texture information of the visible light image; perform gray histogram statistical processing on the visible light image to obtain the contrast information of the visible light image; determine the clarity corresponding to the visible light image based on the texture information and the contrast information.
[0123] Specifically, first, perform texture analysis on the visible light image based on the gray change situation within the local neighborhood of the image to construct a texture complexity matrix , where represent the pixel coordinate points in the image of the local texture intensity. Preferably, the texture intensity can be calculated using the mode distribution frequency or gradient response amplitude after encoding with the local binary pattern. Perform a normalized mean statistic on the texture intensity within the entire image area to obtain the global texture index of the image:
[0124]
[0125] wherein, represents the image texture complexity score, and respectively represent the width and height of the image, respectively represent the pixel coordinates in the horizontal direction and vertical direction of the image.
[0126] Subsequently, perform contrast evaluation based on the grayscale histogram of the visible light image to construct the image grayscale intensity histogram , wherein represents the grayscale value, represents the pixel frequency corresponding to the grayscale value. Calculate the standard deviation of this histogram:
[0127]
[0128] wherein, represents the grayscale contrast score of the image, represents the average grayscale value of the image, represents the total number of pixels in the image.
[0129] Finally, perform weighted fusion on the image texture complexity score and the grayscale contrast score to obtain the image quality score corresponding to the visible light image, that is, the clarity corresponding to this visible light image.
[0130] In some embodiments, the clarity corresponding to the visible light image can also be determined through the following technical process: divide the visible light image into multiple local image regions; determine the local gradient mean corresponding to each local image region, perform weighted fusion on the local gradient means, and determine the clarity corresponding to the visible light image based on the fusion result.
[0131] Among them, the local gradient mean can represent the average of the gradient intensities of all pixel points in a certain local image area in the image, which is used to describe the severity of the brightness change in this area, and further reflect the clarity of the edge details. Preferably, the local gradient intensity can be calculated by performing a first-order gradient operator, such as the Sobel operator, on each pixel point. In a specific implementation, the visible light image can be divided into several image sub-regions of the same size. For the pixel points in each sub-region, the gray-scale gradients in the horizontal and vertical directions are calculated respectively using the Sobel operator, and the gradient amplitude of each pixel point is obtained accordingly; the gradient amplitudes of all pixel points in this sub-region are averaged to obtain the local gradient mean of this region; then, all local gradient means are weighted and fused according to the weights set by the regional position or brightness distribution to obtain a fused gradient index, and finally this fused index is used as the clarity score of the image to measure whether the current visible light image meets the abnormal recognition requirements.
[0132] In some embodiments, feature fusion is performed on the first image feature, the second image feature, and the third image feature to obtain a multi-modal fusion feature, which specifically includes: determining the fusion weights corresponding to the first image feature, the second image feature, and the third image feature according to the environmental information corresponding to the visible light image; performing weighted fusion processing on the first image feature, the second image feature, and the third image feature based on the fusion weights to obtain a multi-modal fusion feature.
[0133] Among them, by dynamically generating fusion weights based on the environmental information corresponding to the visible light image, the system can adaptively select the optimal feature source, reasonably allocate the weights of visible light, infrared, and lidar features under different lighting, visibility, or imaging conditions, so as to improve the adaptability of multi-modal feature fusion and the accuracy of abnormal recognition. In the specific implementation process, the fusion weights corresponding to the first image feature, the second image feature, and the third image feature can be determined according to the illumination intensity information, visibility information, and wind speed information corresponding to the acquisition of the visible light image. For example, in an environmental condition with high illumination intensity, good visibility, and low wind speed, the fusion weight of the first image feature is preferentially increased, and the fusion weights of the second image feature and the third image feature are appropriately reduced; when the illumination intensity is low or the visibility decreases, the fusion weight of the second image feature is correspondingly increased to enhance the discriminant ability for heat source targets; when the wind speed increases, the fusion weight of the third image feature is enhanced.
[0134] In some embodiments, after obtaining the abnormal recognition result, warning can also be performed based on the following technical steps, which specifically include: extracting the visible light image corresponding to the abnormal area according to the abnormal recognition result; visually displaying the visible light image, the abnormal recognition result, and the position information of the abnormal area; generating a corresponding warning response instruction according to the abnormal type corresponding to the abnormal recognition result.
[0135] In specific implementation, after anomaly recognition is completed, visible light alignment data corresponding to the anomaly area in terms of time and space is obtained, and an image segment covering the anomaly area is extracted based on the visible light alignment data to construct a visible light image corresponding to the anomaly area. Based on the above-mentioned visible light image, the target category information and the corresponding confidence information included in the anomaly recognition result are labeled, and combined with the three-dimensional position information of the anomaly area, a visual image interface with recognition results and spatial annotations is constructed, so that the anomaly area is highlighted and category-annotated in the image. Preferably, the distribution of trajectory points of the target at historical and predicted times can also be superimposed on the interface to assist in judging its movement direction and trend. Further, according to the anomaly type characterized by the anomaly recognition result, a response rule corresponding to this type is called, and a warning response instruction is generated. The warning response instruction can include response level, response method, and rescue strategy information. Among them, the response level is used to distinguish the severity of the event, the response method is used to specify the output form of the warning content (such as interface pop-up window, sound and light prompt, or remote notification), and the linkage strategy is used to control subsequent response actions such as drone tracking, rescue equipment startup, or command terminal synchronization.
[0136] In some embodiments, after obtaining the anomaly recognition result, the following technical steps can also be performed for precise drone patrol, specifically including: determining the spatial position information corresponding to the anomaly target based on the anomaly recognition result; generating a flight path adjustment instruction based on the spatial position information and the current position of the drone; adjusting the flight path of the drone based on the flight path adjustment instruction to make the drone approach the anomaly target; and updating the flight path adjustment instruction based on the real-time spatial relationship between the anomaly target and the drone to optimize the observation angle and cruising radius of the drone.
[0137] When the recognition result output by the anomaly recognition model indicates that the target belongs to a preset target category, such as a person falling into the water or an abnormal ship, the three-dimensional spatial position information corresponding to the anomaly target is extracted from the lidar alignment data, and this position is used as the target navigation reference point. Further, the flight attitude parameters and spatial position information of the drone at the current moment are obtained, a spatial vector relationship between the anomaly target and the drone is constructed, and an initial flight path adjustment instruction is generated based on the principle of the shortest flight distance and safety approach constraints. Execute the flight path adjustment instruction to dynamically guide the drone to deflect the current flight path and correct the attitude, so that its flight direction gradually deviates towards the anomaly target space area. During the flight, continuously update the spatial position of the anomaly target and the current pose data of the drone, calculate the relative azimuth and distance information between the two in real time, and optimize the adjustment instruction based on strategies such as multi-angle coverage priority, constant radius approach, or low-energy stable hover, so that the drone can track and approach the anomaly target and achieve the best perspective coverage while maintaining observation stability, thereby improving the patrol response ability and image acquisition quality in a high-dynamic target environment.
[0138] Furthermore, Figure 4 Schematically shows a schematic diagram of the system architecture of a water area safety inspection system based on multi-modal feature analysis according to some embodiments of the present disclosure, which is applied to the safety inspection and early warning response of a target reservoir.
[0139] Among them, the inspection system consists of a public platform, an application layer, a service layer, a data access layer, a persistence layer, and a security management module. The public platform module integrates function entrances such as a console, inspection trajectory, real-time image, data analysis, early warning setting, information recording, and remote monitoring. Users can control the drone based on a preset data collection strategy through the console to collect visible light images, infrared images, and lidar data corresponding to the target water area. The relevant image data is displayed through the real-time image module for the data analysis module to perform edge detection and region segmentation processing to generate a set of candidate regions containing the target contour. The system application layer includes sub-modules such as inspection monitoring, anomaly warning, data analysis, image recognition, and device management. The specific functions interact with the service layer through the WS protocol, HTTP protocol, and FTP protocol to complete image processing, data analysis, anomaly detection, data collection, and early warning notification services.
[0140] In the service layer, the data collection service is used to collect visible light images, infrared images, and lidar data corresponding to the target water area. The image processing service is used to perform preprocessing operations on the visible light images and extract the first image features. The anomaly detection service is used to extract heat source features and perform heat anomaly screening processing based on the infrared alignment data to screen out heat anomaly regions from the set of candidate regions. The data analysis service extracts three-dimensional spatial features and motion state features from the lidar alignment data and performs anomaly recognition on the heat anomaly regions based on the three-dimensional spatial features and motion state features. The early warning analysis service is used to generate corresponding early warning response instructions according to the anomaly type corresponding to the anomaly recognition result. The data access layer supports data writing, deletion, update, and query to realize the dynamic management of inspection data. The data is written into the reservoir information database, inspection record database, early warning information database, device status database, and monitoring database in the persistence layer through the access interface, supporting the traceability analysis of historical inspection results and early warning responses. The security management module provides a unified authentication, unified authorization, unified audit, and unified permission management mechanism to ensure the data security and operation compliance during the system operation. This system architecture completely covers the entire process from perception data collection, multi-modal fusion analysis to anomaly recognition and response. Through the collaborative work of different functional layers, it realizes high-precision, low-latency, and multi-angle intelligent safety inspection in complex water area environments.
[0141] Figure 5Schematically shown is a schematic diagram of a monitoring interface of a water area safety inspection system based on multimodal feature analysis according to some embodiments of the present disclosure. It is applied to the real-time inspection and abnormal early warning scenario design of the reservoir area. This interface mainly includes four functional modules: aerial intelligent inspection, image recognition and analysis, real-time reservoir monitoring, and early warning information release. Among them:
[0142] The aerial intelligent inspection module is used to display the real-time progress and data processing situation of the drone inspection task. Among them, the data shows that the current coverage rate of the target reservoir area by the drone is 25%, the number last month was 63, and the number this month is 114, reflecting the current inspection frequency and time distribution; the early warning information release data shows that the current number of pushed early warning information is 170, and it was 141 last month, reflecting the abnormal discovery ability in the inspection and recognition process; the image recognition and analysis index shows that the system has processed 886 images this month, and the recognition efficiency has increased by 422% compared with last month. The early warning information release module provides the real-time early warning status of the current reservoir area. The figure shows that the current early warning level is a first-level early warning, and the completion degree is shown in the form of a circular chart as 60%. This result is the response level generated by the system based on the three-dimensional space characteristics and motion state characteristics of the targets in the abnormal area for abnormal recognition and combined with preset conditions.
[0143] The image recognition and analysis module further classifies the early warning events in the form of a table, including three columns: early warning level, trigger condition, and associated matters. Among them, the early warning levels are divided into first level, second level, and third level, and the corresponding trigger conditions include personnel falling into the water, abnormal vessels, and abnormal floating objects respectively. Each type of abnormal situation corresponds to specific associated matters. For example, personnel falling into the water is associated with the location of the personnel, abnormal vessels are associated with the location of the vessels, and abnormal floating objects are associated with the type of floating objects. It can be automatically generated by the image segmentation and feature recognition module and presented in a structured manner on the interface. The reservoir real-time monitoring module is a time series graph, which is used to show the change trend of different-level early warning events in the past year. In the figure, the occurrence numbers of the first-level early warning, second-level early warning, and third-level early warning are respectively plotted in the form of a line graph for each month, reflecting the system's ability to continuously perceive and respond to the safety situation of the target water area. The chart data can be obtained from the system's extraction of heat source characteristics and screening of thermal anomalies for the candidate area set, and combined with the target contour motion trajectory extracted from the lidar alignment data to complete abnormal recognition and level division.
[0144] It should be noted that although the steps of the methods in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in that specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution, etc.
[0145] In addition, in the present exemplary embodiment, a water area safety inspection system based on multi-modal feature analysis is also provided. Referring to Figure 6 As shown in Figure 6 , the water area safety inspection system 600 based on multi-modal feature analysis includes: a data acquisition module 610, a data alignment module 620, a contour extraction module 630, a thermal anomaly screening module 640, and an anomaly recognition module 650. Among them:
[0146] The data acquisition module 610 can be used to control the drone to collect visible light images, infrared images, and lidar data corresponding to the target water area based on a preset data acquisition strategy;
[0147] The data alignment module 620 can be used to perform time synchronization and spatial alignment on the visible light image, infrared image, and lidar data to obtain visible light aligned data, infrared aligned data, and lidar aligned data;
[0148] The contour extraction module 630 can be used to perform edge detection and region segmentation processing on the visible light aligned data in response to the clarity of the visible light image being higher than the quality reference value, and generate a set of candidate regions containing the target contour;
[0149] The thermal anomaly screening module 640 can be used to perform heat source feature extraction and thermal anomaly screening processing on the infrared aligned data corresponding to the set of candidate regions, and screen out thermal anomaly regions from the set of candidate regions;
[0150] The anomaly recognition module 650 can be used to perform three-dimensional spatial feature extraction and motion state feature extraction on the lidar aligned data corresponding to the thermal anomaly region, and perform anomaly recognition on the thermal anomaly region based on the three-dimensional spatial features and motion state features.
[0151] The specific details of each module of the above water area safety inspection system based on multi-modal feature analysis have been described in detail in the corresponding water area safety inspection method based on multi-modal feature analysis, so they will not be elaborated here.
[0152] It should be noted that although several modules or units of the water area safety inspection system based on multi-modal feature analysis are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0153] In addition, in the exemplary embodiment of the present disclosure, an electronic device capable of implementing the above water area safety inspection method based on multi-modal feature analysis is also provided.
[0154] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware embodiment, a complete software embodiment (including firmware, microcode, etc.), or an embodiment combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.
[0155] Reference will now be made to Figure 7 to describe the electronic device 700 according to an embodiment of the present disclosure. Figure 7 The illustrated electronic device 700 is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present disclosure.
[0156] As Figure 7 illustrated, the electronic device 700 is presented in the form of a general-purpose computing device. The components of the electronic device 700 may include, but are not limited to: at least one of the above-mentioned processing units 710, at least one of the above-mentioned storage units 720, a bus 730 connecting different system components (including the storage unit 720 and the processing unit 710), and a display unit 740.
[0157] Wherein, the storage unit stores program code, and the program code can be executed by the processing unit 710, so that the processing unit 710 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0158] The storage unit 720 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 721 and / or a cache storage unit 722, and may further include a read-only storage unit (ROM) 723.
[0159] The storage unit 720 may further include a program / utility 724 having a set (at least one) of program modules 725. Such program modules 725 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data, and the implementation of a network environment may be included in each or some combination of these examples.
[0160] The bus 730 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.
[0161] The electronic device 700 may also communicate with one or more external devices 770 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 700, and / or communicate with any device that enables the electronic device 700 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 750. Moreover, the electronic device 700 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 760. As shown in the figure, the network adapter 760 communicates with other modules of the electronic device 700 through the bus 730. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in combination with the electronic device 700, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0162] Through the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, and the software product can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0163] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above method of the present specification is stored. In some possible embodiments, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of the present specification.
[0164] Reference Figure 8 As shown, a program product 800 for implementing the above-mentioned water area safety inspection method based on multi-modal feature analysis according to an embodiment of the present disclosure is described. It may adopt a portable compact disc read-only memory and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, the readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
[0165] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0166] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0167] The program code contained on the readable medium may be transmitted with any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0168] The program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).
[0169] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed, for example, synchronously or asynchronously in multiple modules.
[0170] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solution according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a portable hard drive, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0171] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0172] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A water area safety inspection method based on multi-modal feature analysis, characterized in that, Including: Controlling the drone to collect visible light images, infrared images, and lidar data corresponding to the target water area based on a preset data collection strategy; Performing time synchronization and spatial alignment on the visible light images, infrared images, and lidar data to obtain visible light aligned data, infrared aligned data, and lidar aligned data; In response to the clarity of the visible light image being higher than the quality reference value, performing edge detection and region segmentation processing on the visible light aligned data to generate a set of candidate regions containing the target contour; Performing heat source feature extraction and thermal anomaly screening processing on the infrared aligned data corresponding to the set of candidate regions, and screening out thermal anomaly regions from the set of candidate regions; Performing three-dimensional spatial feature extraction and motion state feature extraction on the lidar aligned data corresponding to the thermal anomaly regions, and performing anomaly recognition on the thermal anomaly regions based on the three-dimensional spatial features and the motion state features; Screening out thermally normal regions from the set of candidate regions, and matching the target contour corresponding to the thermally normal region with a reference contour; In response to the matching result being a human body, determining the three-dimensional spatial features corresponding to the target contour according to the lidar aligned data corresponding to the target contour; Performing anomaly recognition on the thermally normal regions based on the three-dimensional spatial features; Determining the clarity corresponding to the visible light image; In response to the clarity being lower than the quality reference value, respectively performing feature extraction on the visible light aligned data, infrared aligned data, and lidar aligned data to obtain a first image feature, a second image feature, and a third image feature; Determining the fusion weights corresponding to the first image feature, the second image feature, and the third image feature according to the environmental information corresponding to the visible light image; Performing weighted fusion processing on the first image feature, the second image feature, and the third image feature based on the fusion weights to obtain a multi-modal fusion feature; Performing anomaly recognition based on the multi-modal fusion feature.
2. The water area safety inspection method based on multi-modal feature analysis according to claim 1, wherein, The determining the clarity corresponding to the visible light image includes: Performing texture analysis on the visible light image to obtain the texture information of the visible light image; Performing grayscale histogram statistical processing on the visible light image to obtain the contrast information of the visible light image; Determining the clarity corresponding to the visible light image based on the texture information and the contrast information.
3. The water area safety inspection method based on multimodal feature analysis according to claim 1, wherein The determining the clarity corresponding to the visible light image includes: Performing region division on the visible light image to obtain a plurality of local image regions; Determining the local gradient mean corresponding to each local image region, performing weighted fusion on the local gradient means, and determining the clarity corresponding to the visible light image based on the fusion result.
4. The water area safety inspection method based on multimodal feature analysis according to claim 1, characterized in that Before controlling the drone based on a preset data collection strategy, it further includes: Determining the data collection strategy corresponding to the drone according to the current environmental information of the target water area.
5. The method for water area safety inspection based on multimodal feature analysis according to claim 4, characterized in that, The determining the data collection strategy corresponding to the drone according to the current environmental information of the target water area includes: Obtaining the current light intensity information, visibility information, and wind speed information of the target water area; Determine the corresponding data acquisition strategy for the drone based on the light intensity information, visibility information, and wind speed information; Among them, the data acquisition strategy includes the acquisition frequency and acquisition range corresponding to each type of data.
6. The water area safety inspection method based on multi-modal feature analysis according to claim 1, wherein The process of performing heat source feature extraction and thermal anomaly screening on the infrared alignment data corresponding to the candidate region set, and screening out thermal anomaly regions from the candidate region set includes: Extract multi-scale heat source features of the infrared alignment data corresponding to the candidate region set at multiple scales; Based on the multi-scale heat source features, determine the temperature distribution characteristics corresponding to each candidate region; Based on the temperature distribution characteristics, extract the heat source morphology characteristics corresponding to each candidate region; Based on the heat source morphology characteristics, screen out the thermal anomaly regions with irregular edge characteristics or concentrated heat source characteristics from the candidate region set.
7. The method for water area safety inspection based on multi-modal feature analysis according to claim 1, wherein The process of performing three-dimensional spatial feature extraction and motion state feature extraction on the lidar alignment data corresponding to the thermal anomaly region includes: Based on the lidar alignment data, extract the three-dimensional structure features and surface curvature features corresponding to the thermal anomaly region, where the three-dimensional structure features include one or more of size information, volume information, bounding box compactness, and spatial information; Based on the lidar alignment data collected at different times, extract the multi-time trajectory point set corresponding to the target contour in the thermal anomaly region; Based on the multi-time trajectory point set corresponding to the target contour, fit the motion trajectory curve, and extract the motion state features based on the motion trajectory curve.
8. The method for water area safety inspection based on multimodal feature analysis according to claim 7, wherein, The process of performing anomaly recognition on the thermal anomaly region based on the three-dimensional spatial features and the motion state features includes: Input the three-dimensional structure features, the surface curvature features, and the motion state features into the anomaly recognition model to obtain the anomaly recognition result; In response to the anomaly recognition result being a preset target category, predict the motion trajectory of the target contour based on the multi-time trajectory point set.
9. The method for water area safety inspection based on multi-modal feature analysis according to claim 8, characterized in that It also includes: Extract the visible light image corresponding to the anomaly region according to the anomaly recognition result; Perform visual display on the visible light image, the anomaly recognition result, and the position information of the anomaly region; Generate a corresponding early warning response instruction according to the anomaly type corresponding to the anomaly recognition result.
10. The method for water area safety inspection based on multi-modal feature analysis according to claim 8, characterized in that, It also includes: Based on the anomaly recognition result, determine the spatial position information corresponding to the anomaly target; Generate a flight path adjustment instruction based on the spatial position information and the current position of the drone; Based on the flight path adjustment instruction, adjust the flight path of the drone to make the drone approach the anomaly target; Based on the real-time spatial relationship between the anomaly target and the drone, update the flight path adjustment instruction to optimize the observation angle and cruising radius of the drone.
11. A water area safety inspection system based on multimodal feature analysis, which is used to execute the water area safety inspection method based on multimodal feature analysis described in any one of claims 1 to 10, and is characterized in that, It includes: A data acquisition module for controlling the drone to collect visible light images, infrared images, and lidar data corresponding to the target water area based on a preset data acquisition strategy; A data alignment module, configured to perform time synchronization and spatial alignment on the visible light image, infrared image, and lidar data to obtain aligned visible light data, aligned infrared data, and aligned lidar data; A contour extraction module, configured to perform edge detection and region segmentation processing on the aligned visible light data in response to the clarity of the visible light image being higher than a quality reference value, and generate a set of candidate regions containing the target contour; A thermal anomaly screening module, configured to perform heat source feature extraction and thermal anomaly screening processing on the aligned infrared data corresponding to the set of candidate regions, and screen out thermal anomaly regions from the set of candidate regions; An anomaly recognition module, configured to perform three-dimensional spatial feature extraction and motion state feature extraction on the aligned lidar data corresponding to the thermal anomaly regions, and perform anomaly recognition on the thermal anomaly regions based on the three-dimensional spatial features and the motion state features.
Citation Information
Patent Citations
Intelligent security and protection monitoring method and system based on image recognition
CN119495054A
Method and device for identifying unmanned aerial vehicle based on motion features of high-point camera holder
CN119625633A