Inspection robot risk area identification method based on multi-modal fusion perception
By employing multimodal fusion sensing technology and utilizing time alignment, spatial registration, and extended Kalman filter algorithms, the problem of recognition accuracy and stability of multimodal inspection in high dynamic environments in existing technologies has been solved, achieving early warning and stable risk area identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUZHOU VOCATIONAL UNIVERSITY (SUZHOU OPEN UNIVERSITY)
- Filing Date
- 2026-01-27
- Publication Date
- 2026-04-21
AI Technical Summary
Existing multimodal inspection technologies struggle to simultaneously achieve accurate identification, early warning capabilities, and output stability in highly dynamic and highly disruptive industrial inspection scenarios. They suffer from issues such as difficulty in quantifying modal reliability, lack of recursive modeling for risk states, and susceptibility to false alarms triggered by short-term noise.
By acquiring visible light images, thermal imaging images, laser point cloud distance measurements, and acoustic signal measurements, multimodal observation records are generated. Time alignment and spatial registration are performed, confidence weights are calculated, and an extended Kalman filter algorithm is used to construct a recursive risk estimation model. The risk estimate and rate of change are output. By combining consistency judgment and connected component merging for M consecutive sampling periods, stable risk area identification results and early warning levels are generated.
It reduces the risk of missed detection due to single sensor distortion in complex environments, suppresses false alarms caused by occlusion, thermal drift and sparse echo, forms early warning, improves the accuracy and stability of risk area identification, and weakens the impact of isolated grids triggered by short-term noise.
Smart Images

Figure CN121901684A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of multi-sensor fusion perception, risk area identification, and safety early warning for inspection robots, and particularly to a method for risk area identification of inspection robots based on multi-modal fusion perception. Background Technology
[0002] With the increasing digitalization and intelligence of industrial sites, inspection operations in high-risk scenarios such as power stations, petrochemical plants, mine tunnels, warehousing and logistics facilities, and urban underground utility tunnels are gradually shifting from manual inspections to mobile robot inspections. Among existing technologies, visible light images have advantages in target recognition and structural texture extraction; thermal imaging images are more sensitive to temperature anomalies and hidden heat sources; laser point cloud distance measurements can provide stable geometric scale and obstacle spatial structure information; and acoustic signal measurements have high sensitivity to dynamic anomalies such as abnormal noises, leaks, and friction. Meanwhile, in order to transform multi-source observations into navigation-friendly... Risk region representation for safety decision invocation often employs methods such as rasterization, risk cost map generation, and threshold judgment after fusion, supplemented by target detection networks, semantic segmentation, point cloud clustering, or rule-based multi-condition judgment to form a technology for perceiving, fusing, and labeling risk regions. With the development of computing power platforms and embedded sensors, multimodal fusion is gradually moving from simple feature splicing or voting strategies to fusion methods that consider sensor reliability, spatiotemporal registration errors, and uncertainty propagation, enabling inspection robots to still have a certain risk identification capability under complex environmental conditions such as lighting changes, dust and smoke, thermal disturbances, and occlusion.
[0003] However, existing technologies still have significant shortcomings: First, multimodal fusion often relies on equal or static weights, failing to explicitly incorporate factors such as signal-to-noise characteristics, occlusion levels, temperature drift, and echo sparsity of each mode under different operating conditions into the fusion weights. This results in a distorted mode still being forcibly included in the fusion process, leading to false alarms or missed alarms. Second, most risk area determinations rely on the fusion results of a single frame or single moment, lacking a recursive estimation mechanism for the risk evolution process. Faced with situations such as short-term personnel movement, transient hotspots on equipment, echo flypoints, and instantaneous noise impacts, instantaneous misjudgments or delayed warnings are prone to occur, making it difficult to generate early warnings. Third, existing methods typically... Using threshold segmentation or simple connected component processing as the final region output lacks consistency constraints and stability criteria across sampling periods, making it difficult to suppress isolated grids triggered by short-term noise. This leads to boundary jitter in risk areas and unstable level determination, which in turn affects the robot's obstacle avoidance strategy and the scheduling system's selection of risk handling priorities. Fourth, although some solutions attempt to introduce temporal networks or post-processing smoothing, they often only perform statistical filtering at the algorithm level and fail to establish a clear physical-statistical correlation between multimodal observations and risk state evolution. These shortcomings make it difficult for existing technologies to simultaneously ensure identification accuracy, early warning capability, and output stability in highly dynamic and highly interfered industrial inspection scenarios.
[0004] In summary, existing multimodal inspection risk identification technologies generally suffer from problems such as difficulty in quantifying modal reliability, lack of recursive modeling of risk states, and susceptibility to false alarms triggered by short-term noise, causing fluctuations in regional and risk level. The core problem addressed by this invention is: how to form a unified grid feature expression through time alignment and spatial registration under multimodal observation conditions, introduce confidence weights corresponding to each modality, and combine a recursive risk estimation model based on the extended Kalman filter algorithm to output risk estimates and risk change rates, thereby forming an early warning grid set. At the same time, stable risk area identification results and early warning levels are obtained through consistency judgment and connected component merging over M consecutive sampling periods. Summary of the Invention
[0005] The purpose of this section is to outline some aspects of the embodiments of the present invention and to briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this section, the abstract and title of the invention. Such simplifications or omissions shall not be used to limit the scope of the present invention.
[0006] In view of the aforementioned existing problems, the present invention is proposed.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for identifying risk areas of an inspection robot based on multimodal fusion perception, comprising: The inspection robot collects visible light images, thermal images, laser point cloud distance measurements, and acoustic signal measurements of the working environment within the inspection area, and generates multimodal observation records. The multimodal observation records are subjected to temporal alignment and spatial registration to obtain a multimodal raster feature set, and the confidence weights corresponding to each modality are calculated. The multimodal raster feature set and the confidence weight are input into the recursive risk estimation model constructed based on the extended Kalman filter algorithm, and the risk estimate and risk change rate are output to generate an early warning raster set. The consistency determination is performed on the early warning grid set within M consecutive sampling periods. Grid cells that meet the continuous consistency determination threshold are retained and connected components are merged. The risk area identification result and its corresponding early warning level are output.
[0008] In a second aspect, the present invention provides a computer device, comprising: One or more processors; The memory stores operable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, including the process described above for the inspection robot risk area identification method based on multimodal fusion perception.
[0009] Thirdly, the present invention provides a computer-readable medium for storing software, the software including instructions executable by one or more computers, the instructions causing the one or more computers to perform operations, the operations including the process of the aforementioned inspection robot risk area identification method based on multimodal fusion perception.
[0010] The beneficial effects of this invention are as follows: By acquiring visible light images, thermal imaging images, laser point cloud distance measurements, and acoustic signal measurements to generate multimodal observation records, this invention obtains a complementary observation basis for the inspection area's working environment, reducing the risk of missed detections caused by single-sensor distortion; by performing time alignment and spatial registration on the multimodal observation records to obtain a multimodal raster feature set and calculating confidence weights, each modality participates in fusion according to reliability, suppressing false alarms caused by occlusion, thermal drift, and echo sparsity; by inputting the multimodal raster feature set and confidence weights into a recursive risk estimation model constructed based on the extended Kalman filter algorithm, the invention outputs risk estimates and risk change rates, generating an early warning raster set, transforming risk from static discrimination to recursive prediction, thus forming an early warning; by performing consistency determination and connected component merging on the early warning raster set within M consecutive sampling periods, the invention outputs risk area identification results and early warning levels, weakening isolated rasters triggered by short-term noise, and improving the stability and usability of area boundaries and levels. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating the risk area identification method for inspection robots based on multimodal fusion perception, as shown in this invention. Detailed Implementation
[0012] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0013] Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort should fall within the scope of protection of this invention.
[0014] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0015] According to an embodiment of the present invention, in combination Figure 1 The flowchart shown illustrates a method for identifying risk areas in inspection robots based on multimodal fusion perception, which specifically includes the following steps: S1. Utilize the inspection robot to collect visible light images, thermal images, laser point cloud distance measurements, and acoustic signal measurements of the working environment within the inspection area, and generate a multimodal observation record. Note that the following points should be noted in this step: S1.1 At the start of each preset sampling period, the visible light camera, thermal imaging camera, lidar and acoustic sensor on the inspection robot are triggered to complete synchronous sampling at the same sampling timestamp, respectively obtaining visible light image, thermal imaging image, lidar point cloud distance measurement value and acoustic signal measurement value.
[0016] In a preferred embodiment, the visible light camera, thermal imaging camera, lidar, and acoustic sensor mounted on the inspection robot are all connected to the same time reference source. The time reference source is generated by a hardware timer on the main control board of the inspection robot. The hardware timer increments at a fixed frequency and outputs a sampling beat signal. In this embodiment, the preset sampling period is 200 milliseconds. The sampling beat signal arrives at the beginning of each preset sampling period to form a trigger condition.
[0017] To ensure that all modal data have the same sampling timestamp, when the sampling beat signal arrives, the main control board sends synchronous trigger pulses to the visible light camera, thermal imaging camera, lidar, and acoustic sensor respectively, and reads the hardware timer count value as the sampling timestamp at the same moment the synchronous trigger pulse is sent. When a sensor has internal exposure or internal scanning delay, its intra-frame timestamp is read from the data frame returned by the sensor, and the difference between the intra-frame timestamp and the sampling timestamp is corrected. The corrected time is still written back to the sampling timestamp field, so that all modal data are clustered around the same sampling timestamp.
[0018] As an example, a visible light image includes: pixel array data, image resolution field, exposure time field, gain field, lens focal length field, and camera intrinsic parameter calibration number; wherein, the pixel array data can be a 1920×1080 three-channel image frame, the exposure time field is 8 milliseconds, and the gain field is 6 dB.
[0019] As an example, the thermal imaging image includes: temperature array data, emissivity setting field, ambient reflectance temperature field, thermal image resolution field, and thermal imaging camera intrinsic parameter calibration number; wherein, the temperature array data can be a 320×240 temperature matrix, the emissivity setting field is 0.95, and the ambient reflectance temperature field is 25 degrees Celsius.
[0020] As an example, the laser point cloud distance measurement values include: distance sequence, echo intensity sequence, azimuth sequence, and lidar frame number output by scan line; wherein, the distance sequence is in meters, the number of points per frame is 6400, and the echo intensity sequence corresponds one-to-one with the distance sequence.
[0021] As an example, the acoustic signal measurements include: a multi-channel waveform sampling sequence, a sampling rate field, a channel gain field, and an array geometry number; wherein, the multi-channel waveform sampling sequence has 4 channels, the sampling rate field is 48 kHz, and the single-cycle sampling window is 200 milliseconds.
[0022] S1.2 Perform distortion correction and scale normalization on visible light and thermal imaging images, outlier removal on laser point cloud distance measurements, bandpass filtering on acoustic signal measurements, and write sampling timestamps and sensor identifiers to the processing results to obtain multimodal observation records.
[0023] It should be noted that the distortion correction of visible light images includes: locating the radial and tangential distortion coefficients from the camera intrinsic calibration numbers, performing anti-distortion remapping on the pixel coordinates, and outputting the remapped pixel array as the corrected visible light image.
[0024] It should be noted that the distortion correction of the thermal imaging image includes: locating the distortion coefficient of the thermal imaging lens from the intrinsic parameter calibration number of the thermal imaging camera and performing geometric resampling on the temperature array, and outputting the resampled temperature array as the corrected thermal imaging image.
[0025] Furthermore, the corrected visible light image is scaled to 640×360 and a normalized visible light image is output using a fixed interpolation method. The corrected thermal imaging image is scaled to 320×240 and a normalized thermal imaging image is output while keeping the temperature numerical scale unchanged. At the same time, a unified resolution field and interpolation method field are written to both, so that subsequent raster mapping has a consistent pixel scale basis.
[0026] It should be noted that the outlier removal process for laser point cloud distance measurements includes: calculating the number of neighboring points and the neighborhood distance statistics for each point within each lidar frame, with an example of 20 neighboring points; when the average distance from a point to its neighborhood is greater than the average distance of the entire frame plus twice the standard deviation, the point is marked as an outlier and removed from the distance sequence, and the remaining points form the cleaned point cloud distance measurements.
[0027] It should be noted that the bandpass filtering of the acoustic signal measurement values includes: selecting a bandpass frequency band from 300 Hz to 3000 Hz, and performing finite-length impulse response filtering on the multi-channel waveform sampling sequence. The filter order is 201 in the example. The filtered waveform forms the purified acoustic signal measurement value, while the sampling rate field is kept unchanged and written into the filter frequency band field.
[0028] Furthermore, the sampling timestamp and sensor identifier are written into each processing result; where, for example, CAM-VIS corresponds to the visible light mode, CAM-TH corresponds to the thermal imaging mode, LIDAR-RNG corresponds to the point cloud distance mode, and MIC-ARR corresponds to the acoustic mode; and within the same record entry, they are spliced and written in the order of normalized visible light image, normalized thermal imaging image, purified point cloud distance measurement value, and purified acoustic signal measurement value to obtain a multimodal observation record.
[0029] Preferably, in step S1, four types of data—visible light, thermal imaging, point cloud distance, and acoustic signals—are collected at the same sampling timestamp. This allows the illumination texture information, temperature anomaly information, spatial geometric information, and sound source anomaly information at the same moment to enter the subsequent alignment and registration process in parallel. Compared to schemes that rely solely on a single vision or a single point cloud, step S1 can still retain at least one stable observation channel in scenarios such as smoke and dust obstruction, weak light glare, heat source interference, and sparse local echoes. Furthermore, it binds each modal data to the same record entry using the sampling timestamp and sensor identifier, reducing the risk of misjudgment caused by cross-modal mismatch. Compared to the existing method of asynchronous sampling followed by coarse alignment by frame number, this embodiment writes the sampling timestamp with a unified time reference source and performs delay correction during the sampling stage. This ensures that subsequent time alignment no longer depends on unstable frame number intervals, thereby reducing raster feature jitter and mismatch rate.
[0030] S2. Perform temporal alignment and spatial registration on the multimodal observation records to obtain a multimodal raster feature set, and calculate the confidence weight corresponding to each modality. Note that the following points should be noted in this step: S2.1. Based on the sampling timestamp, perform time alignment on the visible light image, thermal imaging image, laser point cloud distance measurement value and acoustic signal measurement value to obtain the aligned observation set.
[0031] Specifically, in this embodiment, time alignment is performed using the sampling timestamp as the primary index. For each multimodal observation record, the intra-frame timestamps of the visible light mode, thermal imaging mode, point cloud distance mode, and acoustic mode are uniformly converted into millisecond counts within the record, and the time deviation is calculated with the sampling timestamp. When the absolute value of the time deviation of a certain mode is less than or equal to 10 milliseconds, the data of that mode is included in the aligned observation set. When the time deviation is greater than 10 milliseconds, the measurement values of the two adjacent records of that mode are read and interpolated measurement values are generated by linear interpolation. The interpolated measurement values are written into the aligned observation set, and an interpolation identifier is written into the mode field so that the stability change caused by interpolation can be identified in the subsequent weight calculation stage.
[0032] S2.2. Using the sensor calibration extrinsic parameter matrix, the laser point cloud distance measurements in the aligned observation set are converted into a point cloud coordinate set, and the point cloud coordinate set is projected onto a unified grid coordinate system to obtain a grid point cloud feature set. At the same time, based on the sensor calibration extrinsic parameter matrix, the visible light image and thermal imaging image in the aligned observation set are mapped to a unified grid coordinate system to obtain a grid image set. Based on the direction of arrival estimation results of the acoustic signal measurement values and the direction intersection of the point cloud coordinate set, the grid cell corresponding to the sound source is determined and a grid acoustic feature set is generated.
[0033] In this embodiment, the unified grid coordinate system takes the chassis coordinate system of the inspection robot as the origin, the forward direction is the positive X-axis of the grid, and the left direction is the positive Y-axis of the grid; the grid resolution is 0.2 meters, the grid coverage is 20 meters in front and 10 meters to the left and right, and the corresponding number of grid rows and columns is 100×100; the sensor calibration extrinsic parameter matrix is obtained and stored offline. The calibration method is as follows: planar targets and corner targets are placed in the calibration site, the inspection robot collects point clouds and images in different postures, and calculates the extrinsic parameters of the visible light camera, thermal imaging camera, and lidar through corner matching and planar constraints, and writes the calculation results into the sensor calibration extrinsic parameter matrix.
[0034] When converting the laser point cloud distance measurements in the aligned observation set into a point cloud coordinate set, the azimuth and elevation angles of the scan line in the lidar intrinsic parameters are read, and each distance measurement value is combined with its corresponding azimuth and elevation angles to convert it into a three-dimensional point in the lidar coordinate system. Then, the extrinsic parameter matrix of the lidar to the chassis is read, and the three-dimensional points are converted into a point cloud coordinate set in the chassis coordinate system. When projecting the point cloud coordinate set to a unified grid coordinate system, the projection coordinates of each three-dimensional point on the chassis horizontal plane are taken, and converted into grid row and column indices according to the grid resolution. When the number of points falling in a certain grid cell is greater than 0, the number of points, the distance to the nearest point, the height of the highest point, and the height variance fields are written to the grid cell to form a grid point cloud feature set.
[0035] When mapping visible light images and thermal imaging images to a unified grid coordinate system, the extrinsic parameter matrix and intrinsic parameter matrix of the corresponding camera to the chassis are read. For each grid cell, the three-dimensional coordinates of the center point of the grid cell in the chassis coordinate system are obtained, and a three-dimensional grid center point is generated according to the ground height. The three-dimensional grid center point is projected onto the imaging plane of the visible light camera and the imaging plane of the thermal imaging camera respectively to obtain the pixel position. A fixed window is taken near the corresponding pixel position, with a window size of 5×5 in an example. The average brightness, average gradient magnitude, and texture contrast of the visible light image are calculated, and the average temperature, maximum temperature, and average temperature gradient of the thermal imaging image are written into the grid cell to form a grid image set. The above method of back-projecting the grid cell to the pixel position and then taking the window for statistics makes the image information correspond one-to-one with the grid cell, avoiding spatial mismatch caused by using the whole frame image as global features.
[0036] When generating the grid acoustic feature set, the direction of arrival (DOA) of the purified acoustic signal measurement values is first estimated. The DOA estimation method is as follows: the cross-correlation peak delay of each channel pair is calculated on the multi-channel waveform sampling sequence, and the azimuth angle of the sound source is determined by combining the array geometric number. Then, the azimuth angle of the sound source is converted into the horizontal ray direction using the chassis coordinate system of the inspection robot as a reference. The directional intersection point is then determined by combining the point cloud coordinate set: points with an angle of less than 5 degrees to the horizontal ray direction are selected from the point cloud coordinate set, and the point closest to the chassis is selected as the directional intersection point. The directional intersection point is projected onto a unified grid coordinate system and the grid row and column index is located. The fields of acoustic band energy, spectral peak frequency, and spectral flatness are written into the grid cell to form the grid acoustic feature set. When the directional intersection point selection set is empty, the acoustic feature is written into the unlocated mark and the modal stability index is reduced in the subsequent weight calculation.
[0037] S2.3. Perform feature calculation and stitching on the grid point cloud feature set, grid image set and grid acoustic feature set according to grid unit to obtain multimodal grid feature set.
[0038] In a preferred embodiment, when performing feature calculation and stitching by grid cell, for each grid cell, the following fields are read: number of points, nearest point distance, highest point height, and height variance from the grid point cloud feature set; mean brightness, mean gradient magnitude, mean temperature, and maximum temperature from the grid image set; and energy and spectral flatness within the acoustic band from the grid acoustic feature set. After performing dimension normalization on the above fields, they are stitched together in a fixed order to obtain the multimodal grid feature vector of that grid cell, which is mathematically expressed as follows: in, For the first Multimodal raster feature vectors of individual raster units; The point count normalization value is obtained by dividing the number of points in the grid cell by the maximum number of points in a single grid cell, where the maximum number of points in a single grid cell is 200. This is the normalized value of the nearest point distance, obtained by linearly scaling the nearest point distance from 0 meters to 20 meters. This is the normalized value of the highest point height, obtained by linearly scaling the highest point height from 0 meters to 3 meters. The height variance normalization value is obtained by scaling the height variance according to the offline statistical upper limit, which is 0.25 square meters. This is the normalized value of the average brightness, obtained by scaling the average brightness by 0 to 255. This is the normalized value of the gradient magnitude mean, obtained by scaling the gradient magnitude mean by the upper limit of offline statistics; This is the normalized value of the temperature mean, obtained by linearly scaling the temperature mean from 0 degrees Celsius to 120 degrees Celsius. This is the normalized value of the maximum temperature, obtained by linearly scaling the maximum temperature from 0 degrees Celsius to 120 degrees Celsius. This is the normalized value of the energy within the acoustic band, obtained by scaling the energy within the band according to the upper limit of offline statistics; This is the spectral flatness normalization value, which is directly written in according to the spectral flatness range of 0 to 1; Generate all raster cells sequentially The features are then grouped according to the raster row and column indices to obtain a multimodal raster feature set.
[0039] S2.4 Calculate the observation residual statistic and sampling stability index for each mode in the aligned observation set; where the observation residual statistic is the mean square difference of the raster feature subset of the mode in two adjacent sampling periods, and the sampling stability index is the signal-to-noise ratio of the measured values of the mode in two adjacent sampling periods; when generating unnormalized weights for each mode, the sampling stability index is used as the positive coefficient, the observation residual statistic is used as the negative coefficient, and a zero-prevention constant (such as 0.000001) is introduced to bias the observation residual statistic before participating in the negative weighting; normalize the unnormalized weights of each mode by summation to obtain the confidence weights corresponding to each mode.
[0040] Specifically, each mode includes: the visible light mode corresponding to the visible light image acquired by the visible light camera, the thermal imaging mode corresponding to the thermal imaging image acquired by the thermal imaging camera, the point cloud distance mode corresponding to the laser point cloud distance measurement value acquired by the lidar, and the acoustic mode corresponding to the acoustic signal measurement value acquired by the acoustic sensor. The confidence weights for each mode include visible light confidence weight, thermal imaging confidence weight, point cloud distance confidence weight, and acoustic confidence weight.
[0041] In a preferred embodiment, the difference mean square value is based on the difference sequence of the same modal raster feature subset within two adjacent sampling periods. For each raster cell, the difference between the feature value of the current sampling period and the feature value of the previous sampling period is calculated, and then the difference is squared and averaged over all raster cells to obtain the difference mean square value of the modality over the adjacent period pair. When some raster cells have missing values in a certain period, the raster cell is not included in the average and is synchronously written into the missing value count.
[0042] In a preferred embodiment, the signal-to-noise ratio is based on the power ratio of the measured values of the mode in two adjacent sampling periods. For each sampling period, the signal power and noise power are calculated within a fixed time window. The signal power is the sum of the spectral power within the passband, and the noise power is the sum of the spectral power outside the passband. The ratio of the two is then calculated and averaged for two adjacent sampling periods to obtain the sampling stability index of the mode. When an interpolation marker or an unlocated marker exists in a certain sampling period, the sampling stability index corresponding to that period is reduced by 0.7.
[0043] Preferably, in step S2, the four-modal observations are first time-aligned using the sampling timestamp, and then the observations from different coordinate systems are unified into a unified grid coordinate system using the sensor calibration extrinsic parameter matrix. This allows each grid cell to simultaneously carry texture, temperature, spatial geometry, and sound source information, thereby providing a clear spatial orientation for subsequent risk estimation. Simultaneously, in step S2.4, the difference mean square value between adjacent periods reflects the jitter degree of the grid feature of the modality, and the signal-to-noise ratio reflects the stability of the measurement value of the modality, thus forming a confidence weight that adaptively adjusts with scene changes. Compared with the fixed weight fusion or weighting based solely on sensor type in the prior art, in this embodiment, under disturbances such as rain, fog, backlight, strong heat sources, sparse echoes, and mechanical noise, the weights can be updated with changes in the difference mean square value and signal-to-noise ratio, reducing the impact of single-modal degradation on the risk area identification results.
[0044] S3. Input the multimodal raster feature set and confidence weights into the recursive risk estimation model constructed based on the extended Kalman filter algorithm, output the risk estimate and risk change rate, and generate an early warning raster set. Note that the following should be noted in this step: S3.1 Generate observation vectors for the multimodal raster feature set according to the raster cells, and perform weighted fusion of the observation vectors according to the confidence weights to obtain weighted observation vectors.
[0045] In a preferred embodiment, when generating observation vectors for the multimodal raster feature set by raster unit, the multimodal raster feature vector of each raster unit is used as the basis, and divided into visible photon vectors, thermal imaging vectors, point cloud distance vectors, and acoustic vectors according to the modal source; wherein, examples of visible photon vectors include normalized values of brightness mean and gradient magnitude mean; examples of thermal imaging vectors include normalized values of temperature mean and maximum temperature; examples of point cloud distance vectors include normalized values of nearest point distance and number of points; and examples of acoustic vectors include normalized values of energy within the acoustic band and spectral flatness.
[0046] Furthermore, when weighting and fusing the observation vectors according to the confidence weights, each modality sub-vector is multiplied by the corresponding confidence weight, and the four weighted sub-vectors are concatenated in a predetermined order to obtain the weighted observation vector of the grid cell. When a certain modality is missing in the grid cell, the modality sub-vector is set to zero and the confidence weight of the modality is reduced by 0.5, so that the weighted observation vector does not make a false contribution to the missing modality.
[0047] S3.2. Within the current sampling period, based on the risk estimate and risk change rate of the previous sampling period, the predicted risk estimate and predicted risk change rate of the current sampling period are recursively predicted to obtain the predicted state quantity.
[0048] It should be noted that the recursive prediction in this embodiment is executed independently on each grid cell. For a certain grid cell, the risk estimate and risk change rate of the previous sampling period have been stored in the previous sampling period. In the current sampling period, the preset sampling period is first read and converted to seconds, and the preset sampling period is 0.2 seconds. Then, the predicted risk estimate is generated by adding the risk estimate of the previous sampling period to the product of the risk change rate of the previous sampling period and the preset sampling period, and the predicted risk change rate is taken as the risk change rate of the previous sampling period or decayed by 0.98. This forms the predicted state quantity.
[0049] S3.3. Input the weighted observation vector and the predicted state quantity into the recursive risk estimation model constructed based on the extended Kalman filter algorithm to recursively update it, and obtain the risk estimate and risk change rate of the current sampling period.
[0050] As an example, the recursive risk estimation model in this embodiment adopts the form of extended Kalman filtering, defining a state vector and an observation vector for each grid cell, and giving expressions for the state transition relationship and the observation relationship: in, For the first The state vector for each sampling period; For the first The risk estimate for each sampling period, with a value range of 0 to 1 (example). For the first Risk change rate over each sampling period This is the state transition function; This is the process noise vector; For observation; For the observed noise vector; This is the value in seconds corresponding to the preset sampling period; for example, it is 0.2. This is the risk change rate decay coefficient, with an example of 0.98; This is the weighted observation vector; The first of the weighted observation vectors One component; The number of components in the weighted observation vector; For observation bias term; The observation mapping coefficients can be obtained through offline calibration. The offline calibration method involves collecting a raster sample set containing both risk and non-risk events, fitting the weighted observation vectors of the samples to manually labeled risk levels, and then obtaining the results. and ; It is a saturation function, and its output is limited to the range of 0 to 1.
[0051] For example, in the recursive update of the extended Kalman filter, the following prediction and update process is adopted: in, To predict the covariance matrix; Update the covariance matrix for the previous period; Let be the Jacobian matrix of the state transition function with respect to the state vector; Let be the process noise covariance matrix. In this example, we take a diagonal matrix with diagonal elements of 0.002 and 0.01, respectively. Kalman gain; The equivalent linearization matrix of the observation function on the state or observation mapping, in this example, can be obtained from the linearization slope of the saturation function and... Combined to obtain; To observe the noise covariance, an example value of 0.02 is used; To predict state variables; To update the state variables; For the reason The predicted observations obtained; It is the identity matrix; by In and Read out the risk estimate and the rate of change of risk respectively.
[0052] S3.4. Perform threshold comparison and sign consistency screening on the risk change rate of the current sampling period, select grid cells whose risk change rate is greater than the early warning change rate threshold (e.g., 0.15 per second) and whose increase direction is consistent with the risk estimate, and obtain the early warning grid set.
[0053] In a preferred embodiment, for each grid cell, the risk change rate of the current sampling period is first compared with the early warning change rate threshold, and grid cells with a risk change rate greater than the early warning change rate threshold are selected. Then, sign consistency screening is performed: the risk estimate of the previous sampling period is read and compared with the risk estimate of the current sampling period. When the risk estimate of the current sampling period is greater than the risk estimate of the previous sampling period and the risk change rate of the current sampling period is positive, the grid cell is included in the early warning grid set. When the risk estimate of the current sampling period is less than or equal to the risk estimate of the previous sampling period, it is not included in the early warning grid set, so as to reduce false triggering caused by short-term noise with a high change rate but no risk increase.
[0054] It should be noted that step S3 first performs weighted fusion of grid observations with confidence weights, so that the observation input is modulated by the stability of each mode; then, recursive prediction and recursive update form a continuous sequence of risk estimates and risk change rates, so that risk judgment is not limited to single-frame threshold triggering, but has two types of criteria: level magnitude and upward trend; on this basis, the early warning grid set is obtained by screening with the early warning change rate threshold and sign consistency, so that the early warning triggering is bound to the risk growth trend; compared with the existing technology that only performs static classification of fused features and outputs risk labels, this embodiment outputs the risk change rate and incorporates it into the early warning criteria, so it is more sensitive to early signs such as rapid temperature rise, rapid approach, and sudden abnormal noise, while the recursive structure is less sensitive to random noise.
[0055] S4. Perform consistency determination on the early warning grid set over M consecutive sampling periods, retain grid cells that meet the continuous consistency determination threshold, merge connected components, and output the risk area identification result and its corresponding early warning level. Note that the following should be noted in this step: S4.1. In each sampling period, the early warning grid set is written into the sliding count buffer, and the number of occurrences of each grid cell in M consecutive sampling periods (e.g., 10) is accumulated to obtain the consistency count value.
[0056] In a preferred embodiment, the sliding counting buffer is constructed as a circular queue with a queue length of 10 consecutive sampling periods. Each queue element of the sliding counting buffer stores a frame of binary raster image, which is generated from an early warning raster set: when a raster cell belongs to the early warning raster set, the raster cell is written with 1, otherwise it is written with 0; in each sampling period, the binary raster image of the current sampling period is enqueued; when the queue length exceeds 10, the earliest enqueued binary raster image is dequeued; for each raster cell, the binary value of the raster cell position is accumulated in the queue, and the accumulation result is written into the consistency count value map to obtain the consistency count value.
[0057] S4.2. Perform threshold comparison on the consistency count values, select grid cells whose consistency count values are greater than or equal to the continuous consistency judgment threshold (such as 7), and obtain a stable early warning grid set.
[0058] Specifically, the consistency count value is compared grid by grid: when the consistency count value of a grid cell is greater than or equal to 7, the grid cell is included in the stable early warning grid set; when the consistency count value is less than 7, it is not included in the stable early warning grid set.
[0059] It should be noted that the purpose of the threshold comparison mentioned above is to ensure that only grid cells that trigger early warnings in at least 7 out of the last 10 sampling periods are included in the stable early warning grid set, thereby distinguishing between occasional triggering and continuous triggering.
[0060] S4.3 Perform grid connectivity merging on the stable early warning grid set to obtain the risk area identification results, and generate early warning levels based on the hierarchical mapping relationship between consistency count values and risk estimates.
[0061] Specifically, the methods for generating early warning levels include: For each risk area in the risk area identification results, calculate the regional risk estimate and the regional consistency count. The regional risk estimate is the maximum of the risk estimates of all grid cells within the risk area, and the regional consistency count is the maximum of the consistency counts of all grid cells within the risk area. Then, compare the regional risk estimate and the regional consistency count with the regional risk threshold (e.g., 0.75) and the regional consistency count threshold (e.g., 8), respectively. When the estimated regional risk value is greater than or equal to the regional risk threshold and the regional consistency count value is greater than or equal to the regional consistency count threshold, the early warning level is high. When the estimated regional risk value is greater than or equal to the regional risk threshold and the regional consistency count value is less than the regional consistency count threshold, or when the estimated regional risk value is less than the regional risk threshold and the regional consistency count value is greater than or equal to the regional consistency count threshold, the early warning level is medium. When the estimated regional risk value is less than the regional risk threshold and the regional consistency count value is less than the regional consistency count threshold, the early warning level is low.
[0062] It should be noted that the merging of grid connected components in this embodiment is performed according to the adjacency rule, which is 8-adjacency. Specifically: the stable early warning grid set is scanned grid by grid. When a certain grid cell has not yet been assigned a region number and the grid cell belongs to the stable early warning grid set, the grid cell is used as the seed to perform breadth expansion. All grid cells connected to the seed under the 8-adjacency rule are grouped into the same region number to obtain a connected component. After scanning all grid cells, multiple connected components are obtained and are respectively used as risk areas in the risk area identification results.
[0063] For example, the risk area identification results include at least: the area number of each risk area, the set of grid cells contained in the area, the bounding rectangle of the area, the centroid coordinates of the area, the area area, and the area risk estimate and the area consistency count; wherein, the centroid coordinates of the area are obtained by weighted average of the coordinates of the center points of the grid cells, and the area area is obtained by multiplying the number of grid cells by the square of the grid resolution. For example, if the grid resolution is 0.2 meters, the area of a single grid cell is 0.04 square meters.
[0064] Offline collection of non-risk period data and known risk event data in the inspection area; generation of risk estimate sequence according to steps S1 to S3; and statistical analysis of the 95th percentile of the regional risk estimate for non-risk periods, using this percentile as a candidate value for the regional risk threshold; then statistical analysis of the regional consistency count distribution of the known risk event data, and selection of the minimum threshold that makes the false negative rate less than 2% as a candidate value for the regional consistency count threshold.
[0065] In a specific application scenario, an inspection robot is inspecting a cable trench area in a substation. The grid coverage is 100×100 pixels, the grid resolution is 0.2 meters, the preset sampling period is 200 milliseconds, the value for M consecutive sampling periods is 10, the continuity consistency judgment threshold is 7, the area risk threshold is 0.75, and the area consistency count threshold is 8. During the inspection, a cable joint is found to be locally overheated and accompanied by a slight discharge noise 6 meters in front of the robot. The thermal imaging mode shows a maximum temperature of 92 degrees Celsius in this area, the visible light mode shows abnormal texture contrast in this area, and the acoustic... The modal analysis in this region indicates an increase in in-band energy and a spectral peak frequency concentrated around 1.2 kHz, while the point cloud distance mode indicates no change in spatial structure. After weight calculation in step S2, the thermal imaging confidence weight example is 0.42, the acoustic confidence weight example is 0.28, the visible light confidence weight example is 0.18, and the point cloud distance confidence weight example is 0.12. After recursive update in step S3, the risk estimate of a certain grid cell in this region increases from 0.63 to 0.81, with a risk change rate of 0.22 per second, meeting the early warning change rate threshold and entering the early warning grid set.
[0066] After running for 10 more sampling cycles in the above scenario, the sliding count buffer statistics show that: the maximum consistency count value of several grid cells in risk area A is 9, and the maximum regional risk estimate is 0.82; the maximum consistency count value in risk area B is 8, and the maximum regional risk estimate is 0.71; the maximum consistency count value in risk area C is 6, and the maximum regional risk estimate is 0.78. According to the early warning level generation rules, in risk area A, the regional risk estimate is greater than or equal to 0.75 and the regional consistency count value is greater than or equal to 8, so the early warning level is judged as high; in risk area B, the regional risk estimate is less than 0.75 and the regional consistency count value is greater than or equal to 8, so the early warning level is judged as medium; in risk area C, the regional risk estimate is greater than or equal to 0.75 but the regional consistency count value is less than 8, so the early warning level is judged as medium.
[0067] For example, the output risk area identification results can be recorded as follows: the centroid coordinates of risk area A are (6.2 meters, 1.1 meters), and the area is 0.72 square meters; the centroid coordinates of risk area B are (4.8 meters, -0.6 meters), and the area is 0.40 square meters; the centroid coordinates of risk area C are (7.0 meters, -2.2 meters), and the area is 0.28 square meters, with each area accompanied by an early warning level field.
[0068] It should be noted that step S4 uses the consistency count value within M consecutive sampling periods to screen the early warning grid set for temporal consistency, making it difficult for instantaneous noise triggers to accumulate to the continuous consistency judgment threshold, thereby reducing false alarms; then, grid connected component merging aggregates discrete grids into interpretable spatial risk regions, so that the output result has region boundary, centroid and area information; compared with the existing technology of directly outputting scattered alarms after single-period threshold triggering, this embodiment introduces two constraints, consistency count and connected component merging, in the output layer, changing the alarm object from scattered points to regions, and jointly mapping it to the early warning level with the region risk estimate and the region consistency count value, thereby improving alarm stability.
[0069] Other aspects disclosed in the embodiments of the present invention also provide a computer device including one or more processors and a memory.
[0070] The memory is used to store operable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, including the flow of the risk area identification method for inspection robots based on multimodal fusion perception in the foregoing embodiments, especially... Figure 1 The flowchart of the method is shown.
[0071] Other aspects disclosed in the embodiments of the present invention also propose a computer-readable medium for storing software including instructions executable by one or more computers, which, upon execution, cause the one or more computers to perform operations including the flow of the inspection robot risk area identification method based on multimodal fusion perception of the foregoing embodiments, particularly... Figure 1 The flowchart of the method is shown.
[0072] It should be recognized that embodiments of the present invention may be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non-transitory computer-readable storage medium.
[0073] The method can be implemented using standard programming techniques, including a non-transitory computer-readable storage medium configured with a computer program in the computer program, wherein the storage medium is configured such that the computer operates in a specific and predefined manner.
[0074] Each program can be implemented in a high-level procedural or object-oriented programming language to communicate with the computer system; however, if required, the program can be implemented in assembly or machine language.
[0075] In any case, the language can be either compiled or interpreted.
[0076] Furthermore, for this purpose, the program can run on programmed application-specific integrated circuits.
[0077] The processes described herein (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that commonly executes on one or more processors. The computer program includes a plurality of instructions executable by one or more processors.
[0078] Furthermore, the method can be implemented in any suitable computing platform, including but not limited to personal computers, minicomputers, mainframes, workstations, networked or distributed computing environments, standalone or integrated computer platforms, or in communication with charged particle tools or other imaging devices.
[0079] Various aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether portable or integrated into a computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer, and when the storage medium or device is read by the computer, it can be used to configure and operate the computer to perform the processes described herein.
[0080] Furthermore, machine-readable code, or parts thereof, can be transmitted via wired or wireless networks.
[0081] When such media includes instructions or programs that combine with a microprocessor or other data processor to implement the steps described above, the invention described herein includes these and other different types of non-transitory computer-readable storage media.
[0082] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for identifying risk areas in an inspection robot based on multimodal fusion perception, characterized in that, include: The inspection robot collects visible light images, thermal images, laser point cloud distance measurements, and acoustic signal measurements of the working environment within the inspection area, and generates multimodal observation records. The multimodal observation records are subjected to temporal alignment and spatial registration to obtain a multimodal raster feature set, and the confidence weights corresponding to each modality are calculated. The multimodal raster feature set and the confidence weight are input into the recursive risk estimation model constructed based on the extended Kalman filter algorithm, and the risk estimate and risk change rate are output to generate an early warning raster set. The consistency determination is performed on the early warning grid set within M consecutive sampling periods. Grid cells that meet the continuous consistency determination threshold are retained and connected components are merged. The risk area identification result and its corresponding early warning level are output.
2. The method for risk area identification of inspection robots based on multimodal fusion perception according to claim 1, characterized in that, The method for generating multimodal observation records includes: At the start of each preset sampling period, the visible light camera, thermal imaging camera, lidar and acoustic sensor on the inspection robot are triggered to complete synchronous sampling at the same sampling timestamp, respectively obtaining visible light image, thermal imaging image, lidar point cloud distance measurement value and acoustic signal measurement value. Distortion correction and scale normalization are performed on the visible light image and the thermal imaging image, outlier removal is performed on the laser point cloud distance measurement value, bandpass filtering is performed on the acoustic signal measurement value, and the sampling timestamp and sensor identifier are written into each processing result to obtain the multimodal observation record.
3. The method for risk area identification of inspection robots based on multimodal fusion perception according to claim 2, characterized in that, The method for obtaining a multimodal raster feature set includes: Based on the sampling timestamp, time alignment is performed on the visible light image, the thermal imaging image, the laser point cloud distance measurement value, and the acoustic signal measurement value to obtain an aligned observation set; The laser point cloud distance measurements in the aligned observation set are converted into a point cloud coordinate set using a sensor calibration extrinsic parameter matrix, and the point cloud coordinate set is projected onto a unified grid coordinate system to obtain a grid point cloud feature set. Simultaneously, the visible light image and the thermal imaging image in the aligned observation set are mapped to the unified grid coordinate system based on the sensor calibration extrinsic parameter matrix to obtain a grid image set. Based on the direction of arrival estimation results of the acoustic signal measurements and the direction intersection of the point cloud coordinate set, the corresponding grid cell for the sound source is determined, and a grid acoustic feature set is generated. The raster point cloud feature set, the raster image set, and the raster acoustic feature set are processed and stitched together according to raster units to obtain the multimodal raster feature set.
4. The method for risk area identification of inspection robots based on multimodal fusion perception according to claim 3, characterized in that, The method for calculating the confidence weights corresponding to each mode includes: For the aligned observation set, the observation residual statistic and sampling stability index of each mode are calculated respectively; wherein, the observation residual statistic is the difference mean square value of the raster feature subset of the mode in two adjacent sampling periods, and the sampling stability index is the signal-to-noise ratio of the measured value of the mode in two adjacent sampling periods; when generating unnormalized weights for each mode, the sampling stability index is used as a positive coefficient, the observation residual statistic is used as a negative coefficient, and a zero-prevention constant is introduced to bias the observation residual statistic before participating in the negative weighting; the unnormalized weights of each mode are normalized by summation to obtain the confidence weights corresponding to each mode.
5. The method for risk area identification of inspection robots based on multimodal fusion perception according to claim 4, characterized in that, The modes include: the visible light mode corresponding to the visible light image acquired by the visible light camera, the thermal imaging mode corresponding to the thermal imaging image acquired by the thermal imaging camera, the point cloud distance mode corresponding to the laser point cloud distance measurement value acquired by the lidar, and the acoustic mode corresponding to the acoustic signal measurement value acquired by the acoustic sensor. The confidence weights corresponding to each mode include visible light confidence weight, thermal imaging confidence weight, point cloud distance confidence weight, and acoustic confidence weight.
6. The method for risk area identification of inspection robots based on multimodal fusion perception according to claim 1, characterized in that, The method for outputting risk estimates and risk change rates to generate an early warning grid set includes: The observation vectors are generated by grid cells for the multimodal grid feature set, and the observation vectors are weighted and fused according to the confidence weights to obtain weighted observation vectors. Within the current sampling period, the predicted risk estimate and predicted risk change rate of the current sampling period are recursively predicted based on the risk estimate and risk change rate of the previous sampling period to obtain the predicted state quantity. The weighted observation vector and the predicted state quantity are input into the recursive risk estimation model constructed based on the extended Kalman filter algorithm for recursive updating, so as to obtain the risk estimate and risk change rate of the current sampling period. The risk change rate of the current sampling period is compared with a threshold and screened for sign consistency. Grid cells with risk change rates greater than the early warning change rate threshold and in the same direction of increase as the risk estimate are selected to obtain the early warning grid set.
7. The method for risk area identification of inspection robots based on multimodal fusion perception according to claim 6, characterized in that, The method for generating the risk area identification results and their corresponding early warning levels includes: In each sampling period, the early warning grid set is written into the sliding count buffer, and the number of occurrences of each grid unit in M consecutive sampling periods is accumulated to obtain a consistency count value. The consistency count values are compared with a threshold, and grid cells with consistency count values greater than or equal to the continuous consistency determination threshold are selected to obtain a stable early warning grid set. The stable early warning grid set is subjected to grid connectivity merging to obtain risk area identification results, and the early warning level is generated according to the hierarchical mapping relationship between the consistency count value and the risk estimate value.
8. The method for risk area identification of inspection robots based on multimodal fusion perception according to claim 7, characterized in that, The method for generating the early warning level includes: For each risk area in the risk area identification results, a regional risk estimate and a regional consistency count are calculated; wherein the regional risk estimate is the maximum value of the risk estimates of each grid cell within the risk area, and the regional consistency count is the maximum value of the consistency counts of each grid cell within the risk area; and the regional risk estimate and the regional consistency count are compared with the regional risk threshold and the regional consistency count threshold, respectively. When the estimated regional risk value is greater than or equal to the regional risk threshold and the regional consistency count value is greater than or equal to the regional consistency count threshold, the early warning level is high. When the estimated regional risk value is greater than or equal to the regional risk threshold and the regional consistency count value is less than the regional consistency count threshold, or when the estimated regional risk value is less than the regional risk threshold and the regional consistency count value is greater than or equal to the regional consistency count threshold, the early warning level is medium. When the estimated regional risk value is less than the regional risk threshold and the regional consistency count value is less than the regional consistency count threshold, the early warning level is low.
9. A computer device, characterized in that, include: One or more processors; The memory stores operable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations, including the flow of the inspection robot risk area identification method based on multimodal fusion perception as described in any one of claims 1 to 8.
10. A computer-readable medium for storing software, characterized in that: The software includes instructions executable by one or more computers, which cause the one or more computers to perform operations, including the process of the inspection robot risk area identification method based on multimodal fusion perception as described in any one of claims 1 to 8.