Multi-modal fusion perception robot dog inspection slope disaster risk assessment method and device and storage medium

Through the robot dog patrol method of multimodal fusion perception, multimodal sensors are used to obtain data and build a three-dimensional voxel structure to identify disaster types on complex rock and soil slopes, solving the problem of inaccurate risk assessment in the existing technology, and achieving efficient and accurate risk assessment.

CN120494535APending Publication Date: 2025-08-15CHANGAN UNIV +1

Patent Information

Application Number
CN202510991084.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-18
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing technology is difficult to fully consider the geological disaster risk factors of complex geotechnical slopes, resulting in insufficient accuracy and reliability of risk assessment.

Method used

The robot dog patrol method of multimodal fusion perception is adopted to obtain underground media reflected signals, surface three-dimensional point clouds, infrared heat maps, multi-spectral images and slope surface RGB images by configuring multimodal sensors, construct three-dimensional voxel structures and perform data fusion, identify disaster types and calculate risk indexes.

Benefits of technology

It improves the accuracy and reliability of geological disaster risk assessment, can quickly and efficiently identify multiple potential risk factors and calculate a reasonable risk index.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494535A_ABST
    Figure CN120494535A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal fusion sensing robot dog inspection slope disaster risk assessment method and device and a storage medium, and relates to the field of slope disaster assessment, and the method comprises the steps: obtaining multi-modal sensing data; performing space-time alignment processing on the multi-modal sensing data and then converting the multi-modal sensing data into a voxel coordinate system; performing feature extraction on the multi-modal sensing data after time-space alignment, and mapping each extracted feature to a unified voxel unit to form a three-dimensional voxel structure containing each extracted feature; constructing a three-dimensional multi-modal data fusion model; identifying disaster types based on the three-dimensional multi-modal data fusion model, wherein the disaster types comprise the ground surface crack length, the underground cavity volume, the water seepage point number, the local collapse and bulging area, the vegetation degradation area and the slope gradient; and calculating a risk index based on the identified disaster type in combination with the association degree of the disaster type. By adopting the evaluation method provided by the invention, rapid and efficient evaluation of slope disasters can be realized, and the method has relatively good accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of slope disaster risk assessment, and in particular to a method, device and storage medium for assessing slope disaster risk using a robot dog patrol with multimodal fusion perception. Background Art

[0002] A slope is a slope of soil or rock with a certain height and gradient, formed by natural or human factors. The safety and stability of a slope are directly related to the safety and economic efficiency of a project. Therefore, geological hazard risk assessment of slopes is of great significance in engineering practice.

[0003] Traditional slope hazard risk assessments typically use conventional sensors, such as inclination sensors, strain gauges, and displacement sensors, to monitor slope displacement, stress, tilt, and other changes in real time for risk assessment. However, slope hazard risk assessments vary widely, and the primary hazard-causing factors vary across different geological types. Current slope hazard risk assessment methods fail to fully consider potential risk factors, making it difficult to ensure the accuracy and reliability of hazard risk assessments for complex geological structures.

[0004] Therefore, a method, device and storage medium for assessing slope disaster risks by a robot dog inspection with multimodal fusion perception are needed to at least partially solve the above technical problems. Summary of the Invention

[0005] In view of this, an embodiment of the present invention provides a method, device and storage medium for assessing slope disaster risks by a robot dog inspection with multimodal fusion perception, so as to solve at least one of the problems in the prior art.

[0006] In a first aspect, an embodiment of the present invention provides a method for assessing slope hazard risk using a robot dog patrol using multimodal fusion perception. The method comprises: A robot dog equipped with a multimodal sensor inspects the target slope and obtains multimodal sensing data, including underground medium reflection signals, surface 3D point clouds, infrared thermal images, multispectral images, and slope surface RGB images. The target slope is divided into regular grid cells of a set size, the terrain voxels and voxel coordinate system are divided, and the multimodal sensor data are uniformly converted to the voxel coordinate system after spatiotemporal alignment processing; Perform feature extraction on the spatiotemporally aligned multimodal sensor data, and map the corresponding extracted features to the same voxel unit to form a three-dimensional voxel structure containing the extracted features; The 3D voxel structure is vectorized along the 3D coordinate axis and fused direction by direction to build a 3D multimodal data fusion model; Identify disaster types based on a 3D multimodal data fusion model, including surface crack length, underground cavity volume, number of water seepage points, local collapse and heaving area, vegetation degradation area, and slope gradient; The risk index is calculated based on the identified disaster types and the degree of correlation between the disaster types; the larger the risk index value, the higher the slope disaster risk.

[0007] In a second aspect, an embodiment of the present invention further provides a multimodal fusion perception robot dog inspection slope disaster risk assessment device, the assessment device comprising: a memory for storing computer-executable instructions; The processor is used to implement the evaluation method of the above technical solution when executing the computer executable instructions stored in the memory.

[0008] In a third aspect, an embodiment of the present invention further provides a storage medium storing computer instructions, wherein the computer instructions are used to enable the computer to execute the evaluation method of the above technical solution.

[0009] According to the assessment method of the present invention, the underground medium reflection signal, the surface three-dimensional point cloud, the infrared thermal map, the multispectral image and the slope surface RGB image are obtained by multimodal sensors. After feature extraction and spatiotemporal alignment processing, a three-dimensional voxel structure and a three-dimensional multimodal data fusion model are constructed to identify the types of disasters including the length of surface cracks, the volume of underground cavities, the number of seepage points, the area of local collapse and bulging, the area of vegetation degradation and the slope gradient. Finally, the risk index is calculated. This assessment method takes into account multiple potential risk factors in various dimensions, thereby improving the accuracy and reliability of geological hazard risk assessment.

[0010] Additional advantages, objects, and features of the present invention will be set forth in part in the following description and will become apparent to those skilled in the art upon examination of the following or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by the structures particularly pointed out in the description and drawings.

[0011] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other purposes that can be achieved by the present invention will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are intended to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention. The components in the drawings are not drawn to scale, but are merely for the purpose of illustrating the principles of the present invention. To facilitate the illustration and description of certain portions of the present invention, corresponding portions in the drawings may be exaggerated, that is, may be larger than other components in an exemplary device actually manufactured according to the present invention. In the drawings: Figure 1 is a flow chart of an evaluation method according to an embodiment of the present invention; Figure 2 is a schematic diagram of an evaluation device according to an embodiment of the present invention; Figure 3 FIG. 1 is a schematic diagram of an evaluation system according to an embodiment of the present invention. DETAILED DESCRIPTION

[0013] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.

[0014] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.

[0015] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.

[0016] It should also be noted that, unless otherwise specified, the term "connection" herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.

[0017] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0018] First, refer to Figure 1 The following describes a method 100 for assessing slope hazard risk using a robot dog patrol using multimodal fusion perception according to an embodiment of the present application. Figure 1 As shown, the evaluation method 100 may include steps S110 to S160, which are specifically as follows: In step S110, a robot dog equipped with a multimodal sensor inspects the target slope and obtains multimodal sensing data, including underground medium reflection signals, surface three-dimensional point clouds, infrared thermal maps, multispectral images, and slope surface RGB images.

[0019] In step S120 , the target slope is divided into regular grid cells of a set size, terrain voxels and voxel coordinate systems are divided, and the multimodal sensing data are subjected to spatiotemporal alignment processing and then uniformly converted to the voxel coordinate system.

[0020] In step S130 , feature extraction is performed on the spatiotemporally synchronized multimodal sensing data, and each corresponding extracted feature is mapped to the same voxel unit to form a three-dimensional voxel structure containing each extracted feature.

[0021] In step S140 , the three-dimensional voxel structure is vectorized along the three-dimensional coordinate axis direction and fused direction by direction to construct a three-dimensional multimodal data fusion model.

[0022] In step S150, the disaster type is identified based on the three-dimensional multimodal data fusion model, including the length of surface cracks, the volume of underground cavities, the number of seepage points, the area of local collapse and heaving, the area of vegetation degradation, and the slope gradient.

[0023] In step S160, a risk index is calculated based on the identified disaster type and in combination with the degree of association between the disaster types; a larger value of the risk index indicates a higher slope disaster risk.

[0024] In an embodiment of the present application, first, a robot dog equipped with a multimodal sensor inspects a target slope to obtain multimodal sensing data, including underground medium reflection signals, surface three-dimensional point clouds, infrared thermal maps, multispectral images, and slope surface RGB images; the target slope is divided into regular grid units of a set size, and terrain voxels and voxel coordinate systems are divided. The multimodal sensing data are subjected to spatiotemporal alignment and then uniformly converted to the voxel coordinate system; then, feature extraction is performed on the spatiotemporally synchronized multimodal sensing data, and the corresponding extracted features are mapped to the same A three-dimensional voxel structure containing each extracted feature is formed into a voxel unit; the three-dimensional voxel structure is vectorized along the three-dimensional coordinate axis and fused direction by direction to construct a three-dimensional multimodal data fusion model; then, based on the three-dimensional multimodal data fusion model, the disaster type is identified, including the length of surface cracks, the volume of underground cavities, the number of seepage points, the area of local collapse and bulging, the area of vegetation degradation, and the slope gradient; finally, based on the identified disaster type and the degree of correlation between the disaster types, the risk index is calculated; the larger the risk index value, the higher the slope disaster risk.

[0025] From the description of the above process, it can be seen that according to the evaluation method 100 of the embodiment of the present application, the underground medium reflection signal, the surface three-dimensional point cloud, the infrared thermal map, the multispectral image and the slope surface RGB image are obtained through the multimodal sensor, and then after feature extraction and spatiotemporal alignment processing, a three-dimensional voxel structure and a three-dimensional multimodal data fusion model are constructed to identify potential risk factors in various dimensions, including the length of surface cracks, the volume of underground cavities, the number of seepage points, the type of disaster including local collapse and bulging area, the area of vegetation degradation area and the slope gradient, and finally the risk index is calculated, thereby improving the accuracy and reliability of geological disaster risk assessment.

[0026] The contents of the above steps of the evaluation method 100 according to the embodiment of the present application will be described in detail below.

[0027] In an embodiment of the present application, in step S110, a robot dog equipped with a multimodal sensor inspects the target slope to obtain multimodal sensing data, including underground medium reflection signals, surface three-dimensional point clouds, infrared thermal maps, multispectral images and slope surface RGB images.

[0028] Specifically, to acquire multimodal sensing data, including subsurface reflection signals, 3D surface point clouds, infrared thermal maps, multispectral images, and RGB images of the slope surface, a robot dog is equipped with multimodal sensors. These sensors may include a geological radar (GPR) mounted on the robot's base, a 3D laser radar (LiDAR) mounted on its top, an infrared thermal imager and a multispectral camera mounted on its front gimbal, and an RGB camera mounted on its front end. During inspections of target slopes, the robot dog uses geological radar (GPR) to perform non-contact reflection imaging of the slope's shallow subsurface structure, acquiring subsurface reflection signals that can be used to identify structural anomalies such as slip surfaces, cavities, and faults. 3D laser radar (LiDAR) collects 3D surface point cloud data. Infrared thermal images are generated by the infrared thermal imager and can be used to monitor slope temperature distribution and water seepage areas, identifying thermal anomalies. Multispectral images are acquired by the multispectral camera to monitor vegetation health and surface weathering, assisting in identifying potential risk areas. High-definition RGB images of the slope surface are collected by an RGB camera for crack identification and image enhancement.

[0029] The robot dog boasts high terrain adaptability and load capacity, enabling stable movement, obstacle traversal, and autonomous navigation on complex slopes. It can adopt both mammal-like and wheeled locomotion modes. The robot dog can automatically switch between a quadrupedal bionic gait and wheeled rolling mode to meet mobility requirements in diverse terrain conditions.

[0030] In an embodiment of the present application, in step S120, the target slope is divided into regular grid units of a set size, terrain voxels and voxel coordinate systems are divided, and the multimodal sensor data are uniformly converted to the voxel coordinate system after time-space alignment processing.

[0031] Specifically, for example, the target slope space can be divided into regular grid units of 5cm×5cm×5cm, and the terrain voxels and voxel coordinate systems (voxels) are divided. The coordinates of each multimodal sensor are mapped to the voxel coordinate system (voxel) through the global coordinate system (world). Each voxel is indexed by its spatial Indicated by number.

[0032] Since each multimodal sensor data is collected based on its own local coordinate system, the multimodal sensor data is spatially misaligned and requires spatiotemporal alignment processing.

[0033] Existing technologies can be used to perform spatiotemporal alignment of multimodal sensor data. For example, a posture and positioning module can be included in the robot dog to obtain posture and latitude and longitude coordinate information. This posture and positioning module includes an IMU (inertial navigation unit) and RTK module.

[0034] The spatiotemporal alignment of multimodal sensor data may refer to: A timestamp synchronization mechanism is used to unify the acquisition frequency of multimodal sensors and achieve time synchronization of multimodal sensor data. For example, internal hardware is configured to synchronize time using a GPS-disciplined rubidium atomic clock. A GPS-disciplined rubidium atomic clock is a high-precision time synchronization device that combines the high-precision time reference of GPS satellite signals with the short-term stability of a rubidium atomic clock. GPS provides long-term accurate time, while the rubidium atomic clock maintains high-precision time output even when GPS signals are unavailable.

[0035] All sensor clocks are synchronized to a common time source: a GPS-disciplined rubidium atomic clock. Each sensor applies a unified timestamp to data collected, ensuring consistent temporal reference for data from different devices. Data collected by different sensors is precisely timestamped. For example, LiDAR generates 10 data packets per second, each with a nanosecond-accurate timestamp; GPS generates one data packet per second, also with a precise timestamp. These timestamps are based on the same clock source, ensuring precise alignment.

[0036] At the same time, based on the attitude and longitude and latitude coordinate information obtained by the attitude and positioning module, coordinate unification is performed to achieve spatial alignment and correction.

[0037] In an embodiment of the present application, in step S130 , feature extraction is performed on the spatiotemporally synchronized multimodal sensing data, and each corresponding extracted feature is mapped to the same voxel unit to form a three-dimensional voxel structure including each extracted feature.

[0038] Specifically, the corresponding extracted features include radar reflection coefficient, dielectric constant, point cloud geometric curvature, infrared temperature value, multispectral vegetation index, soil moisture and RGB image texture.

[0039] The 3D surface point cloud can be filtered to remove noise and compress the data. Poisson reconstruction is then used to generate a mesh. The point cloud coordinates are mapped to a 3D voxel grid, and the geometric curvature of each point cloud is calculated. The specific method involves selecting a set number of points within the neighborhood of each point cloud, calculating the spatial distribution covariance matrix of these points, and solving for their eigenvalues to estimate the geometric curvature characteristics of the point cloud.

[0040] Extract the near-infrared and red bands from the multispectral image and calculate the multispectral vegetation index based on the vegetation index formula , and combined with infrared temperature values , using the empirical regression model to calculate soil moisture : in, and They represent the near-infrared band and the red light band respectively. is the reference temperature of the drying area. 、 It is the empirical coefficient fitted according to the current soil type, which can be calibrated on site or obtained from literature. The optimal unbiased interpolation method based on spatial autocorrelation is used to combine soil moisture with The data were converted into a spatially continuous field and mapped to a three-dimensional voxel grid based on image projection to construct a multispectral voxel structure with vegetation and humidity attributes.

[0041] Collect the wave velocity of the reflected signal from the underground medium and obtain the radar reflection coefficient based on the collected wave velocity information and dielectric constant : in, is the maximum amplitude of the reflected wave from the target interface, is the incident wave amplitude. is the speed of light. is the two-way propagation time of the signal reflected in the underground medium. For depth.

[0042] Then, the GPR one-dimensional reflection profile is converted into a spatially distributed data type (position × depth). The specific method is to form a trajectory coordinate sequence through RTK recording of the GPR survey line position, and convert the calculated reflection coefficient and dielectric constant The position of the trajectory is mapped. The time domain is converted to depth based on the calibrated propagation velocity of the electromagnetic wave in the rock being measured. Three-dimensional data is generated along the trajectory based on the horizontal position and depth data, generating voxel structure data with GPR attributes.

[0043] Among them, the infrared temperature value The RGB image textures are directly obtained from the infrared thermal map and the slope surface RGB image, respectively. Based on global spatiotemporal alignment, the infrared thermal map features are projected onto a 3D voxel grid. The infrared temperature values are mapped to the voxel structure coordinates, generating voxel structure data with temperature properties. Regarding the RGB image texture, the RGB texture image is projected onto the Mesh surface via UV expansion. Mask R-CNN is used to segment the image regions and map them to the corresponding Mesh triangles. For each marked Mesh triangle face, the corresponding region is labeled with an RGB value. Based on its spatial coordinates and distribution range, it is mapped to the voxel structure, generating voxel structure data with RGB image texture properties.

[0044] Finally, the 3D voxel structure is defined as: Where, is an RGB image texture. That is, each 3D voxel structure generated by fusion includes radar reflectivity, dielectric constant, point cloud geometric curvature, infrared temperature value, multispectral vegetation index, soil moisture and RGB image texture.

[0045] In an embodiment of the present application, in step S140 , the three-dimensional voxel structure is vectorized along the three-dimensional coordinate axis direction and fused direction by direction to construct a three-dimensional multimodal data fusion model.

[0046] Specifically, after mapping the five types of sensor data into a three-dimensional voxel structure, deep fusion is performed based on direction vector decomposition. Taking voxels as units, each modal feature is mapped along the three-dimensional coordinate axis. 、 、 The directions are vectorized and fused direction by direction to finally form a unified spatial fusion feature representation.

[0047] Specifically, for each voxel, the feature data collected by the five types of sensors (i.e., LiDAR, RGB camera, infrared thermal imager, multispectral camera, GPR) are first decomposed into corresponding three-dimensional coordinate axes 、 、 Vector components in three directions. LiDAR geometric curvature and normal, RGB image texture, infrared temperature value, multispectral vegetation index The radar reflection coefficient and dielectric constant of soil moisture, GPR, and other data are converted into three-dimensional directional vectors through spatial mapping, thereby constructing the multimodal directional feature distribution of voxels in three-dimensional space. Each voxel obtains the directional distribution characteristics of the five types of data in three-dimensional space.

[0048] Subsequently, a deep fusion network can be used Respectively 、 、 The multi-modal information in each direction is deeply fused using multi-layer nonlinear mapping and residual structure.

[0049] The fusion process in each direction includes feature dimensionality upgrade, that is, projecting the input vector to a higher dimension through one or more layers of full connection / convolution, so that the network can capture richer cross-modal feature interactions to enhance the expressive ability, including mixed interactions between modal features, that is, in high-dimensional space, different modal components interact and share information; and dimensionality reduction and compression combined with residuals, that is, compressing the results back to the appropriate dimension, while retaining the original information through residual connections to prevent feature loss caused by excessive compression, and ensure the complete transmission and efficient coupling of directional features.

[0050] Finally, the constructed three-dimensional multimodal data fusion model , defined as: in, Indicates the Voxel in direction The direction vector on . 、 、 、 and Represent the corresponding extracted features of multimodal sensor data in the Voxel, direction The characteristic components on , for example, 、 、 、 and The corresponding extracted features of the multimodal sensor data of LiDAR, RGB camera, infrared thermal imager, multispectral camera, and GPR can be represented in sequence. Voxel, direction The characteristic components on . Indicates that through deep fusion network right Comprehensive feature vector after deep fusion.

[0051] In an embodiment of the present application, in step S150, the disaster type is identified based on a three-dimensional multimodal data fusion model, including the length of surface cracks, the volume of underground cavities, the number of seepage points, the area of local collapse and bulging, the area of vegetation degradation, and the slope gradient.

[0052] Specifically, based on the 3D multimodal data fusion model, the fusion feature data Based on the current data element Its spatially adjacent data elements If the color difference gradient exceeds the preset threshold, it is considered that there is a potential image crack and the data is marked as a crack risk. At the same time, the local geometric curvature of the data is detected. ,like If the curvature threshold is exceeded, it is marked as a curvature mutation area, and further marked as a potential geometric crack. The potential crack areas obtained by color gradient calculation and geometric curvature detection are projected separately, and the spatial overlap of the two in the data space is calculated. When the spatial overlap exceeds the set threshold, it is confirmed as a crack, and the single crack length is output based on the three-dimensional point cloud data of the crack and the RGB fusion map information. , to obtain the length of all surface cracks .

[0053] The cavity is mostly filled with air or loose materials, and the dielectric constant It is obviously lower than the surrounding rock and soil. At the same time, the interface between the cavity and the rock and soil will cause the radar reflection coefficient Based on the above characteristics, the underground cavity characteristics are initially screened. A convolutional neural network (such as 1D-CNN+BiGPU) is used to extract the reflection wave characteristics (amplitude, frequency, waveform similarity) of the underground medium reflection signal, and to segment and identify the hyperbolic diffraction wave as a cavity sign. The abnormal area is projected along the scanning trajectory using the path trajectory and depth parameters to obtain the underground cavity volume. .

[0054] Based on the multispectral vegetation index, the area where the multispectral vegetation index is less than the set value is output according to the data coordinate information , the cumulative area of vegetation degradation area .

[0055] When extracting the features of the ground 3D point cloud, select the target point cloud and a set number of point clouds in its neighborhood, and calculate the geometric curvature of the target point cloud. and normal vector . Define The slope of a point cloud is the angle between the normal vector and the normal vector of the horizontal plane, that is, Calculate the average slope value of all point clouds in the detection area to obtain the slope of the regional slope .

[0056] Based on the 3D multimodal data fusion model, the fusion feature data Based on this, the temperature gradient, humidity gradient and humidity gradient direction between each data element and its adjacent data elements are calculated to obtain the potential water seepage area. The potential water seepage area is projected onto the horizontal surface to obtain a set of two-dimensional water seepage projection point sets. The connectivity analysis of the water seepage projection point sets is performed to divide the points adjacent or continuously distributed in space into several water seepage blocks, and the cumulative number of water seepage points is obtained. .

[0057] Based on the 3D multimodal data fusion model, the normal vector is extracted And calculate the normal angle difference between each data element and its adjacent data element in the continuous area of the slope .like And height difference , then the initial screening mark is an abnormal terrain area. After obtaining the potential abnormal terrain area, a secondary screening is performed to further improve the recognition accuracy and reduce false detection. The potential abnormal terrain area is determined by analyzing the shadow enhancement, strong reflective area, local color discontinuity of the RGB image data and calculating the temperature gradient in the adjacent infrared data and the background area that meets the set difference requirements. According to the coordinate information in the fusion data of each abnormal terrain area, the area of a single abnormal terrain area is output, and the local collapse and bulge area is accumulated. .

[0058] In an embodiment of the present application, in step S160, a risk index is calculated based on the identified disaster type and the degree of correlation between the disaster types, wherein a larger value of the risk index indicates a higher slope disaster risk.

[0059] Specifically, the risk index is calculated , specifically: in, . The risk parameters, numbered from small to large, are the length of surface cracks, volume of underground cavities, number of water seepage points, area of local collapse and bulge, area of vegetation degradation, and slope gradient after normalization, as shown in Table 1. is the risk parameter item weight. and They are the associated enhancement items, and are the correlation coefficients respectively.

[0060] Table 1 Furthermore, the risk index can be divided into risk levels and response strategies can be formulated, see Table 2.

[0061] Table 2 Based on the above description, according to the evaluation method of the embodiment of the present application, by obtaining five types of multimodal sensor data and then constructing a three-dimensional voxel structure and a three-dimensional multimodal data fusion model, potential risk factors in various dimensions are identified, and finally the risk index is calculated, thereby improving the accuracy and reliability of geological hazard risk assessment.

[0062] refer to Figure 2 The embodiment of the present application further provides an evaluation device 200 for implementing the evaluation method 100 according to the embodiment of the present application. The evaluation device 200 includes a processor 210 and a memory 220. The evaluation device 200 may include one or more processors 210 and one or more memories 220. The memory 220 stores an executable program executed by the processor 210. When the executable program is executed by the processor 210, the processor 210 executes the evaluation method 100 according to the embodiment of the present application described above.

[0063] The processor 210 may be a central processing unit (CPU) or other processing units having data processing capabilities and / or instruction execution capabilities.

[0064] The memory 220 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), a hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 210 may execute the program instructions to implement the client functions (implemented by the processor) and / or other desired functions described in the embodiments of the present application described herein. Various applications and various data, such as various data used and / or generated by the applications, may also be stored in the computer-readable storage medium.

[0065] The evaluation device 200 may also include an input device and an output device, and these components are interconnected via a bus system and / or other forms of connection mechanisms. Figure 2 The components and structure of the evaluation device 200 shown are merely exemplary and non-limiting. The evaluation device 200 may also have other components and structures as needed.

[0066] The input device may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, etc. In addition, the input device may also be any interface for receiving information.

[0067] The output device may output various information (eg, images or sounds) to the outside (eg, a user), and may include one or more of a display, a speaker, etc. In addition, the output device may also be any other device with an output function.

[0068] Illustratively, the example evaluation device 200 for implementing the evaluation method 100 according to an embodiment of the present application can be applied to terminal devices (such as mobile phones), tablet computers, laptop computers, ultra-mobile personal computers (UMPCs), handheld computers, netbooks, personal digital assistants (PDAs), wearable devices (such as smart watches, smart glasses, or smart helmets), augmented reality (AR), virtual reality (VR) devices, smart home devices, car computers, and other electronic devices. The embodiments of the present application do not impose any restrictions on this.

[0069] Those skilled in the art can understand the specific operations of the evaluation device 200 for implementing the evaluation method 100 according to the embodiment of the present application in combination with the contents described above. For the sake of brevity, the specific details are not repeated here, and only some main operations of the processor 210 are described.

[0070] In one embodiment of the present application, when the executable program is executed by the processor 210, the processor 210 performs the following steps: A robot dog equipped with multimodal sensors inspects a target slope and acquires multimodal sensor data, including underground medium reflection signals, surface 3D point clouds, infrared thermal maps, multispectral images, and RGB images of the slope surface. The target slope is divided into regular grid cells of a set size, and terrain voxels and voxel coordinate systems are divided. The multimodal sensor data is then spatially and temporally aligned and uniformly converted to a voxel coordinate system. Feature extraction is performed on the spatiotemporally synchronized multimodal sensor data, and each extracted feature is mapped to the same voxel unit to form a 3D voxel structure containing all the extracted features. The 3D voxel structure is vectorized along the 3D coordinate axis and fused direction by direction to construct a 3D multimodal data fusion model. Based on the 3D multimodal data fusion model, the disaster type is identified, including the length of surface cracks, the volume of underground cavities, the number of seepage points, the area of local collapse and heave, the area of vegetation degradation, and the slope gradient. A risk index is calculated based on the identified disaster type and the degree of correlation between the disaster types. A larger risk index indicates a higher slope disaster risk.

[0071] The above exemplary shows the evaluation method 100 according to the embodiment of the present application. Figure 3 An evaluation system 300 provided in another aspect of an embodiment of the present application is described.

[0072] Reference Figure 3 The following describes an example evaluation system 300 for implementing the evaluation method of an embodiment of the present application. The evaluation system 300 may include a data acquisition module 310, a conversion module 320, a 3D voxel module 330, a fusion model construction module 340, a recognition module 350, and a calculation module 360. The data acquisition module 310 is used to obtain multimodal sensing data from the target slope inspection by a robot dog equipped with a multimodal sensor, including underground medium reflection signals, surface three-dimensional point clouds, infrared thermal maps, multispectral images and slope surface RGB images.

[0073] The conversion module 320 is used to divide the target slope into regular grid units of a set size, divide the terrain voxels and the voxel coordinate system, perform spatiotemporal alignment processing on the multimodal sensor data, and uniformly convert them into the voxel coordinate system.

[0074] The three-dimensional voxel module 330 is used to perform feature extraction on the spatiotemporally synchronized multimodal sensing data, and map the corresponding extracted features to the same voxel unit to form a three-dimensional voxel structure containing the extracted features.

[0075] The fusion model construction module 340 is used to vectorize the three-dimensional voxel structure along the three-dimensional coordinate axis direction and fuse them direction by direction to construct a three-dimensional multimodal data fusion model.

[0076] The identification module 350 is used to identify the disaster type based on the three-dimensional multimodal data fusion model, including the length of surface cracks, the volume of underground cavities, the number of seepage points, the area of local collapse and heaving, the area of vegetation degradation and the slope gradient.

[0077] The calculation module 360 is used to calculate the risk index based on the identified disaster type and the degree of correlation between the disaster types; wherein the larger the risk index value, the higher the slope disaster risk.

[0078] The evaluation system 300 proposed in the embodiment of the present invention can realize rapid and efficient evaluation of slope hazards with good accuracy.

[0079] In addition, according to an embodiment of the present application, the present application further provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, it is used to perform the corresponding steps of the evaluation method 100 of the embodiment of the present application. The storage medium may include, for example, a memory card of a smartphone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.

[0080] In addition, according to an embodiment of the present application, the present application also provides a computer program product, including computer instructions, which implement the steps of the evaluation method 100 of the embodiment of the present application when executed by a processor.

[0081] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely illustrative and are not intended to limit the scope of the present application. Various changes and modifications may be made therein by those skilled in the art without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as required by the appended claims.

[0082] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0083] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units described is merely a logical function division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another device, or ignoring or not performing some features.

[0084] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.

[0085] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbols placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.

[0086] The above description is merely a specific embodiment or illustration of a specific embodiment of the present application, and the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. The scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A multimodal fusion perception robot dog inspection slope disaster risk assessment method, characterized by: The evaluation method includes: A robot dog equipped with a multimodal sensor inspects the target slope and obtains multimodal sensing data, including underground medium reflection signals, surface 3D point clouds, infrared thermal images, multispectral images, and slope surface RGB images. The target slope is divided into regular grid cells of a set size, the terrain voxels and voxel coordinate system are divided, and the multimodal sensor data are uniformly converted to the voxel coordinate system after spatiotemporal alignment processing; Perform feature extraction on the spatiotemporally aligned multimodal sensor data, and map the corresponding extracted features to the same voxel unit to form a three-dimensional voxel structure containing the extracted features; The 3D voxel structure is vectorized along the 3D coordinate axis and fused direction by direction to build a 3D multimodal data fusion model; Identify disaster types based on a 3D multimodal data fusion model, including surface crack length, underground cavity volume, number of water seepage points, local collapse and heaving area, vegetation degradation area, and slope gradient; The risk index is calculated based on the identified disaster types and the degree of correlation between the disaster types; the larger the risk index value, the higher the slope disaster risk.

2. The evaluation method according to claim 1, wherein: The feature extraction of the spatiotemporally synchronized multimodal sensing data specifically includes: Poisson reconstruction is used to generate a wireless grid for the three-dimensional point cloud of the ground surface. The point cloud coordinates are mapped to the three-dimensional voxel grid, and the geometric curvature of each point cloud is calculated. Extract the near-infrared and red bands from the multispectral image and calculate the multispectral vegetation index based on the vegetation index formula , and combined with infrared temperature values , using the empirical regression model to calculate soil moisture , in, and represent the near-infrared band and the red light band respectively; is the reference temperature of the drying area; 、 is the empirical coefficient fitted according to the current soil type; Collect the wave velocity of the reflected signal from the underground medium and obtain the radar reflection coefficient based on the collected wave velocity information and dielectric constant , in, is the maximum amplitude of the reflected wave from the target interface, is the incident wave amplitude; is the speed of light; is the two-way propagation time of the reflected signal in the underground medium, for depth; Among them, the infrared temperature value and RGB image textures are directly obtained from infrared thermal images and slope surface RGB images, respectively.

3. The evaluation method according to claim 1, wherein: Constructed three-dimensional multimodal data fusion model , defined as: in, 、 、 Represent the three-dimensional coordinate axes, Indicates the Voxel in direction The direction vector on 、 、 、 and Represent the corresponding extracted features of multimodal sensor data in the Voxel, direction The characteristic components of Indicates that through deep fusion network right Comprehensive feature vector after deep fusion.

4. The evaluation method according to claim 2, wherein: The identification of disaster types based on the three-dimensional multimodal data fusion model specifically includes: Based on a 3D multimodal data fusion model, the boundary of potential crack areas obtained through color gradient calculation and geometric curvature detection is projected separately, and the spatial overlap of the two in the data space is calculated. When the spatial overlap exceeds a set threshold, it is confirmed as a crack and the length of the crack is calculated to obtain the length of all surface cracks. A convolutional neural network is used to extract the reflection wave characteristics of the underground medium reflection signal, and to segment and identify the hyperbolic diffraction waves that serve as the hallmark of the cavity. The path trajectory and depth parameters are used to project the abnormal area along the scanning trajectory to obtain the underground cavity volume. Based on the multispectral vegetation index, the area of the region where the multispectral vegetation index is less than the set value is output according to the data coordinate information, and the area of the vegetation degradation area is accumulated; When extracting the features of the three-dimensional point cloud of the ground surface, the target point cloud and a set number of point clouds in its neighborhood are selected, the geometric curvature and normal vector of the target point cloud are calculated, the slope is defined as the angle between the normal vector and the normal vector of the horizontal plane, and the slope values of all point clouds in the detection area are averaged to obtain the slope of the regional slope; Based on a three-dimensional multimodal data fusion model, the temperature gradient, humidity gradient, and humidity gradient direction between each data element and its adjacent data elements are calculated to obtain potential water seepage areas. The potential water seepage areas are projected onto the horizontal surface to obtain a set of two-dimensional water seepage projection points. Connectivity analysis is performed on this set of water seepage projection points, and spatially adjacent or continuously distributed points are divided into several water seepage blocks, and the number of water seepage points is accumulated. Based on the three-dimensional multimodal data fusion model, the normal vector is extracted and the normal angle difference between each data element and its adjacent data element in the continuous area of the slope is calculated to obtain potential abnormal terrain areas. The potential abnormal terrain areas are determined by analyzing the shadow enhancement, strong reflective areas, local color discontinuity of the RGB image data, and calculating the temperature gradient in the adjacent infrared data and the background area that meets the set difference requirements. The area of each abnormal terrain area is output based on the coordinate information in the fused data, and the local collapse and bulge areas are accumulated.

5. The evaluation method according to claim 1, wherein: The calculated risk index , specifically: in, , The risk parameters are the length of surface cracks, volume of underground cavities, number of water seepage points, area of local collapse and bulge, area of vegetation degradation, and slope gradient after normalization. is the risk parameter item weight, and They are the associated enhancement items, and are the correlation coefficients respectively.

6. The evaluation method according to claim 2, wherein: The corresponding extracted features include radar reflection coefficient, dielectric constant, point cloud geometric curvature, infrared temperature value, multispectral vegetation index, soil moisture and RGB image texture; Among them, the three-dimensional voxel structure is defined as: Where, 、 、 are the spatial index of each voxel respectively; It is an RGB image texture.

7. The evaluation method according to claim 1, wherein: The robot dog is equipped with a multimodal sensor, wherein the multimodal sensor includes a geological radar set at the bottom of the robot dog, a three-dimensional laser radar set at the top of the robot dog, an infrared thermal imager and a multispectral camera set at the front gimbal of the robot dog, and an RGB camera set at the front end of the robot dog.

8. The evaluation method according to claim 7, characterized in that It also includes a posture and positioning module set to the robot dog, for obtaining posture and longitude and latitude coordinate information; the posture and positioning module includes an IMU and RTK module; The spatiotemporal alignment of multimodal sensor data specifically refers to: Adopting the timestamp synchronization mechanism, unifying the acquisition frequency of multi-modal sensors, and achieving time synchronization of multi-modal sensor data; Coordinate unification is performed based on the attitude and longitude and latitude coordinate information obtained by the attitude and positioning module.

9. A multi-modal fusion perception robot dog inspection slope disaster risk assessment device, characterized in that: The evaluation device comprises: a memory for storing computer-executable instructions; A processor, configured to implement the evaluation method according to any one of claims 1 to 8 when executing the computer-executable instructions stored in the memory.

10. A storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the evaluation method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Whole set of technical equipment for dike danger patrol based on bionic robot dog

    CN114997359A

  • Highway high slope landslide disaster early warning method based on evaluation function training

    CN117058845A

  • Urban water affair inspection device based on bionic robot dog

    CN117901132A

  • Slope inspection robot autonomous exploration three-dimensional point cloud mapping algorithm, storage medium and slope inspection robot

    CN120244945A

Cited By

  • Road slope modeling method based on unmanned aerial vehicle inspection route

    CN120931853A

  • Multi-source data fusion-based northern spring corn unmanned aerial vehicle image waterlogging extraction method

    CN120974146A

  • Vehicle-mounted and mobile platform ground disaster multi-mode sensing, alarming and risk avoiding method and system

    CN121148105A

  • Automatic inspection sensing method and system based on three-dimensional point cloud data

    CN121305051A

  • Slope intelligent inspection early warning method and system based on time-space-multi-mode fusion

    CN121330858A