Multi-modal fusion perception and semantic generation method for underground dynamic degradation environment

By building a multimodal perception platform in a dynamic degraded environment in underground space and applying joint calibration and deep learning models, the problems of multimodal perception applications are solved and noise interference are achieved, and efficient and accurate environmental monitoring and perception effects are achieved.

CN120014392APending Publication Date: 2025-05-16TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411872479.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing technology has limited multimodal perception applications in dynamic degraded environments in underground space, and its computing power exceeds the edge computing processing capabilities of unmanned devices. The environment complexity causes a single modal perception algorithm to fail to complete the perception task, and dynamic dust generates a lot of noise that affects the positioning accuracy.

Method used

The multimodal fusion perception and semantic generation method is adopted, and the multimodal perception platform is built, and the multimodal data features are extracted using deep learning algorithms, and the semantic information is generated by combining perception results and timestamp information, and it is released in real time through the ROS communication mechanism.

Benefits of technology

It realizes reliable acquisition and processing of multimodal data in dynamic degraded underground environments, reduces noise interference, improves the timeliness and effectiveness of environmental monitoring, and provides accurate perception and three-dimensional positioning capabilities that are not affected by dynamic lighting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014392A_ABST
    Figure CN120014392A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal fusion perception and semantic generation method for an underground dynamic degradation environment, and relates to the technical field of computer vision, and the method comprises the steps: constructing a multi-modal perception platform; carrying out joint calibration on the multi-modal sensing platform; carrying out multi-modal data acquisition on a to-be-detected target through the multi-modal sensing platform subjected to joint calibration; extracting data features of the multi-modal data through a multi-modal target sensing model based on a deep learning algorithm; determining a sensing result of the to-be-detected target according to the data features of the multi-modal data; combining the sensing result of the to-be-detected target with the timestamp information to generate semantic information; through an ROS communication mechanism, semantic information is published in real time, and real-time visualization of a perception result is realized. According to the method, reliable perception capability can still be provided in a resource-limited and communication-limited environment, so that the perception precision and semantic comprehension capability in an underground space environment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a multimodal fusion perception and semantic generation method for an underground dynamically degraded environment. Background Art

[0002] Multimodal fusion perception is a technology that integrates multiple different types of information and data sources, aiming to improve the ability to understand and recognize the environment or objects. Semantic generation refers to the process of automatically creating descriptive information or knowledge based on perception results and contextual information. The generated semantic information usually includes the target's category, status, and relationship with the environment.

[0003] In the dynamically degraded environment of underground space, traditional sensing technology has encountered multiple challenges. Accurate positioning can monitor the changes in underground space in real time, promptly discover potential safety hazards (such as landslides, water seepage, etc.), and issue early warnings to ensure the safety of workers. Therefore, accurate positioning of underground space is of great significance for ensuring safety, achieving effective navigation, enhancing environmental understanding, improving system robustness and promoting technological progress.

[0004] However, the current multimodal perception research is limited in its application in the dynamic degradation environment of underground space, mainly because most multimodal perception algorithms have too high requirements on computing power, which exceeds the processing capacity of unmanned equipment in edge computing scenarios, and the complexity of the environment makes it impossible for single-modal perception algorithms to complete perception tasks. At the same time, these algorithms have large bandwidth requirements for data sharing. When the communication link is blocked, it is difficult to achieve real-time detection and visualization, resulting in the loss of sensor data features, which in turn increases the difficulty of positioning and navigation. In addition, under the influence of dynamic dust, visible light, night vision and laser LiDAR sensors will generate a lot of noise, resulting in the point cloud information provided by the LiDAR being insufficient to provide accurate posture constraints, thus affecting the accuracy of positioning. Summary of the invention

[0005] In order to solve the technical problems that the application of multimodal perception research in the existing technology is limited in the dynamic degradation environment of underground space, resulting in the inability of single-modal perception algorithms to complete perception tasks, making it difficult to achieve real-time detection and visualization, resulting in the loss of sensor data features, and thus increasing the difficulty of positioning and navigation, and under the influence of dynamic dust, visible light, night vision and laser LiDAR sensors will generate a large amount of noise, resulting in the point cloud information provided by the laser radar being insufficient to provide accurate posture constraints, thereby affecting the accuracy of positioning, the present invention provides a multimodal fusion perception and semantic generation method for underground dynamic degradation environment.

[0006] The technical solution provided by the embodiment of the present invention is as follows:

[0007] First aspect

[0008] An embodiment of the present invention provides a multimodal fusion perception and semantic generation method for an underground dynamic degradation environment, comprising:

[0009] S1: Build a multimodal perception platform;

[0010] S2: Joint calibration of multimodal perception platforms;

[0011] S3: Collect multimodal data of the target to be detected through the multimodal perception platform after joint calibration;

[0012] S4: Extract data features of multimodal data through a multimodal target perception model based on deep learning algorithm;

[0013] S5: Determine the perception result of the target to be detected according to the data characteristics of the multimodal data;

[0014] S6: Combining the perception result of the target to be detected with the timestamp information to generate semantic information;

[0015] S7: Through the ROS communication mechanism, semantic information is published in real time to achieve instant visualization of perception results.

[0016] Second aspect

[0017] An embodiment of the present invention provides a multimodal fusion perception and semantic generation system for an underground dynamic degradation environment, comprising:

[0018] processor;

[0019] A memory having computer-readable instructions stored thereon, which, when executed by a processor, implements the multimodal fusion perception and semantic generation method for an underground dynamically degraded environment as described in the first aspect.

[0020] The third aspect

[0021] An embodiment of the present invention provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the multimodal fusion perception and semantic generation method for an underground dynamic degradation environment as described in the first aspect is implemented.

[0022] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0023] In the present invention, a multimodal perception platform is constructed by combining multiple sensors to create a comprehensive multimodal perception platform, and then the multimodal perception platform is jointly calibrated to make multimodal data acquisition more reliable. The data of multimodal data is extracted by a multimodal target perception model based on deep learning, thereby reducing noise interference. Finally, the perception results and timestamp information are combined to generate rich semantic information, and the semantic information is released in real time through the ROS communication mechanism, further improving the timeliness and effectiveness of environmental monitoring. This method takes advantage of the multimodal sensor and builds a comprehensive perception system through the fusion of thermal infrared, night vision, visible light and laser LiDAR sensors. The system can provide accurate perception without dynamic illumination, and through the multimodal target perception model constructed by the deep learning algorithm, reduce the limitation of dust on the perception task, realize the three-dimensional positioning of the target, obtain the category information, and suppress the influence of dust on the laser LiDAR sensor, and generate semantic information according to the timestamp information. In addition, the system also publishes the perception results in real time through the ROS communication mechanism, realizes the real-time transmission and processing of data, improves the accuracy of positioning and navigation, provides reliable perception capabilities and realizes real-time detection and visualization. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0025] Figure 1 A flowchart of a multimodal fusion perception and semantic generation method for an underground dynamic degradation environment provided by an embodiment of the present invention;

[0026] Figure 2 A schematic diagram of a framework of a multimodal fusion perception and semantic generation method for an underground dynamic degradation environment provided by an embodiment of the present invention;

[0027] Figure 3 A schematic diagram of a framework of a multimodal target perception model provided by an embodiment of the present invention;

[0028] Figure 4 A schematic diagram of the structure of a multimodal fusion perception and semantic generation system for an underground dynamically degraded environment provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The technical solution of the present invention is described below in conjunction with the accompanying drawings.

[0030] In the embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "example" in the present invention should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the word "example" is intended to present the concept in a specific way. In addition, in the embodiments of the present invention, the meaning expressed by "and / or" can be both, or it can be either of the two.

[0031] In order to make the technical problems, technical solutions and advantages to be solved by the present invention more clear, a detailed description will be given below with reference to the accompanying drawings and specific embodiments.

[0032] Reference Manual Attached Figure 1 , showing a flow chart of a multimodal fusion perception and semantic generation method for an underground dynamically degraded environment provided by an embodiment of the present invention.

[0033] The embodiment of the present invention provides a method for multimodal fusion perception and semantic generation of an underground dynamically degraded environment. The method can be implemented by a multimodal fusion perception and semantic generation device for an underground dynamically degraded environment. The multimodal fusion perception and semantic generation device for an underground dynamically degraded environment can be a terminal or a server. The processing flow of the method for multimodal fusion perception and semantic generation of an underground dynamically degraded environment can include the following steps:

[0034] S1: Build a multimodal perception platform.

[0035] Among them, a multimodal perception platform refers to a system that integrates multiple sensors and data processing technologies. It aims to simultaneously obtain information from different modalities, such as visible light, infrared, lidar, etc., to enhance environmental understanding and perception capabilities. By building a multimodal perception platform that integrates multiple sensors and uses night vision, visible light, thermal infrared and laser LiDAR technology, it can effectively respond to multiple challenges in the complex environment of underground space.

[0036] In a possible implementation, S1 specifically includes:

[0037] S101: Combining the characteristics of night vision sensors and visible light sensors, an adaptive light switching sensor based on light-sensitive device triggering is constructed.

[0038] Among them, night vision sensors can work in low light conditions and generate images by amplifying weak visible light or capturing thermal radiation using infrared technology. They are usually used in military, security, monitoring and night driving scenarios. Visible light sensors refer to sensors that can detect and image visible light bands (usually 380nm to 750nm). They are widely used in ordinary cameras, smart phones and various visual systems to provide high-resolution and true color images. Photosensors are used to sense changes in the intensity of ambient light and can automatically adjust the working status of other sensors according to lighting conditions. For example, when the light is strong, the night vision function is turned off and imaging is performed using visible light sensors. Adaptive light switching sensors refer to the system's function of automatically selecting the most suitable sensor for data collection according to the current ambient light conditions, aiming to optimize the perception effect and extend the service life of the equipment.

[0039] S102: Introduce a thermal infrared sensor into the adaptive light switching sensor to distinguish whether the target is a living organism.

[0040] Among them, thermal infrared sensor refers to a sensor that can detect thermal radiation (infrared radiation) emitted by objects. It is usually used to identify temperature changes and is widely used in security, medical imaging and night monitoring. Thermal infrared sensors can work in a completely dark environment to help identify heat sources, such as living organisms or machine equipment.

[0041] S103: Introduce a laser LiDAR sensor into the adaptive light switching sensor to obtain three-dimensional information of the target.

[0042] Among them, LiDAR (Light Detection and Ranging) sensor refers to the use of laser beams for distance measurement, creating high-precision three-dimensional point cloud data by emitting lasers and receiving reflected light. LiDAR is widely used in autonomous driving, geographic mapping and environmental modeling, and can provide detailed spatial information. Target three-dimensional information refers to the three-dimensional spatial data of the target object collected by the sensor, including its shape, size, position and relative motion. This information is very important for understanding the geometric characteristics of the target and conducting further analysis (such as path planning and collision detection).

[0043] S104: Build a multimodal perception platform based on adaptive light switching sensors, thermal infrared sensors and laser LiDAR sensors.

[0044] It should be noted that the combination of the characteristics of night vision sensors, visible light sensors, thermal infrared sensors and laser LiDAR sensors enables the system to work efficiently in complex underground environments. The adaptive light switching sensor can automatically adjust the working mode according to the real-time lighting conditions to ensure effective data collection under different lighting conditions. The introduction of thermal infrared sensors can not only identify living organisms in the dark, but also enhance the detection ability of potential threats. The three-dimensional information acquisition capability of laser LiDAR enables the system to accurately construct a three-dimensional model of the underground space and provide a detailed understanding of the environment. The combination of this series of technologies has greatly improved the robustness, accuracy and intelligence of the perception system in dynamic and complex environments, and can effectively support application scenarios such as security monitoring, emergency response and automated navigation.

[0045] In a possible implementation, the adaptive illumination switching sensor specifically includes: a visible light camera, an infrared transmitter, and a photosensor.

[0046] When the ambient light intensity detected by the light sensor is greater than or equal to 10 lux, the infrared transmitter is automatically turned off and the visible light data in the scene is captured using the visible light camera.

[0047] When the ambient light intensity detected by the light-sensitive sensor is less than 10 lux, the infrared transmitter is automatically turned on, and the infrared transmitter and the visible light camera are used to capture night vision data in the scene.

[0048] It should be noted that the automatic adjustment function of the photosensor improves the working efficiency and data collection quality of the adaptive light switching sensor under different lighting conditions. This intelligent light switching mechanism not only improves the adaptability and response speed of the sensor, but also effectively extends the service life of the equipment and improves the perception ability in complex and dynamic environments.

[0049] Reference Manual Attached Figure 2 , showing a framework schematic diagram of a multimodal fusion perception and semantic generation method for an underground dynamic degradation environment provided by an embodiment of the present invention.

[0050] Figure 2In the paper, a multimodal perception system is constructed. The multimodal perception system includes a thermal infrared sensor, a laser LiDAR sensor, and an adaptive light switching sensor composed of a visible light sensor, an infrared light emitter and a photosensor. The multimodal perception system provides thermal infrared images, point cloud data, and visible light images or night vision images. The multimodal target perception model receives the data provided by the adaptive light switching sensor and the thermal infrared sensor as input data, outputs the category information, positioning information and detection timestamp information of the data target, and extracts the corresponding depth information in the depth map by mapping. The above information can be used to generate semantic information, and the semantic information is shared through ROS communication, and finally a dynamically updated visual representation in three-dimensional space is performed in rviz.

[0051] S2: Joint calibration of multimodal perception platforms.

[0052] It should be noted that by using the calibration tool, the system can effectively convert the data coordinates of the laser LiDAR sensor into the visible light coordinate system, thereby achieving seamless data fusion. This precise calibration not only improves the measurement accuracy of the sensor, but also reduces the data inconsistency caused by sensor errors and enhances the system's ability to understand the environment.

[0053] In a possible implementation manner, S2 is specifically:

[0054] Use the Calibration tools to obtain the intrinsic and extrinsic parameters of the sensor in the multimodal perception platform and the rotation matrix that converts the coordinates of the laser LiDAR sensor into the visible light coordinate system.

[0055] Among them, Calibration tools are software and hardware tools used to calibrate and adjust sensors, cameras or other measuring devices. The main purpose of these tools is to ensure that the output data of the device is accurate and reliable, thereby improving the measurement accuracy. The intrinsic parameters of the sensor refer to the inherent characteristic parameters of the sensor itself, such as focal length, aperture and resolution of the image sensor, etc. These parameters affect the quality and accuracy of sensor imaging. The extrinsic parameters of the sensor refer to the position and orientation of the sensor in the coordinate system, including the rotation and displacement information of the sensor relative to other sensors or objects, which is crucial for realizing the alignment and fusion of multi-sensor data. The visible light coordinate system refers to the coordinate system based on the image generated by the visible light sensor, which is used for image processing and feature extraction.

[0056] It should be noted that calibration tools are an important part of ensuring the performance of multimodal perception systems. Through accurate parameter estimation and correction, the sensor data quality and system reliability are improved.

[0057] S3: Collect multimodal data of the target to be detected through the jointly calibrated multimodal perception platform.

[0058] Among them, the multimodal perception platform after joint calibration refers to a calibrated perception platform, which is composed of multiple sensors and can effectively collect and fuse data from different modalities in a unified coordinate system. The target to be detected refers to the object that needs to be identified, located or analyzed during the perception process, which can be an object, a person or any entity that needs to be monitored.

[0059] In a possible implementation, the multimodal data includes: thermal infrared images, visible light or night vision images, and point cloud data provided by laser LiDAR.

[0060] Reference Manual Attached Figure 3 , showing a schematic diagram of the framework of a multimodal target perception model provided by an embodiment of the present invention.

[0061] Figure 3 middle, I vis represents the visible light image, I ir represents the thermal infrared image, f νis ' and f ir 'represent the convolution layer operations for extracting features from visible light images and thermal infrared images, Concat represents the concatenation operation, and F fusion represents the fusion feature map, F mid represents the intermediate layer feature map, F multi-scale represents a multi-scale feature map, FPN represents a feature pyramid network, FC represents a fully connected network layer, and P c represents the predicted category probability of the target to be detected, B represents the predicted bounding box coordinates of the target to be detected, and softmax represents the softmax activation function.

[0062] Visible light images and thermal infrared images are input, and convolution operations are performed on the features of the input visible light images and thermal infrared images respectively. The extracted feature maps are passed through several convolution layers to extract deeper features, and the features of the visible light image and the thermal infrared image are fused and connected. The normalized visible light feature map and the thermal infrared feature map are fused into a fused feature map through feature splicing operations. The fused feature map is then processed through a feature pyramid network to extract features of different scales and obtain a multi-scale feature map. Based on the multi-scale feature map, the final classification and positioning of the multi-scale feature map is achieved through the fully connected layer.

[0063] It should be noted that the multimodal target perception model uses the convolutional neural network in deep learning to build a multimodal target perception model for the needs of feature fusion and target detection in the multimodal perception system, and realizes the efficient fusion and target detection of the multimodal perception system data; the network is especially targeted at the dynamic degradation environment of underground space, and can integrate the data of adaptive light switching sensors and thermal infrared sensors, and use multi-scale feature fusion strategies to enhance the feature representation and generalization capabilities of the model, while reducing the noise data generated by dynamic dust. This innovative method can accurately predict the target location information, providing an efficient and reliable technical solution for the safety monitoring, navigation and automated operation of underground space.

[0064] S4: Extract data features of multimodal data through a multimodal target perception model based on deep learning algorithm.

[0065] Among them, the deep learning algorithm is a machine learning technology based on artificial neural networks. It automatically extracts and learns data features through a multi-layer network structure and is widely used in image recognition, speech recognition and other fields. The multimodal target perception model refers to a deep learning model that integrates multiple sensor data (such as images, point clouds, etc.). It aims to achieve target recognition and understanding by analyzing these multimodal data. Data features refer to key information or characteristics extracted from the collected data. These features can help the model understand the content and structure of the data and are the basis for subsequent processing and analysis.

[0066] It should be noted that data feature extraction through a multimodal target perception model based on a deep learning algorithm has significant advantages. This model can automatically process multimodal data from different sensors and make full use of the characteristics of each sensor to improve the accuracy and efficiency of target recognition. The deep learning algorithm can extract high-dimensional features from complex data and capture information that is difficult to obtain through traditional methods, making the system more adaptable when dealing with complex scenarios. In addition, by utilizing the fusion of multimodal data, the model can maintain good performance under different environmental conditions and enhance the ability to perceive the target.

[0067] In a possible implementation, the multimodal data includes: thermal infrared images, visible light or night vision images, and point cloud data provided by laser LiDAR.

[0068] In a possible implementation, S4 specifically includes:

[0069] S401: Uniform sampling of point cloud data:

[0070] V=[v x ,v y ,v z ]

[0071]

[0072] Where V represents the voxel size of the target to be detected, (v x ,v y ,v z ) represents the size of the voxel of the target to be detected in each dimension, p i,j,k represents the representative point of all points in the voxel of the target to be detected, n represents the total number of voxels in the target to be detected, x represents the target to be detected, S i,j,k Represents the point set within the i, j, k voxel of the target to be detected.

[0073] S402: Convert the representative point coordinate system to the camera coordinate system:

[0074] P cam =R·P pc +T

[0075] Among them, P cam represents the representative point camera coordinate system, R represents the rotation matrix, T represents the translation vector, P pc Represents the coordinate system of the representative point cloud.

[0076] S403: Convert the representative point camera coordinate system into a pixel coordinate system to obtain a depth map of the target to be detected:

[0077]

[0078] Depth(u,v)=Z cam

[0079] Among them, (X cam ,Y cam ,Z cam ) represents the coordinate information of the representative point camera coordinate system, (x norm ,y norm ) represents the coordinate information of the representative point camera coordinate system projected onto the normalized image plane, K represents the parameter matrix within the camera, Depth(u,v)=Z cam Indicates that the depth value corresponding to each pixel (u, v) is Z cam , Depth(u,v) represents the depth map of the target to be detected.

[0080] It should be noted that the point cloud data provided by the laser LiDAR sensor is reduced in data volume through uniform sampling, and 10 percent of the original point cloud is retained for subsequent algorithm calculations, thereby reducing the noise generated by dust in the point cloud data and the complexity of the original point cloud data.

[0081] In a possible implementation, S4 specifically includes:

[0082] S404: Extract features of thermal infrared image and visible light or night vision image through the same convolution operation respectively:

[0083] F vis =Conv k,s (I vis ),F ir =Conv k,s (I ir )

[0084] Among them, I vis Indicates visible light or night vision images, I ir Represents thermal infrared image, Conv k,s represents a convolution operation, and the kernel size of the convolution layer is k and the step length is s, F vis Indicates visible light or night vision feature map, F ir Represents thermal infrared signature.

[0085] S405: Normalize and ReLU activate the visible light or night vision feature map and the thermal infrared feature map respectively:

[0086] F ν ' is =BN[ReLU(F νis )],F ir ′=BN[ReLU(F ir )]

[0087] Among them, BN represents normalization processing, ReLU represents ReLU activation function, and F ν ' is represents the normalized visible light or night vision feature map, F ir ′ represents the normalized thermal infrared feature map.

[0088] S406: Fusion of normalized visible light or night vision feature map and normalized thermal infrared feature map:

[0089] F fusion =Concat(F' vis ,F' ir )

[0090] Among them, Concat represents the concatenation operation, F fusion Represents the fused feature map.

[0091] S407: Extracting the intermediate layer feature information of the fused feature map:

[0092] F mid =Conv k,s (F fusion )

[0093] Among them, Fmid Represents the intermediate layer feature map.

[0094] S408: Fusion of different scale features of the intermediate layer feature map through feature pyramid network:

[0095] F multi-scale =FPN(F mid )

[0096] Among them, F multi-scale represents a multi-scale feature map, and FPN represents a feature pyramid network.

[0097] S409: Reduce the dimension of the multi-scale feature map through convolution operation, and use the fully connected layer to predict the bounding box category and position of the multi-scale feature map:

[0098] P c =softmax[FC(F multi-scale )]

[0099] B=FC(F multi-scale )

[0100]

[0101] L=(1-λ)L c +λL b

[0102] Among them, FC represents the fully connected network layer, P c represents the predicted category probability of the target to be detected, B represents the predicted bounding box coordinates of the target to be detected, and L c Represents the category loss of the target to be detected, y i Represents the true category label of the target to be detected, represents the predicted probability of the i-th target among the targets to be detected, ∑ represents the summation symbol, log represents the natural logarithm function, L b represents the bounding box loss of the target to be detected, smooth represents the smoothing function, B i represents the predicted bounding box coordinates of the i-th target in the target to be detected, b i represents the true bounding box coordinates of the i-th target in the target to be detected, L represents the total loss function, and λ represents the learnable parameter.

[0103] S4010: Mapping the predicted bounding box coordinates and predicted category of the target to be detected into the corresponding depth map, and extracting the depth information of the center point of the bounding box in the corresponding depth map to obtain the positioning information of the target to be detected.

[0104] It should be noted that through effective feature extraction, normalization, feature fusion and multi-scale feature integration, efficient processing of multimodal data is achieved, and the accuracy and robustness of target detection are significantly improved. This series of advantages provides strong support for practical applications such as autonomous driving, monitoring systems and robot navigation, and can achieve more efficient and accurate perception capabilities in complex environments.

[0105] S5: Determine the perception result of the target to be detected according to the data features of the multimodal data.

[0106] Among them, the perception result refers to the identification and positioning information about the target to be detected obtained after data feature extraction and analysis, which is usually used to indicate the category, status or other important information of the target.

[0107] It should be noted that determining the perception results of the target to be detected based on the data characteristics of multimodal data has significant advantages. This process utilizes the information provided by different sensors and can comprehensively analyze the multi-dimensional characteristics of the target to ensure the accuracy and comprehensiveness of recognition. In complex environments, the fusion of multimodal data can effectively reduce the errors caused by the limitations of a single sensor, allowing the system to maintain stable performance in dynamic and complex situations. In addition, accurate perception results can provide a reliable basis for subsequent decision-making and support real-time monitoring and emergency response.

[0108] S6: Combine the perception result of the target to be detected with the timestamp information to generate semantic information.

[0109] Among them, perception results refer to information about the target to be detected based on multimodal data analysis, usually including key information such as the target category, status and location. Timestamp information refers to the mark that records the specific time of data collection or event occurrence, which is usually used to synchronize and track data changes at different time points, and provide a time reference for analysis. Semantic information refers to information with specific meaning, which is usually generated through data processing and analysis, and can help the system understand the state of the environment or target.

[0110] It should be noted that by combining the perception results of the target to be detected with the timestamp information to generate semantic information, the system's ability to understand and respond to the environment is significantly enhanced. This process can convert static perception data into dynamic semantic information with time series characteristics, providing the context for target identification and positioning. This combination not only improves the richness and accuracy of information, but also enables the system to track changes in the target in real time and adapt to the needs of the dynamic environment.

[0111] In a possible implementation, S6 specifically includes:

[0112] S601: Construct a semantic information framework according to the predicted category of the target to be detected, the positioning information of the target to be detected and the corresponding timestamp information.

[0113] Among them, the predicted category refers to the classification result of the target to be detected by the system based on previous data analysis and model inference. The positioning information refers to the position data of the target to be detected in three-dimensional space, including the coordinates of the target, bounding box and other information. The positioning information helps the system understand the actual position of the target in the environment and its relationship with other objects. The semantic information framework is a structured information organization method that contains data such as the category, positioning information and timestamp of the target. The semantic information framework combines these elements together to facilitate subsequent analysis and decision-making.

[0114] It should be noted that the semantic information framework is constructed by combining the predicted category, location information and timestamp information of the target to be detected, aiming to provide more comprehensive contextual information for subsequent data analysis and enhance the system's ability to understand and process the environment.

[0115] S602: Based on the semantic information framework, the perception result of the target to be detected is combined with the timestamp information to generate semantic information.

[0116] Among them, semantic information refers to structured data with specific meaning and context generated by combining the perception results and timestamp information of the target to be detected. Based on the semantic information, a comprehensive, accurate and dynamic understanding of the target to be detected can be achieved, providing a solid foundation for subsequent intelligent decision-making and automated response.

[0117] S7: Through the ROS communication mechanism, semantic information is published in real time to achieve instant visualization of perception results.

[0118] Among them, ROS (Robot Operating System) is an open source robot operating system that provides many tools and libraries for robot development and supports the development and integration of multiple robot functions.

[0119] It should be noted that the real-time release of semantic information using the rostopic mechanism of ROS has significant advantages. This process can ensure that the system updates and transmits key information in real time during runtime, enhancing its ability to respond to environmental changes.

[0120] In a possible implementation, S7 specifically includes:

[0121] S701: The semantic information generated based on the semantic information framework is published in real time through the rostopic topic of the robot operating system ROS.

[0122] Among them, rostopic refers to the topic mechanism used for message passing in ROS, which allows different nodes to exchange information by publishing and subscribing to specific topics, and is the basis of ROS communication.

[0123] S702: Subscribe to the corresponding rostopic topic through rviz and receive semantic information.

[0124] Among them, rviz is a visualization tool of ROS, which can display the status, environmental model and data of the robot and its sensors in real time, helping developers monitor and debug the robot system.

[0125] S703: Convert the positioning data in the received semantic information into a three-dimensional label in the point cloud coordinate system through rviz, and mark the predicted category of the target to be detected on the corresponding three-dimensional label.

[0126] Among them, three-dimensional markers refer to visual identifiers used to represent specific targets in three-dimensional space, and are usually used to help users identify and track targets.

[0127] S704: Extract and display the point cloud data corresponding to the predicted time according to the timestamp information, so as to realize a dynamically updated visual representation of the target to be detected.

[0128] It should be noted that through the rviz tool, developers can visualize the published semantic information in real time, monitor the dynamic state and position of the target to be detected, and thus improve the operability and debugging efficiency of the system. This real-time publishing and visualization capability enables the system to adapt quickly in complex and dynamic environments, helping to enhance the ability to identify potential threats and respond to emergencies. In addition, dynamically updated semantic information can provide a more accurate basis for subsequent decision-making, thereby optimizing the overall performance of the system and improving the intelligence level of automated applications. This advantage is particularly important in the fields of autonomous driving, intelligent monitoring, and robot navigation, and can achieve more efficient and reliable operations.

[0129] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0130] In the present invention, a multimodal perception platform is constructed by combining multiple sensors to create a comprehensive multimodal perception platform, and then the multimodal perception platform is jointly calibrated to make multimodal data acquisition more reliable. The data of multimodal data is extracted by a multimodal target perception model based on deep learning, thereby reducing noise interference. Finally, the perception results and timestamp information are combined to generate rich semantic information, and the semantic information is released in real time through the ROS communication mechanism, further improving the timeliness and effectiveness of environmental monitoring. This method takes advantage of the multimodal sensor and builds a comprehensive perception system through the fusion of thermal infrared, night vision, visible light and laser LiDAR sensors. The system can provide accurate perception without dynamic illumination, and through the multimodal target perception model constructed by the deep learning algorithm, reduce the limitation of dust on the perception task, realize the three-dimensional positioning of the target, obtain the category information, and suppress the influence of dust on the laser LiDAR sensor, and generate semantic information according to the timestamp information. In addition, the system also publishes the perception results in real time through the ROS communication mechanism, realizes the real-time transmission and processing of data, improves the accuracy of positioning and navigation, provides reliable perception capabilities and realizes real-time detection and visualization.

[0131] Reference Manual Attached Figure 4 , showing a structural schematic diagram of a multimodal fusion perception and semantic generation system for an underground dynamic degradation environment provided by the present invention.

[0132] The present invention also provides a multimodal fusion perception and semantic generation system 20 for underground dynamic degradation environment, which is applied to the multimodal fusion perception and semantic generation method for underground dynamic degradation environment, comprising:

[0133] Processor 201.

[0134] The memory 202 stores computer-readable instructions. When the computer-readable instructions are executed by the processor 201, the multimodal fusion perception and semantic generation method of the underground dynamic degradation environment as described in the method embodiment is implemented.

[0135] The multimodal fusion perception and semantic generation system 20 for underground dynamic degradation environment provided by the present invention can execute the above-mentioned multimodal fusion perception and semantic generation method for underground dynamic degradation environment and achieve the same or similar technical effects. To avoid repetition, the present invention will not go into details.

[0136] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0137] In the present invention, a multimodal perception platform is constructed by combining multiple sensors to create a comprehensive multimodal perception platform, and then the multimodal perception platform is jointly calibrated to make multimodal data acquisition more reliable. The data of multimodal data is extracted by a multimodal target perception model based on deep learning, thereby reducing noise interference. Finally, the perception results and timestamp information are combined to generate rich semantic information, and the semantic information is released in real time through the ROS communication mechanism, further improving the timeliness and effectiveness of environmental monitoring. This method takes advantage of the multimodal sensor and builds a comprehensive perception system through the fusion of thermal infrared, night vision, visible light and laser LiDAR sensors. The system can provide accurate perception without dynamic illumination, and through the multimodal target perception model constructed by the deep learning algorithm, reduce the limitation of dust on the perception task, realize the three-dimensional positioning of the target, obtain the category information, and suppress the influence of dust on the laser LiDAR sensor, and generate semantic information according to the timestamp information. In addition, the system also publishes the perception results in real time through the ROS communication mechanism, realizes the real-time transmission and processing of data, improves the accuracy of positioning and navigation, provides reliable perception capabilities and realizes real-time detection and visualization.

[0138] It should be understood that the processor in the embodiment of the present invention may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0139] It should also be understood that the memory in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0140] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When a computer instruction or computer program is loaded or executed on a computer, a process or function according to an embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a tape), an optical medium (for example, a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state hard disk.

[0141] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. A and B can be singular or plural. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship, but it may also indicate an "and / or" relationship. Please refer to the context for specific understanding.

[0142] In the present invention, "at least one" means one or more, and "more than one" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0143] It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0144] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0145] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described equipment, devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0146] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0147] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0148] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0149] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods of various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc., various media that can store program codes.

[0150] An embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the multimodal fusion perception and semantic generation method for an underground dynamically degraded environment as described in the method embodiment is implemented.

[0151] A computer-readable storage medium provided by the present invention can implement the steps and effects of the multimodal fusion perception and semantic generation method of the underground dynamic degradation environment of the above-mentioned method embodiment. To avoid repetition, the present invention will not go into details.

[0152] The beneficial effects brought about by the technical solution provided by the embodiment of the present invention include at least:

[0153] In the present invention, a multimodal perception platform is constructed by combining multiple sensors to create a comprehensive multimodal perception platform, and then the multimodal perception platform is jointly calibrated to make multimodal data acquisition more reliable. The data of multimodal data is extracted by a multimodal target perception model based on deep learning, thereby reducing noise interference. Finally, the perception results and timestamp information are combined to generate rich semantic information, and the semantic information is released in real time through the ROS communication mechanism, further improving the timeliness and effectiveness of environmental monitoring. This method takes advantage of the multimodal sensor and builds a comprehensive perception system through the fusion of thermal infrared, night vision, visible light and laser LiDAR sensors. The system can provide accurate perception without dynamic illumination, and through the multimodal target perception model constructed by the deep learning algorithm, reduce the limitation of dust on the perception task, realize the three-dimensional positioning of the target, obtain the category information, and suppress the influence of dust on the laser LiDAR sensor, and generate semantic information according to the timestamp information. In addition, the system also publishes the perception results in real time through the ROS communication mechanism, realizes the real-time transmission and processing of data, improves the accuracy of positioning and navigation, provides reliable perception capabilities and realizes real-time detection and visualization.

[0154] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

[0155] There are a few points to note:

[0156] (1) The drawings of the embodiments of the present invention only relate to the structures related to the embodiments of the present invention, and other structures may refer to the general design.

[0157] (2) For the sake of clarity, in the drawings used to describe the embodiments of the present invention, the thickness of the layers or regions is exaggerated or reduced, that is, these drawings are not drawn according to the actual scale. It is understood that when an element such as a layer, film, region or substrate is referred to as being "on" or "under" another element, the element may be "directly" "on" or "under" the other element or there may be intermediate elements.

[0158] (3) In the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other to obtain new embodiments.

[0159] The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. The protection scope of the present invention shall be based on the protection scope of the claims.

Claims

1. A multimodal fusion perception and semantic generation method for underground dynamic degradation environment, characterized in that: include: S1: Build a multimodal perception platform; S2: jointly calibrating the multimodal perception platform; S3: Collect multimodal data of the target to be detected through the multimodal perception platform after joint calibration; S4: extracting data features of the multimodal data through a multimodal target perception model based on a deep learning algorithm; S5: Determine the perception result of the target to be detected according to the data features of the multimodal data; S6: Combining the perception result of the target to be detected with the timestamp information to generate semantic information; S7: The semantic information is published in real time through the ROS communication mechanism to achieve instant visualization of the perception results.

2. The multimodal fusion perception and semantic generation method of underground dynamic degradation environment according to claim 1 is characterized in that: The S1 specifically includes: S101: Combining the characteristics of night vision sensors and visible light sensors, an adaptive light switching sensor based on light-sensitive device triggering is constructed; S102: Introducing a thermal infrared sensor into the adaptive light switching sensor to distinguish whether the target is a living organism; S103: Introducing a laser LiDAR sensor into the adaptive illumination switching sensor to obtain three-dimensional information of the target; S104: Constructing the multimodal perception platform based on the adaptive light switching sensor, the thermal infrared sensor and the laser LiDAR sensor.

3. The multimodal fusion perception and semantic generation method of underground dynamic degradation environment according to claim 2 is characterized in that: The adaptive illumination switching sensor specifically includes: a visible light camera, an infrared transmitter and a photosensor; When the ambient light intensity detected by the light sensor is greater than or equal to 10 lux, the infrared transmitter is automatically turned off, and the visible light camera is used to capture visible light data in the scene; When the ambient light intensity detected by the light sensor is less than 10 lux, the infrared transmitter is automatically turned on, and the infrared transmitter and the visible light camera are used to capture night vision data in the scene.

4. The multimodal fusion perception and semantic generation method of underground dynamic degradation environment according to claim 1 is characterized in that: The multimodal data includes: thermal infrared images, visible light or night vision images, and point cloud data provided by laser LiDAR.

5. The multimodal fusion perception and semantic generation method of underground dynamic degradation environment according to claim 4 is characterized in that: The S4 specifically includes: S401: uniformly sampling the point cloud data: V=[v x ,v y ,v z ] Where V represents the voxel size of the target to be detected, (v x ,v y ,v z ) represents the size of the voxel of the target to be detected in each dimension, p i,j,k represents the representative point of all points in the voxel of the target to be detected, n represents the total number of voxels in the target to be detected, x represents the target to be detected, S i,j,k Represents the point set within the i, j, kth voxel of the target to be detected; S402: Convert the representative point coordinate system into a camera coordinate system: P cam =R·P pc +T Among them, P cam represents the representative point camera coordinate system, R represents the rotation matrix, T represents the translation vector, P pc Represents the coordinate system of the representative point cloud; S403: Convert the representative point camera coordinate system into a pixel coordinate system to obtain a depth map of the target to be detected: Depth(u,v)=Z cam Among them, (X cam ,Y cam ,Z cam ) represents the coordinate information of the representative point camera coordinate system, (x norm ,y norm ) represents the coordinate information of the representative point camera coordinate system projected onto the normalized image plane, K represents the parameter matrix within the camera, Depth(u,v)=Z cam Indicates that the depth value corresponding to each pixel (u, v) is Z cam , Depth(u,v) represents the depth map of the target to be detected.

6. The multimodal fusion perception and semantic generation method of underground dynamic degradation environment according to claim 1 is characterized in that: The S4 specifically includes: S404: extracting features of the thermal infrared image and the visible light or night vision image respectively through the same convolution operation: F vis =Conv k,s (I vis ),F ir =Conv k,s (I ir ) Among them, I vis Indicates visible light or night vision images, I ir Represents thermal infrared image, Conv k,s represents a convolution operation, and the kernel size of the convolution layer is k and the step length is s, F vis Indicates visible light or night vision feature map, F ir Represents thermal infrared signature; S405: performing normalization and ReLU activation operations on the visible light or night vision feature map and the thermal infrared feature map respectively: F ν ′ is =BN[ReLU(F νis )],F ir ′=BN[ReLU(F ir )] Among them, BN represents normalization processing, ReLU represents ReLU activation function, and F ν ' is represents the normalized visible light or night vision feature map, F ir ′ represents the normalized thermal infrared feature map; S406: Fusion of the normalized visible light or night vision feature map and the normalized thermal infrared feature map: F fusion =Concat(F' vis ,F' ir ) Among them, Concat represents the concatenation operation, F fusion Represents the fused feature map; S407: Extracting the intermediate layer feature information of the fused feature map: F mid =Conv k,s (F fusion ) Among them, F mid Represents the intermediate layer feature map; S408: Fusion of different scale features of the intermediate layer feature map through a feature pyramid network: F multi-scale =FPN(F mid ) Among them, F multi-scale represents a multi-scale feature map, and FPN represents a feature pyramid network; S409: reducing the dimension of the multi-scale feature map by a convolution operation, and predicting the bounding box category and position of the multi-scale feature map using a fully connected layer: P c =softmax[FC(F multi-scale )] B=FC(F multi-scale ) L=(1-λ)L c +λL b Among them, FC represents the fully connected network layer, P c represents the predicted category probability of the target to be detected, B represents the predicted bounding box coordinates of the target to be detected, and L c Represents the category loss of the target to be detected, y i Represents the true category label of the target to be detected, represents the predicted probability of the i-th target among the targets to be detected, ∑ represents the summation symbol, log represents the natural logarithm function, L b represents the bounding box loss of the target to be detected, smooth represents the smoothing function, B i represents the predicted bounding box coordinates of the i-th target in the target to be detected, b i represents the true bounding box coordinates of the i-th target in the target to be detected, L represents the total loss function, and λ represents the learnable parameter; S4010: Mapping the predicted bounding box coordinates and predicted category of the target to be detected into a corresponding depth map, and extracting the depth information of the center point of the bounding box in the corresponding depth map to obtain the positioning information of the target to be detected.

7. The multimodal fusion perception and semantic generation method of underground dynamic degradation environment according to claim 1 is characterized in that: The S6 specifically includes: S601: constructing a semantic information framework according to the predicted category of the target to be detected, the positioning information of the target to be detected and the timestamp information corresponding to each other; S602: Based on the semantic information framework, the perception result of the target to be detected is combined with the timestamp information to generate semantic information.

8. The multimodal fusion perception and semantic generation method of underground dynamic degradation environment according to claim 1 is characterized in that: S7 specifically includes: S701: The semantic information generated based on the semantic information framework is published in real time through the rostopic topic of the robot operating system ROS; S702: subscribing to a corresponding rostopic topic through rviz and receiving the semantic information; S703: converting the positioning data in the received semantic information into a three-dimensional mark in the point cloud coordinate system through rviz, and marking the predicted category of the target to be detected on the corresponding three-dimensional mark; S704: Extracting and displaying the point cloud data corresponding to the predicted time according to the timestamp information, so as to realize a dynamically updated visual representation of the target to be detected.

9. A multimodal fusion perception and semantic generation system for underground dynamic degradation environment, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the multimodal fusion perception and semantic generation method for an underground dynamically degraded environment as described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the multimodal fusion perception and semantic generation method for an underground dynamic degradation environment as described in any one of claims 1 to 8 is implemented.