Three-dimensional reconstruction-based intelligent multi-modal fusion sensing backpack with body and sensing method

By integrating multi-source sensors into the robot backpack and performing 3D reconstruction, the problem of insufficient multimodal fusion in the robot's perception system was solved, enabling the robot to autonomously perceive and make safe decisions in complex environments, and improving the flexibility and reliability of task execution.

CN121374728APending Publication Date: 2026-01-23娄浩哲
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511707509.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-01-23

AI Technical Summary

Technical Problem

Existing robot perception systems are inadequate in terms of the diversity of sensing modalities, the reconfigurability of hardware platforms, and the fusion mechanism of multi-source information, and cannot meet the needs of embodied intelligent robots for comprehensive perception and autonomous adaptation in complex, open, and dynamic scenarios.

Method used

Design an embodied intelligent multimodal fusion perception backpack based on 3D reconstruction, integrating multiple source sensors such as vision, touch, and smell, and realizing the unified fusion of multimodal data through multi-level data processing units to generate a 3D scene model for environmental recognition, path planning, and safety decision-making.

Benefits of technology

It enhances the robot's autonomous perception and interaction capabilities in complex environments, achieves efficient integration of multimodal information and safe obstacle avoidance, and is suitable for various scenarios such as inspection, search and rescue, and transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121374728A_ABST
    Figure CN121374728A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent multi-modal fusion sensing backpack based on three-dimensional reconstruction and a sensing method. The sensing knapsack comprises a knapsack body, a mechanical arm, a control cabin, a host end, a visual sensor installed on the knapsack body, a thermal imaging camera used for collecting environment and object surface temperature distribution, a touch sensor array integrated at the end of the mechanical arm of the knapsack, and a multi-stage data processing unit. The sensing method comprises the following steps: integrating a visual sensor, a depth sensor, a thermal imaging camera, a tactile array sensor and an olfactory sensor on the sensing backpack; vision and depth information is obtained, three-dimensional reconstruction of the surrounding environment is achieved, temperature distribution and hazardous gas sensing are achieved in combination with thermal imaging and an olfactory sensor, and therefore a high-temperature or dangerous area is recognized and avoided. Meanwhile, the characteristics of a carried object or an external contact target are recognized and distinguished through force sense sensing data, and the autonomous sensing and interaction capacity of a robot bearing system in a complex environment is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot intelligent perception and operation, and in particular to a somatic intelligent multi-modal fusion perception backpack based on three-dimensional reconstruction and a perception method. BACKGROUND

[0002] With the rapid development of artificial intelligence, robotics technology and wearable smart equipment, there is an increasing demand for robot systems with autonomous environmental interaction capabilities to perform tasks such as inspection, rescue, operation and material transportation in complex and dynamic environments. Such tasks usually require robots to have high environmental adaptability, multi-modal perception capability and real-time decision support. Although traditional fixed perception modules or robotic arm type perception devices have been able to achieve environmental recognition and basic operation functions to some extent, there are still significant technical bottlenecks in spatial flexibility, multi-source information fusion capability and cross-modal environmental understanding, which limit the comprehensive performance of robots in real complex scenarios.

[0003] Firstly, the perception methods of existing robot systems are generally single, most of which rely on a few sensing modalities such as visual sensors or inertial measurement units (IMU), lacking the ability to cooperatively collect and analyze multiple physical quantities such as temperature, humidity, mechanical touch, gas concentration, etc. For example, in high-temperature operation or vibration operation environments, it is difficult to accurately determine the surface temperature or structural stability of equipment relying only on visual information, and the lack of tactile feedback also makes the robot unable to perceive the force and surface material when performing grasping or contact tasks, thereby affecting the operation precision and system safety. This lack of perception dimension directly restricts the adaptability and reliability of robots in extreme or special working conditions.

[0004] Secondly, in terms of system architecture, there is currently no detachable intelligent perception backpack platform designed specifically for somatic intelligent robots with high compatibility and reconfigurability. Existing perception modules are mostly attached to end effector mechanisms such as robotic arms and dexterous hands, or integrated into the robot body, resulting in limited perception field and inflexible deployment, making it difficult to quickly adjust the perception configuration according to the scene requirements in tasks. Especially in applications such as material transportation and field exploration that require robots to frequently switch working modes, there is a lack of a plug-and-play intelligent terminal that supports multi-modal sensor expansion, limiting the quick response and functional reuse capability of robots in cross-scene tasks.

[0005] Furthermore, at the level of sensory information fusion and understanding, traditional robot systems often employ loosely coupled or post-fusion data processing methods. Different modalities such as vision, inertia, and tactile data are typically integrated asynchronously at the feature layer or decision layer, lacking a deep fusion mechanism under a unified spatiotemporal reference. This makes it difficult for the system to achieve effective complementarity and redundancy verification of multi-source information in dynamic environments. Especially during robot movement, inconsistencies in sensor sampling frequencies, coordinate systems, and delays can easily lead to perceptual conflicts and state estimation drift, severely impacting the accuracy and robustness of understanding complex terrains and multi-object interaction scenarios.

[0006] In summary, existing robot perception systems still have significant shortcomings in terms of the diversity of sensing modalities, the reconfigurability of hardware platforms, and the fusion mechanism of multi-source information. These shortcomings fail to fully meet the urgent needs of embodied intelligent robots for comprehensive perception, autonomous adaptation, and intelligent decision-making in complex, open, and dynamic scenarios. Therefore, a novel intelligent sensing carrying system and method are urgently needed. This system should integrate multiple physical quantity sensing units, support flexible disassembly and expansion, and possess efficient and unified multimodal perception fusion capabilities. This would enhance the overall performance and intelligence level of robots under all-terrain and multi-task conditions. Summary of the Invention

[0007] Purpose of the invention: To address the shortcomings of existing dexterous robotic arms or traditional robots, such as single perception modality, limited environmental understanding, and insufficient spatial perception coverage, this invention proposes an embodied intelligent multimodal fusion perception backpack and perception method based on 3D reconstruction. By integrating multi-source sensors into the robotic backpack, it achieves multimodal fusion perception of environmental geometric features, temperature distribution, gas concentration, and mechanical state, constructing an intelligent perception platform with spatial understanding and obstacle avoidance capabilities. Specifically, through the collaborative perception of visual, tactile, and olfactory information, it achieves 3D reconstruction, temperature distribution mapping, and tactile response fusion. Furthermore, by integrating the perception backpack with the robotic arm, it enriches the environmental perception capabilities of the integrated system, enhancing its autonomous identification, path planning, and safety decision-making capabilities in complex environments. It is applicable to various scenarios such as inspection, search and rescue, transportation, and intelligent interaction.

[0008] Technical solution: The present invention is a three-dimensional reconstruction-based embodied intelligent multimodal fusion sensing backpack, which includes a multi-backpack body, a robotic arm, a control cabin, a host terminal, a visual sensor installed on the shell of the backpack body, a thermal imaging camera for collecting environmental and object surface temperature distribution, a tactile sensor array integrated into the end of the backpack's deployable robotic arm, and a multi-level data processing unit.

[0009] The multi-level data processing unit includes a first-level preprocessing module located at the end of the robotic arm, a second-level integration module inside the backpack's core control cabin, and a third-level fusion processing module on the host side. The first-level preprocessing module filters and denoises the real-time data from the tactile sensor array at the end of the robotic arm, using an algorithm combining median filtering to remove environmental interference signals, and then transmits the preprocessed data to the second-level integration module via a high-speed serial bus.

[0010] The second-level integrated module establishes a multi-threaded data processing framework: on the one hand, it receives mechanical sensing data from the first-level preprocessing module, and on the other hand, it simultaneously acquires raw data from visual sensors, depth sensors, thermal imaging cameras, and olfactory sensors; it extracts point cloud models based on visual and depth data, reconstructs 3D object models in a virtual environment, interpolates and amplifies thermal imaging data, performs feature matching on thermal imaging data and point cloud models to map temperature information onto the 3D object model, quantifies concentration and determines thresholds for olfactory sensor data, and establishes a unified timestamp synchronization mechanism to align different modal data according to acquisition time and generate standardized data frames.

[0011] The third-level fusion processing module takes the standardized data frames transmitted by the second-level integration module as input, and based on the data frames containing the three-dimensional scene model, completes advanced decision-making tasks such as target recognition, danger zone determination, and path planning; it then converts the decision results into standardized control commands and sends them to the embodied intelligent device for execution.

[0012] The visual sensors include an RGB camera and a depth sensor.

[0013] The embodied intelligent multimodal fusion perception method based on 3D reconstruction of this invention includes the following steps:

[0014] (1) Connect the embodied intelligent device to the sensing backpack through an external interface to realize the function of power supply of the machine backpack and to perceive and identify external objects.

[0015] (2) Integrate a visual sensor, a depth sensor, a thermal imaging camera, a tactile array sensor and an olfactory sensor into a machine-sensing backpack.

[0016] (3) Based on visual and depth sensor data, further perform 3D reconstruction of the surrounding environment and target objects to obtain an environment model and target model containing geometric shape and spatial position, which together form a 3D scene model. The process is as follows:

[0017] Based on visual and depth data acquired by visual and depth sensors, point cloud models of embodied intelligent devices and external objects are obtained. Points in the point cloud model are represented as follows: ,in Let be the coordinates of each point in the point cloud in a spatial rectangular coordinate system; the point cloud model is fused with RGB image feature points to assign color values ​​to the corresponding point cloud data points, resulting in a 3D scene model of the environment in which the smart device and object are located. This 3D scene model is composed of a 3D Gaussian ellipsoid, represented as:

[0018]

[0019] in, Represents the set of Gaussian spheres. The set of the center points of the three-dimensional Gaussian ellipsoid. These are the coordinates of the center point of each three-dimensional Gaussian ellipsoid. The set of covariance matrices of a three-dimensional Gaussian ellipsoid. It is the covariance matrix of each three-dimensional Gaussian ellipsoid, representing the scaling ratio and rotation angle of each Gaussian sphere in space; This is a set of transparency parameters for a three-dimensional Gaussian ellipsoid. Is each three-dimensional Gaussian ellipsoid in Transparency in color space It is the set of spherical harmonic coefficients of a three-dimensional Gaussian ellipsoid. It is the spherical harmonic coefficient of each three-dimensional Gaussian ellipsoid, representing the color of the Gaussian ellipsoid when viewed from a certain angle.

[0020] (4) Register and fuse the thermal imaging data with the 3D reconstruction results to map the temperature distribution onto the 3D scene model. Combine this with tactile and olfactory sensor data to generate a 3D scene model that includes geometry, mechanics, gas concentration, and temperature distribution. The process is as follows:

[0021] The images acquired by the thermal imaging camera are interpolated to obtain a thermal imaging output image with the same resolution as the RGB image. Then, the thermal pixel coordinates [x, y] and the depth value d are converted into three-dimensional temperature points, represented as:

[0022]

[0023] In the formula, This represents a temperature point in three-dimensional space, where K is a camera intrinsic parameter. For camera pose; feature extraction (maximum and minimum values, position coordinates) is performed on the haptic array data. The extraction process is as follows:

[0024] For tactile arrays, firstly, the first... Spatial extremum detection is performed using frame tactile array data; pressure values ​​are traversed across M×N sensing units. Where M represents the robotic hand index, N represents the number of finger joints, and the maximum value of that frame is recorded. minimum value Introducing an effective contact threshold (This effective contact threshold is set based on 1.5 times the standard deviation of the sensor noise floor.) If there is no valid solution, then it is determined that there is no effective solution, and the extreme values ​​are all recorded as 0; if If the value is true, then it is a valid value. Record the coordinates of the tactile sensor corresponding to this valid contact threshold. and the coordinates of the minimum value corresponding to the tactile sensor It integrates geometric / temperature information to find the coordinates of the corresponding region in the 3D scene model, and assigns the weight of the object at the corresponding coordinates.

[0025] The thermal imaging data is registered and fused with the 3D reconstruction results. Specifically, the 3D temperature points output by the thermal imaging are used to assign temperature values ​​to the corresponding 3D Gaussian ellipsoid points, mapping the temperature onto the 3D scene model. Combined with tactile and olfactory sensor data, a 3D scene model incorporating geometry, mechanics, gas concentration, and temperature distribution is generated. The spatial model constituting this 3D scene model is then represented as follows:

[0026] In the formula, Represents a 3D scene model. It is a set of three-dimensional Gaussian spheres. The set of weights of external objects. For the set of gas concentration distributions, It is a set of three-dimensional temperature distribution points.

[0027] (5) An operation method for generating a machine-perceptive backpack based on the generated 3D scene model is used. During operation or carrying tasks, the backpack automatically avoids high-temperature or dangerous areas based on thermal imaging results, and performs safe path planning and object carrying. In recognition tasks, data from tactile and inertial measurement units are used to distinguish objects that look the same but have different masses or materials. The process is as follows:

[0028] First, external objects are initially perceived using an RGB camera, thermal imaging camera, depth camera, and olfactory sensor. Based on temperature and gas concentration, high-temperature hazardous objects and safe objects are distinguished. The sensing backpack then drives an internal robotic arm to approach a safe object. Upon reaching the safe object, a dexterous hand is activated to grasp it. The weight information of the object is obtained using a tactile sensor integrated into the dexterous hand, ultimately achieving target object recognition. The recognition process is as follows:

[0029] First, a target object feature library is constructed, containing visual features (color histogram, texture features, 3D size parameters), weight features (standard weight range, weight deviation threshold), and environmental adaptation features (safe temperature range, hazardous gas-free association tags). Each known object in the feature library is labeled with a unique identification ID and multi-dimensional feature baseline values. Real-time data collected by multiple sensors is preprocessed to extract the object's RGB image color distribution, surface texture details, and 3D length, width, and height dimensions. This is combined with the actual weight calculated by the dexterous hand tactile sensor (after gravity calibration and gripping force cancellation processing), as well as the previously confirmed safe temperature value and hazardous gas-free environmental characteristics, to form a real-time feature vector. A weighted cosine similarity algorithm is used for matching to obtain a similarity score. :

[0030]

[0031] Among them, visual feature weights Weight feature weight Environmental adaptation feature weights Calculate the similarity between the real-time feature vector and the baseline feature values ​​of each known object in the feature library.

[0032] In step (6), a three-level matching threshold is set for similarity:

[0033] (a) When the similarity is ≥92%, the corresponding object's identification ID and name are directly output for accurate identification.

[0034] (b) When 75% ≤ similarity < 92%, a secondary feature acquisition is initiated. By slightly rotating the object with a dexterous hand, visual features and pressure distribution features from different perspectives are acquired. After recalculating the similarity, the optimal matching result is output.

[0035] (c) When the similarity is < 75%, it is determined to be an unknown object. The complete feature vector is recorded and stored in the feature library for later update. At the same time, the prompt "unknown object + core feature information (weight, size, color)" is output.

[0036] (6) The operation method generated in step (5) is sent to the embodied intelligent device. During the rotation or movement of the embodied intelligent device's viewpoint, the 3D scene model is dynamically adjusted based on the real-time collected visual, tactile, and olfactory data to perform closed-loop control and operation optimization. The dynamic adjustment process is as follows: Record The maximum number of elements in each set:

[0037] Let S be a set of Gaussian spheres, and let S be a 3D scene model. For the set of gas concentration distributions, This is a set of three-dimensional temperature distribution points; when the viewpoint of the embodied intelligent device rotates or moves, a new set of Gaussian spheres is obtained according to steps (3)-(4). The set of external object weights Gas concentration distribution set Three-dimensional temperature distribution point set Meanwhile, the data set lost during the rotation includes the Gaussian sphere set. The set of external object weights Gas concentration distribution set Three-dimensional temperature distribution point set Based on the newly added and deleted sets, the updated 3D scene model is as follows:

[0038]

[0039] remember , , , The updated 3D scene model is then represented as: .

[0040] Working Principle: This invention relates to an embodied intelligent multimodal fusion sensing backpack and method based on 3D reconstruction, applicable to robots, robotic arms, and drones. By integrating a high-density tactile array, depth sensor, thermal imaging camera, and olfactory sensor onto the backpack, a multimodal sensing system with environmental perception and intelligent interaction capabilities is formed. The sensing methods include: acquiring visual and depth information to achieve 3D reconstruction of the surrounding environment; combining thermal imaging and olfactory sensors to detect temperature distribution and hazardous gases, thereby identifying and avoiding high-temperature or dangerous areas; and simultaneously identifying and distinguishing the characteristics of carried objects or external contact targets through force sensing data, significantly improving the autonomous perception and interaction capabilities of the robot carrying system in complex environments.

[0041] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0042] (1) This invention solves the limitation of traditional mobile robot carrying systems that mainly rely on single visual perception and are difficult to fully acquire environmental physical state and thermal information. By integrating a machine perception backpack into a traditional robot or robotic arm, and integrating a thermal imaging camera and a high-density tactile array sensor into the perception backpack, a multimodal environmental perception system with wide coverage and sensitive response is constructed.

[0043] (2) In the process of 3D reconstruction and scene understanding, the present invention realizes the data fusion of visual sensors, depth sensors, thermal imaging cameras, tactile sensors and olfactory sensors, which significantly improves the comprehensive perception and information integration capabilities of the machine perception backpack in complex environments.

[0044] (3) The present invention enables the robotic backpack and robotic arm or robot integrated system to have temperature field recognition and safe obstacle avoidance capabilities, and physical attribute identification capabilities based on multi-source signals, thereby performing highly reliable auxiliary operation tasks in complex task scenarios such as disaster relief (high temperature, smoke), industrial inspection (precision equipment, hot components) and daily services (high temperature kitchen environment, handling of fragile objects). Attached Figure Description

[0045] Figure 1 This is an external structural diagram of the embodied intelligent multimodal fusion sensing backpack based on 3D reconstruction according to the present invention;

[0046] Figure 2 This is a diagram of the internal structure of the embodied intelligent multimodal fusion perception backpack based on 3D reconstruction according to the present invention.

[0047] Figure 3 This is a flowchart of the embodied intelligent multimodal fusion perception method based on 3D reconstruction of the present invention. Detailed Implementation

[0048] like Figure 1 and Figure 2 As shown, the embodied intelligent multimodal fusion perception backpack based on three-dimensional reconstruction of the present invention includes an RGB camera 11, a depth sensor 15, a thermal imaging camera 12, an olfactory sensor 13, an external interface 14, and a tactile sensor 26 integrated on the first dexterous hand 21 and the second dexterous hand 23 at the top of the first robotic arm 27 and the second robotic arm 28.

[0049] The first dexterous hand 21 and the second dexterous hand 23 are efficient and anthropomorphic hardware platforms; the RGB camera 11 outputs a resolution of 1280×960 and a frame rate of 30fps; the tactile sensor 26 is integrated into an array with an array sampling accuracy of milligram level; the thermal imaging camera 12 acquires data at a resolution of 160×120 and outputs it at a resolution of 1280×960 through an interpolation algorithm; the remaining structures include a heat dissipation device 16, an integrated graphics card 25, a dual-arm blocking device 22, and a dual-arm support plate 24.

[0050] This invention relates to an embodied intelligent multimodal fusion perception method based on 3D reconstruction, comprising the following steps:

[0051] (1) Connect the embodied intelligent device to the machine-sensing backpack through an external interface to enable the machine backpack to function as a power source and to perceive and identify external objects.

[0052] (2) The robot backpack adopts a modular slot design to integrate sensor components: the vision sensor and the thermal imaging camera are coaxially mounted on the upper part of the backpack; the depth sensor is embedded in the middle layer of the backpack, and the lens is equipped with an anti-fog coating and a dust cover; the tactile array sensor adopts a flexible patch structure and is distributed in the robotic arm area at the top of the robotic arm; the olfactory sensor is integrated below the depth sensor and is equipped with a multi-channel gas adsorption module, which can detect a variety of common dangerous gases at the same time; all sensors are connected to the backpack main control board through a standardized interface, which supports hot-swapping and replacement, and the sensor modules can be flexibly added or removed according to the task requirements.

[0053] (3) Based on visual and depth sensor data, point cloud models of the embodied intelligent device and external objects are obtained. The points in the point cloud model are represented as follows: ,in Let be the coordinates of each point in the point cloud in a spatial rectangular coordinate system; the point cloud model is fused with RGB image feature points to assign color values ​​to the corresponding point cloud data points, resulting in a 3D scene model of the environment in which the smart device and object are located. The 3D scene model is composed of a 3D Gaussian ellipsoid and is represented as:

[0054]

[0055] in, Represents the set of Gaussian spheres. The set of the center points of the three-dimensional Gaussian ellipsoid. These are the coordinates of the center point of each three-dimensional Gaussian ellipsoid. The set of covariance matrices of a three-dimensional Gaussian ellipsoid. It is the covariance matrix of each three-dimensional Gaussian ellipsoid, representing the scaling ratio and rotation angle of each Gaussian sphere in space; This is a set of transparency parameters for a three-dimensional Gaussian ellipsoid. Is each three-dimensional Gaussian ellipsoid in Transparency in color space It is the set of spherical harmonic coefficients of a three-dimensional Gaussian ellipsoid. It is the spherical harmonic coefficient of each three-dimensional Gaussian ellipsoid, representing the color of the Gaussian ellipsoid when viewed from a certain angle.

[0056] (4) The 160×120 resolution image acquired by the thermal imaging camera is interpolated to obtain a thermal imaging output image with the same resolution as the 1280×960 RGB image. Then, the thermal pixel coordinates [x, y] and the depth value d are converted into three-dimensional temperature points. , is represented as:

[0057]

[0058] In the formula, This represents a temperature point in three-dimensional space, where K is a camera intrinsic parameter. The camera pose is determined; the haptic array data is subjected to median filtering, and the filtering result is fused with geometric / temperature information.

[0059] Thermal imaging data is registered and fused with a 3D scene model. Specifically, the 3D temperature points output by thermal imaging are used to assign temperature values ​​to corresponding 3D Gaussian ellipsoid points, mapping the temperature onto the 3D scene model. Combined with tactile and olfactory sensor data, a 3D scene model is generated that includes geometry, mechanics, gas concentration, and temperature distribution. The spatial model constituting this 3D scene model is then represented as follows:

[0060]

[0061] In the formula, Represents a 3D scene model. It is a set of three-dimensional Gaussian spheres. The set of weights of external objects. For the set of gas concentration distributions, It is a set of three-dimensional temperature distribution points.

[0062] (5) The operation method for generating a machine-perceived backpack based on the generated 3D scene model is as follows: First, the external objects are fused and perceived as shown in steps (2)-(4) based on the RGB camera, thermal imaging camera, depth camera and olfactory sensor. The safety index is set according to the temperature information and gas concentration information. Used to distinguish between high-temperature hazardous objects and safe objects, as follows:

[0063] In the formula, Mark whether it is high temperature. This indicates whether the concentration of toxic gas exceeds the standard. When determining whether an object or area is a dangerous object or a safe object.

[0064] The sensing backpack drives the internal robotic arm to approach a safe object. When the robotic arm reaches the safe object, it drives the dexterous hand to perform a grasping task. The tactile sensors integrated into the dexterous hand obtain the object's weight information, ultimately achieving the recognition of the target object. The recognition process is as follows:

[0065] First, a target object feature library is constructed, containing visual features (color histogram, texture features, 3D size parameters), weight features (standard weight range, weight deviation threshold), and environmental adaptation features (safe temperature range, hazardous gas-free association tags). Each known object in the feature library is labeled with a unique identification ID and multi-dimensional feature baseline values. Real-time data collected by multiple sensors is preprocessed to extract the object's RGB image color distribution, surface texture details, and 3D length, width, and height dimensions. This is combined with the actual weight calculated by the dexterous hand tactile sensor (after gravity calibration and gripping force cancellation processing), as well as the previously confirmed safe temperature value and hazardous gas-free environmental characteristics, to form a real-time feature vector. A weighted cosine similarity algorithm is used for matching to obtain a similarity score. :

[0066]

[0067] Among them, visual feature weights Weight feature weight Environmental adaptation feature weights Calculate the similarity between the real-time feature vector and the baseline feature values ​​of each known object in the feature library.

[0068] In step (5), a three-level matching threshold is set for similarity:

[0069] (a) When the similarity is ≥92%, the corresponding object's identification ID and name are directly output for accurate identification.

[0070] (b) When 75% ≤ similarity < 92%, a secondary feature acquisition is initiated. By slightly rotating the object with a dexterous hand, visual features and pressure distribution features from different perspectives are acquired. After recalculating the similarity, the optimal matching result is output.

[0071] (c) When the similarity is < 75%, it is determined to be an unknown object. The complete feature vector is recorded and stored in the feature library for later update. At the same time, the prompt "unknown object + core feature information (weight, size, color)" is output.

[0072] (6) The operation method generated in step (5) is sent to the embodied intelligent device. During the rotation or movement of the embodied intelligent device's viewpoint, the 3D scene model is dynamically adjusted based on the real-time collected visual, tactile, and olfactory data to perform closed-loop control and operation optimization. The dynamic adjustment process is as follows: The maximum number of elements in each set:

[0073] Let S be a set of Gaussian spheres, and let S be a 3D scene model. For the set of gas concentration distributions, This is a set of three-dimensional temperature distribution points; when the viewpoint of the embodied intelligent device rotates or moves, a new set of Gaussian spheres is obtained according to steps (3)-(4). The set of external object weights Gas concentration distribution set Three-dimensional temperature distribution point set Meanwhile, the data set lost during the rotation includes the Gaussian sphere set. The set of external object weights Gas concentration distribution set Three-dimensional temperature distribution point set Based on the newly added and deleted sets, the updated 3D scene model is as follows:

[0074]

[0075] remember , , , The updated 3D scene model is then represented as: .

Claims

1. A three-dimensional reconstruction-based embodied intelligent multimodal fusion sensing backpack, characterized in that: It includes a backpack body, a robotic arm, a control cabin, a main unit, a vision sensor mounted on the backpack body, a thermal imaging camera that collects environmental and object surface temperature distribution, a tactile sensor array integrated into the end of the robotic arm on the backpack body, and a multi-level data processing unit. The multi-level data processing unit includes a first-level preprocessing module located at the end of the robotic arm, a second-level integration module located in the control cabin, and a third-level fusion processing module located on the host computer.

2. The embodied intelligent multimodal fusion sensing backpack based on 3D reconstruction according to claim 1, characterized in that: The first-level preprocessing module filters and denoises the real-time data from the tactile sensor array at the end of the robotic arm, removes environmental interference signals, and transmits the preprocessed data to the second-level integrated module via a high-speed serial bus.

3. The embodied intelligent multimodal fusion sensing backpack based on 3D reconstruction according to claim 1, characterized in that: The second-level integrated module builds a multi-threaded data processing framework: while receiving mechanical sensing data from the first-level preprocessing module, it simultaneously collects raw data from the vision sensor, depth sensor, thermal imaging camera, and olfactory sensor; it extracts point cloud models based on the vision and depth data, reconstructs a three-dimensional object model in the virtual environment, interpolates and amplifies the thermal imaging data, performs feature matching on the thermal imaging data and point cloud model to map the temperature onto the three-dimensional object model, and quantifies the concentration and makes threshold judgments on the olfactory sensor data. Establish a timestamp synchronization mechanism to align data from different modalities according to their acquisition time and generate standardized data frames.

4. The embodied intelligent multimodal fusion sensing backpack based on 3D reconstruction according to claim 1, characterized in that: The third-level fusion processing module takes the standardized data frames transmitted by the second-level integration module as input, and performs target recognition, danger zone determination, and path planning decisions based on the data frames containing the three-dimensional scene model; it then converts the decision results into standardized control commands and sends them to the embodied intelligent device for execution.

5. A method for embodied intelligent multimodal fusion perception based on 3D reconstruction, characterized in that: The embodied intelligent multimodal fusion perception backpack based on 3D reconstruction as described in claim 1 is implemented, wherein the perception method includes the following steps: (1) Connect the embodied intelligent device to the sensing backpack through an external interface, power the sensing backpack, and perform sensing and recognition of external objects; (2) Integrate a visual sensor, a depth sensor, a thermal imaging camera, a tactile array sensor, and an olfactory sensor into the sensing backpack; (3) Based on visual and depth sensors, visual and depth data are acquired to obtain point cloud models of the embodied intelligent device and external objects; the point cloud models are fused with RGB image feature points, and color values ​​are assigned to the corresponding point cloud data points to obtain a three-dimensional scene model of the environment in which the intelligent device and objects are located. The three-dimensional scene model is composed of a three-dimensional Gaussian ellipsoid and is represented as: ; in, Represents the set of Gaussian spheres. The set of the center points of the three-dimensional Gaussian ellipsoid. These are the coordinates of the center point of each three-dimensional Gaussian ellipsoid. The set of covariance matrices of a three-dimensional Gaussian ellipsoid. It is the covariance matrix of each three-dimensional Gaussian ellipsoid, representing the scaling ratio and rotation angle of each Gaussian sphere in space; This is a set of transparency parameters for a three-dimensional Gaussian ellipsoid. Is each three-dimensional Gaussian ellipsoid in Transparency in color space It is the set of spherical harmonic coefficients of a three-dimensional Gaussian ellipsoid. is the spherical harmonic coefficient of each three-dimensional Gaussian ellipsoid, representing the color of the Gaussian ellipsoid when viewed from a certain angle; (4) The images acquired by the thermal imaging camera are interpolated to obtain a thermal imaging output image with the same resolution as the RGB image. Then, the thermal pixel coordinates [x, y] and the depth value d are converted into three-dimensional temperature points, represented as: ; In the formula, Let K be the temperature point in three-dimensional space, and K be the camera's intrinsic parameter. The camera pose is determined; features are extracted from the tactile array data and fused with geometric and temperature information; and the weight of the object in the corresponding region is assigned in the 3D scene model. Thermal imaging data is registered and fused with 3D reconstruction results, and combined with tactile and olfactory sensor data to generate a 3D scene model that includes geometry, mechanics, gas concentration, and temperature distribution. ; In the formula, Represents a 3D scene model. It is a set of three-dimensional Gaussian spheres. The set of weights of external objects. For the set of gas concentration distributions, It is a set of three-dimensional temperature distribution points; (5) First, based on the RGB camera, thermal imaging camera, depth camera and olfactory sensor, the external object is initially perceived in steps (2) to (4), and a safety index is set according to the temperature and gas concentration. To distinguish between high-temperature hazardous objects and safe objects, when the sensing backpack drives the robotic arm to reach a safe object, it uses the tactile sensors integrated into the dexterous hand to obtain the weight of the object, and then identifies the target object. (6) During the rotation or movement of the viewpoint, the embodied intelligent device dynamically adjusts the 3D scene model based on real-time collected visual, tactile, temperature, and olfactory concentration data. The process is as follows: The maximum number of elements in each set: ; in, Let S be a set of Gaussian spheres, and let S be a 3D scene model. For the set of gas concentration distributions, This is a set of three-dimensional temperature distribution points; when the viewpoint of the embodied intelligent device rotates or moves, a new set of Gaussian spheres is obtained according to steps (3)-(4). The set of external object weights Gas concentration distribution set Three-dimensional temperature distribution point set Meanwhile, the data set lost during the rotation includes the Gaussian sphere set. The set of external object weights Gas concentration distribution set Three-dimensional temperature distribution point set Based on the newly added and deleted sets, the updated 3D scene model is as follows: ; remember , , , The updated 3D scene model is as follows: .

6. The embodied intelligent multimodal fusion perception method based on three-dimensional reconstruction according to claim 5, characterized in that: In step (2), the visual sensor and the thermal imaging camera are coaxially mounted on the upper part of the sensing backpack, the depth sensor is embedded in the middle layer of the sensing backpack, and the lens is equipped with an anti-fog coating and a dust cover; the tactile array sensor is distributed in the robotic arm area; the olfactory sensor is integrated below the depth sensor and configured with a multi-channel gas adsorption module to detect a variety of common hazardous gases at the same time.

7. The embodied intelligent multimodal fusion perception method based on three-dimensional reconstruction according to claim 5, characterized in that: In step (4), the process of feature extraction from the tactile array data is as follows: First, the first... Spatial extremum detection is performed using frame tactile array data; pressure values ​​are traversed across M×N sensing units. Where M represents the robotic hand index and N represents the number of finger joints. And record the maximum value of that frame. minimum value Introducing an effective contact threshold ,like If so, it is determined that there is no valid solution, and the maximum value is recorded as 0; if If the value is true, then it is a valid value. Record the tactile sensor coordinates corresponding to the valid value. and the coordinates of the minimum value corresponding to the tactile sensor It integrates geometric / temperature information to find the coordinates of the corresponding region in the 3D scene model and assigns the weight of the object at the corresponding coordinates.

8. The embodied intelligent multimodal fusion perception method based on three-dimensional reconstruction according to claim 7, characterized in that: In step (4), when the similarity is ≥92%, the corresponding object's identification ID and name are output; When the similarity is 75% ≤ similarity < 92%, a secondary feature acquisition is initiated. The object is rotated by a dexterous hand to collect additional visual features and pressure distribution features from different perspectives. After recalculating the similarity, the optimal matching result is output. When the similarity is less than 75%, it is determined to be an unknown object. The complete feature vector of the unknown object is recorded and stored in the feature library for later updating.

9. The embodied intelligent multimodal fusion perception method based on three-dimensional reconstruction according to claim 5, characterized in that: In step (5), the process of identifying the target object is as follows: First, a target object feature library containing visual features, weight features, and environmental adaptation features is constructed. Each known object in the feature library is labeled with a unique identification ID and multi-dimensional feature benchmark values. After preprocessing, the real-time data collected by multiple sensors extracts the RGB image color distribution, surface texture details, and three-dimensional dimensions of the object. Combined with the actual weight calculated by the dexterous hand tactile sensor, as well as the confirmed safe temperature value and environmental features without dangerous gases, a real-time feature vector is formed. The weighted cosine similarity algorithm is used for matching to obtain a similarity score. : ; in, For visual feature weights, Weights for weight features To adapt the feature weights to the environment, calculate the similarity between the real-time feature vector and the baseline feature values ​​of each known object in the feature library.

10. The embodied intelligent multimodal fusion perception method based on three-dimensional reconstruction according to claim 5, characterized in that: In step (5), a safety index is set based on temperature and gas concentration. To distinguish between high-temperature hazardous objects and safe objects: ; In the formula, Mark whether it is high temperature. This indicates whether the concentration of toxic gas exceeds the standard. When determining whether an object or area is a dangerous object or a safe object.