Intelligent fire-fighting robot fusing multi-modal embodied intelligence and application control method thereof

By integrating multi-modal embodied intelligent firefighting robots with multispectral vision, lidar and other sensors and edge intelligence hubs, the problems of single perception and lengthy response in firefighting systems have been solved. This enables deep understanding of the fire scene and autonomous decision-making, thereby improving firefighting efficiency and safety.

CN122172659APending Publication Date: 2026-06-09HANGZHOU TAIXIAO TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU TAIXIAO TECHNOLOGY CO LTD
Filing Date
2026-02-09
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing fire protection systems suffer from limited perception dimensions, lengthy response chains, and high personnel safety risks. Firefighting robots lack in-depth understanding and dynamic response capabilities, making them unable to collaborate with building fire protection systems.

Method used

The multimodal embodied intelligent firefighting robot integrates a multispectral vision head, solid-state lidar, multi-gas sensor array, and acoustic and positioning modules. Combined with an edge intelligent hub and a highly mobile execution unit, it achieves a closed loop of perception, decision-making, and execution, and performs collaborative operations through a multi-agent distributed task allocation algorithm.

Benefits of technology

It has achieved a deep understanding of the fire scene and autonomous decision-making, improved fire fighting efficiency and safety, and can operate autonomously in complex environments, freeing up manpower and enabling system collaboration and continuous evolution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122172659A_ABST
    Figure CN122172659A_ABST
Patent Text Reader

Abstract

The application provides a kind of intelligent fire-fighting robot of fusion multi-modal embodied intelligence and its application control method, it aims at deeply fusing multi-modal perception fusion, hierarchical reinforcement learning, whole body motion planning, multi-agent collaboration and other frontier technologies, designs and realizes a kind of mobile fire duty robot (fusion multi-modal embodied intelligent intelligent fire-fighting robot and its application control method) with embodied intelligence, promotes the paradigm change of fire operation from "manual experience driving" to "machine autonomous intelligence". The mobile fire duty robot with embodied intelligence is a new generation of fire-fighting equipment form created by deeply fusing artificial intelligence, robotics and control theory. It realizes the leap from "perception-control" separation to "perception-cognition-decision-action" integration, is a key step to promote fire rescue to high intelligence and autonomy, and has great social value and application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent fire protection technology, and in particular to an intelligent fire-fighting robot that integrates multimodal embodied intelligence, its application control method, electronic equipment, and computer-readable storage medium. Background Technology

[0002] The core of fire safety lies in "early prevention, extinguishing small fires, and rapid rescue." However, the current fire protection system mainly relies on a dual model of "fixed facilities + manual duty," and its limitations are becoming increasingly apparent against the backdrop of increasingly complex building structures, diversified risks, and high labor costs.

[0003] 1.1 Systemic defects of the traditional fire duty system The perception dimensions are limited and passive: Fixed smoke and heat detectors rely on threshold alarms and cannot distinguish between fire types (such as electrical fires and oil fires), and their response to smoldering fires and rapidly flowing fires is slow. Two-dimensional video surveillance relies on human eyes, which is susceptible to visual fatigue and misjudgment, and it lacks depth and temperature information.

[0004] The response chain is lengthy and inefficient: from discovery, confirmation, reporting to deployment, the traditional process takes several minutes, missing the golden window for initial fire suppression (usually <3 minutes). Fixed sprinkler systems provide indiscriminate coverage, which can easily cause secondary water damage and cannot target the core of the fire source.

[0005] High personnel safety risks: Firefighters entering unknown fire scenes characterized by high temperatures, dense smoke, and potential collapse face enormous threats to their lives. Manual reconnaissance is inefficient and makes it difficult to quickly obtain comprehensive information about the interior of the fire scene.

[0006] 1.2 Technical Bottlenecks of Existing Firefighting Robots Most firefighting robots currently on the market are "remotely controlled mobile platforms with basic functions," far from being true intelligent agents. "Remote-controlled extension" rather than "embodied intelligence": It relies heavily on back-end operators and faces problems such as video delay, complex operation, and poor sense of presence. The operator's burden is extremely heavy, and it is difficult for one person to operate multiple devices at the same time.

[0007] Superficial perception, lack of understanding: Sensors are piled up (visible light, simple thermal imager) but lack deep fusion and semantic understanding, making it impossible to recognize "what fire is, where it spreads, and what dangerous objects are nearby".

[0008] Rigid behavior and inability to adapt: ​​The "patrol-discover-spray" logic based on preset procedures is unable to cope with dynamic fire scenes (sudden changes in fire intensity, falling obstacles, multiple fire sources) and complex tasks (demolition, search and rescue, coordination).

[0009] The system is isolated and unable to coordinate: it operates as a standalone unit and lacks real-time intelligent interaction and task coordination with the building's fire protection system (alarm control panel, smoke exhaust valve), other robots, and the command center. Summary of the Invention

[0010] To address the technical problems existing in the prior art, the present invention provides the following technical solution: On the one hand, an intelligent firefighting robot integrating multimodal embodied intelligence is provided, including: The embodied sensing unit includes a multispectral vision head, a solid-state lidar, a multi-gas sensor array, and an acoustic and positioning module, used for high-density acquisition of geometric, temperature, gas, and acoustic information of the environment; The edge intelligence hub, including the main computing platform and real-time control unit, is used to run multimodal perception fusion, neural implicit cognition, hierarchical decision-making, and whole-body motion planning algorithms; The highly mobile execution unit includes an omnidirectional mobile chassis and a multi-degree-of-freedom working arm. The working arm integrates a six-dimensional force sensor at its end and supports the rapid replacement of modular end tools. Collaborative communication and energy units, including multimode communication gateways and smart energy systems; The embodied perception unit, the highly mobile execution unit, and the collaborative communication and energy unit are all connected to the edge intelligence hub to form a closed loop of perception, decision-making, and execution.

[0011] Preferably, the edge intelligence hub operates a multimodal embodied perception and mapping model, which is a neural implicit representation-based network. Its input is a spatiotemporal query point (x, y, z, t), and its output is the color c, volume density σ, temperature T, semantic feature vector s, and instantaneous velocity v of the point. This model is used to construct and update a four-dimensional environmental representation containing geometric, semantic, and physical attributes online. Wherein, (x, y, z): the three-dimensional spatial coordinates (world coordinate system) of the query point. t: Query time (time stamp relative to the start of the task). Introducing time t enables the model to express dynamic changes.

[0012] Preferably, the edge intelligence hub also runs a hierarchical deep reinforcement learning decision-maker, which includes a high-level policy network and a low-level policy network; the high-level policy network outputs high-level skill instructions and target states every c steps based on the observation vector output by the four-dimensional environment representation; the low-level policy network receives the observation vector and the target state at each time step and outputs the intention of continuous control instructions at the low level.

[0013] Preferably, the edge intelligent hub also operates a whole-body motion model predictive controller. Based on the high-level skill instructions, the intent of the low-level control instructions, and the dynamic constraints provided by the four-dimensional environmental representation, the controller solves a finite-time-domain optimization problem in each control cycle to generate optimal control instructions that satisfy dynamic constraints, actuator amplitude limits, state ranges, and collision avoidance constraints, and sends them to the high-mobility execution unit.

[0014] Preferably, the objective function of the optimization problem of the whole-body motion model predictive controller is: , Definitions of each item: N: Prediction time domain length, representing the state and input that MPC predicts for the next N steps within each control cycle; Predict the system state vector at the k-th step in the time domain ( p is the state dimension, which typically includes q and ,Right now T represents the transpose, describing the motion state (position, velocity) of the system at step k. Similarly; : The reference state vector at step k ( (This is generated by the high-level intent and trajectory generator, and represents the ideal state that the system should track.) Similarly;

[0015] : The control input vector at step k ( ), that is, the actuator output (or its increment), is the optimization variable for MPC; Q: State error weight matrix ( (positive definite symmetric), the diagonal elements represent the importance of the corresponding state error; R: Control input weight matrix ( (positive definite symmetric), the diagonal elements represent the penalty intensity of the corresponding control input; P: Terminal state error weight matrix ( (positive definite symmetric), used to penalize the last step of prediction in the time domain ( The state error is controlled to ensure that the terminal state is close to the reference. Weighted norm, defined as , is used to quantize the "cost" of vector z, where the cost is the error or the input size.

[0016] Preferably, the collaborative communication and energy unit runs a multi-agent distributed online task allocation algorithm. The algorithm is based on a consensus-based bundled auction mechanism, which enables multiple robots to exchange task bids and winning lists through a communication network, thereby achieving a consensus on task allocation in a distributed manner and maximizing the overall net utility of the system.

[0017] Preferably, in the execution of the multi-agent distributed online task allocation algorithm, the robot For the task bid It equals the task value minus the execution cost, that is: , Among them, the definitions are as follows: :robot For the task The bid; :Task Value; :robot Execute the task The cost; During the auction phase, the robot calculates... Evaluate the net benefit (value - cost) of performing the task; the higher the bid, the greater the net benefit for the robot to perform the task and the higher the demand; the robot will include the task with the highest bid in the winning list. The list and corresponding bids are broadcast for subsequent conflict resolution; bots resolve allocation conflicts by comparing their bids for the same task, with the highest bidder winning the task. Consensus Phase: When the robot Received from neighbor After receiving the information, conflict resolution is performed by the instruction function (robot). with neighbors Between the tasks Conflict resolution indicator variable Determining the bid: If both parties are interested in the same task... If a bid is placed, then compare bids. (j) and (j) The highest bidder wins the task, and the lowest bidder removes it from their list; through multiple rounds of communication, all robots eventually reach a consensus on task allocation; the expression formula for the indicator function is: , in, This is an indicator function (it takes the value 1 if the condition is true, and 0 otherwise). , They are respectively , For the task The bid (i.e.) When on the winning list, ); Definitions of each item: : and Regarding the task Conflict resolution indicator variables; : Indicator function, its value is 1 when the internal condition is true, and 0 otherwise; : For the task Bids (from the auction phase) ); : For the task The bid; like and All for the same task If a bid is made (task conflict), then through Determine if the bid is high or low: like ( ),but Keep this task. Need to Removed from the list of winners; like ( ),but Keep this task. Need to be removed ; Through exchange With this type of information, the robot gradually resolves conflicts and eventually reaches a consensus on task allocation; the robot resolves conflicts by exchanging this information.

[0018] Preferably, the embodied perception unit is connected to the edge intelligent hub via a high-speed image bus and a medium-to-low-speed device bus; the edge intelligent hub is connected to the high-mobility execution unit via a real-time Ethernet bus; and the collaborative communication and energy unit is connected to the edge intelligent hub via a high-speed peripheral component interconnection bus.

[0019] Preferably, the modular end-effector includes a variable flow fire extinguishing nozzle, a breaching tool set, and a multi-functional gripper.

[0020] On the other hand, an application control method based on the above-described intelligent fire-fighting robot is provided, comprising the following steps: Step 1: Environmental Prior Learning and Autonomous Mapping: After the robot is deployed for the first time, it autonomously explores the environment and uses the multimodal embodied perception and mapping model to fuse multi-source sensor data online, constructing and storing a high-precision neural implicit map containing geometric and semantic attributes; Step 2: Routine Autonomous Inspection and Multimodal Anomaly Perception: The robot receives inspection tasks, plans its path based on the hierarchical deep reinforcement learning decision-maker, and moves autonomously. At the same time, it analyzes the perception data in real time through the multimodal embodied perception and mapping model to detect flames, abnormal temperatures, smoke characteristics, and abnormal gas concentrations. Step 3: Close-up reconnaissance, situational understanding and autonomous decision-making: The robot moves to the vicinity of the anomaly point and performs a multi-angle fine scan. It generates a structured fire situation report through the multimodal embodied perception and mapping model. The hierarchical deep reinforcement learning decision-maker autonomously decides to initiate fire extinguishing, fire control or rescue operations based on the report. Step 4: Multi-robot collaboration and dexterous operation: Multiple robots negotiate and allocate tasks through the multi-agent distributed online task allocation algorithm. Each robot plans a collaborative, collision-free motion trajectory based on the whole-body motion model predictive controller and controls the multi-degree-of-freedom working arm to perform fire extinguishing, demolition, or search and rescue operations. Step 5: Post-event handling, data feedback and model evolution: After the operation is completed, the robot performs an ember scan, generates an action data package and uploads it to the cloud platform. The cloud platform uses this data to train and optimize the core algorithm model offline and sends the updated model parameters to the robot.

[0021] On the other hand, an electronic device is provided, comprising: a processor; and a memory storing computer-readable instructions, which, when executed by the processor, implement the method described above.

[0022] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction is stored therein, the at least one instruction being loaded and executed by a processor to implement the above method.

[0023] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: 1. From “passive threshold response” to “active scene understanding”: Based on the neural implicit representation of MEM-Net, the robot constructs a four-dimensional internal model of the environment that includes geometric, semantic and physical attributes, realizing a deep understanding of “what, why and how” the fire scene is, surpassing the simple “exceeding the standard alarm” mode of traditional sensors.

[0024] 2. From “Remote Control Tool” to “Autonomous Intelligent Agent”: The H-DRL hierarchical decision-making framework endows robots with long-term task planning and dynamic decision-making capabilities. Like an experienced commander, it can decompose tasks, assess risks, and adjust strategies, truly achieving “unmanned” autonomous operation in complex and dynamic environments, thus freeing up human resources.

[0025] 3. From "Wheeled Platform" to "Hollow Manipulator": The WB-MPC full-body motion control system unifies the planning of the chassis and robotic arm, achieving a balance between high mobility and stability in unstructured terrain, as well as dexterous and force-controlled tool manipulation. The robot can not only "move" but also "operate well," possessing the physical capabilities to perform delicate tasks such as demolition and valve closure.

[0026] 4. From "Individual Combat" to "Swarm Intelligence": The DOTA distributed collaborative algorithm enables multi-robot systems to possess self-organizing and flexible collaborative capabilities. Like a swarm of bees, they can dynamically allocate tasks based on the global situation and collaboratively complete complex firefighting and rescue tasks that a single robot cannot accomplish, resulting in a qualitative improvement in the overall system efficiency and robustness.

[0027] 5. From "Factory-Determined" to "Continuous Evolution": The collaborative learning loop between the "edge and cloud" enables the entire system to learn from every real-world test and simulation, continuously optimizing its perception, decision-making, and control models. This means that the capabilities of firefighting robots will grow over time, adapting to new building types and new fire scenarios, becoming a true "expert system."

[0028] 6. Comprehensive Enhancement of Firefighting Efficiency and Safety: This solution frees firefighters from dangerous and repetitive tasks, allowing them to focus on higher-level command and decision-making. Through early and accurate detection, rapid autonomous response, and efficient collaborative operations, fire losses can be minimized while maximizing the safety of personnel, representing the inevitable future direction of smart firefighting. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of the hardware system composition provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the software system architecture provided in an embodiment of the present invention; Figure 3This is a schematic diagram of the software and hardware interaction control logic provided in the embodiments of the present invention; Figure 4 This is a schematic diagram of the implementation process of the application control method provided in the embodiments of the present invention. Detailed Implementation

[0031] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0032] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0033] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0034] In this embodiment of the invention, sometimes a subscript such as W1 may be mistakenly written as a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0035] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0036] I. The Introduction and Inevitability of the Embodied AI Paradigm Embodied intelligence emphasizes that advanced intelligence originates from the perception-action loop generated by an agent's continuous interaction with the physical environment through its body (the robot itself). Introducing embodied intelligence into firefighting robots aims to create an intelligent entity with "human-like" environmental awareness, autonomous decision-making, and dexterous operation capabilities. Its core characteristics are: 1. Active perception and scene understanding: Actively explore through multimodal sensors to construct an internal model of the environment with physical and semantic attributes.

[0037] 2. Value-based autonomous decision-making: Based on internal models and mission objectives (firefighting, rescue, property protection), plan and dynamically adjust action sequences in real time.

[0038] 3. Body-environment interaction and learning: Through the execution-feedback loop, learn how to use your own "body" (robotic arm, mobile chassis) to interact with the environment more effectively, and accumulate experience to optimize future behavior.

[0039] This solution aims to deeply integrate cutting-edge technologies such as multimodal perception fusion, hierarchical reinforcement learning, whole-body motion planning, and multi-agent collaboration to design and implement a truly embodied intelligent mobile fire-fighting robot (an intelligent fire-fighting robot integrating multimodal embodied intelligence and its application control methods), promoting a paradigm shift in fire-fighting operations from "human experience-driven" to "machine autonomous intelligence".

[0040] II. Overall Hardware System Architecture for Intelligent Anti-Robot System Integrating Multimodal Embodied Intelligence 2.1 Hardware System Composition The system adopts an integrated design of "perception-thinking-action-collaboration", and the hardware architecture is as follows: Figure 1 As shown, it comprises four core units: The embodied sensing unit, acting as the robot's "senses," collects high-density, multi-dimensional environmental physical information (geometry, temperature, gas, sound) and achieves precise self-positioning.

[0041] 2.1.1 The embodied perception unit includes: 1. Multispectral vision head: RGB-D camera + mid-wave infrared thermal imager (MWIR), rigidly connected and calibrated.

[0042] 2. Solid-state lidar (LiDAR): 128 lines, 360° horizontal field of view.

[0043] 3. Multi-gas sensor array: CO2, CO, VOC, O2 sensors, distributed arrangement.

[0044] 4. Acoustics and positioning module: microphone array, high-precision IMU, UWB indoor positioning tag / RTK-GNSS dual antenna.

[0045] 2.2.2 The edge intelligence hub, acting as the robot's "brain," runs all AI algorithms, completing the entire process from raw data fusion to advanced decision generation. This includes: (1) Main computing platform: CPU+GPU heterogeneous architecture (such as NVIDIA Jetson AGX Orin), computing power ≥200TOPS (INT8). (2) Real-time control unit: based on FPGA or high-performance MCU, running 1kHz servo loop. (3) High-speed storage: NVMe SSD, used to cache perception data and operation logs.

[0046] 2.1.3 The highly mobile actuator, acting as the robot's "body," provides flexible movement in complex terrain and dexterous manipulation of objects in a fire. This includes: (1) Omnidirectional mobile chassis: Four-wheel independent drive steering (4WIS / 4WID) or omnidirectional wheels, obstacle crossing height ≥200mm, with active suspension. (2) Multi-degree-of-freedom working arm: 6-7 degree-of-freedom humanoid robotic arm, load ≥15kg, with a six-dimensional force sensor integrated at the end. (3) Modular end tools: quick-change interface supports variable flow fire extinguishing nozzles, demolition tool sets (hydraulic shears, spreaders), and multi-functional grippers.

[0047] 2.1.4 Collaborative communication and energy unit, ensuring real-time and reliable transmission of information flow between the robot and external systems, as well as energy supply for long-term autonomous operation. Includes: (1) Multimode communication gateway: supports TSN (Time Sensitive Networking), 5G (uRLLC slicing), and Wi-Fi 6. (2) Smart energy system: high energy density lithium battery, wireless charging receiver module, and smart BMS (Battery Management System).

[0048] 2.2 Hardware-to-Hardware Data Communication Connection Methods Seamless communication between hardware units is achieved through a highly reliable industrial bus and real-time Ethernet: the multispectral vision head, lidar, and other devices of the embodied sensing unit are directly connected to the main computing platform of the edge intelligent hub through the MIPI-CSI-2 interface, and the gas sensor array and acoustic module aggregate data through the RS-485 bus; the edge intelligent hub and the highly mobile execution unit use the EtherCAT real-time bus (cycle ≤1ms) to transmit control commands, and at the same time realize the coordinated motion control of the chassis and the robotic arm through the CANopen protocol; the collaborative communication and energy unit is interconnected with the main computing platform through the PCIe 4.0 interface to provide unified communication scheduling and energy management for the entire system, forming a closed-loop data link of "perception-decision-execution".

[0049] To ensure the real-time performance, determinism, and reliability of the data stream, an example architecture using a layered hybrid network is described below: 1) High-speed internal bus of the sensor: Connect: RGB-D camera, infrared thermal imager → CoaXPress 2.0 or GigE Vision → Edge Intelligence Hub.

[0050] Features: Provides ultra-high-speed (>25 Gbps), low-latency image data streams for real-time vision processing.

[0051] 2). Real-time control network: Connection: Edge intelligent hub (real-time control unit) → EtherCAT daisy chain → chassis hub motor drivers, robotic arm joint servo drivers, end effector I / O modules.

[0052] Features: Microsecond-level synchronization cycle (typically 1ms), enabling high-precision, coordinated motion control of all actuators. Employing a "fly-read / fly-write" mechanism, the master station broadcasts control data, and each slave station synchronously reads and transmits its status back.

[0053] 3). Medium and low speed devices and state networks: Connection: LiDAR, gas sensor, IMU, BMS, ultrasonic loop → CAN FD → edge intelligent hub (or via gateway conversion).

[0054] Features: High reliability, used for transmitting periodic or event data such as point clouds, sensor readings, and system status.

[0055] 4) Coordination and Command Network: Internally: The computing units are interconnected through TSN switches to allocate priorities and bandwidth guarantees for different service flows (such as sensing flow, control flow, and cooperative signaling flow).

[0056] Externally: Access to the fire command cloud platform and other intelligent agents via 5G / Wi-Fi 6. Utilizing ROS 2 over DDS middleware, it enables topic publishing / subscription and service calls, supporting distributed, high-real-time communication.

[0057] III. Software and Algorithm Architecture and Core Model Software systems are the carriers of embodied intelligence, and their architecture is as follows: Figure 2 As shown, it follows a progressive logic of "perception-cognition-decision-planning-control" and is tightly coupled with hardware. The architecture consists of the following five layers: 1. Multimodal perception fusion layer: Integrates multi-source data such as vision, LiDAR, and gas sensing, and constructs the original environmental perception tensor through spatiotemporal registration and feature fusion algorithms.

[0058] 2. Neural Implicit Cognitive Layer: Based on MEM-Net (a multimodal embodied perception and mapping model built on memory-enhanced networks), a four-dimensional environmental representation is constructed to achieve unified modeling of geometric structure, semantic attributes (fire source, obstacles, hazardous materials) and physical state (temperature field, gas concentration).

[0059] 3. Hierarchical Decision Layer: The H-DRL (Hierarchical Deep Reinforcement Learning) framework is adopted. The upper policy network is responsible for task planning (such as fire extinguishing priority ranking), and the lower execution network generates specific action sequences (such as path selection and tool switching).

[0060] 4. Full-body motion planning layer: Through the WB-MPC (whole-body model predictive control) algorithm, the motion trajectory of the mobile chassis and the robotic arm is optimized in a coordinated manner to ensure high mobility and operational accuracy in complex terrain.

[0061] 5. Real-time control execution layer: Based on FPGA+MCU heterogeneous architecture, it realizes 1kHz high bandwidth servo control, supports force control operation (such as demolition force feedback) and smooth motion transition.

[0062] The core algorithm principles of the software system will be described in detail below.

[0063] (I) Multimodal Embodied Perception and Mapping Model (MEM-Net) Objective: To fuse multi-source heterogeneous sensor data from robots to generate a four-dimensional spatiotemporal environment representation (4D Scene Representation) centered on the robot and incorporating geometric, semantic, and physical attributes. This is the cognitive foundation for realizing embodied intelligence.

[0064] Input and preprocessing: Synchronous acquisition: RGB image Depth map D, infrared temperature matrix Laser point cloud Odometer IMU data.

[0065] Preprocessing: Spatiotemporal alignment (based on calibration results and hardware synchronization signals), coordinate system aligned with robot base coordinate system.

[0066] Model structure and operating mechanism: MEM-Net is an online incremental neural implicit representation model, the core of which is training a multilayer perceptron (MLP). Its input is a spatiotemporal query point (x, y, z, t), and its output is the multimodal properties of that point: , Definitions of each term in the formula: (x, y, z): The three-dimensional spatial coordinates (world coordinate system) of the query point; t: Query time (time stamp relative to the start of the task). Introducing time t enables the model to express dynamic changes (such as the spread of fire, the opening and closing of doors). c = (r, g, b): The RGB color value of this point; σ: Volume Density, representing the probability of matter existing at that point, used for subsequent rendering and geometry reconstruction; T: Absolute temperature at that point (°C); s ∈ RK: A K-dimensional semantic feature vector, which learns information such as object category (wall, door, flame, gas cylinder, person) through training; The instantaneous three-dimensional velocity of this point (for a dynamic object) is initially zero.

[0067] Operating mechanism: 1) Encoding and Mapping: The query point (x, y, z) is encoded using positional encoding. Mapping to a higher-dimensional space allows MLPs to better learn high-frequency details:

[0068] ], γ(p): Position encoding function, outputting a feature vector of dimension 2L; : Input three-dimensional spatial coordinates (x, y, z); L: Number of encoding layers, controlling the resolution of high-frequency features; , , ..., Exponentially increasing frequency coefficients enable multi-scale feature encoding; Operating mechanism: By encoding the input coordinate p with sine and cosine functions of different frequencies, the low-dimensional spatial coordinates are mapped to the high-dimensional feature space, enabling the neural network to learn the positional dependencies and high-frequency detailed features of the input data.

[0069] 2). Network inference: Input the encoded spatiotemporal points into the MLP. Through forward propagation of the network, the multiple attributes of the point are output.

[0070] 3). Volume rendering and map generation: RGB-D rendering: Multiple points are sampled along the sensor projection line, and the color and depth of each pixel are synthesized using a volume rendering formula. The loss is calculated by comparing this data with sensor observations, and backpropagation is used to optimize the network parameters θ. The volume rendering formula is: Ĉ(r) = Σi=1N*Ti· (1 - exp(-σiδi)) · ci, where: Ĉ(r): the color of the rendered pixel corresponding to ray r; Ti = exp(-Σj=1i-1σjδj): the transmittance of the i-th sampling point, representing the proportion of light that is not absorbed when it travels from the starting point to the i-th point; σi: the volume density of the i-th sampling point (unit: m-1), characterizing the medium's ability to absorb / scatter light; δi: the distance between adjacent sampling points (unit: m); ci: the color value (RGB three channels) of the i-th sampling point.

[0071] Operating mechanism: N points are sampled from near to far along the sensor ray r. The final rendering result is synthesized by accumulating the contribution value of each sampled point to the pixel color. This process simulates the attenuation law of light in a medium, realizing the mapping from a 3D scene to a 2D image.

[0072] Semantic / Temperature Map Generation: Also using volume rendering, semantic segmentation maps, temperature field heatmaps, and occupancy grid maps can be rendered from any viewpoint. For example, pixel categories can be obtained by processing the rendered semantic feature vector 's' using the argmax operation.

[0073] The model consumes raw data streams from the embodied perception unit in real time. As the core of perception and layer building, its output implicit representation is shared by all subsequent modules. For example, the decision module can query... To obtain the "category and temperature of the object 5 meters ahead"; the planning module can query fθ to obtain occupancy information and the speed v of dynamic obstacles.

[0074] MEM-Net constructs a unified, dense, and queryable model of the environment, enabling robots not only to "see" the environment but also to "imagine" and "understand" it from any angle, serving as the cornerstone for subsequent advanced cognition and planning. Its online update capability allows it to adapt to dynamically changing fire scenes.

[0075] (ii) Hierarchical Deep Reinforcement Learning Decision Maker (H-DRL) 1. Objective: Based on the internal environment model built by MEM-Net, autonomously generate and optimize long-term behavioral strategies to complete complex fire-fighting tasks (such as "inspecting Zone B and dealing with all fire sources").

[0076] 2. Problem Modeling: The robot's decision-making process is modeled as a partially observable Markov decision process (POMDP). The specific modeling process is as follows: 1) State Space (S) Definition: Includes environmental state and robot body state. The environmental state covers fire source location / temperature field distribution / obstacle dynamics (such as collapse risk) / hazardous material type. The robot state includes pose (x,y,z,θ), battery level (0-100%), tool state (water gun / demolition tool / mechanical gripper), and task progress (fire extinguishing / search and rescue / reconnaissance completion rate); 2) Motion Space (A) Design: Divided into high-level strategic actions (task switching / target selection) and low-level execution actions (chassis movement speed (v,ω) / robotic arm joint angle / tool ​​switching quantity), it contains a total of 27-dimensional continuous-discrete hybrid motion space; Action A can be divided into two layers: High-level skills Abstract actions lasting c steps (e.g., c=50, corresponding to 5 seconds), such as Navigate To (the robot uses SLAM-based environmental mapping and real-time localization to autonomously plan paths and move to target locations (e.g., fire source coordinates, trapped personnel location, or designated task area), supporting dynamic obstacle avoidance and terrain adaptation), Inspect (using multimodal sensors (RGB-D camera, infrared thermal imager, gas detector) to perform omnidirectional scanning of target objects, extracting their geometric dimensions, surface temperature, material properties, and semantic categories (e.g., identifying them as "liquefied gas tank," "electrical distribution box," or "trapped personnel"), generating structured descriptive information), Extinguish (adaptively selecting extinguishing methods (water gun / dry powder / foam) based on fire source type (electrical fire / oil fire / solid fire), adjusting nozzle angle and flow rate through visual servo control of the robotic arm, and achieving precise extinguishing of the core fire area by combining temperature field feedback), and Open. Door (identifies door type via vision (hinged door / sliding door / fire door), calls on a robotic arm to perform corresponding operations (rotate handle / press push rod / break lock), supports force control sensing to prevent door jamming or structural damage), Fetch Tool (autonomously selects specified equipment (hydraulic shears / life detector / thermal imager) from the tool library according to current task requirements (such as demolition / rescue / inspection), completes the grabbing, replacement and locking process through the end effector, supports tool status self-check).

[0077] Low-level actions The specific control commands executed at each time step, such as chassis speed. (Rotational speed and angular velocity in the xy direction), robotic arm joint increment Δq, tool commands (switch, flow rate).

[0078] 3). Construction of the observation model (O|S,A): Based on the MEM-Net neural implicit representation, the multimodal sensor data (RGB-D image / laser point cloud / thermal imaging / gas concentration) is rendered into a 238-dimensional observation vector. The system includes a semantic segmentation map (128×128×8 channels), a temperature field heatmap (64×64 resolution), and robot body state parameters; observation O: at time t, the observation vector consists of the local semantic map and the robot's own state (position, charge, tool state), acquired by the robot through MEM-Net. Since the robot cannot obtain a globally perfect state, its observations... The rendering and understanding of the local environment from MEM-Net is achieved through the following specific acquisition process: 1. Multimodal sensor data acquisition: The multispectral vision head (RGB-D image + infrared thermal imaging) of the embodied perception unit, 128-line LiDAR point cloud, distributed gas sensing array (CO2 / CO / VOC concentration), and acoustic positioning module (microphone array + IMU) are used to synchronously acquire raw data at a frequency of 10Hz; 2. Spatiotemporal registration and feature extraction: Through timestamp alignment (±0.5ms accuracy) and spatial calibration (extrinsic parameter matrix correction), multi-source data are mapped to a unified coordinate system to extract low-level features such as color texture, depth contour, thermal field distribution, and gas concentration gradient; 3. MEM-Net dynamic modeling: The neural implicit network constructs a four-dimensional environmental representation based on historical time-series data (sliding window 10s), and renders in real time the semantic segmentation map (fire source / obstacle / hazardous material classification), temperature field heat map (accuracy ±1℃), and three-dimensional occupancy grid (resolution 0.1m×0.1m×0.1m) from the current perspective; 4. Observation state encapsulation: The rendering results are fused with the robot's body state (pose, battery level, tool state) to form an observation vector containing 238-dimensional features. This serves as the input for the Hierarchical Decision Layer (H-DRL).

[0079] 4). Transition model (S'|S,A) learning: Temporal difference network (TDN) is used to predict dynamic changes in the environment, and LSTM network is used to capture state transition probabilities such as fire spread (based on heat conduction equation) and obstacle movement (such as the trajectory of falling objects); 5) Reward Function (R(S,A)) Design: The reward R is carefully designed to guide the agent to learn effective skills, taking into account the task benefits (firefighting efficiency + search and rescue success rate), resource consumption (power loss + water consumption) and safety costs (collision risk + high temperature damage). , Definitions of each term in the formula: Large rewards for completing sub-tasks (e.g., +10 for reaching the target point, +100 for extinguishing a fire). Penalties for entering high-risk areas (high temperature, toxic gas areas); Penalties for violating physical constraints (collision, exceeding limits); , : These are the coefficients.

[0080] For example, a sparse reward mechanism is used here: R = 100 × (number of fires extinguished) - 0.1 × (distance traveled) - 50 × (number of collisions) - 10 × (duration of temperature exceeding the limit).

[0081] 3. Algorithm and Training Mechanism (HIRO Variant): The HIRO (Data-Efficient Hierarchical Reinforcement Learning) framework is adopted, which includes two policy networks (hierarchical policy networks): High-level strategy π_high: Every c steps, based on the current observation Output a high-level skill and a specific target state (For example, every 50 steps (5 seconds) based on observations) Output abstract skills (such as NavigateTo) and target state. (e.g., coordinate features), optimize long-term task planning through MDP. Objective Provides direction for underlying strategies.

[0082] The underlying strategy π_low: at each time step t, receives observations. and high-level goals Output underlying actions The PPO algorithm outputs continuous control commands (such as chassis speed) to optimize short-term action sequences; its goal is to maximize the cumulative reward over the next c steps from the current moment.

[0083] Offline Experience Replay and Target Recalibration: HIRO's key innovation lies in its ability to effectively utilize offline experience by employing a Priority Experience Replay Pool (PER) to store old experiences. , , , The sampling weights are dynamically adjusted according to TD-error. During training, the old experience stored in the replay buffer ( , , , Based on the improved underlying strategy, a more suitable high-level target will be determined through backpropagation. Then use ( , , , This allows for updating high-level strategies. It addresses the discrepancy between the target and the actual trajectory, increasing sample utilization by 300%, which significantly improves sample efficiency.

[0084] 4. Layered training process: Pre-training: Complete 106 unsupervised training steps in a digital twin environment to initialize basic motion control capabilities; Fine-tuning: Adopting a course-based learning strategy, the difficulty is gradually increased from simple tasks (fixed-point navigation) to complex tasks (multi-fire source coordinated fire extinguishing), with 5×104 episodes trained for each level of task; Online optimization: After deployment, the strategy is continuously updated through incremental learning. The network is fine-tuned once every 1,000 practical experiences are accumulated to avoid catastrophic amnesia.

[0085] Applications are as follows: The π_high and π_low policy networks are deployed on the GPU of the edge intelligence hub.

[0086] At the decision-making level, the H-DRL module receives observations from MEM-Net. π_high output skills and For the underlying layer, π_low outputs the initial underlying action intent. This intent is then passed to the planning layer for refinement and feasibility checks.

[0087] Technical effects: H-DRL enables robots to "think" like humans, breaking down grand goals into actionable skill sequences and dynamically adjusting plans based on environmental feedback. For example, if trapped people are discovered on the way to the main fire, the robot can autonomously decide to prioritize rescue efforts.

[0088] (III) Whole Body Motion Model Predictive Controller (WB-MPC) 1. Objective: To transform the abstract motion intent output by H-DRL (such as "move to (x,y) and aim the water gun at the base of the flame") into a smooth, collision-free motion trajectory that is coordinated between the chassis and the robotic arm, is dynamically feasible, and is optimized in real time to cope with dynamic changes.

[0089] 2. Control Model and Optimization Problem: The robot's chassis and robotic arm are considered as a unified floating base dynamic system. Its continuous-time dynamic equations can be simplified as follows:

[0090] Definitions of each item: q: Generalized coordinate vector, including chassis pose ( These are the longitudinal and lateral positions and the heading angle, respectively, and the joint angles of the robotic arm (e.g., longitudinal, lateral, and heading angles). ), which describes the overall configuration of the system.

[0091] M(q): System inertia matrix ( (where n is the generalized coordinate dimension), depends on q, is symmetric and positive definite, and reflects the resistance characteristics of the system's mass distribution to acceleration.

[0092] : The combined vector of non-inertial forces (Coriolis force, centripetal force) and gravity term ( ), where the Coriolis force / centripetal force depends on q and (Velocity), the gravity term depends only on q (configuration), and the overall description is the resultant force of non-inertial forces and gravity.

[0093] S: Driver selection matrix ( (m is the number of actuators), a binary matrix (or real matrix), used to select which generalized coordinates are driven by actuators (such as chassis hub motors). Robotic arm joint motor drive ).

[0094] : Actuator output vector ( This includes the driving force / torque of the hub motor and the torque of the robotic arm joint motor.

[0095] Contact point Jacobian matrix ( Assuming the ground is in planar contact, it is usually taken as ,correspond (Contact motion in the direction of contact), describing the relationship between the velocity at the contact point and the generalized coordinate velocity ( , (Location of the contact point).

[0096] Ground contact force vector ( When in planar contact ,correspond The ground reaction force / moment in the direction is a constraint reaction force generated by the interaction between the ground and the chassis.

[0097] External disturbance vector ( For example, the recoil and wind resistance of a robot carrying a water gun are unexpected forces / torques input from outside the system.

[0098] Operating mechanism: This equation describes the force balance of the floating base system (chassis + robotic arm): left: It is the inertial force (the product of acceleration and mass). It is the resultant force of non-inertial force and gravity, and the sum of the two is the "resistance" term of the system's motion.

[0099] right: The driving force / torque (active input) provided to the actuator. It contributes to the generalized force of ground contact force (constraint reaction force, ensuring that the chassis does not penetrate the ground). External disturbances (passive inputs), the sum of which constitutes the "dynamic" term of the system's motion.

[0100] The essence of the equation is that the resultant force of the system's inertial force, non-inertial force, and gravity is equal to the resultant force of the actuator driving force, ground constraint reaction force, and external disturbance, and follows the Newton-Euler laws.

[0101] 3. WB-MPC solves a finite-time domain (e.g., the next 1.5 seconds) optimization problem in each control cycle (e.g., 50ms).

[0102] (1) Construct the objective function: , Definitions of each item: N: Length of the prediction time domain (e.g., (corresponding to a 1.5-second cycle, with each step being 50ms), indicating that MPC predicts the state and input for the next N steps within each control cycle.

[0103] Predict the system state vector at the k-th step in the time domain ( p is the state dimension, which typically includes q and ,Right now T represents the transpose, describing the motion state (position, velocity) of the system at step k. Similarly.

[0104] : The reference state vector at step k ( ), driven by higher-level intent (such as "reaching the target pose") The ideal state that the system should track is generated by the trajectory generator (such as polynomial trajectory, spline trajectory) and the trajectory generator. Similarly.

[0105] : The control input vector at step k ( ), that is, the actuator output (or its increment), is the optimization variable for MPC.

[0106] Q: State error weight matrix ( (positive definite symmetric), the diagonal elements represent the importance of the corresponding state error (e.g., increasing the weight of the x-direction error to emphasize the priority of longitudinal position tracking).

[0107] R: Control input weight matrix ( (positive definite symmetric), the diagonal elements represent the penalty intensity of the corresponding control input (such as increasing the weight of motor torque, suppressing excessive torque output, and avoiding actuator overload).

[0108] P: Terminal state error weight matrix ( (positive definite symmetric), used to penalize the last step of prediction in the time domain ( The state error is reduced to ensure that the terminal state is close to the reference (e.g., P takes the solution of the Riccati equation to ensure closed-loop stability).

[0109] Weighted norm, defined as , used to quantize the “cost” (error or input size) of vector z.

[0110] (2) Operating mechanism: The core of the objective function is to balance three objectives: State tracking accuracy: The deviation between the actual state and the reference state is penalized to make the system track the reference trajectory as closely as possible.

[0111] Controlling input smoothness: Punish excessive control inputs (such as sudden changes in motor torque) to prevent actuator damage or system vibration.

[0112] Terminal state consistency: Ensure that the prediction time domain ends ( The state is close to the reference to prevent "short-sighted" optimization (focusing only on recent errors and ignoring long-term stability).

[0113] Overall objective: To obtain a set of optimal control inputs by minimizing the total cost J. Only input the first step The process is applied to the system, and then repeated in the next control cycle (rolling optimization).

[0114] (3) Constraints MPC optimization must meet the following constraints to ensure the feasibility and safety of control inputs: (3.1) Discretized dynamic constraints:

[0115] in It is a discrete approximation of the continuous dynamic equations (such as Euler's forward method). , To control the cycle; or the Runge-Kutta method), describing how the state changes with the input.

[0116] Operating mechanism: Guarantee the predicted state sequence It conforms to the dynamic characteristics of robots, meaning that the state change at each step follows physical laws (such as "input"). This will lead to Become (”).

[0117] (3.2) Actuator limiting constraint:

[0118] in ( ), ( These represent the minimum and maximum outputs of the actuator (e.g., motor torque limit). N·m, N·m).

[0119] Operating mechanism: Ensure that the control input is within the physical limits of the actuator to avoid overload damage (such as motor burnout due to excessive torque).

[0120] (3.3) State range constraints:

[0121] in ( ), ( These are the allowable ranges for system states (e.g., chassis position constraints). m, m; Robotic arm joint angle constraints: , ).

[0122] Operating mechanism: Ensure that the robot's state is within a safe or feasible range (e.g., the robotic arm does not exceed the joint limits, and the chassis does not drive out of the working area).

[0123] (3.4) Collision avoidance inequality constraints:

[0124] in These are constraint functions used to describe safety requirements (such as collision avoidance, joint constraints, etc.). For example, a collision avoidance constraint:

[0125] In the formula It is the distance between the robot (such as the chassis, the end effector of the robotic arm) and the nearest obstacle (perceived by sensors (LiDAR, cameras)). It is a safety threshold (e.g., 0.5m).

[0126] Operating mechanism: Safety requirements are transformed into mathematical conditions to ensure that the predicted state sequence will not lead to collisions. or violate other security rules (such as) and (To ensure that the joint angle is within the limit).

[0127] 4. Overall operating mechanism of the whole-body motion model predictive controller: WB-MPC (Whole-Body MPC) performs the following steps within each control cycle (e.g., 50ms): State observation: Acquiring the current state through sensors (encoders, IMUs) (Actual state); Reference generation: A reference state for the next N steps is generated by a high-level intent and trajectory generator. ; Optimization Solution: Under the premise of satisfying discrete dynamics constraints, actuator amplitude limiting constraints, state range constraints, and collision avoidance constraints, minimize the objective function J to obtain the optimal control input sequence. ; Input application: Only the first step of the optimal sequence is input. Send the data to the actuators (hub motors, articulated motors) to control the robot's movement; Rolling update: Enter the next control cycle and repeat the above steps (using the new current state). Replace the old (Predict the time domain to roll forward one step). Through this "prediction-optimization-rolling" mechanism, WB-MPC can handle the complex dynamics of floating base systems (chassis + robotic arm coupling) while satisfying multiple constraints (actuator limits, safety obstacle avoidance), achieving high-precision and high-safety motion control.

[0128] The interaction is as follows: WB-MPC runs on the real-time control unit or high-performance CPU core of the edge intelligence hub; Receive the reference trajectory from the decision layer and real-time obstacle distance information from MEM-Net (used to construct collision constraints g_collision); Solved low-level control instructions The data is sent directly to each motor driver via the EtherCAT bus.

[0129] Technical benefits: WB-MPC achieves dynamic balance, active obstacle avoidance, and disturbance resistance throughout the entire body. For example, when moving on rough terrain, it can coordinate the movement of wheels and robotic arms to maintain stability; when using water cannons to extinguish fires, it can actively compensate for the impact of recoil on aiming accuracy.

[0130] (iv) Multi-agent distributed online task allocation algorithm (DOTA) 1. Objective: When multiple robots work together, achieve online, distributed, optimal or near-optimal allocation of tasks (such as fire source points and reconnaissance areas) to maximize overall fire extinguishing efficiency (U).

[0131] 2. Problem Modeling (Distributed Constraint Optimization Problem, DCOP): There are N robots (intelligent agents) and M tasks. Each task... There is a value (Related to fire size and risk level). Each robot Execute the task Costs are required (Related to distance and required resources). The goal is to find an allocation scheme φ: → Maximize the total utility of all robots in completing the task: total net utility (U), while satisfying the capability constraints of each robot.

[0132] Total utility formula (optimization objective): The total net utility (U) of a multi-robot system equals the sum of the values ​​of all tasks minus the total cost of all robots performing their assigned tasks, i.e.: , in, For task allocation scheme, Indicates assignment to the robot A set of tasks.

[0133] Definitions of each item: (U): Total net utility of the system (optimization objective); (M): Number of tasks; (N): Number of robots; :Task Value; :robot Execute the task The cost; Task allocation scheme (mapping each task to a robot); Assigned to the robot A set of tasks.

[0134] Operating Mechanism: (U) is the core optimization objective of the algorithm, aiming to maximize the difference between the total task value and the total execution cost by rationally allocating tasks. Because the total task value... With a fixed value (assuming all tasks need to be assigned), maximizing (U) is equivalent to minimizing the total execution cost, ultimately maximizing the overall fire suppression efficiency.

[0135] 3. Algorithm Mechanism: Consistency-Based Bundling Algorithm (CBBA) Variant: CBBA is a decentralized, distributed algorithm based on market auction principles, consisting of two alternating phases: 1. Auction Phase: Each robot Maintain a list of winning tasks and corresponding bid ; The robot is capable of performing all unassigned tasks. Calculate a bid , For robots Unassigned but capable of performing tasks bid It equals the task value minus the execution cost, that is , Definitions of each item: :robot For the task The bid; :Task The value (related to the size of the fire and the level of risk). :robot Execute the task The cost (related to distance and required resources).

[0136] Operating mechanism: During the auction phase, the robot calculates... Evaluate the net benefit (value - cost) of performing the task. A higher bid indicates a greater net benefit and higher demand for the task. The robot includes the highest-bid task in its winning list. The list and corresponding bids are broadcast for subsequent conflict resolution.

[0137] The robot broadcasts its own (information) through the communication network. , ).

[0138] 2. Consensus Phase: When the robot Received from neighbor After receiving the information, conflict resolution is performed by the instruction function (robot). with neighbors Between the tasks Conflict resolution indicator variable Determine if the bid is high or low: If both parties are working on the same task If a bid is placed, then compare bids. (j) and (j) The highest bidder wins the task, and the lowest bidder removes it from their list.

[0139] Through multiple rounds of communication, all robots eventually reached a consensus on task allocation.

[0140] The expression formula for the indicator function is:

[0141] in, This is an indicator function (it takes the value 1 if the condition is true, and 0 otherwise). , They are respectively , For the task The bid (i.e.) When on the winning list, ).

[0142] Definitions of each item: : and Regarding the task Conflict resolution indicator variables; : Indicator function, its value is 1 when the internal condition is true, and 0 otherwise; : For the task Bids (from the auction phase) ); : For the task The bid.

[0143] Operating mechanism: During the consensus phase, if and All for the same task If a bid is made (task conflict), then through Determine if the bid is high or low: like ( ),but Keep this task. Need to Removed from the list of winners; like ( ),but Keep this task. Need to be removed .

[0144] Through exchange The robot uses this type of information to gradually resolve conflicts and ultimately reach a consensus on task allocation. The robot resolves conflicts by exchanging this information.

[0145] System interaction: The algorithm module is located in the cooperative communication layer.

[0146] Cost per robot The computation relies on the map information provided by its MEM-Net and the state of its own as evaluated by H-DRL.

[0147] Low-latency information exchange is achieved through the 5G / Wi-Fi network of the collaborative communication and energy unit in the form of multicast or multicast (CBBA communication volume is small).

[0148] Technical benefits: Enables self-organization and flexible collaboration within a multi-robot system. When a new fire occurs or a robot malfunctions and exits, the remaining robots can be quickly redistributed, resulting in extremely strong overall system robustness.

[0149] IV. Hardware-Software Interaction Process and Control Logic The robot's operation manifests as a tightly coupled closed loop of "perception-decision-planning-execution," and its interaction process is as follows: Figure 3 As shown: 1. Sensing data: Hardware triggering: A global hardware synchronization signal triggers all sensors (camera, LiDAR, IMU) to collect data simultaneously.

[0150] Drive and Transmission: Raw data is aggregated to the edge intelligent hub via a high-speed bus (CoaXPress, CAN FD).

[0151] Software processing: MEM-Net receives the data stream, performs online neural rendering and map updates, and generates the latest environmental observations. and the implicit model available for querying .

[0152] 2. Decision-making and planning: Decision generation: H-DRL module receives The high-level strategy π_high assesses the current task progress; if new skills are needed, it outputs [the necessary information]. and The underlying strategy π_low is combined. and Output the initial instructions for the desired underlying actions.

[0153] Motion planning: WB-MPC receives desired commands and dynamic environmental constraints from MEM-Net. It is based on the latest robot state. (From encoder and IMU), solve the optimization problem to generate optimal and safe joint / wheel control commands. .

[0154] Collaborative Negotiation (Parallel): The DOTA module runs continuously, exchanging task status and bids with other robots through the communication network, dynamically adjusting its own task list, and feeding back to H-DRL as decision input.

[0155] 3. Control Execution: Command issuance: The control commands calculated by WB-MPC (specifically the target position / torque of each motor) are accurately issued to each servo drive via the EtherCAT bus at a period of 1ms.

[0156] Hardware execution: Servo drivers drive motors, which in turn move the wheels and robotic arm. Six-dimensional force sensors and other components provide real-time data feedback.

[0157] Closed-loop feedback: Encoder, IMU, and force sensor data are transmitted back to the real-time control unit of the edge intelligent hub via EtherCAT and CAN FD, forming a closed-loop control to ensure accurate tracking of WB-MPC commands and monitor anomalies (such as excessive collision force).

[0158] 4. Exception handling and learning loop: If the control loop detects a serious anomaly (such as a collision or motor overheating), it immediately triggers a hardware safety stop and sends an interrupt signal to the decision-making level.

[0159] Daily operational data (observations, actions, rewards) are recorded and packaged periodically.

[0160] When the robot returns to the charging dock or when the network is idle, incremental data is encrypted and uploaded to the cloud. The cloud uses a larger-scale simulation environment and historical data to perform offline reinforcement learning and training on the policy network of H-DRL and the representation capabilities of MEM-Net, generating better model parameters, and then securely distributing updates to each robot via OTA (Over-The-Air) technology to achieve continuous evolution.

[0161] V. Application Control Methods like Figure 4 As shown, the steps are as follows: 5.1 Step 1: Environmental Prior Learning and Autonomous Mapping Implementation Process: After the robot's initial deployment, the operator initiates mapping mode via advanced commands (such as "learn the entire floor"). The robot autonomously explores the unknown environment, while MEM-Net integrates LiDAR point cloud and visual data online to construct a high-precision neural implicit map that includes geometry and semantics (initial recognition of doors, windows, and fire hydrants through a pre-trained model). During this process, the robot autonomously attempts to interact with typical objects (such as gently pushing a door to determine if it can be opened), enriching its understanding of physical interaction attributes.

[0162] Technical principle: By leveraging the online learning capabilities of MEM-Net and the robot's proactive exploration behavior, simultaneous localization and mapping (SLAM) is achieved, and the map is endowed with rich semantic and physical attributes.

[0163] Technical effects: It forms the robot's "innate memory" of the working environment. This map serves as the spatial cognition basis for all subsequent autonomous tasks, supporting precise navigation and task planning.

[0164] 5.2 Step Two: Routine Autonomous Inspection and Multimodal Anomaly Detection Implementation process: (1). Task assignment: The command center or system automatically generates periodic inspection tasks (such as "inspect area A along the preset route every 2 hours").

[0165] (2). Autonomous navigation: H-DRL decomposes the task into a series of NavigateTo skills. WB-MPC controls the robot to move along the optimal path, and MEM-Net continuously updates the environmental information along the way.

[0166] (3). Anomaly detection: Real-time analysis by MEM-Net in both mobile and stationary observations: Visual / infrared: Unplanned heat sources (>60°C and continuing to rise), flame spectral characteristics, smoke texture.

[0167] Gases: CO / VOC concentrations accumulate abnormally in poorly ventilated areas.

[0168] Acoustics: Detects abnormal sound patterns such as glass shattering and short-circuit explosions.

[0169] (4). Once an anomaly with a confidence level higher than the threshold is detected, the location is immediately marked, the inspection is interrupted, and the fire confirmation mode is entered.

[0170] Technical principle: By combining the multimodal perception capabilities of MEM-Net with the task management of H-DRL, the upgrade from "fixed-point monitoring" to "mobile panoramic scanning" is achieved.

[0171] Technical effects: Greatly expands monitoring coverage and significantly improves the probability of early detection and location accuracy of fires (especially smoldering fires).

[0172] 5.3 Step Three: Close-range reconnaissance, situational understanding, and autonomous decision-making Implementation process: (1). The robot moves to a safe position near the anomaly point, adjusts its posture, and uses the multispectral vision head mounted on the robotic arm to perform multi-angle, close-range scanning.

[0173] (2). MEM-Net performs detailed rendering and inference, and outputs a structured fire scene report: flame three-dimensional volume, temperature gradient field, spread direction prediction, identification and status assessment of nearby key targets (chemicals, electrical boxes, potential trapped personnel).

[0174] (3). H-DRL makes independent decisions based on this report: If it is a small, initial fire with no significant risk, the decision is made to initiate "autonomous firefighting".

[0175] If the fire is large or the risks are complex, the decision should be to implement "control first, then extinguish" (such as spraying flame retardant between the fire point and nearby combustibles to create an isolation zone) or "prioritize rescue," while sending a detailed report to the command center and requesting coordination and support.

[0176] Technical principle: The dense and quantifiable environmental understanding provided by MEM-Net enables H-DRL to make decisions based on rich situational information rather than simple rules.

[0177] Technical benefits: Providing firsthand intelligence comparable to that of on-site reconnaissance personnel before firefighters arrive, enabling them to make preliminary and reasonable emergency responses, thus gaining the initiative for subsequent operations.

[0178] 5.4 Step Four: Multi-machine Collaboration and Embossed Dexterous Operation Implementation process: (1) Task allocation: The command center or robots can quickly negotiate and allocate tasks through the DOTA algorithm. For example: Robot A (heavy) is responsible for suppressing the main fire from the front; Robot B (dexterous) is responsible for breaking down ventilation or closing remote valves; Robot C is responsible for transporting fire extinguishing agents or search and rescue.

[0179] (2). Cooperative motion: When planning their own trajectories, each WB-MPC robot treats each other as dynamic obstacles through communication, so as to achieve collision-free cooperative approach and work station.

[0180] (3). Dexterous work: Fire extinguishing: The robot H-DRL controls the robotic arm, which dynamically adjusts the spray angle, distance and flow rate based on the three-dimensional position and temperature field of the flame root fed back by MEM-Net, to achieve "precise spray cooling".

[0181] Demolition: Based on the recognition of the door structure (hinge position) by MEM-Net, WB-MPC controls the robotic arm to operate the demolition tool with appropriate force and angle.

[0182] Search and rescue: Where permitted, use robotic arms to gently remove obstacles or deliver breathing masks to trapped individuals.

[0183] Technical principles: DOTA enables system-level task optimization; WB-MPC enables individual-level secure collaboration; H-DRL and MEM-Net achieve closed-loop functionality and tool flexibility.

[0184] Technical effects: Achieving "swarm intelligence" doubles operational efficiency; robots transform from "mobile gun emplacements" into "multi-functional firefighters" capable of handling complex firefighting and rescue scenarios.

[0185] 5.5 Step Five: Post-event processing, data feedback, and model evolution Implementation process: (1). After the open flame is extinguished, the robot uses an infrared thermal imager to scan the embers to ensure there is no re-ignition point.

[0186] (2). Generate a complete action report and data package for this task (including key decision points, environmental changes, and control effects).

[0187] (3). Autonomously return to the charging dock for resupply.

[0188] (4) The system automatically encrypts and uploads the data to the cloud-based fire protection AI platform. The platform uses this real "clinical data" to conduct large-scale offline reinforcement learning training in a digital twin simulation environment, optimizing the policy network of H-DRL, the feature extraction capabilities of MEM-Net, etc.

[0189] (5) The new model parameters after training and validation are updated to the front-line robot cluster via OTA through a secure channel.

[0190] Technical principle: Construct a closed loop of "edge execution - cloud learning" to transform every practical experience into nourishment for system evolution.

[0191] Technical benefits: The robot system has lifelong learning capabilities. Its fire extinguishing strategies and environmental awareness will be continuously optimized with the accumulation of practical experience, becoming more and more intelligent with use.

[0192] in conclusion: The mobile fire-fighting robot with embodied intelligence described in this solution is not a simple improvement on existing technologies, but a new generation of fire-fighting equipment created by deeply integrating cutting-edge achievements in artificial intelligence, robotics, and control theory. It represents a leap from the separation of "perception and control" to the integration of "perception, cognition, decision-making, and action," a crucial step towards highly intelligent and autonomous fire-fighting operations, and possesses significant social value and application prospects.

[0193] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An intelligent firefighting robot integrating multimodal embodied intelligence, characterized in that, include: The embodied sensing unit includes a multispectral vision head, a solid-state lidar, a multi-gas sensor array, and an acoustic and positioning module, used for high-density acquisition of geometric, temperature, gas, and acoustic information of the environment; The edge intelligence hub, including the main computing platform and real-time control unit, is used to run multimodal perception fusion, neural implicit cognition, hierarchical decision-making, and whole-body motion planning algorithms; The highly mobile execution unit includes an omnidirectional mobile chassis and a multi-degree-of-freedom working arm. The working arm integrates a six-dimensional force sensor at its end and supports the rapid replacement of modular end tools. Collaborative communication and energy units, including multimode communication gateways and smart energy systems; The embodied perception unit, the highly mobile execution unit, and the collaborative communication and energy unit are all connected to the edge intelligence hub to form a closed loop of perception, decision-making, and execution.

2. The intelligent firefighting robot according to claim 1, characterized in that, The edge intelligence hub operates a multimodal embodied perception and mapping model, which is a neural implicit representation-based network. Its input is a spatiotemporal query point (x, y, z, t), and its output is the color c, volume density σ, temperature T, semantic feature vector s, and instantaneous velocity v of that point. This model is used to construct and update a four-dimensional environmental representation that includes geometric, semantic, and physical attributes online. Here, (x, y, z) represents the three-dimensional spatial coordinates (world coordinate system) of the query point. t: Query time (time stamp relative to the start of the task). Introducing time t enables the model to express dynamic changes.

3. The intelligent firefighting robot according to claim 2, characterized in that, The edge intelligence hub also runs a hierarchical deep reinforcement learning decision-maker, which includes a high-level policy network and a low-level policy network. The high-level policy network outputs high-level skill instructions and target states every c steps based on the observation vector output by the four-dimensional environment representation. The low-level policy network receives the observation vector and the target state at each time step and outputs the intention of continuous control instructions at the lower level.

4. The intelligent firefighting robot according to claim 3, characterized in that, The edge intelligent hub also runs a whole-body motion model prediction controller. Based on the high-level skill instructions, the intent of the low-level control instructions, and the dynamic constraints provided by the four-dimensional environmental representation, the controller solves a finite-time domain optimization problem in each control cycle to generate optimal control instructions that satisfy dynamic constraints, actuator amplitude limits, state ranges, and collision avoidance constraints, and sends them to the high-mobility execution unit.

5. The intelligent firefighting robot according to claim 4, characterized in that, The objective function of the optimization problem of the whole-body motion model predictive controller is: , Definitions of each item: N: Prediction time domain length, representing the state and input that MPC predicts for the next N steps within each control cycle; Predict the system state vector at the k-th step in the time domain ( p is the state dimension, which typically includes q and ,Right now T represents the transpose, describing the motion state (position, velocity) of the system at step k. Similarly; : The reference state vector at step k ( (This is generated by the high-level intent and trajectory generator, and represents the ideal state that the system should track.) Similarly; : The control input vector at step k ( ), that is, the actuator output (or its increment), is the optimization variable for MPC; Q: State error weight matrix ( (positive definite symmetric), the diagonal elements represent the importance of the corresponding state error; R: Control input weight matrix ( (positive definite symmetric), the diagonal elements represent the penalty intensity of the corresponding control input; P: Terminal state error weight matrix ( (positive definite symmetric), used to penalize the last step of prediction in the time domain ( The state error is controlled to ensure that the terminal state is close to the reference. Weighted norm, defined as , is used to quantize the "cost" of vector z, where the cost is the error or the input size.

6. The intelligent firefighting robot according to claim 1, characterized in that, The collaborative communication and energy unit operates a multi-agent distributed online task allocation algorithm. The algorithm is based on a consensus-based bundled auction mechanism, which enables multiple robots to exchange task bids and winning lists through a communication network, thereby achieving a consensus on task allocation in a distributed manner and maximizing the overall net utility of the system.

7. The intelligent firefighting robot according to claim 6, characterized in that, In the execution of the multi-agent distributed online task allocation algorithm, the robot For the task bid It equals the task value minus the execution cost, that is: , Among them, the definitions are as follows: :robot For the task The bid; :Task Value; :robot Execute the task The cost; During the auction phase, the robot calculates... Evaluate the net benefit (value - cost) of performing the task; the higher the bid, the greater the net benefit for the robot to perform the task and the higher the demand; the robot will include the task with the highest bid in the winning list. The list and corresponding bids are broadcast for subsequent conflict resolution; bots resolve allocation conflicts by comparing their bids for the same task, with the highest bidder winning the task. Consensus Phase: When the robot Received from neighbor After receiving the information, conflict resolution is performed by the instruction function (robot). with neighbors Between the tasks Conflict resolution indicator variable Determining the bid: If both parties are interested in the same task... If a bid is placed, then compare bids. (j) and (j) The highest bidder wins the task, and the lowest bidder removes it from their list; through multiple rounds of communication, all robots eventually reach a consensus on task allocation; the expression formula for the indicator function is: , in, This is an indicator function (it takes the value 1 if the condition is true, and 0 otherwise). , They are respectively , For the task The bid (i.e.) When on the winning list, ); Definitions of each item: : and Regarding the task Conflict resolution indicator variables; : Indicator function, its value is 1 when the internal condition is true, and 0 otherwise; : For the task Bids (from the auction phase) ); : For the task The bid; like and All for the same task If a bid is made (task conflict), then through Determine if the bid is high or low: like ( ),but Keep this task. Need to Removed from the list of winners; like ( ),but Keep this task. Need to be removed ; Through exchange With this type of information, the robot gradually resolves conflicts and eventually reaches a consensus on task allocation; the robot resolves conflicts by exchanging this information.

8. The intelligent firefighting robot according to claim 1, characterized in that, The embodied perception unit is connected to the edge intelligent hub via a high-speed image bus and a medium-to-low-speed device bus; the edge intelligent hub is connected to the high-mobility execution unit via a real-time Ethernet bus; and the collaborative communication and energy unit is connected to the edge intelligent hub via a high-speed peripheral component interconnection bus.

9. The intelligent firefighting robot according to claim 1, characterized in that, The modular end-effector tools include variable flow fire extinguishing nozzles, breaching tool sets, and multi-functional grippers.

10. An application control method for the intelligent firefighting robot according to any one of claims 1-9, characterized in that, Includes the following steps: Step 1: Environmental Prior Learning and Autonomous Mapping: After the robot is deployed for the first time, it autonomously explores the environment and uses the multimodal embodied perception and mapping model to fuse multi-source sensor data online, constructing and storing a high-precision neural implicit map containing geometric and semantic attributes; Step 2: Routine Autonomous Inspection and Multimodal Anomaly Perception: The robot receives inspection tasks, plans its path based on the hierarchical deep reinforcement learning decision-maker, and moves autonomously. At the same time, it analyzes the perception data in real time through the multimodal embodied perception and mapping model to detect flames, abnormal temperatures, smoke characteristics, and abnormal gas concentrations. Step 3: Close-up reconnaissance, situational understanding and autonomous decision-making: The robot moves to the vicinity of the anomaly point and performs a multi-angle fine scan. It generates a structured fire situation report through the multimodal embodied perception and mapping model. The hierarchical deep reinforcement learning decision-maker autonomously decides to initiate fire extinguishing, fire control or rescue operations based on the report. Step 4: Multi-robot collaboration and dexterous operation: Multiple robots negotiate and allocate tasks through the multi-agent distributed online task allocation algorithm. Each robot plans a collaborative, collision-free motion trajectory based on the whole-body motion model predictive controller and controls the multi-degree-of-freedom working arm to perform fire extinguishing, demolition, or search and rescue operations. Step 5: Post-event handling, data feedback and model evolution: After the operation is completed, the robot performs an ember scan, generates an action data package and uploads it to the cloud platform. The cloud platform uses this data to train and optimize the core algorithm model offline and sends the updated model parameters to the robot.