Intelligent multi-mode patrol sensing method, system and equipment and storage medium

Through the combination of large language model and OWL framework, natural language understanding and multi-modal data fusion are realized, which solves the problem of low accuracy in dynamic task understanding and abnormal recognition in the existing technology, improves the stability and navigation success rate of the inspection system, and reduces operating costs.

CN120337934APending Publication Date: 2025-07-18QINGDAO UNIV

Patent Information

Application Number
CN202510486506.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The prior art cannot realize dynamic task understanding based on natural language, and the lack of efficient multimodal data fusion of perception systems leads to low abnormal recognition accuracy and difficulty in dealing with complex tasks or emergencies.

Method used

The large language model and OWL framework are used to combine multi-agent collaboration and multi-modal expert model integration to realize natural language understanding, autonomous task planning, intelligent navigation and abnormal event recognition, through natural language instruction analysis, task decomposition and allocation, multi-modal information collection and fusion, abnormal identification and event response, result feedback and task closed loop.

Benefits of technology

It improves the stability and scalability of task execution, enhances the system's navigation success rate and task completion rate in an unstructured environment, and reduces operating costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337934A_ABST
    Figure CN120337934A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of multi-mode sensing fusion, and relates to an intelligent multi-mode patrol sensing method, system and equipment and a storage medium, which are used for carrying out automatic patrol, real-time sensing and abnormity warning in a complex environment. Comprising the steps of natural language instruction analysis, task decomposition and distribution, intelligent agent scheduling and execution, multi-modal information collection and fusion, anomaly recognition and event response, result feedback and task closed loop, and through the integration of an OWL framework and a DeepSeek-R1 large model, a multi-modal data fusion technology and an NL-SLAM autonomous navigation algorithm, the multi-modal data fusion algorithm and the multi-modal data fusion technology are integrated. The problems that in the prior art, dynamic task understanding based on natural languages cannot be achieved, a sensing system lacks efficient multi-modal data fusion, so that the anomaly recognition precision is low, and complex tasks or emergencies are difficult to deal with are solved, various abnormal conditions in a complex environment are effectively dealt with, and the method is suitable for popularization and application. The stability and expansibility of task execution are improved, and the navigation success rate and the task completion rate of the system in an unstructured environment are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention belongs to the field of multimodal perception fusion, and particularly relates to a multimodal perception method for an embodied intelligence-based system to perform automatic inspection, real-time perception, and anomaly warning in a complex environment. Background Art:

[0002] With the rapid development of industrial automation and artificial intelligence technologies, inspection technologies relying on movable devices such as robots are increasingly widely used in fields such as electricity, chemical industry, and transportation. Most of the existing automated inspection technologies rely on robot platforms and preset paths for inspection, collecting data through sensors and performing simple analysis. With the improvement of robot hardware performance, more and more systems have started to use SLAM technology (simultaneous localization and mapping) to achieve navigation and path planning, and use visual sensors and infrared thermal imaging technologies for basic anomaly detection. However, the existing technologies still have certain limitations in natural language understanding, task autonomous planning and execution, and the accuracy and efficiency of multimodal data fusion. Especially in a dynamic environment, the system reacts slowly and is difficult to meet the requirements of complex inspection tasks. The existing technologies usually rely on manually preset inspection paths and task instructions, unable to achieve task scheduling and execution based on natural language, and difficult to quickly adjust the execution plan according to on-site requirements; data fusion usually relies on simple parallel processing or traditional computer vision algorithms, failing to fully utilize the spatio-temporal correlations between various sensors, resulting in errors in anomaly recognition and slow reaction speed in a complex environment; unable to perform effective scheduling in multi-task concurrency or complex tasks; the navigation system performs path planning based on a fixed map and cannot adjust actions according to real-time feedback. Especially when facing dynamic obstacles or temporary task instructions, the navigation accuracy and execution efficiency are low.

[0003] In the prior art, Chinese Patent CN119106101A discloses a multi-agent collaborative power inspection method, system, device, and storage medium, which relates to the field of power inspection technology. The method includes using a pre-established prompt word agent to generate a prompt according to a user service instruction; using a pre-established language understanding agent to parse the user service instruction from the prompt; according to the parsed user service instruction, using a pre-established routing agent to perform intelligent routing and decision-making among multiple service agents, and completing the specific task process through a pre-established service agent. However, this technical solution is limited to analyzing and generating prompts for user instructions and languages using pre-established prompt words and languages, and cannot achieve dynamic task understanding based on natural language.

[0004] Chinese Patent CN113190002A discloses a method for an inspection robot of high-speed railway box girders to achieve automatic inspection, including the following steps: S1, generating a map by the radar SLAM algorithm; S2, positioning the inspection robot and planning a travel path on the map; S3, confirming the motion control scheme of the inspection robot; S4, controlling the inspection robot to move; S5, collecting the track movement amount of the inspection robot; S6, detecting internal hidden dangers of the high-speed railway box girder and obtaining the absolute positions of the hidden dangers. However, this technical solution relies on a preset map for autonomous navigation, the path planning is relatively fixed, and manual intervention is required in case of complex situations.

[0005] Chinese Patent CN119693824A discloses an intelligent inspection method for an unmanned aerial vehicle (UAV) used for energy facility inspection. The UAV is equipped with a high-definition camera and a thermal sensor to inspect energy facilities, and visually captured images and thermal imaging data are collected in real time and transmitted to the server side. The server uses a variety of recognition training and deep learning models to process, analyze, classify risks, and detect anomalies for the data. The UAV uses simultaneous localization and mapping technology to obtain three-dimensional information, construct a three-dimensional map, and uses the RRT algorithm to plan the inspection path to achieve multiple functions. Deep learning technology is used to identify potential risks to achieve early warning and automatic obstacle avoidance. The edge computing unit processes data in real time and transmits it to the ground control center for further analysis and prediction of risk trends. However, the overall system architecture of this technical solution is too centralized, lacking the collaboration and task decomposition capabilities among intelligent agents, and it is difficult to handle complex tasks or emergencies.

[0006] The above-mentioned existing technologies have problems such as being unable to achieve dynamic task understanding based on natural language, the lack of efficient multi-modal data fusion in the perception system resulting in low anomaly recognition accuracy, and being difficult to handle complex tasks or emergencies. After the inventor's search and analysis, there is no existing technology that discloses a patrol and perception method that uses a large language model and an OWL framework and realizes natural language understanding, autonomous task planning, intelligent navigation, and anomaly event recognition through multi-agent collaboration and multi-modal expert model fusion. Therefore, inventing an embodied intelligent multi-modal patrol and perception method, system, device, and storage medium can improve the deficiencies of the existing technology, achieve dynamic integrated processing of multi-source information, effectively handle various anomalies in complex environments, improve the stability and scalability of task execution, and enhance the completion rate of complex environment tasks. Summary of the Invention:

[0007] The object of the present invention is to overcome the shortcomings of the prior art. Based on the improvement of the prior art, it is designed to provide a patrol perception method that utilizes large language models and the OWL framework and realizes natural language understanding, autonomous task planning, intelligent navigation, and abnormal event recognition through multi-agent collaboration and multi-modal expert model fusion. In particular, it is an embodied intelligent multi-modal patrol perception method, system, device, and storage medium, which solves the problems existing in the prior art, such as the inability to achieve dynamic task understanding based on natural language, the lack of efficient multi-modal data fusion in the perception system resulting in low abnormal recognition accuracy, and the difficulty in coping with complex tasks or emergencies.

[0008] To achieve the above object, the present invention provides an embodied intelligent multi-modal patrol perception method, including the following steps:

[0009] S1: Natural language instruction parsing: Obtain the inspection task instructions input in natural language form, perform semantic parsing on the natural language description of the inspection task instructions through the integrated large language model, extract key inspection targets, objects, area information, and task logic, and generate a structured task description;

[0010] S2: Task decomposition and allocation: Decompose the structured task through the lightweight scheduling framework OWL, divide it into several subtasks, and allocate each subtask to the corresponding agent; the agents include a navigation agent, a perception agent, a detection agent, a monitoring agent, an image analysis agent, and an acoustic analysis agent;

[0011] S3: Agent scheduling and execution: Each agent executes and completes the corresponding navigation, collection, and analysis subtasks by calling sensors, algorithms, or hardware interfaces in the preset tool pool;

[0012] S4: Multi-modal information collection and fusion: Input the multi-source heterogeneous data collected after completing the collection subtask into the MOE (Mixture of Experts) multi-modal fusion model for analysis and fusion after preprocessing; the MOE model includes multiple dedicated expert sub-models such as an image recognition expert, an infrared analysis expert, and an acoustic anomaly expert, and dynamically allocates computing resources and expert weights according to the current task scenario and data characteristics through a gating network to achieve spatio-temporal registration, weighted fusion, and rapid judgment of multi-modal information;

[0013] S5: Abnormal recognition and event response: Comprehensively judge the fused multi-modal data through the expert output fusion mechanism to identify whether there is abnormal behavior; if abnormal behavior is identified, perform any one or a combination of the following operations: generate an alarm signal, upload it to the remote management platform, notify the user, update the subsequent inspection path or task priority;

[0014] S6: Result feedback and task closed-loop: Transmit the execution results, recognition results, and path status of all subtasks back to the large model decision-making center, which performs global summarization and result generation; after detecting environmental changes or new instructions, re-enter a new round of task scheduling and path update processes.

[0015] The tasks executed by the agent described in step S3 of the present invention specifically include the following steps:

[0016] S301: Use the navigation agent to combine with the NL-SLAM algorithm to perform path planning for the target area, use the NL-SLAM algorithm to combine the natural language parsing results with the environmental map, realize the mapping from language intention to navigation path, and dynamically sense environmental changes during the inspection process to perform dynamic obstacle avoidance and route adjustment;

[0017] S302: Use the perception agent to control the sensors carried by the inspection device to collect data on the target environment; the sensors include visible light cameras, infrared thermal imagers, audio-video devices, and lidar.

[0018] The method for preprocessing multi-source heterogeneous data described in step S4 of the present invention includes:

[0019] S401: Perform data cleaning, use deletion, filling, and marking methods to handle missing values, use threshold filtering, binning smoothing, and model correction to handle outliers, and perform duplicate data deletion to solve problems of noise, missing values, outliers, and duplicates in the data;

[0020] S402: Align the modes through unit conversion, perform redundancy processing through correlation analysis, and perform data union through multi-directional splicing to integrate multi-source data and solve problems of mode conflicts, redundancy, and entity matching;

[0021] S403: Normalize the data through min-max scaling, discretize the data using binning and one-hot encoding, and vectorize unstructured data using text processing, image processing, and time series data to convert the data into a unified format suitable for analysis;

[0022] S404: Perform data reduction through dimensionality reduction, data sampling, and feature selection to reduce the data scale and improve processing efficiency;

[0023] S405: Perform multi-modal data fusion through early fusion, late fusion, and hybrid fusion for semantic association and collaborative analysis of heterogeneous data;

[0024] S406: Unify the storage format and centrally store multi-source data, recording the data source, collection time, version, and preprocessing steps.

[0025] In step S5 involved in the present invention, abnormal behaviors include fire, abnormal sounds, personnel intrusion, and abnormal equipment operation.

[0026] The present invention also relates to an embodied intelligent multi-modal inspection and perception system, which includes a control layer, an execution layer, and an output layer. The control layer includes a large language model and an OWL framework. The execution layer includes intelligent agents and a tool pool. The output layer includes an alarm and feedback system. Among them, an MOE module is integrated in the intelligent agents; the large language model is used to receive natural language instructions, perform task decomposition and understanding, and generate an executable task execution plan to guide the intelligent agents to schedule the tool pool to complete corresponding tasks; the OWL framework is used to coordinate the collaboration between different intelligent agents and the tool pool; the intelligent agents include a navigation intelligent agent, a perception intelligent agent, a detection intelligent agent, a monitoring intelligent agent, an image analysis intelligent agent, and an acoustic analysis intelligent agent, which are used to execute specific tasks such as path navigation, equipment detection, and environmental monitoring. The intelligent agents call their respective tools for specific operations to complete the assigned subtasks. The MOE module is used to receive data collected by sensors and perform data fusion and anomaly detection to identify abnormal situations in the environment; the tool pool is used for the intelligent agents to call the required hardware devices, including sensors, cameras, navigation modules, and inspection devices. The sensors among them include visible light cameras, infrared thermal imagers, audio-video devices, and lidar. The sensors and cameras are used to collect visual, thermal, and sound information of the environment. The navigation module is located inside the inspection device and is used to calculate the best path through the NL-SLAM navigation algorithm and control the movement of the robot according to user instructions and system planning. The inspection device is used to move in the target environment and is equipped with sensors, cameras, and a navigation module; the alarm and feedback system is used for the response after anomaly detection. When the MOE module identifies an abnormal behavior, it triggers an alarm mechanism and feeds it back to the user through the cloud platform or performs emergency processing.

[0027] The large language model involved in the present invention is the DeepSeek-R1 large model.

[0028] The present invention also relates to an electronic device, which includes a processor and a memory. The processor is used to execute a computer program stored in the memory to implement the embodied intelligent multi-modal inspection and perception method.

[0029] The present invention also relates to a computer-readable storage medium, which stores at least one instruction. When the at least one instruction is executed by a processor, the embodied intelligent multi-modal inspection and perception method is implemented.

[0030] Compared with the prior art, the present invention has the following beneficial effects: (1) It solves the problem that the existing system cannot understand natural language instructions, improves the intelligent scheduling ability of the inspection task, enables the system to autonomously generate tasks based on natural language input and complete dynamic scheduling and execution; (2) It improves the fusion efficiency of multi-modal sensor data and the accuracy of anomaly recognition, realizes the dynamic integrated processing of multi-source information, and effectively responds to various abnormal situations in complex environments; (3) It overcomes the problem of poor collaborative ability of the centralized system architecture module, establishes an intelligent agent collaborative system with unified scheduling and distributed execution capabilities, and improves the stability and scalability of task execution; (4) It enhances the autonomy and adaptability of the navigation module, realizes path planning and dynamic adjustment by integrating natural language understanding and environmental perception, and enhances the navigation success rate and task completion rate of the system in unstructured environments; (5) By replacing manual inspection with intelligent inspection, it can effectively reduce the maintenance costs of personnel and equipment, reduce unexpected expenditures on equipment failures and maintenance, and thus significantly reduce the operating costs. Brief Description of the Drawings:

[0031] Figure 1 It is a flowchart showing the process of the embodied intelligent multi-modal inspection and perception method involved in the present invention. Detailed Embodiments:

[0032] The present invention will be further described below with reference to the drawings and embodiments.

[0033] Embodiment 1:

[0034] This embodiment relates to an embodied intelligent multi-modal inspection and perception method, as Figure 1 shown, including the following steps:

[0035] S1: Natural language instruction parsing: Obtain the inspection task instruction input in natural language form, perform semantic parsing on the natural language description of the inspection task instruction through an integrated large language model, extract key inspection targets, objects, area information and task logic, and generate a structured task description;

[0036] S2: Task decomposition and allocation: Decompose the structured task through the lightweight scheduling architecture OWL, divide it into several subtasks, and allocate each subtask to the corresponding intelligent agent; the intelligent agents include a navigation intelligent agent, a perception intelligent agent, a detection intelligent agent, a monitoring intelligent agent, an image analysis intelligent agent, and an acoustic analysis intelligent agent;

[0037] S3: Intelligent agent scheduling and execution: Each intelligent agent calls sensors, algorithms or hardware interfaces in the preset tool pool to execute and complete the corresponding navigation, collection and analysis subtasks;

[0038] S4: Multimodal Information Acquisition and Fusion: The multi-source heterogeneous data collected after completing the acquisition subtasks is preprocessed and then input into the MOE (Mixture of Experts) multimodal fusion model for analysis and fusion; the MOE model includes multiple dedicated expert sub-models such as image recognition experts, infrared analysis experts, and acoustic anomaly experts, and dynamically allocates computing resources and expert weights according to the current task scenario and data features through a gating network to achieve spatio-temporal registration, weighted fusion, and rapid judgment of multimodal information;

[0039] S5: Anomaly Recognition and Event Response: The fused multimodal data is comprehensively judged through an expert output fusion mechanism to identify abnormal behaviors such as fires, abnormal sounds, personnel intrusion, and abnormal equipment operation; if an abnormal behavior is identified, perform any one or a combination of the following operations: generate an alarm signal, upload it to a remote management platform, notify the user, update the subsequent inspection path or task priority;

[0040] S6: Result Feedback and Task Closed-Loop: The execution results, recognition results, and path status of all subtasks are transmitted back to the large model decision-making center, and it performs global summarization and result generation; after identifying environmental changes or new instructions, re-enter a new round of task scheduling and path update processes.

[0041] In this embodiment, the tasks performed by the intelligent agent in step S3 specifically include the following steps:

[0042] S301: Use the navigation intelligent agent to perform path planning on the target area in combination with the NL-SLAM algorithm, and use the NL-SLAM algorithm to combine the natural language parsing results with the environmental map to achieve the mapping from language intention to navigation path, and dynamically perceive environmental changes during the inspection process to perform dynamic obstacle avoidance and route adjustment;

[0043] S302: Use the perception intelligent agent to control the sensors carried by the inspection device to collect data on the target environment; the sensors include visible light cameras, infrared thermal imagers, audio-video devices, and lidar.

[0044] In this embodiment, the method for preprocessing multi-source heterogeneous data in step S4 includes:

[0045] S401: Perform data cleaning, use deletion methods, filling methods, and marking methods to process missing values, use threshold filtering, bin smoothing, and model correction to process outliers, and perform duplicate data deletion to solve problems of noise, missing, outliers, and duplicates in the data;

[0046] S402: Align modes through unit conversion, perform redundancy processing through correlation analysis, and perform data union through multi-directional splicing to integrate multi-source data and solve problems of mode conflicts, redundancy, and entity matching;

[0047] S403: Normalize the data through min-max scaling, discretize the data using binning and one-hot encoding, vectorize the unstructured data using text processing, image processing, and time-series data, and convert the data into a unified format suitable for analysis;

[0048] S404: Reduce the data scale and improve the processing efficiency by performing data reduction through dimensionality reduction, data sampling, and feature selection;

[0049] S405: Perform multimodal data fusion through early fusion, late fusion, and hybrid fusion for semantic association and collaborative analysis of heterogeneous data;

[0050] S406: Unify the storage format and centrally store the multi-source data, recording the data source, collection time, version, and preprocessing steps.

[0051] This embodiment also relates to an embodied intelligent multi-modal inspection and perception system, which includes a control layer, an execution layer, and an output layer. The control layer includes the DeepSeek-R1 large model and the OWL framework. The execution layer includes agents and a tool pool. The output layer includes an alarm and feedback system. Among them, the MOE module is integrated in the agents; the DeepSeek-R1 large model is used to receive natural language instructions, perform task decomposition and understanding, and generate an executable task execution plan to guide the agents to schedule the tool pool to complete corresponding tasks; the OWL framework is used to coordinate the cooperation between different agents and the tool pool; the agents include navigation agents, perception agents, detection agents, monitoring agents, image analysis agents, and acoustic analysis agents, which are used to execute specific tasks such as path navigation, equipment detection, and environmental monitoring. The agents call their respective tools for specific operations to complete the assigned subtasks. For example, the navigation agent is responsible for path planning and obstacle avoidance, and the detection agent uses multi-modal sensors for anomaly detection. The MOE module is used to receive the data collected by the sensors and perform data fusion and anomaly detection. It dynamically routes through a gating network and can intelligently select the most suitable expert model to process the data from each sensor to identify abnormal situations in the environment, such as fires and equipment failures; the tool pool is used for the agents to call the required hardware devices, including sensors, cameras, navigation modules, and inspection devices. The sensors among them include visible light cameras, infrared thermal imagers, audio-video devices, and lidar. The sensors and cameras are used to collect visual, thermal, and sound information of the environment. The navigation module is located inside the inspection device and is used to calculate the best path and control the movement of the robot according to user instructions and system planning through the NL-SLAM navigation algorithm. It can perform real-time path planning and dynamically adjust the walking route according to environmental changes. The inspection device is used to move in the target environment and is equipped with sensors, cameras, and a navigation module; the alarm and feedback system is used for the response after anomaly detection. When the MOE module identifies abnormal behavior, it triggers the alarm mechanism and feedbacks to the user through the cloud platform or performs emergency processing.

[0052] In the embodied intelligent multi-modal inspection and perception system involved in this embodiment, the DeepSeek-R1 large model decomposes and distributes tasks to each agent through the OWL framework. The task execution of each agent is scheduled and coordinated by the OWL framework. These agents call the hardware devices in the tool pool according to the task requirements. For example, the navigation module controls the inspection device to move, and the multi-modal sensors collect environmental data; all the collected data is fused and analyzed by the MOE module. If an anomaly is identified, the alarm and feedback system will trigger an alarm and report it to the user.

[0053] This embodiment also relates to an electronic device, including a processor and a memory. The processor is used to execute the computer program stored in the memory to implement the embodied intelligent multi-modal inspection and perception method.

[0054] This embodiment also relates to a computer-readable storage medium storing at least one instruction, and when the at least one instruction is executed by a processor, the embodied intelligent multi-modal inspection and perception method is implemented.

[0055] Embodiment 2:

[0056] This embodiment relates to the specific working methods of the embodied intelligent multi-modal inspection and perception method, system, device and storage medium described in Embodiment 1, including the following steps:

[0057] 1. Natural language instruction parsing and task decomposition: The user inputs a natural language task instruction through voice or text. The input content should be in Chinese or English. The DeepSeek-R1 large model is used for natural language understanding, parsing the user's intention and generating the corresponding task decomposition. According to the task goal, the system automatically decomposes the task into several subtasks and assigns them to the corresponding agents for execution through the OWL framework; the system can accurately understand and parse the user's instruction and convert it into a specific task executable by the machine;

[0058] 2. Generation of autonomous navigation path based on NL-SLAM: It supports navigation tasks from 10 meters to 800 meters, the navigation map accuracy error ≤ ±10 cm, and supports dynamic obstacle avoidance; through the NL-SLAM algorithm, the system generates the optimal navigation path according to the user's task goal and the environmental map. The system updates the map in real time during navigation, avoids dynamic obstacles, and ensures that the robot can reach the specified target safely and smoothly; the robot can autonomously plan the walking route and dynamically adjust the path according to environmental changes; the navigation success rate in a complex environment ≥ 93.7%, the obstacle avoidance response delay ≤ 1 second, significantly improving the execution efficiency of the inspection task;

[0059] 3. Multi-modal data collection and synchronous processing: The collection conditions are that the video sensor ≥ 15 fps, the infrared image update frequency ≥ 5 Hz, the acoustic wave sensor sampling rate ≥ 44.1 kHz, and the synchronization requirements are that the time synchronization error < 50 ms and the space synchronization error < 2 cm; start the multi-modal sensors carried by the robot, including high-definition cameras, infrared thermal imagers, lidar, microphone arrays, etc., to collect environmental data; all the collected data are aligned in time and space through the synchronization module to ensure no error during data fusion; it can realize multi-angle and multi-dimensional perception of the environment, including visual, thermal and sound data, etc., providing comprehensive information support for subsequent anomaly detection. The system can obtain 360° seamless environmental information coverage, and the synchronization and accuracy of environmental data meet the requirements of subsequent processing;

[0060] 4. Multimodal Data Fusion and Anomaly Recognition: The MOE model contains at least 6 experts in different fields, the call latency of each expert model is < 200ms, and the dynamic routing frequency of the gating network is ≥ 1Hz; The collected multimodal data (images, thermal imaging, sound waves, etc.) is input into the MOE (Mixture of Experts) module; The MOE module dynamically adjusts the weights of the expert models through the gating network, and selects the most suitable expert model for data analysis based on the characteristics of the data; The expert models perform anomaly detection on the data according to their respective specialties. For example, the visual expert identifies abnormal objects, and the thermal imaging expert judges fires, etc.; It can efficiently and accurately identify abnormal events in the environment (such as fires, intrusions, equipment failures, etc.), and provide real-time feedback through the system. The anomaly recognition accuracy is ≥ 95.3%, and the response latency is ≤ 1 second, significantly improving the system's adaptability to complex environments and the response speed to abnormal events;

[0061] 5. Task Feedback and Dynamic Decision Adjustment: The system automatically classifies abnormal events into three levels: minor, severe, and urgent, and takes corresponding handling measures according to the event level. The response time for decision adjustment is ≤ 200ms; After detecting an abnormal event, the system automatically takes corresponding measures according to the severity of the event. For minor anomalies, the system will record them and continue to execute the current task. For severe anomalies, the system issues an alarm and evaluates whether to suspend the current task for emergency handling. For urgent anomalies (such as fires), the system will automatically interrupt the task and re-plan the path to guide the robot away from the dangerous area; The decision feedback mechanism of the system is evaluated based on the DeepSeek-R1 large model to ensure the timeliness and reasonableness of the response; The system can quickly make decisions according to the level and urgency of the event, ensuring the smooth completion and safety of the task;

[0062] 6. Inspection Result Sorting and Cloud Upload: The data upload requirement is that the upload rate is ≥ 5Mbps, and data encryption uses the AES-256 standard to ensure data security. The data upload latency is < 2 seconds; The system sorts information such as task execution logs, anomaly recognition results, and environmental data into standard data formats (such as JSON, XML) and uploads them to the cloud platform. The cloud platform is responsible for data backup, review, and cross-platform sharing. Users can view historical task data and anomaly handling records through the platform; It realizes the whole-process recording and data archiving of inspection tasks, facilitating subsequent task review and information sharing; Cloud upload ensures the integrity and timeliness of the data, and the system can optimize future task execution based on historical data.

[0063] Example 3:

[0064] This embodiment relates to a specific operation example of the embodied intelligent multi-modal inspection and perception method, system, device, and storage medium described in Embodiment 1 as a power plant safety inspection system. The inspection system in this embodiment includes, but is not limited to, robots and robotic dogs. As a high-risk industrial environment, the power plant requires high-precision and high-timeliness inspections of the equipment operating status. Traditional manual inspections are inefficient and pose safety hazards. This embodiment realizes automated operations through the embodied intelligent inspection system, and the specific content is as follows:

[0065] Operation steps:

[0066] Step 1: The user inputs an inspection task through natural language (such as "Check the operating status of the equipment in Area A of the power plant and confirm no abnormalities"). The DeepSeek-R1 large model in the system analyzes the user's intention and decomposes the task into multiple subtasks, including "Visit Area A of the power plant", "Monitor the equipment status", and "Data collection and analysis";

[0067] Step 2: Combining with the NL-SLAM navigation algorithm, the system generates an optimal path to enable the robot to autonomously navigate within the power plant safely and efficiently; The robotic dog walks based on a preset map and avoids obstacles in real time;

[0068] Step 3: The robot collects the environmental data around the equipment in real time through integrated multi-modal sensors such as infrared thermal imaging, acoustic imaging, and high-definition visible light cameras, conducts temperature, sound, and video monitoring, and synchronously uploads the data to the cloud;

[0069] Step 4: The system analyzes through multi-modal data fusion technology (MOE architecture) to identify equipment failures or abnormalities (such as abnormal temperature, abnormal sound, etc.), and issues an alarm in a timely manner to ensure that the staff can respond quickly.

[0070] Operation result: The system successfully detected the abnormal temperature and overloaded operation of the equipment in Area A of the power plant, issued an alarm in a timely manner, and avoided potential safety accidents. The abnormal detection rate of this system reached 98.5%, and the efficiency was increased by 3 times compared with traditional manual inspections.

[0071] Embodiment 4:

[0072] This embodiment relates to a specific operation example of the embodied intelligent multi-modal inspection and perception method, system, device, and storage medium described in Embodiment 1 as a chemical plant safety inspection system. The inspection tasks in chemical plants involve chemical reaction equipment, storage facilities, etc., posing extremely high safety hazards; Traditional inspections are often difficult to carry out continuously and efficiently, and there are significant safety risks. This embodiment realizes automated operations through the embodied intelligent inspection system, and the specific content is as follows:

[0073] Operation steps:

[0074] Step 1: The user inputs the inspection requirements through natural language instructions (such as "Check the gas leakage risk in Area B of the chemical plant and confirm the normal operation of gas sensors"). The system parses the instructions through the DeepSeek-R1 model and identifies two subtasks: "gas leakage detection" and "gas sensor status check";

[0075] Step 2: The robot autonomously plans the route according to the task requirements, conducts environmental positioning and path planning through the NL-SLAM algorithm, and real-time perceives the changes in the surrounding environment;

[0076] Step 3: The system uses an infrared thermal imaging camera to detect the temperature change on the surface of the equipment, and a sonic sensor to monitor the sound of gas leakage. It combines the gas sensor data for real-time analysis to achieve comprehensive detection and identification;

[0077] Step 4: When the system detects abnormal gas leakage or equipment failure, it immediately issues an alarm and makes decision feedback according to the situation (such as switching to standby equipment, shutting down the main power supply, etc.).

[0078] Operation result: In the application in Area B of the chemical plant, the system successfully identified two gas leakage incidents, gave early warnings, and avoided potential explosion hazards. Compared with manual inspection, the response time of the system was shortened by 40%, and the abnormal identification accuracy rate reached 97.8%.

[0079] Example 5:

[0080] This example involves a specific operation instance of the embodied intelligent multimodal inspection and perception method, system, device, and storage medium described in Example 1 as a high-speed rail and subway safety inspection system. Public transportation facilities such as high-speed rail and subway have extremely high requirements for safety inspection. Traditional manual inspection has problems such as long cycle and inability to cover all aspects. This example realizes automated operation through the embodied intelligent inspection system, and the specific content is as follows:

[0081] Operation steps:

[0082] Step 1: The user inputs instructions through voice or text, such as "Check whether the electrical equipment on Subway Line A is operating normally". The DeepSeek-R1 large model parses the task and identifies subtasks such as "check electrical equipment" and "detect the status of power facilities";

[0083] Step 2: The robot uses the NL-SLAM algorithm to conduct autonomous navigation in the subway track and station, and real-time generates the optimal path, avoiding obstacles and going to the specified equipment for inspection;

[0084] Step 3: The robot collects the status data (such as temperature, vibration, image, etc.) of the electrical equipment through multimodal sensors (such as visible light cameras, infrared sensors, etc.), and combines the MOE architecture for data fusion to identify equipment abnormalities (such as overheating, circuit failures, etc.);

[0085] Step 4: If the system detects an anomaly, it immediately issues an alarm through the cloud platform and guides subway staff to take corresponding maintenance measures. The system can also quickly provide solutions to similar problems based on historical data, improving the processing efficiency.

[0086] Operation result: During a high-speed rail safety inspection, the system successfully detected an overheating anomaly in the power equipment and promptly triggered the cooling system to start, avoiding equipment damage. Compared with traditional manual inspections, the system not only increased the inspection efficiency by 2.5 times but also achieved an accuracy rate of 99.2% for anomaly detection.

Claims

1. An embodied intelligence multi-modal inspection and perception method, characterized in that, It includes the following steps: S1: Natural language instruction parsing: Obtain the inspection task instructions input in natural language form, perform semantic parsing on the natural language description of the inspection task instructions through an integrated large language model, extract key inspection targets, objects, area information, and task logic, and generate a structured task description; S2: Task decomposition and allocation: Decompose the structured task through the lightweight scheduling architecture OWL, divide it into several subtasks, and allocate each subtask to the corresponding agent; the agents include a navigation agent, a perception agent, a detection agent, a monitoring agent, an image analysis agent, and an acoustic analysis agent; S3: Agent scheduling and execution: Each agent executes and completes the corresponding navigation, acquisition, and analysis subtasks by calling sensors, algorithms, or hardware interfaces in the preset tool pool; S4: Multi-modal information collection and fusion: Input the multi-source heterogeneous data collected after completing the acquisition subtask into the MOE (Mixture of Experts) multi-modal fusion model for analysis and fusion after preprocessing; the MOE model includes multiple dedicated expert sub-models such as an image recognition expert, an infrared analysis expert, and an acoustic anomaly expert, and dynamically allocates computing resources and expert weights according to the current task scenario and data characteristics through a gating network to achieve spatio-temporal registration, weighted fusion, and rapid judgment of multi-modal information; S5: Abnormality recognition and event response: Comprehensively judge the fused multi-modal data through the expert output fusion mechanism to identify whether there is abnormal behavior; If abnormal behavior is recognized, perform any one or a combination of the following operations: generate an alarm signal, upload it to the remote management platform, notify the user, update the subsequent inspection path or task priority; S6: Result feedback and task closed-loop: Transmit the execution results, recognition results, and path status of all subtasks back to the large model decision-making center, and it performs global summarization and result generation; after recognizing environmental changes or new instructions, re-enter a new round of task scheduling and path update process.

2. The embodied intelligent multi-modal inspection and perception method according to claim 1, wherein: The tasks executed by the agents described in step S3 specifically include the following steps: S301: Use the navigation agent to combine the NL-SLAM algorithm to perform path planning for the target area, and use the NL-SLAM algorithm to combine the natural language parsing results with the environmental map to realize the mapping from language intention to navigation path, and dynamically sense environmental changes during the inspection process for dynamic obstacle avoidance and route adjustment; S302: Use the perception agent to control the sensors carried by the inspection device to collect data on the target environment; the sensors include a visible light camera, an infrared thermal imager, an audio-video instrument, and a lidar.

3. The embodied intelligent multi-modal inspection and perception method according to claim 1, characterized in that: The method for preprocessing multi-source heterogeneous data described in step S4 includes: S401: Perform data cleaning, perform missing value processing using the deletion method, filling method, and marking method, perform outlier processing using threshold filtering, bin smoothing, and model correction, and perform duplicate data deletion to solve noise, missing, abnormal, and duplicate problems in the data; S402: Align patterns through unit conversion, perform redundancy processing through correlation analysis, conduct data union through multi-directional splicing, integrate multi-source data, and solve pattern conflicts, redundancy, and entity matching problems; S403: Normalize data through min-max scaling, discretize data using binning and one-hot encoding, vectorize unstructured data using text processing, image processing, and time-series data, and convert the data into a unified format suitable for analysis; S404: Perform data reduction through dimensionality reduction, data sampling, and feature selection to reduce the data scale and improve processing efficiency; S405: Conduct multi-modal data fusion through early fusion, late fusion, and hybrid fusion for semantic association and collaborative analysis of heterogeneous data; S406: Unify the storage format and centrally store multi-source data, recording the data source, collection time, version, and preprocessing steps.

4. The embodied intelligent multi-modal inspection and perception method according to claim 1, characterized in that: In step S5, abnormal behaviors include fires, abnormal sounds, personnel intrusion, and abnormal device operation.

5. An embodied intelligent multi-modal inspection and perception system, characterized in that, It includes a control layer, an execution layer, and an output layer. The control layer includes a large language model and an OWL framework. The execution layer includes agents and a tool pool. The output layer includes an alarm and feedback system. Among them, the agent integrates an MOE module; the large language model is used to receive natural language instructions, decompose and understand tasks, and generate an executable task execution plan to guide the agent to schedule the tool pool to complete corresponding tasks; the OWL framework is used to coordinate the collaboration between different agents and the tool pool; the agents include navigation agents, perception agents, detection agents, monitoring agents, image analysis agents, and acoustic analysis agents, which are used to execute specific tasks such as path navigation, device detection, and environmental monitoring. The agents call their respective tools for specific operations to complete the assigned subtasks. The MOE module is used to receive data collected by sensors, perform data fusion and anomaly detection, and identify abnormal situations in the environment; the tool pool is used for agents to call the required hardware devices, including sensors, cameras, navigation modules, and inspection devices. The sensors among them include visible light cameras, infrared thermal imagers, audio imagers, and lidar. The sensors and cameras are used to collect visual, thermal, and sound information of the environment. The navigation module is located inside the inspection device and is used to calculate the optimal path and control the movement of the robot through the NL-SLAM navigation algorithm according to user instructions and system planning. The inspection device is used to move in the target environment and is equipped with sensors, cameras, and a navigation module; the alarm and feedback system is used for the response after anomaly detection. When the MOE module identifies an abnormal behavior, it triggers the alarm mechanism and feeds it back to the user through the cloud platform or conducts emergency processing.

6. The embodied intelligent multi-modal inspection and perception system according to claim 5, characterized in that: The large language model described above is the DeepSeek-R1 large model.

7. An electronic device, characterized in that, It includes a processor and a memory. The processor is used to execute the computer program stored in the memory to implement the embodied intelligent multi-modal inspection and perception method as described in claims 1-4.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, it implements the embodied intelligent multimodal inspection and perception method as described in claims 1-4.

Citation Information

Patent Citations

  • Method for realizing automatic inspection of high-speed rail box girder inspection robot

    CN113190002A

  • Multi-agent cooperative electric power inspection method, system and device and storage medium

    CN119106101A

  • Unmanned aerial vehicle intelligent inspection method for energy facility inspection

    CN119693824A

Cited By

  • Unmanned aerial vehicle expressway intelligent inspection method and system based on patrol requirements

    CN120690202A

  • Fusion method and device suitable for sound image long-distance identification system

    CN120873983A

  • Unmanned aerial vehicle road inspection task intelligent planning management system

    CN121032150A

  • Video inspection method and device, server and storage medium

    CN121099006A

  • Electric power scene full-chain ubiquitous dynamic sensing body-equipped intelligent inspection operation method and system

    CN121279323A