A multi-modal fusion intelligent closed-loop control method and system

By using a multimodal fusion intelligent closed-loop control method and system, the problems of single perception mode and isolated execution channel in existing technologies are solved. It achieves highly robust recognition and multilingual interaction in complex scenarios, adapts to the deployment needs of multiple scenarios, and has the collaborative capabilities of an intelligent assistant.

CN122131603APending Publication Date: 2026-06-02SIYUAN HUANYU (QINGDAO) INTERNATIONAL TRADING CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SIYUAN HUANYU (QINGDAO) INTERNATIONAL TRADING CO LTD
Filing Date
2026-03-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing intelligent control technologies suffer from limited perception methods and insufficient robustness in complex scenarios. Their execution channels are isolated, lack a unified scheduling mechanism, and cannot achieve a closed-loop linkage between multimodal perception and multi-channel execution. They also have weak command interaction capabilities, making them unsuitable for the interactive automated task requirements of upper-level AI agents. Furthermore, their device forms are unclear, making it difficult to adapt to deployments in multiple scenarios.

Method used

A multimodal fusion intelligent closed-loop control method and system are constructed. The system collects the state data of the controlled object through a multimodal perception module, performs weighted fusion processing, and combines cloud API to achieve adaptive adjustment of weights. It supports real-time programming and instruction parsing in multiple languages, and adopts a unified decision engine to automatically match the execution channel, realizing real-time judgment and correction throughout the closed loop, and has the ability to make proactive collaborative decisions.

Benefits of technology

It improves recognition accuracy and robustness in complex scenarios, lowers the operational threshold, enables multilingual interaction and device adaptation, has collaborative capabilities similar to an intelligent assistant, adapts to deployment needs in multiple scenarios, and has low threshold and high robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122131603A_ABST
    Figure CN122131603A_ABST
Patent Text Reader

Abstract

This invention relates to a multimodal fusion intelligent closed-loop control method and system, comprising the following steps: (1) Multimodal information acquisition and fusion recognition: The state data of the controlled object is acquired through a multimodal perception module, wherein the state data includes at least one of visual image data, auditory audio data, and auxiliary sensor data; the state data is weighted and fused, and the fusion algorithm can be implemented by calling a cloud API. The multimodal fusion intelligent closed-loop control method: System-level architecture innovation: Constructing an integrated core architecture of "multimodal perception - real-time judgment - unified scheduling - human-machine collaboration - multi-channel execution - cross-modal closed-loop correction", adopting a miniaturized integrated intelligent control terminal form, the core algorithm can be implemented by calling a cloud API, compatible with a variety of existing mainstream network connection methods, forming a hierarchical difference with existing single-point execution patents, completely avoiding conflicts, and forming an irreplaceable system-level barrier.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, specifically to a multimodal fusion intelligent closed-loop control method and system. Background Technology

[0002] This invention relates to the field of intelligent control and cross-device interaction technology, and in particular to an intelligent closed-loop control scheme that integrates multimodal perception, unified decision-making and scheduling, multi-channel execution, and real-time human-machine collaboration across the entire chain. It is suitable for collaboration with upper-layer AI agents and large-scale model APIs. The core algorithm can be implemented by calling cloud APIs, supports IoT communication functions, and is compatible with various existing mainstream network connection methods for network access. The device is a miniaturized, integrated intelligent control terminal. It supports multilingual real-time programming, multi-interface command input, real-time interruption, modification, and proactive querying. Through real-time analysis, it achieves automated, low-threshold, highly robust, and interactive control of physical devices, touch-screen electronic devices, and computer terminals, possessing full-scene collaborative capabilities similar to an intelligent assistant, and achieving stable control without requiring professional operating experience.

[0003] The core technology chain of this invention follows the basic architecture of visual recognition—algorithm decision-making—terminal execution, and on this basis, forms a full-link intelligent closed-loop control system with multimodal perception, unified scheduling, multi-channel execution, real-time judgment and correction, and closed-loop correction.

[0004] In existing intelligent control technologies, there are already execution schemes for a single controlled object: such as simulating touch control of a mobile phone through a transparent capacitive screen, simulating control of a computer terminal through keyboard and mouse signals, and operating physical buttons / knobs through mechanical structures. The above-mentioned single-point execution schemes are the technical solutions that the applicant intends to apply for patent protection simultaneously. However, existing technologies have several core flaws: First, the perception method is singular, resulting in insufficient reliability and poor robustness in complex scenarios. Second, the execution channels are isolated, lacking a unified scheduling mechanism, making it impossible to achieve "one system controlling the entire domain," and the operational threshold is high. Third, a closed-loop linkage between multimodal perception and multi-channel execution has not been established, and the verification of execution results and error correction lack universality. Fourth, the command interaction capability is weak, supporting only fixed commands or single language input, lacking real-time programming, real-time interruption, and real-time modification capabilities, and unable to dynamically adjust tasks during execution. Fifth, there is a lack of proactive inquiry and collaborative decision-making mechanisms, requiring full human intervention in abnormal scenarios, making it difficult to adapt to the large-scale, interactive, and automated task requirements of upper-level AI agents, unable to achieve proactive collaborative interaction effects similar to intelligent assistants, and unable to balance low-threshold operation with high robustness. Sixth, the algorithm implementation relies on local computing, lacking flexibility and flexible IoT networking capabilities, unable to achieve remote collaboration and data synchronization, and the device form is unclear, making it difficult to adapt to multi-scenario deployment needs.

[0005] There is currently no integrated intelligent control scheme that can simultaneously perform unified automated control of physical devices, touch electronic devices, and computer terminals, and combine multimodal adaptive fusion perception, real-time closed-loop correction, and multilingual human-machine collaborative interaction.

[0006] This invention aims to solve the above-mentioned problems by constructing a system-level architecture of "multimodal perception input → unified decision-making and scheduling → multi-channel execution output → real-time human-machine collaborative interaction → cross-modal closed-loop correction". The core algorithm can be implemented by calling cloud APIs and is compatible with various existing mainstream network connection methods. The device adopts a miniaturized integrated intelligent control terminal form, which complements rather than conflicts with existing single-point execution patents, realizing intelligent collaboration and autonomous interaction, taking into account both low-threshold operation experience and high robust operation performance, while improving algorithm flexibility and device deployment adaptability. Summary of the Invention

[0007] The purpose of this invention is to provide a multimodal fusion intelligent closed-loop control method and system to solve the problems mentioned in the background art.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a multimodal fusion intelligent closed-loop control method and system, comprising the following steps: (1) Multimodal information acquisition and fusion recognition: the state data of the controlled object is acquired through a multimodal perception module, the state data including at least one of visual image data, auditory audio data, and auxiliary sensor data; the state data is weighted and fused, and the fusion algorithm can be implemented by calling a cloud API; the system records the task execution success rate and operation error data under different scenarios, and iteratively optimizes the weight coefficients corresponding to each modality based on the execution effect, dynamically and adaptively adjusts the weights according to the scene interference degree and recognition success rate, and continuously updates to achieve the optimal balance between recognition accuracy and system stability, thereby improving the accuracy, robustness and scene adaptability of fusion recognition. The type of the controlled object is identified as a physical device, a touch electronic device or a computer terminal, and the current state features of the controlled object are extracted to ensure high robustness of recognition; visual recognition is used as the core perception means to extract the visual features, position and state information of the controlled object.

[0009] (2) Multi-channel real-time instruction reception and parsing: The system receives multilingual real-time programming instructions issued by users or upper-layer AI agents through at least one of the voice interface, text interface, mobile application interface, and instant messaging interface. The instructions include task intent, execution logic, and scheduling rules, and support real-time interruption, real-time modification, and incremental programming, simplifying the operation process and lowering the threshold for use. The system parses the instructions through a multilingual natural language understanding module. The parsing algorithm can be implemented by calling cloud APIs to generate standardized task instructions, enabling flexible access to instructions in multiple scenarios and in multiple ways.

[0010] (3) Task parsing and execution channel decision-making: Combining the controlled object type, state characteristics, and standardized task instructions, the algorithm decision is completed through a unified decision engine, decomposing it into atomic operation sequences; the core algorithm of the decision engine can be implemented by calling cloud APIs, improving decision-making efficiency and adaptability, and automatically matching the corresponding execution channel. The execution channels include physical execution channels, touch simulation execution channels, and keyboard and mouse simulation execution channels, which do not require manual intervention for matching, further reducing the operation threshold.

[0011] (4) Standardized instruction issuance and execution: The atomic operation sequence is converted into control instructions for the corresponding execution channel and forwarded to the target execution unit through the multi-channel execution interface. The execution unit completes the terminal execution to realize physical operation, touch simulation operation or keyboard and mouse simulation operation of the controlled object, ensuring the stability and high robustness of the operation. During the execution process, remote transmission of control instructions and state synchronization can be realized through various existing mainstream network connection methods.

[0012] (5) Cross-modal real-time analysis and closed-loop correction: The state data of the controlled object after execution is collected again through the multi-modal perception module and compared with the preset expected results across modes to calculate the operation error and realize the real-time analysis and control of the whole process in a closed loop; the analysis and error calculation algorithm can be implemented by calling the cloud API to improve the accuracy of the analysis; if the error exceeds the preset threshold, the corrected atomic operation sequence and control instructions are regenerated and iterated until the error meets the threshold requirements to improve the robustness of the system; if continuous iteration fails, abnormal scenarios are detected, or doubtful scenarios such as uncertain operation logic, parameter settings, and abnormal causes are encountered in the analysis process, the system immediately initiates intelligent inquiry to the user through the preset instant communication interaction method, receives real-time replies from the user and dynamically adjusts the execution strategy to realize proactive collaborative decision-making and reduce the threshold of manual intervention; if the user does not reply or the instruction is terminated, an abnormal alarm is triggered and the task is terminated; abnormal information can be synchronized to the user's mobile terminal or the upper-level AI intelligent agent through IoT communication methods (card slot Internet access, traditional IoT Internet access, etc.).

[0013] (6) Task closure and result reporting: When the operation error meets the threshold requirements, the task is determined to be completed. The execution result is standardized and fed back to the upper-level AI agent or user. Operation logs, correction data and real-time interaction records are archived to optimize subsequent fusion recognition, decision scheduling and human-machine collaboration logic, and continuously improve the low-threshold operability and high robustness of the system. Feedback and log archiving can be remotely synchronized through a variety of existing mainstream network connection methods and are compatible with various existing network access forms.

[0014] Furthermore, the auxiliary sensor data includes at least one of tactile feedback data, inertial measurement unit attitude data, and pressure sensing data, further enhancing the robustness of multimodal fusion recognition.

[0015] Furthermore, the fusion processing employs a weighted fusion algorithm, which can be implemented by calling a cloud API. This algorithm can dynamically optimize weight parameters based on cloud computing power, dynamically adjust the weights of visual, auditory, and auxiliary sensor data according to the level of scene interference, improve the recognition reliability in complex scenes, enhance the system's robustness, and at the same time eliminate the need for manual weight adjustment, thus reducing the operational threshold.

[0016] Furthermore, the multimodal sensing module adopts a modular design, which can flexibly add new sensor types without reconstructing the acquisition and fusion logic, thereby improving the system's scalability and robustness, while reducing the threshold for equipment upgrades and scenario adaptation.

[0017] Furthermore, when the controlled object is identified as a touch electronic device, the matched execution channel is a touch simulation execution channel, and the control command is a transparent capacitive touch driving command. The control logic can be controlled by simulating capacitance changes to accurately simulate touch operations and ensure the high robustness of touch operations.

[0018] Furthermore, when the controlled object is identified as a physical device, the matched execution channel is the physical execution channel. The control commands can generate high-precision robotic arm drive commands, stepper motor drive commands, and displacement, force, and timing drive commands of mechanical actuators according to the type of physical device components. Among them, the stepper motor drive commands are used to control the rotation of existing knob-like components, and the high-precision robotic arm is used to complete precise pressing and tossing operations. At the same time, it can be combined with analog capacitance change control commands to adapt to touch-screen auxiliary components, ensuring high robustness and adaptability of physical operations.

[0019] Furthermore, when the controlled object is identified as a computer terminal, the matched execution channel is a keyboard and mouse simulation execution channel. The control command is a standard keyboard and mouse communication protocol command, which uses simulated keyboard and mouse signals to achieve control. It does not require intrusion into the computer terminal system, has strong compatibility, and ensures the stability and high robustness of computer terminal operation.

[0020] Furthermore, the abnormal alarms include at least one of the following: device malfunction alarm, no operation response alarm, missing perception data alarm, and interaction timeout alarm, which promptly reminds users to handle abnormalities, reduces operational difficulty, and ensures system robustness.

[0021] Furthermore, the multilingual real-time programming instructions support natural language input in at least two languages, including Chinese, English, Japanese, and Korean. It supports offline multilingual recognition and online large-model collaborative translation and parsing. The parsing algorithm can be implemented by calling cloud APIs, adapting to multilingual usage scenarios, reducing the operational threshold for users of different languages, and improving the robustness of instruction parsing. For niche languages, it supports adaptation through extended interfaces. Based on the translation capabilities of the upper-level large-model API, it can quickly complete instruction parsing and real-time programming adaptation for niche languages ​​without refactoring the core system logic. Only interface parameter configuration and language model adaptation are required to achieve real-time programming, interruption, and modification functions for niche languages, adapting to multi-scenario and multilingual usage needs.

[0022] Furthermore, the real-time interruption includes voice interruption, text command interruption, one-click interruption via mobile application, and instant messaging command interruption. After interruption, it supports at least one of the following operations: incremental modification of task parameters, replanning of execution steps, and termination of task, so as to realize flexible and controllable task execution, reduce the operation threshold, and ensure the robustness of operation.

[0023] Furthermore, the intelligent inquiry includes rule-based inquiry and large-model-based intelligent inquiry. The inquiry content includes confirmation of the cause of the anomaly, selection of operation method, adjustment of target parameters, etc. The response method is consistent with the command input interface to ensure the continuity and convenience of the interaction, further reduce the operation threshold, and improve the system's collaborative robustness. The logic algorithm of intelligent inquiry can be implemented by calling cloud APIs to improve the intelligence and adaptability of the inquiry.

[0024] Furthermore, the IoT communication method is compatible with a variety of existing mainstream network connection methods, including but not limited to SIM card slot internet access (supporting 4G / 5G networks), conventional IoT communication methods (Bluetooth, 2.4G wireless communication, NB-IoT, LoRa, etc.) and wired Ethernet (directly connected to the network cable). The appropriate network access method can be flexibly selected according to the deployment scenario to ensure the stability and flexibility of the network connection and realize remote command transmission, status synchronization and data archiving.

[0025] A smart hardware system for realizing multi-channel unified closed-loop control, the system is a miniaturized integrated smart control terminal, the device is small in size and regular in shape, the shell is made of flame-retardant ABS material with frosted surface treatment, multiple status indicator lights are set on the front (for displaying power, network, and running status), and various adapter interfaces (including network access interface, data transmission interface, etc.) are reserved on the side. All adapter interfaces are compatible with industry general standards to ensure compatibility with different execution units and network devices. Anti-slip pads are set on the bottom for easy deployment and can be adapted to installation in industrial, office and other scenarios; the system includes: (1) Multimodal perception unit: integrating visual sensor, audio sensor and auxiliary sensor, used to collect visual image data, auditory audio data and auxiliary sensor data of the controlled object, and complete data preprocessing and fusion recognition, output the type and status characteristics of the controlled object, and provide a foundation for the low threshold and high robustness operation of the system;

[0026] (2) Real-time Human-Machine Collaborative Interaction Unit: Integrates a multilingual voice acquisition module, a multilingual natural language understanding module, a full-channel instruction receiving module, and an intelligent inquiry module. It possesses intelligent collaborative capabilities, supports real-time programming, real-time interruption, real-time modification, and intelligent user inquiry in multiple languages. It can receive instruction input in various ways, including voice, text, mobile applications, and instant messaging, and simultaneously send inquiry instructions to users and receive replies, achieving proactive interaction across all scenarios. This simplifies the operation process, lowers the usage threshold, and improves the robustness of interaction. Among these features, the multilingual natural language understanding module reserves an extension interface for niche languages. It can be connected to the large model API and used for niche languages ​​through... The system expands its interface to import minority language recognition models, enabling rapid adaptation to various minority languages. After speech-to-text conversion, it performs machine cognition and semantic understanding on the text, achieving accurate parsing of minority language commands. Once adapted, it enables real-time programming, interruption, modification, and intelligent query interaction in minority languages, ensuring scalability and flexibility for multilingual adaptation. The system also supports visual drag-and-drop programming, encapsulating functions such as moving, clicking, pressing, looping, conditional judgment, and error correction into graphical building block components. Users can complete task flow arrangement by dragging and assembling components without writing code, further reducing the operational threshold and meeting the needs of non-professionals to quickly configure automated tasks.

[0027] (3) Unified decision-making and scheduling unit: It communicates with the multimodal perception unit and the real-time human-machine collaborative interaction unit. It has a built-in decision engine, task decomposition module and real-time replanning module. It is used to receive standardized task instructions and real-time interaction instructions, decompose them into atomic operation sequences, automatically match the corresponding execution channels, and dynamically adjust the execution strategy according to real-time interruption and intelligent query response to ensure the flexibility, accuracy and high robustness of task execution. No manual intervention is required for scheduling, which reduces the operation threshold. The core algorithms of the decision engine and task decomposition module can be implemented by calling cloud APIs, relying on cloud computing power to improve decision efficiency and adaptability.

[0028] (4) Multi-channel execution interface unit: It communicates with the unified decision scheduling unit and includes physical execution interface, touch simulation execution interface and keyboard and mouse simulation execution interface. It is used to convert standardized control commands into drive signals that can be recognized by each execution unit to ensure high robustness of operation of multiple types of devices.

[0029] (5) Real-time analysis unit: It communicates with the multimodal perception unit, unified decision scheduling unit, and real-time human-machine collaborative interaction unit respectively. It has built-in error calculation module, correction instruction generation module, and anomaly detection module. The core realizes the whole-process closed-loop real-time analysis and control, and participates in the whole-link of "pre-execution prediction - in-execution monitoring - post-execution verification". It is used to compare the state data before and after execution, calculate errors, generate correction instructions, and detect abnormal scenarios. For uncertain operational details, abnormal causes, parameter adaptation and other doubtful issues that cannot be determined in the analysis process, the intelligent inquiry mechanism is automatically triggered to initiate inquiries to users to ensure accurate analysis and reliable control, realize closed-loop correction, and continuously improve the robustness of the system. The core algorithms of error calculation and anomaly detection can be implemented by calling cloud API to improve the accuracy and efficiency of analysis.

[0030] (6) Communication and storage unit: It communicates with the unified decision-making and scheduling unit, real-time analysis unit, and real-time human-machine collaborative interaction unit to realize standardized data transmission with upper-layer AI intelligent agents, large model APIs, mobile applications, and instant messaging interfaces. At the same time, it archives operation logs, model parameters, correction data, and real-time interaction records to ensure data traceability and system optimization, and further enhances the system's low-threshold and high-robust operation characteristics. The communication and storage unit integrates an IoT communication module, which is compatible with a variety of existing mainstream network connection methods. It can flexibly switch the network access form according to the needs of the scenario to ensure the stability and flexibility of remote communication. It also supports offline storage, and can perform local tasks normally in the absence of network environment. After the network is restored, the data is automatically synchronized.

[0031] Furthermore, the visual sensor is an industrial-grade camera, the audio sensor is a noise-canceling microphone, and the auxiliary sensors include at least one of a tactile sensor, a six-axis inertial measurement unit, and a pressure sensor; the physical execution unit can be adapted to a high-precision robotic arm and a stepper motor, where the high-precision robotic arm is used to complete precise pressing, flicking and other operations, and the stepper motor is used to drive the rotation of knob-like parts. It also supports analog capacitance change control logic, adapts to touch-controlled parts, improves the robustness of multimodal perception and the adaptability of execution parts, and provides support for low-threshold operation.

[0032] Furthermore, the physical execution interface is compatible with high-precision robotic arms, stepper motors, electromagnetic presses, and knob drives; the touch simulation execution interface is compatible with transparent capacitive touch units, supports simulated capacitance change control, and can accurately simulate touch operations; the keyboard and mouse simulation execution interface is compatible with the USBHID standard protocol, ensuring compatibility with multiple types of execution units, improving the robustness of system execution, and lowering the device adaptation threshold.

[0033] Furthermore, the communication and storage unit supports various existing mainstream network connection methods and offline storage, and can operate stably in environments without external network access, ensuring the continuity of real-time interaction and data security, improving system robustness, while eliminating the need for complex network configuration and lowering the barrier to entry; the communication and storage unit integrates various network access interfaces, is compatible with various existing network forms, and achieves multi-scenario network adaptation.

[0034] Furthermore, the instant messaging interface supports standard open platform protocols, enabling access to various instant messaging systems and seamless integration of message sending and receiving, command interaction, and intelligent multilingual queries. It also supports the function of interrupting at any time in multilingual scenarios, adapting to the interaction needs of multiple scenarios such as office and industry, reducing the interaction threshold, and improving the robustness of interaction.

[0035] Furthermore, the intelligent inquiry module supports multiple notification methods such as voice broadcast, text push, mobile application notification, and instant messaging messages, ensuring that users receive and respond in a timely manner, improving interaction response efficiency, reducing the operation threshold, and ensuring the robustness of the interaction.

[0036] Furthermore, the real-time replanning module can quickly adjust the atomic operation sequence, execution channel, and control parameters based on real-time interruption instructions or intelligent query responses during task execution, without restarting the task, thereby improving the flexibility and efficiency of task execution, reducing the operational threshold, and ensuring the robustness of execution.

[0037] Furthermore, the cloud API calls support encrypted transmission to ensure the security of algorithm calls and data interaction. It also supports API interface expansion, allowing the addition of new algorithm call types as needed, thereby improving system scalability. The IoT communication process employs encryption protocols to prevent data leakage and ensure the security of remote control and data synchronization.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] This multimodal fusion intelligent closed-loop control method features: system-level architecture innovation: constructing a core architecture of "multimodal perception + real-time human-machine collaboration + unified scheduling + multi-channel execution + cross-modal closed-loop correction", adopting a miniaturized integrated intelligent control terminal form, and the core algorithm can be implemented by calling cloud APIs, compatible with a variety of existing mainstream network connection methods, forming a hierarchical difference from existing single-point execution patents, completely avoiding conflicts, and forming an irreplaceable system-level barrier;

[0040] Advantages of intelligent human-machine collaboration: It has the proactive interaction capability of an intelligent assistant, supports multilingual real-time programming, full interface command input, real-time interruption / modification / intelligent inquiry, realizes dynamic adjustment of tasks during execution, proactively collaborates with users to make decisions in abnormal scenarios, greatly reduces the cost of human intervention, creates an efficient and intelligent human-machine collaboration experience, and at the same time takes into account low threshold operation and high robust operation.

[0041] Low barrier to entry and high robustness: Multimodal data fusion overcomes the limitations of single-sensor scenarios, achieving accurate identification and reliable control even in complex scenarios such as visual occlusion, noise interference, and device malfunction, demonstrating outstanding robustness. It also simplifies operation processes, supports multiple command input methods, and automates decision-making and scheduling, requiring no specialized operating experience, thus lowering the barrier to entry and adapting to various user groups. The core achieves real-time closed-loop analysis and control throughout the entire process, proactively querying users for scenarios with questionable analysis and adjusting execution strategies based on user commands to further improve control accuracy and robustness, reducing manual intervention costs. The core algorithm calls cloud APIs, eliminating the need for complex local computing power, reducing device hardware costs and operational barriers. Multiple IoT networking methods allow flexible adaptation to different scenarios, eliminating the need for complex network deployments, further lowering the barrier to entry.

[0042] Ecosystem synergy advantages: It achieves standardized integration with upper-layer AI intelligent agents, large model APIs, mobile applications, and instant messaging systems. It ensures the stability of remote collaboration through various IoT communication methods, becoming a unified and interactive execution entry point for AI systems to reach the physical world, electronic devices, and computer terminals. It has low integration thresholds and strong robustness in collaborative operation.

[0043] Highly scalable: The sensing module, execution channel, and interaction interface can all be flexibly expanded. When adding new sensor types, controlled objects, or interaction methods, there is no need to reconstruct the core logic; only the interface needs to be adapted. The cloud API calls support interface expansion, allowing the addition of new algorithm types. It is compatible with multiple existing mainstream network connection methods, significantly reducing the cost of subsequent product iterations, while ensuring that the expanded system still maintains its low-threshold and high-robust characteristics. The device is compact, has a regular shape, and is flexibly deployed, making it suitable for various scenarios such as industrial and office environments.

[0044] Significant commercial value: A single system covers multiple scenarios such as industry, office, and consumption. The device form is clear and easy to deploy. The core algorithm calls cloud APIs to improve flexibility and iteration efficiency. It is compatible with a variety of existing mainstream network connection methods to adapt to the needs of multiple scenarios. The low threshold and high robustness characteristics enhance product adaptability and market competitiveness, and provide core intellectual property support for financing and large-scale product implementation.

[0045] Compliance and controllability: The real-time human-machine collaboration mechanism ensures that all operations are under the full control of the user. The system is only a reasonable extension of the legitimate will of humans. Combining local decision-making and privacy protection design, it eliminates the possibility of unauthorized use from a technical point of view, while ensuring the robustness of compliant operation and lowering the threshold for compliant operation. Cloud API calls and IoT communications are all encrypted to ensure data security and operational compliance. Attached Figure Description

[0046] Figure 1 This is a flowchart of the workflow of the present invention. Detailed Implementation

[0047] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0048] Please see Figure 1 Example 1: Multimodal collaborative closed-loop control of industrial physical equipment (high-precision robotic arm + stepper motor adapter)

[0049] In this embodiment, the intelligent control terminal adopts a miniaturized, integrated form (small in size and regular in shape; the flame-retardant ABS frosted shell can also be made of metal, plastic, or other shells; multiple status indicator lights on the front; various adapter interfaces reserved on the sides; and anti-slip feet on the bottom). Deployed next to the production line control cabinet, it can achieve remote communication with the upper-level AI intelligent agent through various existing mainstream network connection methods (this embodiment uses a 4G card slot for internet access). The core algorithm can be implemented by calling the cloud API (encrypted transmission), ensuring operational efficiency and accuracy. Task instructions: The user issues the command "Start the air compressor, adjust the output pressure to 0.8MPa (adjustable according to actual working conditions)" via Chinese voice, and sets "If the pressure cannot be reached, actively ask whether to adjust the target parameters."

[0050] Execution process:

[0051] Multimodal acquisition and recognition: Visual sensors identify switch positions, knob scales, and pressure gauge pointers; audio sensors collect buzzer sounds to determine the initial standby state of the device; tactile sensors pre-detect switch damping states and knob initial positions, and collect tactile feedback data, which is a type of auxiliary sensor data, further improving the robustness of multimodal fusion recognition; the fusion recognition result is "physical device - air compressor - standby state". The fusion algorithm can be implemented by calling cloud APIs, and the weight parameters can be dynamically optimized based on cloud computing power. The weights of visual, auditory, and auxiliary sensor data can be dynamically adjusted according to the level of scene interference, improving the recognition reliability in complex scenarios and strengthening the system's high robustness. At the same time, no manual weight adjustment is required, reducing the operation threshold and accurately identifying the device status. It also identifies pressure knobs as adjustment components that can be driven by stepper motors and start switches as components that require mechanical pressing, ensuring high robustness of recognition without the need for manual confirmation of device status and adaptation methods.

[0052] Command parsing and decision-making: The real-time human-machine collaborative interaction unit parses Chinese voice commands. The parsing algorithm can be implemented by calling cloud APIs. The unified decision engine breaks it down into "high-precision robotic arm presses the start switch → stepper motor drives the pressure knob to rotate to the 0.8MPa scale", automatically matching the physical execution channel. The robotic arm corresponds to the physical pressing execution interface, and the stepper motor corresponds to the knob drive execution interface. No manual intervention is required throughout the process, and the operation threshold is low.

[0053] Command Issuance and Execution: The physical execution interface sends drive commands to the high-precision robotic arm and stepper motor respectively. The high-precision robotic arm accurately completes the switch operation (control accuracy up to 0.1mm), and the stepper motor drives the pressure knob to rotate according to the preset step distance. The execution process is stable and robust. At the same time, the touch simulation execution interface realizes the control of simulated capacitance changes, which can be adapted to the touch auxiliary buttons on the panel (if any), realize the simulation control of touch operation, improve the adaptation flexibility, and form a collaborative adaptation with the touch simulation execution channel. During the execution, the device connects to the Internet through the 4G card slot and synchronizes the real-time execution status to the upper-level AI intelligent agent and the user's mobile terminal.

[0054] Cross-modal judgment and collaborative correction: The visual system recognizes the pressure gauge pointer to 0.75MPa (adjustable according to actual working conditions), the audio sensor detects the air compressor's operating noise (confirming successful startup), and the tactile sensor detects that the knob is not stuck and the robotic arm resets normally. The real-time judgment unit triggers an intelligent inquiry mechanism. The judgment algorithm can be implemented by calling the cloud API to accurately determine error anomalies and send an inquiry message to the user via instant messaging: "Current pressure is 0.75MPa, which has not reached the preset 0.8MPa. Do you want to adjust the target pressure to 0.75MPa and complete the task?" The user replies "confirm" via instant messaging. The interaction is convenient and has a low threshold. At the same time, the real-time replanning module issues instructions to control the stepper motor to fine-tune the knob to the 0.75MPa scale, realizing error correction and ensuring system robustness. The multimodal perception module adopts a modular design, which can flexibly add sensor types to form a general machine perception and execution nervous system.

[0055] This system is a universal machine perception and execution neural network for all types of devices. It employs a general intelligent architecture, first constructing the system body and prioritizing adaptation to various existing physical devices, regardless of device form, application scenario, or robot body. Subsequently, based on this system body, it can be gradually integrated into various devices to achieve intelligent upgrades. Through a unified perception-neural-execution system, this system endows all adapted and integrated devices with standardized and scalable basic intelligent capabilities. The solution focuses on practicality, simplifies core module dependencies, reduces deployment costs, and covers all the fundamental innovations involved in this patent, forming a complete technical protection system.

[0056] This architecture abandons the complex self-developed brain module and adopts a layered design of "perception-neuron-execution". The core innovation lies in: using bionic nerves as the central nervous system, instead of independently developing a decision-making brain, it directly calls mature APIs from third-party vendors to complete core decision-making functions, balancing stability and deployment efficiency, forming a closed-loop chain of "perception acquisition-neural transmission-API decision-making-execution output". At the same time, each layer of modules adopts a standardized interface design, which can flexibly adapt to different types of sensing hardware, actuators and third-party APIs, improving system scalability. The specific architecture is as follows:

[0057] Perception Layer (Bionic Eye, Extended Innovation): Innovatively realizes full-domain data acquisition of the two-dimensional and three-dimensional world, synchronous acquisition and fusion of multi-modal signals, adapts to the access of various sensing hardware, breaks the limitations of single sensing mode, and provides comprehensive and accurate underlying data support for subsequent processing. This is the basic innovation for realizing full-domain perception.

[0058] The neural center (bionic nerve, core innovation): It is responsible for data transmission, signal analysis, and instruction distribution. It simulates the signal transmission mechanism of biological nerves. The core innovation is that "it does not develop its own decision-making module, but relies on third-party APIs to make decisions", which reduces the difficulty of research and development, shortens the implementation cycle, and ensures the stability of decision-making. This is different from the design of traditional intelligent systems that require the development of a complete brain module. It is the core innovation for the implementation of this system.

[0059] Execution layer (bionic hand, innovation extension): an innovative actuator that adapts to various devices, can quickly respond to API decision commands distributed by the central nervous system, achieve precise end-effector operation, motion control and command output, complete physical interaction and task execution, break the limitations of device form, adapt to devices in all scenarios, and is an important supporting innovation for the general architecture.

[0060] This system, through the aforementioned three-layer integrated architecture and leveraging the mature capabilities of third-party APIs, enables any physical device to quickly acquire the basic intelligent capabilities to understand the two-dimensional and three-dimensional world, autonomously comprehend the spatial environment, make autonomous decisions, and accurately execute operations. This significantly improves deployment efficiency and reduces R&D and deployment costs. This innovative architecture effectively protects the deployment-ready design of "bionic neural network + third-party API," distinguishing it from the architectural patterns of traditional intelligent systems.

[0061] Breaking through the limitations of a single perception modality, this system achieves simultaneous acquisition, registration, and joint modeling of 2D and 3D environmental perception, covering 19 multimodal perception methods. Each method possesses differentiated innovative adaptability advantages and can be flexibly selected according to different scenarios to form a full-domain perception capability. Specific perception solutions are as follows (all are extensions and implementations of basic innovations, covering all innovations at the perception level):

[0062] Monocular vision perception: This innovative approach combines a single camera with algorithms to simultaneously acquire 2D images and estimate 3D spatial information. It features low hardware costs, compatibility with small and lightweight devices, and addresses the technical challenge of achieving 3D perception with a single monocular vision system.

[0063] Binocular vision perception: Innovatively adopts dual-camera stereo matching technology to simultaneously acquire two-dimensional images and three-dimensional spatial structures, balancing accuracy and cost, adapting to medium-precision demand scenarios, and breaking through the bottleneck of fusion between binocular vision and three-dimensional perception.

[0064] Multi-view vision perception: This innovative approach uses an array of three or more cameras combined with stereo vision algorithms to reconstruct high-precision 3D spatial information, addressing the need for high-precision 3D perception and adapting to sophisticated scenarios.

[0065] Linear laser sensing: This innovative technology integrates two-dimensional image features with laser line projection to simultaneously acquire two-dimensional features, three-dimensional contours, height, and deformation information of objects, thereby improving the reliability of sensing in complex lighting scenarios.

[0066] Structured light sensing: This innovative approach combines coded structured light projection with 2D image acquisition to rapidly reconstruct the 3D shape of an object, addressing the technical need for rapid 3D shape acquisition and improving acquisition efficiency.

[0067] Contact probe sensing: This innovative approach combines probe contact acquisition with 2D image positioning to simultaneously obtain 3D coordinates and 2D positioning information, improving positioning accuracy and adapting to precision measurement scenarios.

[0068] Robot Force Control Perception: Innovations utilize force sensors and damping control at the end of the robotic arm to acquire contact force, position, and physical interaction information, supporting two-dimensional and three-dimensional spatial positioning and achieving multimodal fusion of force perception and vision.

[0069] Depth camera perception: Innovatively outputs two-dimensional color images and three-dimensional depth data simultaneously, achieving seamless integration of two-dimensional vision and three-dimensional spatial information, improving integration, and adapting to small and medium-sized smart devices.

[0070] LiDAR perception: The innovative method of registering LiDAR 3D point cloud information with 2D image data enables multimodal fusion and enhances long-distance, large-area environmental perception capabilities.

[0071] Solid-state LiDAR perception: This innovative combination of solid-state LiDAR and 2D vision data enhances perception robustness, solves the problems of large size and insufficient stability of traditional LiDAR, and is suitable for small devices.

[0072] Millimeter-wave radar sensing: This innovative technology integrates distance and location information acquired from millimeter-wave signals with two-dimensional images to achieve all-weather sensing and overcome the limitations of sensing in adverse weather conditions.

[0073] Ultrasonic array sensing: This innovative technology uses multiple ultrasonic probes combined with two-dimensional images to achieve three-dimensional spatial positioning. It is low-cost, highly resistant to interference, and suitable for close-range detection scenarios.

[0074] Inertial Measurement Sensing: This innovative approach uses gyroscopes and accelerometers to acquire attitude information, providing an attitude reference for two-dimensional and three-dimensional perception and solving the problem of perception deviation in mobile devices.

[0075] Visual-inertial fusion perception: Innovatively integrates 2D images and inertial measurement data to improve the accuracy of 2D target tracking and 3D pose estimation, and solve the problem of insufficient perception accuracy in motion.

[0076] Infrared thermal imaging perception: This innovative technology integrates two-dimensional thermal images of an object's temperature distribution with three-dimensional spatial information to achieve multimodal feature perception, overcoming the limitations of conventional vision in detecting concealed, high-temperature targets.

[0077] Spectral perception: This innovative approach uses spectral sensors to acquire two-dimensional distribution information such as the material and composition of objects, and then combines this information with three-dimensional spatial data to create a joint model, achieving a fusion of material recognition and three-dimensional modeling.

[0078] 3D scanning perception: Innovatively, guided by 2D images, it uses structured light, laser, or photogrammetry to complete the overall 3D model and spatial data acquisition of an object, improving the accuracy and efficiency of 3D modeling.

[0079] Proximity sensing: This innovative technology uses proximity sensors, combined with two-dimensional and three-dimensional sensing data, to achieve close-range object detection and distance prediction, thereby improving the device's collision avoidance capabilities and positioning accuracy.

[0080] Tactile Array Perception: This innovative approach uses a flexible tactile sensor array to acquire two-dimensional pressure distribution and three-dimensional contact position information on the contact surface, achieving the fusion of tactile, visual, and three-dimensional perception to enhance device interactivity.

[0081] This embodiment is the core of the system's virtual-real environment compatibility, providing a hardware-level video signal co-source splitting and bypass acquisition scheme. As a key external signal source for the system's multimodal sensing unit, it addresses the technical pain points of traditional software acquisition (screen capture, screen recording, memory reading) such as easy shielding, high host resource consumption, and high latency. Simultaneously, it clarifies the technical roadmap of "ensuring a minimum pure vision experience and future pure vision development," balancing current practical application with long-term innovation. The specific implementation method is as follows:

[0082] Multimodal signal acquisition source (screen image and cursor acquisition)

[0083] In this embodiment, a multimodal signal acquisition scheme based on video signal source splitting and bypass acquisition is provided for the system. As one of the key external signal sources of the multimodal sensing unit of the system, it is used to acquire the screen display and mouse cursor position information in real time. It is the core technical support for the system to achieve compatibility between virtual and real environments and is also an important basic innovation of this patent.

[0084] The specific implementation process of this embodiment is as follows: The original video signal output by the host is processed by the video splitter unit through physical layer one-to-two splitting to form a first video signal and a second video signal. The first video signal is output to the display device to complete the normal screen display; the second video signal, as an independent bypass acquisition signal source, is input to the video acquisition unit, and after analog-to-digital conversion, it forms a digital real-time video stream, which is then sent to the virtual acquisition environment of this system. The virtual acquisition environment parses and processes the real-time video stream, extracts the screen display information, and obtains the coordinate position of the mouse cursor on the screen through image recognition. The screen data and cursor coordinate data are encapsulated into a unified format multimodal sensing signal, which is then used as an external input signal source to access the system's central control unit, providing underlying data support for the system's status monitoring, behavior analysis, and intelligent control.

[0085] This implementation method employs video signal source splitting and bypass acquisition technology. By performing physical layer signal splitting and bypass capture on the original video signal output by the host, one display signal is split into a second independent acquisition signal source without modification, interruption, or intrusion into the host system. This signal is then sent to a virtual acquisition environment for real-time analysis of the image and cursor position, providing highly reliable multimodal perception input for the system. The virtual acquisition environment uses a lightweight image analysis algorithm to further reduce acquisition latency, ensure synchronization with the display image, and adapt to video signals of different resolutions and interfaces, thus improving compatibility.

[0086] This implementation method differs from traditional software-based acquisition methods such as screen capture, screen recording, and memory reading. It directly acquires display data from the signal link level through hardware-level signal offloading and bypass capture. This technology boasts high technical distinctiveness and quickly demonstrates the core advantages of underlying hardware-level control, perfectly matching the overall control positioning of this system. Furthermore, this solution explicitly states that hardware-level video signal offloading is the preferred solution for rapid implementation, guaranteeing support for pure visual acquisition. Future optimization and upgrades will continue towards pure visual acquisition, balancing current feasibility with long-term technological development. The acquisition process achieves parallel processing of display and acquisition, ensuring that the image and cursor information remain synchronized from the same source. It does not consume host computing resources, is not easily blocked by applications, and features low latency, high stability, and high compatibility. This significantly improves the system's perception capabilities and operational reliability in human-computer interaction scenarios, effectively protecting the hardware-level bypass acquisition technology, distinguishing it from traditional software acquisition modes.

[0087] Universal for all devices: It is not limited to robots or device forms, and can be connected to various physical devices (such as automated equipment, mobile terminals, industrial machinery, etc.). There is no need to develop separately for different devices, achieving unified intelligent upgrades and reducing adaptation costs. The innovation lies in breaking the limitations of device form and building a universal intelligent architecture, which is different from traditional targeted intelligent systems.

[0088] Full-domain perception fusion: Simultaneously identify, perceive, and fuse information from the two-dimensional and three-dimensional worlds, covering multiple perception modalities such as vision, laser, radar, force, touch, thermal imaging, spectroscopy, and inertia. The synchronous fusion and joint modeling of multi-modal perception breaks through the limitations of a single perception modality and adapts to various complex scenarios.

[0089] Compatible with both virtual and real environments: It can perceive the real physical world and, through the hardware-level co-source bypass acquisition technology of this embodiment, access virtual environment signals such as screen images and cursors. The unified access and fusion of virtual and real perception signals expands the application boundaries of the system, while clarifying the baseline of pure vision and the future development direction, taking into account both practical application and long-term innovation. Non-intrusive acquisition innovation: The core adopts hardware-level video signal splitting and bypass acquisition technology, which does not intrude on the host system, does not occupy host computing resources, is not blocked by applications, and does not interrupt normal display. It has the characteristics of low latency, high stability, and high compatibility. The implementation of non-intrusive, hardware-level acquisition solves the technical pain points of traditional software acquisition, requires no modification to the original equipment system, and is easy to deploy.

[0090] Innovation in a practical bionic neural architecture: It adopts a closed-loop process of perception-neuron-execution, and calls mature APIs from third-party vendors for core decision-making functions. It does not develop complex brain modules in-house. The lightweight and practical design of the architecture reduces the difficulty of research and development and shortens the implementation cycle, which is different from the complex architecture of traditional intelligent systems.

[0091] High robustness and adaptability: The complementary fusion of multiple sensors can adapt to various actual working conditions such as complex environments, low light, occlusion, long distance, and high-precision positioning. The complementary design of multimodal perception improves the reliability of the system in practical applications, reduces operation and maintenance costs, and adapts to the needs of application in all scenarios.

[0092] 1. Equipment jamming alarm implementation

[0093] The system uses visual sensors to collect the motion trajectory of the equipment's execution components (such as robotic arms and moving mechanisms) in real time. Combined with motion parameters collected by auxiliary sensors (such as encoders and torque sensors), it collaboratively analyzes the equipment's operating status across modalities. If, within three consecutive sampling cycles (the sampling cycle is adjusted according to actual working conditions), the system visually detects no displacement change in the equipment's execution components, and the torque value collected by the auxiliary sensors is ≥ a preset threshold (adjusted according to actual working conditions) and the speed value is ≤ a preset speed (adjusted according to actual working conditions), the system determines that the equipment is stuck and immediately triggers an alarm. The alarm method uses audible and visual alarms (buzzer sound + indicator light flashing), and simultaneously pushes alarm information through a mobile application, clearly marking "equipment stuck" and the suspected stuck location (e.g., "robotic arm joint 3 stuck"), reminding the user to stop the machine promptly and check for obstructions, component wear, and other problems.

[0094] 2. Implementation of alarm for no response during operation

[0095] The system monitors various control commands in real time, including multilingual programming instructions, mobile application operations, and instant messaging commands. It combines auditory sensors (to recognize voice commands), visual sensors (to recognize actions triggered by the user interface), and the system's internal command receiving module to verify command transmission and execution status across modalities. If, after a command is issued, no device action is detected within a preset response time limit (e.g., 1 second, adjustable according to actual operating conditions), the system does not return command execution feedback, and the visual sensor does not capture movement of the device's executing components, and the auditory sensor does not collect the device's operating sound, the system is considered unresponsive and an alarm is triggered. The alarm message simultaneously displays the command content (e.g., "Unresponsive 'Move to coordinates (50,80) (adjustable according to actual operating conditions)' command)" and automatically retryes command transmission once. If there is still no response, the user is prompted to check the command transmission link and the device's power supply status.

[0096] 3. Implementation of alarms for missing perception data

[0097] The system continuously monitors real-time data from vision, hearing, and auxiliary sensors, comparing the data transmission frequency and integrity of each sensor across modalities. If any sensor (such as a vision sensor) fails to transmit data for five consecutive sampling cycles (adjustable according to actual operating conditions), or if the transmitted data is missing or abnormal (e.g., a black screen for vision, or a signal-to-noise ratio ≤20dB for hearing data), and other sensors cannot compensate for this data loss (e.g., auxiliary sensors cannot replace vision sensors in acquiring device location information), it is determined that sensor data is missing, triggering an alarm. The alarm clearly indicates the sensor type with missing data (e.g., "Vision sensor data missing"), and simultaneously activates a backup data acquisition mode (e.g., enabling redundant auxiliary sensors) to maintain basic system operation, reminding the user to check sensor connections, wiring, and device status.

[0098] 4. Implementation of interactive timeout alarms

[0099] For user-system interactions (such as parameter setting, command confirmation, and exception handling feedback), the system presets an interaction time limit (e.g., 30 seconds). It uses visual sensors (to recognize user actions) and auditory sensors (to recognize user voice feedback) to collaboratively assess the interaction status. If, after an interaction begins (e.g., a user initiates a parameter modification operation), no further user actions or voice feedback are detected within the preset time limit, and the system receives no interaction confirmation command, the interaction is considered timed out, triggering an alarm. The alarm uses a pop-up window on the interface combined with a voice prompt (e.g., "Interaction timed out, please complete the operation promptly"). Simultaneously, the current interaction progress is saved to prevent data loss, reminding the user to complete subsequent operations or cancel the interaction as soon as possible.

[0100] All of the above alarms support collaborative correction. After an alarm is triggered, the system automatically records abnormal parameters, abnormal time and related sensor data based on cross-modal analysis results, providing data support for users to troubleshoot problems. After the user handles the abnormality, the system re-analyzes the data through multimodal data. Once the abnormality is confirmed to be resolved, the alarm is automatically terminated and the equipment is restored to normal operation.

[0101] The multimodal sensing module adopts a modular design, with each sensor unit independently packaged and having standardized interfaces. Sensor types can be flexibly added according to actual application scenarios without refactoring the module's core logic. Below are some typical examples of flexible combinations, focusing on supplementing the identification features suitable for precise target locking, continuous tracking, and industrial scenarios, thus meeting the needs of industrial equipment monitoring, target control, and dynamic tracking:

[0102] 1. Basic Scenario Combination: High-definition industrial camera (visual sensor) + auditory sensor (microphone). The high-definition industrial camera supports high-definition imaging, dynamic capture, and detail recognition, and can identify basic target features (such as human body outline, personnel facial features, equipment appearance outline, and equipment identification code). The auditory sensor can recognize voice commands and abnormal noises from equipment operation. Together, they meet the basic multimodal perception and preliminary target recognition needs, and are suitable for basic industrial site monitoring, personnel attendance, and preliminary equipment inspection scenarios, laying the foundation for subsequent precise control.

[0103] 2. Complex Industrial Environment Combination: Building upon the basic combination, infrared and ultrasonic sensors are added. The high-definition industrial camera can capture targets in low-light environments. Combined with the infrared sensor, it can achieve temperature sensing and target recognition in dim and obscured scenes (it can identify human body surface temperature characteristics and equipment heating characteristics, accurately distinguishing between people, industrial equipment, and debris). The ultrasonic sensor is used for distance measurement, obstacle detection, and target movement trajectory prediction. It is suitable for complex industrial scenarios such as outdoor industrial plants and warehouse obscured areas, enabling accurate monitoring of personnel and equipment.

[0104] 3. High-precision industrial scenario combination: On top of the basic combination, a lidar sensor and an inertial measurement unit (IMU) are added. The high-definition industrial camera can capture the subtle features of the target (such as loose screws, personnel operation gestures, and scratches on the surface of the equipment). The lidar improves the spatial positioning accuracy and can capture the target's three-dimensional contour, movement trajectory, and limb movement features (such as personnel operation posture and equipment operation movements) to achieve accurate target positioning and action prediction. The IMU captures changes in equipment posture and is suitable for precise monitoring of industrial equipment and tracking of standardized personnel operation in scenarios with high requirements for positioning, posture perception, and target feature recognition.

[0105] 4. Industrial Safety Monitoring Suite: Building upon the basic suite, this suite adds gas and vibration sensors. High-definition industrial cameras can monitor in real time for personnel violations and abnormal equipment conditions (such as equipment leaks or component detachment). Gas sensors are used to detect the concentration of harmful gases in the industrial environment, identify safety hazards, and assist in identifying gas characteristics left by target activities (such as volatile gases from personal protective equipment or leaked gases from equipment), enabling hazard tracing. Vibration sensors are used to monitor equipment vibration, predict equipment failures, and capture ground vibration characteristics generated by target movement, making it suitable for safety monitoring scenarios in concealed areas and hazardous work areas of industrial plants.

[0106] 5. Industrial Autonomous Cruise and Target Control Combination: Building upon the basic combination, a new millimeter-wave radar sensor and thermal imaging sensor are added. High-definition industrial cameras support wide-area cruise capture and target magnification recognition, accurately identifying details such as employee badges, equipment numbers, and equipment operating status. Millimeter-wave radar enables long-distance target detection and speed measurement in complex industrial environments such as rain, snow, and dense fog, capturing target movement speed, distance, and direction of movement to accurately predict personnel and equipment movement trajectories. The thermal imaging sensor overcomes light limitations, accurately capturing the surface thermal radiation characteristics and body contours of living targets (such as workers) and abnormal heating of industrial equipment. It can also identify differentiated features such as clothing color, body shape, height, and posture, and capture details of human thermal radiation distribution (distinguishing workers from industrial simulation models and debris). This combination is suitable for autonomous cruise in industrial plants, worker management, and equipment anomaly monitoring scenarios, achieving precise target locking and dynamic tracking to ensure comprehensive and accurate monitoring.

[0107] 6. Industrial Anti-interference and Redundancy Protection Combination: Building upon the basic combination, an electromagnetic interference sensor and a backup inertial measurement unit (IMU) are added. The high-definition industrial camera has anti-electromagnetic interference capabilities, enabling it to stably capture target features and transmit image data in strong industrial electromagnetic environments. The electromagnetic interference sensor monitors the surrounding electromagnetic environment in real time and automatically switches to anti-interference mode when strong electromagnetic interference is detected (such as in industrial machine tools or signal shielding scenarios), ensuring uninterrupted acquisition of target recognition features (such as image details, voice, thermal radiation, and body movements). The backup IMU seamlessly takes over when the main sensor fails, ensuring the continuous stability of the sensing system and target recognition function. This meets the core requirements of continuous and stable monitoring, anti-interference, and accurate target locking in industrial scenarios, and is suitable for continuous industrial monitoring scenarios in complex interference environments.

[0108] The above combinations are all directly spliced ​​and adapted through modular interfaces. Adding sensors and related recognition feature adaptations only requires configuring interface parameters and driver adaptation to integrate into the multimodal perception module, enabling flexible expansion of perception capabilities and target recognition accuracy. The high-definition industrial camera, as the core visual perception unit, can achieve detailed capture, accurate recognition, and dynamic monitoring in industrial scenarios. Combined with other sensors, it can accurately capture multi-dimensional target recognition features (covering body shape, odor, movement, thermal radiation, appearance details, etc.), adapting to various industrial target locking, continuous monitoring, and security control scenarios, ensuring accurate target identification and locking, and equipment status monitoring even in complex industrial environments.

[0109] I. Implementation Examples of Multi-Camera Fusion Scenarios

[0110] The fusion process employs a weighted fusion algorithm, specifically:

[0111] Suppose the multiple data streams to be fused are image data collected by camera one. Image data collected by camera 2 Image data collected by the three cameras ( (These are all pixel data corresponding to the image pixel value matrix), and preset weights are assigned to each camera data stream. ), ( ), ( ), where the weights satisfy: + + =1;

[0112] The fusion output result is calculated according to the following formula:

[0113]

[0114] As a specific embodiment, the weights are taken in this embodiment:

[0115] =0.3、 =0.4、 =0.3

[0116] Those skilled in the art can adjust the weight values ​​according to the actual application scenario, such as the viewing angle priority and image clarity of each camera.

[0117] II. Examples of Multi-Feature Fusion Scenarios

[0118] The fusion process employs a weighted fusion algorithm, specifically:

[0119] Let the multiple data streams to be fused be the contour feature data extracted from the image ( ), texture feature data ( ), color feature data ( (These are all data corresponding to the dimensions of the feature vectors), and preset weights are assigned to each feature data path. ), ( ), ( ), where the weights satisfy: + + =1;

[0120] The fusion output result is calculated according to the following formula:

[0121]

[0122] As a specific embodiment, the weights are taken in this embodiment:

[0123] =0.4、 =0.3、 =0.3

[0124] Those skilled in the art can adjust the weight values ​​according to the actual application scenarios, such as the extraction accuracy and importance of each feature.

[0125] III. Example of Image + Control Signal Fusion Scenarios

[0126] The fusion process employs a weighted fusion algorithm, specifically:

[0127] Let the multiple data streams to be fused be image feature data acquired by the camera ( ), equipment control signal data ( ), auxiliary control signal data ( ), (( ) represents the target coordinate data, ), ( (This refers to equipment operating status data), and a preset weight is assigned to each data stream. ), ( ), ( ), where the weights satisfy + + =1;

[0128] The fusion output result is calculated according to the following formula:

[0129]

[0130] As a specific embodiment, the weights are taken in this embodiment:

[0131] =0.4、 =0.3、 =0.3

[0132] Those skilled in the art can adaptively adjust the weight parameters according to actual application scenarios such as image positioning accuracy requirements and device control priorities; or they can determine the optimal weight parameters through autonomous learning and iterative optimization.

[0133] This system can dynamically optimize weight parameters based on cloud computing power. According to the current scene interference level, it can assign adaptive fusion weights to the input data of visual sensors, auditory sensors and auxiliary sensors to achieve robust fusion of multimodal information and decision output.

[0134] In practice, the system evaluates the signal-to-noise ratio, confidence level, and scene interference level of each modality in real time. The cloud server dynamically adjusts the weight coefficients of each modality based on historical data and real-time inference results. To more intuitively illustrate the weight allocation mechanism, the following are weight examples for typical scenarios:

[0135] Normal interference-free scenario

[0136] Visual data weight: 0.6

[0137] Auditory data weight: 0.2

[0138] Auxiliary sensor weight: 0.2

[0139] Scenes with strong light / obstructions and other visual interference

[0140] Visual data weight: 0.3

[0141] Auditory data weight: 0.4

[0142] Auxiliary sensor weight: 0.3

[0143] scenarios with strong noise and other auditory interference

[0144] Visual data weight: 0.7

[0145] Auditory data weight: 0.1

[0146] Auxiliary sensor weight: 0.2

[0147] Complex and highly interfering scenarios (both visual and auditory senses are affected)

[0148] Visual data weight: 0.4

[0149] Auditory data weight: 0.2

[0150] Auxiliary sensor weight: 0.4

[0151] The weights are not limited to the values ​​mentioned above, and the system supports self-learning and iterative optimization based on execution results.

[0152] The cloud updates the above weights in real time according to the scene interference level, enabling the system to maintain stable recognition and control accuracy in different environments. The above weights are only for illustrative purposes, and in actual applications, they can be adaptively adjusted and optimized according to the algorithm model, hardware configuration, and application scenario.

[0153] The multi-language real-time programming instructions support natural language input in at least two languages. The basic languages include Chinese and English, and the extended languages include Japanese, Korean, etc. It adopts an offline multi-language recognition + online large model collaborative translation and parsing architecture, and the parsing algorithm can be implemented by calling the cloud API to adapt to multi-language usage scenarios.

[0154] 1. Examples of basic languages (Chinese, English) The offline end has built-in Chinese and English recognition models, and basic instruction recognition can be completed without a network connection. When connected to the network, accurate parsing and instruction standardization are achieved through the cloud API.

[0155] Example of Chinese instruction: "Move to coordinates (50, 80) (can be adjusted according to the actual working conditions), execute recognition and control", corresponding English meaning: Move to coordinates (50, 80), execute recognition and control. After offline recognition, call the cloud parsing interface to convert it into a standard control instruction executable by the system;

[0156] Example of English instruction: "Move to coordinates (50, 80), execute recognition and control", corresponding Chinese meaning: Move to coordinates (50, 80) (can be adjusted according to the actual working conditions), execute recognition and control. After offline recognition + online parsing, output a standardized task flow homologous to Chinese.

[0157] 1. Examples of extended languages (Japanese, Korean) For Japanese and Korean, parsing is achieved through the extended interface + upper-layer large model translation ability, without reconstructing the core logic of the system. Only configure the interface parameters and language models to adapt.

[0158] Example of Japanese instruction: "移動し、認識と制御を実行せよ", corresponding Chinese meaning: Move to coordinates (50, 80) (can be adjusted according to the actual working conditions), execute recognition and control, corresponding English meaning: Move to coordinates (50, 80), execute recognition and control. The system specifies the language parameter as Japanese through the extended interface and calls the cloud translation and parsing API to convert it into a unified standard instruction;

[0159] Example of a Korean command: “좌표(50,80)로이동하고인식및제어를실행하라”, which means: Move to coordinates (50,80) (can be adjusted according to actual working conditions), execute recognition and control. The corresponding English meaning is: Move to coordinates(50,80), execute uterecognitionandcontrol. The specified language parameter is Korean, which is directly mapped to the same execution logic after translation and parsing, realizing the equivalent execution of multilingual commands.

[0160] The above parsing process reuses a unified core framework, completing the identifier configuration and model loading only at the language adaptation layer, which can quickly support the parsing of multilingual real-time programming instructions.

[0161] Results reporting: The system connects to the Internet via 4G SIM card slot and reports to the upper-level AI agent and user that "the air compressor has started, the pressure has stabilized at 0.75MPa, and the high-precision robotic arm and stepper motor are operating normally." The system archives operation logs, interaction records, and operating parameters of the robotic arm and stepper motor. The log data is synchronized to cloud storage via IoT communication, providing data support for subsequent optimization of the system's low-threshold, high-robustness characteristics and the control precision of the actuators.

[0162] Please see Figure 1 Example 2: Multilingual real-time modification closed-loop control of touch electronic devices

[0163] In this embodiment, the intelligent control terminal is deployed next to the touch tablet. It can achieve real-time communication and interaction with the user's mobile terminal through various existing mainstream network connection methods (this embodiment uses 2.4G IoT networking). The core algorithm can be implemented by calling cloud APIs (encrypted transmission). The device is a standard miniaturized form, small in size and regular in shape. A USB interface is reserved on the side for connecting to the touch simulation execution unit. Multiple status indicator lights on the front display the network and operating status. Task command: The foreign engineer issues the command "Open the production line report and switch to the daily view" via English voice. During execution, the user can interrupt and modify the command to "Switch to the weekly report view" in real time via instant messaging. Execution flow:

[0164] Multimodal acquisition and recognition: When the controlled object is identified as a touch electronic device, the matched execution channel is a touch simulation execution channel. The control command is a transparent capacitive touch driving command, which can control the logic through simulated capacitance changes to accurately simulate touch operations. The visual sensor recognizes the "Report APP" icon on the tablet desktop and the current interface status; the audio sensor collects the operation feedback sound of the touch tablet; the fusion recognition result is "touch electronic device - industrial tablet - desktop status". The fusion algorithm can be implemented by calling the cloud API, which is accurate, robust and requires no manual assistance.

[0165] The multimodal sensing module adopts a modular design, with each sensor unit independently packaged and featuring standardized interfaces. It allows for flexible addition of sensor types based on actual application scenarios without requiring reconstruction of the module's core logic. Below are some typical examples of flexible combinations, with a focus on enhancing screen reading and industrial control panel reading functions, and adding automated monitoring capabilities to meet the needs of industrial equipment monitoring, target control, terminal monitoring, and dynamic tracking.

[0166] 1. Basic Scenario Combination: High-definition industrial camera (visual sensor) + auditory sensor (microphone). The high-definition industrial camera supports high-definition imaging, dynamic capture, detail recognition, and screen / panel reading functions. It can accurately read the button status, numerical display (such as pressure gauge and temperature controller readings), and indicator light status of industrial operation panels. It can also read the display content of industrial control terminal screens, monitoring screens, computer screens, and industrial information screens. It can identify basic target features (such as human outline, facial features, equipment outline, and equipment identification code). The auditory sensor can recognize voice commands and abnormal noises from equipment operation. Together, they meet the basic multimodal perception, preliminary target recognition, and basic screen / panel reading needs. It is suitable for basic industrial site monitoring, personnel attendance, preliminary equipment inspection, and simple terminal monitoring scenarios, laying the foundation for subsequent precise control and automated monitoring functions.

[0167] 2. Complex Industrial Environment Combination: Building upon the basic combination, infrared and ultrasonic sensors are added. The high-definition industrial camera can capture targets in low-light environments. Combined with the infrared sensor, it can achieve temperature sensing and target recognition in dimly lit and obscured scenes (it can identify human body surface temperature characteristics and equipment heating characteristics, accurately distinguishing between human bodies, industrial equipment, and debris). The ultrasonic sensor is used for distance measurement, obstacle detection, and target movement trajectory prediction. It is suitable for complex industrial scenarios such as outdoor industrial plants and warehouse obscured areas, enabling accurate monitoring of personnel and equipment.

[0168] 3. High-precision industrial scenario combination: Building upon the basic combination, a new LiDAR sensor and inertial measurement unit (IMU) are added. The high-definition industrial camera can capture subtle features of the target (such as loose screws, personnel operation gestures, and scratches on the equipment surface), significantly enhancing the accuracy of screen reading and industrial operation panel reading. It can accurately identify small buttons on the panel, small characters on the screen, and parameter values ​​(error ≤1%). It can distinguish different colors and brightness levels of panel indicator lights and accurately read real-time data and alarm information from industrial control terminal screens, computer screens, and industrial information screens. The LiDAR improves spatial positioning accuracy, capturing the target's three-dimensional contour, movement trajectory, and limb movement features (such as personnel operation postures and equipment operation movements) to achieve accurate target positioning and action prediction. The IMU captures changes in equipment posture, making it suitable for scenarios with high requirements for precise monitoring of industrial equipment, tracking of standardized personnel operations, and accurate reading of industrial panels.

[0169] 4. Industrial Safety Monitoring Suite: Building upon the basic suite, gas sensors and vibration sensors are added. High-definition industrial cameras can monitor in real time for personnel violations and abnormal equipment conditions (such as equipment leaks or component detachment). Gas sensors are used to detect the concentration of harmful gases in the industrial environment, identify safety hazards, and help identify gas characteristics left by target activities (such as volatile gases from personal protective equipment or leaked gases from equipment), enabling hazard tracing. Vibration sensors are used to monitor equipment vibration, predict equipment failures, and capture ground vibration characteristics generated by target movement, making them suitable for safety monitoring scenarios in concealed areas and hazardous work areas of industrial plants.

[0170] 5. Industrial Autonomous Cruise and Target Control Combination: Building upon the basic combination, this system adds millimeter-wave radar and thermal imaging sensors, and a high-definition industrial camera to support wide-area cruise capture and target magnification recognition. It significantly enhances screen reading, industrial control panel reading, and automated monitoring functions. It can accurately read real-time data, operating status, and alarm information from multiple industrial control panels, industrial control terminal screens, computer screens, and industrial information screens from a distance, automatically identifying abnormal screen pop-ups and panel fault indicator lights to achieve automated screen monitoring. It can accurately identify detailed features such as employee ID cards, equipment numbers, and equipment operating status. The millimeter-wave radar can achieve long-distance target detection and speed measurement in complex industrial environments such as rain, snow, and dense fog, capturing target movement. The sensor identifies speed, distance, and direction of movement to accurately predict the movement trajectories of personnel and equipment. Overcoming light limitations, the thermal imaging sensor accurately captures the surface thermal radiation characteristics and body contours of living targets (such as workers), as well as abnormal heating of industrial equipment. It can also identify differences in clothing color, body shape, height, and posture, and capture details of human thermal radiation distribution (distinguishing workers from industrial simulation models and debris). Adaptable to autonomous patrols in industrial plants, worker management, equipment anomaly monitoring, and industrial terminal monitoring scenarios, it achieves precise target locking, dynamic tracking, and automated monitoring, ensuring no omissions or misjudgments in monitoring. It efficiently completes core monitoring tasks such as screen viewing, reading, and anomaly identification, meeting the needs of industrial automation monitoring.

[0171] 6. Industrial Anti-interference and Redundancy Protection Combination: Building upon the basic combination, an electromagnetic interference sensor and a backup inertial measurement unit (IMU) are added. The high-definition industrial camera has anti-electromagnetic interference capabilities, enabling it to stably capture target features and transmit image data in strong industrial electromagnetic environments. The electromagnetic interference sensor monitors the surrounding electromagnetic environment in real time and automatically switches to anti-interference mode when strong electromagnetic interference is detected (such as in industrial machine tools or signal shielding scenarios), ensuring uninterrupted acquisition of target recognition features (such as image details, voice, thermal radiation, and body movements). The backup IMU seamlessly takes over when the main sensor fails, ensuring the continuous stability of the sensing system and target recognition function. This meets the core requirements of continuous and stable monitoring, anti-interference, and accurate target locking in industrial scenarios, and is suitable for continuous industrial monitoring scenarios in complex interference environments.

[0172] All of the above combinations are directly spliced ​​and adapted through modular interfaces. With the addition of sensor and recognition feature-related adaptations, only interface parameter configuration and driver adaptation are required to integrate them into the multimodal perception module, enabling flexible expansion of perception capabilities and target recognition accuracy. The high-definition industrial camera, as the core visual perception unit, focuses on enhancing screen reading and industrial operation panel reading functions. It can accurately read panel parameters, data, and alarm information from computer screens, industrial information screens, and industrial control terminal screens. Simultaneously, an automated monitoring function has been added, automatically completing real-time monitoring, data reading, and anomaly identification of industrial control terminal screens, monitoring screens, computer screens, and industrial information screens. This automates core monitoring tasks such as screen viewing, monitoring, reading, and anomaly warning, adapting to industrial terminal monitoring scenarios. Combined with other sensors, it can accurately capture multi-dimensional target recognition features (covering body shape, odor, movement, thermal radiation, appearance details, etc.), adapting to various industrial target locking, continuous monitoring, security control, and terminal monitoring scenarios. This ensures accurate target identification and locking, and equipment status monitoring even in complex industrial environments, effectively improving the automation level of industrial scenarios through automated monitoring.

[0173] Command parsing and decision-making: The real-time human-computer collaborative interaction unit parses English voice commands. The parsing algorithm can be implemented by calling cloud APIs. The unified decision engine breaks it down into "click the report APP icon → wait for loading → click the 'Daily Report' option", matching the touch simulation execution channel. Multi-language support reduces the operation threshold, and the command parsing is robust.

[0174] Command issuance and execution: The touch simulation execution interface issues coordinate touch commands to the transparent capacitive execution unit to complete APP clicks and loading waits. The execution is stable, adapts to the operation requirements of touch tablets, and has high robustness. The execution status is synchronized to the user's mobile terminal via 2.4G IoT network.

[0175] Real-time interruption and replanning: Before execution reaches the "click 'Daily Report' option", the user sends a real-time interruption command "switch to Weekly Report view" via instant messaging; the real-time replanning module immediately adjusts the atomic operation sequence to "click 'Weekly Report' option". The interruption operation is convenient and has a low threshold, and the replanning response is rapid, ensuring execution robustness.

[0176] Cross-modal analysis and correction: From visual recognition to the completion of weekly report view loading, audio collection to feedback sound, and the error meeting the threshold, the analysis algorithm can be implemented by calling the cloud API, resulting in accurate analysis and further enhancing the system's robustness; after the task is completed, the results are reported via 2.4G IoT network connection and the operation log is archived.

[0177] Please see Figure 1 Example 3: Real-time programming closed-loop control of computer terminal (analog keyboard and mouse signal adaptation)

[0178] In this embodiment, the intelligent control terminal is deployed on the office desktop. It can achieve data synchronization with the upper-level AI intelligent agent through various existing mainstream network connection methods (this embodiment uses NB-IoT). The core algorithm can be implemented by calling the cloud API (encrypted transmission). The device is small and has a regular shape. A USB port is reserved on the side for connecting a keyboard and mouse simulation execution unit. Anti-slip feet on the bottom prevent slipping. Multiple status indicator lights on the front distinguish between power, network, and running states. Task command: The user issues a real-time programming command via instant messaging: "Open office document software, create a new text document, name it 'Intelligent Control Financing Report V1.0'. If naming fails, actively ask if you want to change the name." Execution flow:

[0179] Multimodal Acquisition and Recognition: A visual sensor identifies the WPS icon and mouse cursor position on the computer terminal desktop, while an inertial measurement unit detects any displacement or shaking of the computer terminal. This multimodal information is then fused and recognized to determine the device status as "Computer Terminal - Office Use - Desktop Status." The fusion algorithm can be implemented via a cloud API, simultaneously identifying the computer terminal's compatibility with analog keyboard and mouse signal control, ensuring high robustness of the recognition results and eliminating the need for manual verification of device status and control method.

[0180] During the identification process, priority is given to detecting the presence of a signal source from the same screen. If such a signal source is found, it is used directly as the basis for determining the device status, with visual and other sensor information serving only as auxiliary verification. Simultaneously, compatible devices are detected via the computer terminal's 3.5mm headphone jack, Bluetooth, or other common audio / communication interfaces to ensure the system can receive and parse the voice signals transmitted through the corresponding interfaces, as well as analog voice commands sent via a microphone module or wirelessly, achieving stable voice reading and analog voice input.

[0181] When the multimodal sensing module identifies the controlled object as a touch electronic device (including industrial touch panels, touch operation terminals, touch computers, touch all-in-one machines, and touch human-machine interfaces) through a high-definition industrial camera, the system automatically matches the corresponding touch simulation execution channel and generates and outputs transparent capacitive touch drive commands.

[0182] This instruction achieves precise operation based on analog capacitance change control logic: the system generates capacitance coupling change signals in the corresponding coordinate area based on the touch points and operation types (click, swipe, long press, parameter adjustment, function switching) obtained from screen reading and panel recognition. This simulates the capacitance effect generated when a human finger touches the screen, enabling touch electronic devices to recognize and respond to operations without physical contact. It completes automated control actions such as parameter setting, instruction confirmation, data retrieval, interface switching, and abnormal reset of industrial touch interfaces, achieving stable, accurate, and contactless control of various touch-type industrial terminals without manual intervention.

[0183] Command parsing and decision-making: The real-time human-computer collaborative interaction unit parses user commands, and the parsing algorithm can be implemented by calling cloud APIs; the unified decision engine dynamically generates a standardized operation sequence of "double-click the office software icon → click New Document → enter document name → save" based on the parsing results, calls the keyboard and mouse simulation module, automatically matches the keyboard and mouse simulation execution channel, and determines to use the simulated keyboard and mouse signal method to complete the automated control. No manual intervention is required throughout the process, and the operation threshold is low.

[0184] Command Issuance and Execution: The keyboard and mouse simulation execution interface issues standard HID simulated keyboard and mouse signal commands to the execution unit. By simulating keyboard and mouse signals, it simulates manual double-clicking, clicking, text input, and other operations without intruding into the computer terminal system. The execution process is stable, highly compatible, and robust. It can also coordinate with high-precision robotic arms, stepper motors, and simulated capacitor change control. If the computer terminal needs to cooperate with physical button operation, it can be linked with the high-precision robotic arm to complete the task, adapting to multiple scenario control needs. The real-time status during the execution process is synchronized to the user's mobile terminal via NB-IoT Internet of Things networking.

[0185] In this embodiment, a multimodal signal acquisition scheme based on video signal splitting is provided for the intelligent control system. As one of the external signal sources of the system's multimodal perception unit, it is used to acquire the screen display and mouse cursor position information in real time.

[0186] The specific implementation process of this embodiment is as follows:

[0187] The raw video signal output from the host is split into two streams by the video splitter unit, forming a first video signal and a second video signal. The first video signal is output to the display device for normal screen display; the second video signal is input to the video acquisition unit, converted from analog to digital to form a digital real-time video stream, and sent to the system's virtual acquisition environment. The virtual acquisition environment analyzes and processes this real-time video stream, extracting screen display information and obtaining the mouse cursor's coordinates on the screen through image recognition. The screen data and cursor coordinate data are encapsulated into a unified format multimodal sensing signal, which is then used as an external input signal source to connect to the intelligent control system's central control unit, providing underlying data support for system status monitoring, behavior analysis, and intelligent control.

[0188] This implementation achieves parallel processing of display and acquisition through hardware-level signal splitting, keeping the screen and cursor information from the same source and in sync. The acquisition process does not consume host computing resources and is not easily blocked by applications. It features low latency, high stability, and high compatibility, which can improve the perception capability and operational reliability of the intelligent control system in human-computer interaction scenarios.

[0189] This implementation method hijacks and splits the original video signal output by the host at the physical layer. Without modifying, interrupting, or intruding on the host system, it hijacks one display signal into a second independent acquisition signal source and sends it into the virtual acquisition environment for real-time analysis of the screen and cursor position, providing highly reliable multimodal perception input for the intelligent control system.

[0190] Cross-modal analysis and proactive inquiry: When visual recognition fails to input a document name (the system prompts "name already exists"), the auxiliary sensor detects that the simulated keyboard and mouse signal execution unit is functioning normally and the signal transmission is stable. The real-time analysis unit triggers an intelligent inquiry mechanism. The analysis algorithm can be implemented by calling the cloud API to accurately determine the type of anomaly and send an inquiry message to the user via instant messaging: "The document name 'Smart Control Financing Report V1.0' already exists. Do you want to change it to 'Smart Control Financing Report V1.1'?" The user replies "yes" via instant messaging. This proactive inquiry mechanism lowers the threshold for manual intervention while ensuring the system's robustness in dealing with anomalies.

[0191] Correction execution: The real-time replanning module generates a correction sequence of "simulating keyboard and mouse signals to control the cursor to the end of the name → deleting 'V1.0' → entering 'V1.1' → clicking save". After iterative execution, the task is completed. The correction process is automated, with low threshold and strong robustness. The simulated keyboard and mouse signal control is accurate and without deviation.

[0192] Results reporting: Through NB-IoT Internet of Things networking, feedback is sent to users that "the financing report document has been created and named 'Smart Control Financing Report V1.1', and the simulated keyboard and mouse signal control is operating normally." Operation logs, text interaction records, and simulated keyboard and mouse signal transmission parameters are archived, and log data is synchronized to cloud storage, providing data support for subsequent optimization of the system's low-threshold and high-robustness characteristics and the accuracy of simulated keyboard and mouse control.

Claims

1. A multimodal fusion intelligent closed-loop control method and system, applied to the unified automated control of physical devices, touch-screen electronic devices, and computer terminals, characterized in that, Includes the following steps: (1) Multimodal information acquisition and fusion recognition: The state data of the controlled object is acquired through the multimodal perception module. The state data includes at least one of visual image data, auditory audio data and auxiliary sensor data. The state data is fused to identify the controlled object and extract the current state features of the controlled object. (2) Multi-channel real-time instruction reception and parsing: The system receives multilingual real-time programming instructions issued by the user or the upper-level AI agent through at least one of the following interfaces: voice interface, text interface, mobile application interface, instant messaging interface, and visual drag-and-drop programming interface. The instructions include task intent and execution logic. The system parses the instructions through the multilingual natural language understanding module to generate standardized task instructions. (3) Task parsing and execution channel decision: Combining the controlled object type, state characteristics and standardized task instructions, the task is decomposed into atomic operation sequences through a unified decision engine and the corresponding execution channel is automatically matched. The execution channel includes physical execution channel, touch simulation execution channel and keyboard and mouse simulation execution channel. (4) Standardized instruction issuance and execution: Convert atomic operation sequences into control instructions for corresponding execution channels, forward them to the target execution unit through the multi-channel execution interface, and have the execution unit complete the operation on the controlled object; (5) Cross-modal real-time analysis and closed-loop correction: The state data of the controlled object after execution is collected again through the multi-modal sensing module and compared with the preset expected results across modes to calculate the operation error; if the error exceeds the preset threshold, the corrected atomic operation sequence and control instructions are regenerated and iteratively executed until the error meets the threshold requirements; (6) Task closure and result reporting: When the operation error meets the threshold requirements, the task is determined to be completed, the execution result is fed back to the upper-level AI agent or user, and the operation log and correction data are archived.

2. The method according to claim 1, characterized in that, The fusion process described in step (1) employs a weighted fusion algorithm, specifically calculated according to the following formula: in , , These are the feature vectors corresponding to visual image data, auditory audio data, and auxiliary sensor data, respectively; , , The corresponding weight coefficients, and satisfying + + =1; the weighting coefficient is dynamically and adaptively adjusted according to the level of scene interference.

3. The method according to claim 2, characterized in that, The weighting coefficients are preset according to the following scenarios and support self-learning optimization: In normal, interference-free scenarios, the following settings are provided. =0.6、 =0.2、 =0.2; In scenarios where visual interference is present, set... =0.3、 =0.4、 =0.3; In scenarios where hearing is disturbed, set =0.7、 =0.1、 =0.2; In scenarios where both visual and auditory perception are impaired, set =0.4、 =0.2、 =0.4; The system records the success rate and operation error data of task execution in different scenarios, and optimizes the weight coefficients corresponding to each scenario through self-learning iteration to achieve the best balance between recognition accuracy and system stability.

4. The method according to claim 1, characterized in that, The specific implementation of the multilingual natural language understanding module in step (2) includes: setting up a speech-to-text module or calling a third-party speech-to-text API interface to process the audio signals collected by the multimodal perception module into speech-to-text, accurately converting them into standardized text instructions; calling multiple decision models, combined with the multilingual natural language understanding module, to perform joint reasoning on multi-source perception information and the text instructions, outputting multi-path model decision results, and improving the accuracy and robustness of instruction parsing; the multilingual natural language understanding module supports at least two languages ​​from Chinese, English, Japanese, and Korean to meet the multilingual instruction input requirements; for niche languages, by extending the interface to call the API interface of a third-party vendor, the translation and parsing of instructions can be completed with the help of its cloud translation service, without reconstructing the core logic of the system, thus achieving flexible adaptation to niche languages.

5. The method according to claim 1, characterized in that, The specific method by which the unified decision engine decomposes the task into a sequence of atomic operations in step (3) is as follows: Based on the type of the controlled object and the standardized task instructions, the corresponding operation template is matched from the pre-stored operation template library, and the operation template is instantiated into a specific atomic operation sequence. The atomic operation includes at least one of moving, pressing, rotating, clicking and inputting.

6. The method according to claim 1, characterized in that, In step (4): When the controlled object is identified as a touch electronic device, a touch simulation execution channel is matched, and the control command is a transparent capacitive touch driving command used to simulate capacitance change control logic; When the controlled object is identified as a physical device, a physical execution channel is matched, and the control command is a robotic arm drive command, a stepper motor drive command, or a displacement, force, or timing drive command of a mechanical actuator. When the controlled object is identified as a computer terminal, a keyboard and mouse simulation execution channel is matched, and the control command is a standard keyboard and mouse communication protocol command.

7. The method according to claim 1, characterized in that, The specific method for calculating the operational error in step (5) is as follows: The Euclidean distance between the state feature vector collected after execution and the expected result feature vector is calculated. If the distance value is greater than a preset threshold, it is determined that there is an error.

8. The method according to claim 1, characterized in that, Step (5) also includes an exception handling mechanism: If the number of consecutive iteration failures exceeds a preset threshold, or if abnormal scenarios such as device stagnation, unresponsive operation, or missing perception data are detected, an abnormal alarm will be triggered and the task will be terminated; the abnormal alarm includes at least one of device stagnation alarm, unresponsive operation alarm, missing perception data alarm, and interaction timeout alarm.

9. The method according to claim 1, characterized in that, Step (5) also includes an intelligent query mechanism: When encountering uncertain operational logic, parameter settings, or abnormal causes during the analysis process, an intelligent inquiry is initiated to the user via instant messaging to receive the user's real-time response and dynamically adjust the execution strategy; if the user does not respond for more than a preset time limit, an abnormal alarm is triggered and the task is terminated.

10. The method according to claim 1, characterized in that, The real-time programming instructions in step (2) support real-time interruption and real-time modification. The real-time interruption includes at least one of the following: multilingual voice interruption, interruption by inputting text instructions through access to third-party office software or proprietary App, one-click interruption via mobile application and instant messaging instructions, and real-time modification interruption via visual drag-and-drop programming interface.

11. A multimodal fusion intelligent closed-loop control system for implementing the method as described in any one of claims 1-10, characterized in that, include: (1) A multimodal sensing unit, including a visual sensor, an audio sensor and an auxiliary sensor, is used to collect state data of the controlled object, perform fusion processing on the state data, and output the type and state characteristics of the controlled object; (2) Real-time human-computer collaborative interaction unit, including multilingual voice acquisition module, multilingual natural language understanding module, full-channel instruction receiving module, visual drag-and-drop programming module and intelligent inquiry module, used to receive multilingual real-time programming instructions, parse and generate standardized task instructions, and initiate intelligent inquiries to users in abnormal scenarios; (3) A unified decision scheduling unit, including a processor and a memory, wherein the memory stores computer programs for a decision engine, a task decomposition module and a real-time replanning module, and the processor performs the following functions when executing the computer program: combining the controlled object type, state characteristics and standardized task instructions, decomposing the task into a sequence of atomic operations and automatically matching the corresponding execution channel; (4) A multi-channel execution interface unit, including a physical execution interface, a touch simulation execution interface and a keyboard and mouse simulation execution interface, is used to convert atomic operation sequences into control instructions for the corresponding execution channels and forward them to the target execution unit; (5) Real-time analysis unit, including processor and memory, wherein the memory stores computer programs of error calculation module, correction instruction generation module and anomaly detection module, and the processor performs the following functions when executing the computer program: cross-modal comparison of state data before and after execution, calculation of operation error, generation of correction instruction, and detection of abnormal scenarios; (6) Communication and storage unit, including IoT communication module and memory, for data transmission with upper-layer AI intelligent agent, mobile application and instant messaging interface, while archiving operation log and correction data.

12. The system according to claim 11, characterized in that, The visual sensor is an industrial-grade camera, the audio sensor is a noise-canceling microphone, and the auxiliary sensor includes at least one of a tactile sensor, an inertial measurement unit, and a pressure sensor.

13. The system according to claim 11, characterized in that, The IoT communication module supports at least one network connection method among 4G / 5G networks, Bluetooth, 2.4G wireless communication, NB-IoT, LoRa and wired Ethernet, and supports offline storage, automatically synchronizing data after the network is restored.

14. The system according to claim 11, characterized in that, The system adopts a miniaturized, integrated intelligent control terminal form. The shell is made of flame-retardant ABS material with a frosted surface. Multiple status indicator lights are set on the front to display power, network and operating status. Various adapter interfaces are reserved on the side, and anti-slip pads are set on the bottom.