Distributed intelligent multi-machine cooperative operation system and method, deployment device and equipment

The distributed intelligent multi-machine collaborative operating system solves the reliability and flexibility issues of centralized architecture, realizes efficient and autonomous operation of multi-machine collaborative operation, improves the system's fault tolerance and equipment utilization, and supports plug-and-play of devices from multiple vendors.

CN120972664APending Publication Date: 2025-11-18CHONGQING RES INST OF HARBIN UNIV OF TECH +1

Patent Information

Application Number
CN202511080756.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

The low reliability, rigid static task allocation, poor compatibility with heterogeneous devices, and insufficient fault tolerance of the centralized architecture in the existing technology result in poor overall system reliability and flexibility, making it difficult to meet real-time requirements and dynamic production needs.

Method used

It adopts a distributed intelligent multi-machine collaborative operating system, establishes a communication network through the OpenHarmony distributed soft bus, and combines dynamic task allocation algorithm, heterogeneous device adaptation module and fault tolerance and self-healing module to realize real-time data transmission, resource optimization and fault recovery between machines.

Benefits of technology

It enables efficient and autonomous operation of multi-machine collaborative operation, reduces equipment conflict rate and task delay, improves system fault tolerance and equipment utilization, supports plug-and-play of multi-vendor and multi-protocol devices, and reduces the need for manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120972664A_ABST
    Figure CN120972664A_ABST
Patent Text Reader

Abstract

The invention provides a distributed intelligent multi-machine cooperative operation system and method, a deployment device and equipment, belongs to the technical field of industrial automation and intelligent operation system crossing, and aims to solve the problems of low reliability, rigid static task distribution, poor compatibility of heterogeneous equipment and insufficient fault-tolerant capability caused by a centralized architecture in the prior art. The operating system comprises a communication module, a task scheduling module, a heterogeneous equipment adaptation module, a fault-tolerant and self-healing module and a human-computer interaction module. The operating method comprises the following steps: S1, establishing an equipment communication network; s2, equipment state visualization and task instruction issuing; s3, dynamically allocating tasks and scheduling the tasks; and S4, carrying out equipment fault monitoring and self-healing recovery strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a distributed intelligent multi-machine collaborative operating system, method, deployment device and equipment, belonging to the cross-technical field of industrial automation and intelligent operating system. Background Technology

[0002] With the rapid development of Industry 4.0 and intelligent manufacturing, multi-machine collaborative operation has become a key technology for improving production efficiency and accomplishing complex tasks. However, existing technologies face the following bottlenecks: 1. Reliability defects of centralized architecture: Traditional systems rely on a single central controller for task allocation and scheduling (e.g., patent CN110123456A). If the controller fails or communication is interrupted, the entire system collapses. 2. Rigidity of static task allocation: Fixed task allocation strategies cannot adapt to dynamic production needs (e.g., order changes, sudden equipment failures), leading to low resource utilization (see literature "Research on Collaborative Control Technology of Industrial Robots"). 3. Difficulty in coordinating heterogeneous equipment: Industrial robots, AGVs, sensors, and other equipment from different manufacturers have significantly different protocols, making unified scheduling difficult (e.g., patent US20210056789B2). 4. Insufficient fault tolerance: Existing systems lack rapid fault recovery mechanisms, requiring manual intervention after equipment malfunctions, resulting in significant downtime losses.

[0003] Current mainstream solutions have the following limitations: 1. Centralized optimization algorithms: Although they can achieve globally optimal scheduling, they suffer from high computational latency, making it difficult to meet real-time requirements. 2. Fixed communication protocols: Using a single PROFINET or EtherCAT protocol, they cannot be compatible with older equipment or cross-platform hardware. 3. Passive fault-tolerant design: Reliability is improved solely through hardware redundancy, which is costly and lacks flexibility. Summary of the Invention

[0004] To address the problems of low reliability, rigid static task allocation, poor compatibility with heterogeneous devices, and insufficient fault tolerance caused by centralized architecture in the prior art, this invention proposes a distributed intelligent multi-machine collaborative operating system, method, deployment device, and equipment.

[0005] The technical solution adopted by this invention to solve the above problems is: the distributed intelligent multi-machine collaborative operating system proposed in this invention includes: The communication module uses the OpenHarmony distributed soft bus method to establish a distributed communication network between machines for real-time data transmission and synchronization. The task scheduling module adopts a multi-machine device collaboration strategy based on a dynamic task allocation algorithm and optimizes resource allocation in real time. Heterogeneous device adapter module: The heterogeneous device adapter module connects robots, sensors and actuators with different hardware specifications to the distributed intelligent multi-machine collaborative operating system through the OpenHarmony unified interface protocol; The fault tolerance and self-healing module is used to monitor the system's operating status in real time and trigger task migration and redundant resource scheduling when a machine or equipment failure is detected. The human-computer interaction module provides a visual interface and operation command interface based on OpenHarmony, allowing operators to intervene and authorize collaborative strategies in real time.

[0006] Furthermore, the communication module includes: A real-time communication unit based on OpenHarmony time-sensitive networking is used to ensure that the transmission latency of critical task data is less than 10ms. Distributed blockchain nodes use OpenHarmony-based distributed data management capabilities to store device identity information and operation logs, ensuring that the data is immutable; An adaptive channel switching unit is used to dynamically select one of the communication frequency bands among 5G, Wi-Fi 6, and LoRa based on the intensity of environmental interference.

[0007] Furthermore, the heterogeneous device adaptation module includes: The device abstraction layer uses the OpenHarmony driver framework to encapsulate the driver interfaces of different robots into standard APIs. The resource virtualization unit is used to map the computing power, storage space and sensor data of physical devices into a virtual resource pool; Protocol converter, compatible with ROS2, OPC UA and Modbus communication protocols; The resource management unit enables global visibility of device resources through a dynamic registration and discovery mechanism, and allows for rapid retrieval and combined invocation of devices.

[0008] Furthermore, the fault tolerance and self-healing module includes: Redundant equipment pre-deployment strategy unit, used to configure backup machines on the critical path; The fault detection unit uses the OpenHarmony distributed heartbeat mechanism and sensor data analysis to determine whether the device is offline or abnormal. The task migration unit is used to reassign unfinished tasks from faulty devices to healthy devices and restore data consistency. The fault diagnosis unit uses a strategy that combines rule engine and graph reasoning to obtain local fault diagnosis results, regional fault diagnosis results and global fault diagnosis results, and generates a repair priority list.

[0009] Furthermore, the machine-to-machine interaction module includes: An AR augmented reality interface based on OpenHarmony is used to overlay and display the location of machines and equipment, task status, and a 3D map of the environment. The access control unit is used to differentiate the operating permissions of administrators, operators, and visitors; Emergency intervention interface, which supports one-click pause of task or switching to manual control mode; The natural language interaction module integrates speech recognition and semantic parsing to accurately identify user commands.

[0010] Distributed intelligent multi-machine collaborative operation methods include: Step 1: Initialize the distributed intelligent multi-machine collaborative operating system based on the communication module, establish a communication network between all machines and devices, and connect the robot, sensor and actuator to the distributed intelligent multi-machine collaborative operating system and verify their identity information through the heterogeneous device adapter module; Step 2: The operator views the device status, device location, and 3D map of the environment through the human-computer interaction module, issues task instructions, and uses a pre-trained language model to parse the contextual semantics, accurately identify the instructions, and send them to the task scheduling module. Step 3: The task scheduling module divides the collaborative subgroups according to the task requirements. Each subgroup includes at least one master control node and multiple execution nodes. The master control node generates control commands based on the dynamic task allocation algorithm and distributes them to the execution nodes through the communication module. The execution nodes feed back status data to the master control node, forming a closed-loop control link. Step 4: Monitor the system's operating status in real time through the fault tolerance and self-healing module. When a machine or equipment failure is detected, the unfinished tasks of the faulty equipment are reassigned to healthy machines or equipment, and data consistency is restored.

[0011] Furthermore, step 3 specifically includes: Step 3.1: The task scheduling module receives the task instructions from the human-computer interaction module and parses the task objectives and constraints; Step 3.2: Collect real-time status data of each device using OpenHarmony's distributed capabilities, including location, power level, load capacity, and sensor information; Step 3.3: Divide the collaborative subgroups according to the physical location topology of the devices, task priority and deadline, and the matching degree between device capabilities and task requirements. Each subgroup shall include at least one master node and multiple execution nodes. Step 3.4: The master node calculates the optimal coordination strategy and generates a set of device control instructions by invoking the dynamic task allocation algorithm; the dynamic task allocation algorithm specifically includes: The algorithm is based on one or more of the following: Q-Learning algorithm based on reinforcement learning, Nash equilibrium allocation model based on game theory, and multi-objective optimization model based on genetic algorithm. Specifically, Q-Learning algorithm based on reinforcement learning optimizes long-term task efficiency through reward function; Nash equilibrium allocation model based on game theory is used to balance resource competition among multiple machines; and multi-objective optimization model based on genetic algorithm is used to minimize task completion time and energy consumption.

[0012] Step 3.5: Use the target device corresponding to each instruction as the execution node, send the device control instructions to the target device through the communication module, and monitor the progress of the task. Step 3.6: Dynamically adjust the allocation of equipment control commands based on environmental changes or equipment status updates.

[0013] Furthermore, step 4 specifically includes: Step 4.1: Detect equipment malfunctions using the fault detection unit; Step 4.2: When a device malfunction is detected, assess the scope of the malfunction's impact and classify the affected tasks into levels. Step 4.3: Resume task execution by calling the backup equipment configured in the redundant equipment pre-deployment strategy unit or by adopting one of the resource reallocation strategies. The resource reallocation strategy includes the proximity principle, the capability matching principle, and the cost minimization principle. The proximity principle is used to prioritize the physical location of the equipment to take over the task; the capability matching principle is used to select the backup equipment with similar performance to the failed equipment; and the cost minimization principle is used to select the optimal solution by comprehensively considering the calculation time, energy consumption, and cost. Step 4.4: Record the fault log and update the device health score.

[0014] A distributed intelligent multi-machine deployment device includes: The edge server cluster is deployed with a distributed intelligent multi-machine collaborative operating system, and adopts a Kubernetes containerized microservice architecture based on OpenHarmony for elastic module expansion; it uses a load balancing algorithm to dynamically allocate computing resources; and it uses hardware acceleration units to improve the computational efficiency of the dynamic task allocation algorithm. A lightweight communication gateway that connects edge server clusters and machine devices, supporting protocol conversion and data caching; The local database stores device configuration information, historical task data, and algorithm model parameters.

[0015] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement a distributed intelligent multi-machine collaborative operation method.

[0016] The beneficial effects of this invention are: 1. This invention achieves efficient collaboration and autonomous operation of multiple machines in complex industrial scenarios through the deep integration of layered architecture design, dynamic task allocation algorithm, heterogeneous device adaptation technology and intelligent fault tolerance mechanism.

[0017] 2. The distributed intelligent multi-machine collaborative operating system proposed in this invention adopts a layered architecture. The edge layer handles high real-time tasks (such as obstacle avoidance and emergency shutdown), while the cloud focuses on long-term strategy optimization (such as capacity planning and energy efficiency analysis). The two work together to balance efficiency and global optimization.

[0018] 3. This invention employs a hybrid strategy of "offline planning + online adjustment" during the equipment task allocation phase, effectively reducing the conflict rate of equipment task execution and continuously monitoring task execution status and environmental changes. When equipment failure, order changes, or environmental disturbances (such as personnel entering the work area) are detected, the system triggers strategy replanning, significantly reducing the need for manual intervention.

[0019] 4. To enable plug-and-play functionality for devices from multiple vendors and using multiple protocols, this invention designs a heterogeneous device adaptation module. At the same time, the fault tolerance mechanism of this invention covers the entire process of fault prediction, rapid diagnosis, and self-healing recovery, ensuring that the system can still operate stably when devices malfunction, and significantly reducing overall task latency.

[0020] 5. This invention visualizes the equipment status, task queue, and environmental heat map through a human-computer interaction module, facilitating equipment monitoring by operators. It also employs a pre-trained language model to parse the contextual semantics of instructions, achieving an instruction recognition accuracy of no less than 95% even in noisy or mispronounced scenarios. The emergency instruction response time is less than 0.5 seconds. Furthermore, it incorporates a hierarchical access control unit to monitor operational behavior in real time. Upon detecting unauthorized operations, it immediately terminates the instruction and triggers an alarm, ensuring operational traceability and security. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the structure of the distributed intelligent multi-machine collaborative operating system provided by the present invention; Figure 2 A flowchart illustrating the distributed intelligent multi-machine collaborative operation method provided by the present invention; Figure 3 A flowchart illustrating the dynamic task allocation algorithm provided by this invention; Figure 4 A schematic diagram of the structure of the computer device provided by the present invention. Detailed Implementation Specific implementation method one: This system adopts a layered architecture design with "edge-cloud" collaboration, consisting of three main parts: physical device layer, edge computing layer, and cloud management platform. It combines real-time communication network and intelligent algorithms to form an efficient, reliable, and scalable collaborative operation system.

[0023] The physical equipment layer includes industrial robots, AGVs (Automated Guided Vehicles), sensors (such as LiDAR and vision cameras), and various actuators. These devices connect to the system through standardized interfaces, becoming physical execution units for collaborative operation. For example, AGVs are responsible for material transportation, robotic arms perform precision assembly, and sensors collect environmental data in real time to support dynamic decision-making.

[0024] The edge computing layer is the core processing unit of the system, deployed in an edge server cluster close to the devices. This layer includes a dynamic task allocation engine, device driver adapters, and a real-time communication gateway. The dynamic task allocation engine, based on a hybrid optimization algorithm and machine learning model, generates and optimizes collaborative strategies in real time; the device driver adapter unifies the hardware interfaces of heterogeneous devices into standardized commands through protocol conversion and capability abstraction; the real-time communication gateway supports the convergence of multiple protocols, including Time-Sensitive Networking (TSN), 5G, and Industrial Ethernet, ensuring that critical data is transmitted with a latency of less than 10 milliseconds. The distributed design of the edge layer avoids the single point of failure risk of a centralized architecture, while significantly reducing cloud load through localized computing.

[0025] The cloud management platform is responsible for global resource coordination and long-term optimization. Its core modules include a digital twin engine and a big data analytics module. The digital twin engine constructs a virtual simulation environment, maps the operating status of physical devices in real time, and rehearses task strategies to verify feasibility. The big data analytics module mines historical task data, optimizes algorithm parameters, and generates device health prediction models. The cloud and edge layers achieve bidirectional data synchronization through an encrypted channel, ensuring optimal resource allocation from a global perspective while supporting rapid autonomous decision-making at the edge layer.

[0026] The technical advantage of this layered architecture lies in the fact that the edge layer handles high real-time tasks (such as obstacle avoidance and emergency shutdown), while the cloud focuses on long-term strategy optimization (such as capacity planning and energy efficiency analysis). The two work together to balance efficiency and global optimization. For example, when a sudden order is inserted, the edge layer can quickly adjust the AGV path, while the cloud synchronously updates the production plan and reassigns it.

[0027] like Figure 1 As shown, the structure of the distributed intelligent multi-machine collaborative operating system described in this embodiment includes: The module includes a communication module, a task scheduling module, a heterogeneous device adaptation module, a fault tolerance and self-healing module, and a human-computer interaction module.

[0028] The communication module establishes a distributed communication network between machines based on OpenHarmony distributed soft bus technology, supporting real-time data transmission and synchronization. It includes: a real-time communication unit based on OpenHarmony time-sensitive networking to ensure that the transmission latency of critical task data is less than 10ms; distributed blockchain nodes that use OpenHarmony-based distributed data management capabilities to store device identity information and operation logs, ensuring that the data is tamper-proof; and an adaptive channel switching unit to dynamically select one of the communication frequency bands among 5G, Wi-Fi 6, and LoRa according to the intensity of environmental interference.

[0029] The task scheduling module generates multi-machine collaboration strategies based on a dynamic task allocation algorithm and optimizes resource allocation in real time. The heterogeneous device adaptation module connects robots, sensors, and actuators with different hardware specifications to the system through the OpenHarmony unified interface protocol. To achieve plug-and-play functionality for devices from multiple vendors and using multiple protocols, this implementation design includes a heterogeneous device adaptation module covering four main functions: protocol conversion, capability virtualization, device encapsulation, and resource pooling management. Protocol Converter: This layer abstracts away hardware differences through a simplified interface. The system incorporates parsers for mainstream industrial protocols such as ROS 2, OPC UA, and Modbus, converting heterogeneous commands into a unified JSON format. For example, when robotic arm A (supporting ROS 2 protocol) and robotic arm B (supporting Modbus protocol) need to perform the same grasping action, the adapter parses their original commands into standardized JSON commands, such as {"action": "grasp", "position": [x,y,z], "force": 20N}, ensuring that upper-layer algorithms do not need to consider underlying hardware differences. Furthermore, the protocol conversion layer supports dynamic plugin loading. When new devices are connected, the cloud automatically sends the corresponding driver to the edge layer, achieving "zero-configuration" access.

[0030] Resource virtualization unit: This unit abstracts the capabilities of physical devices into composable services. For example, parameters such as the load capacity of an AGV, the precision level of a robotic arm, and the resolution of a camera are mapped to service description files and stored in an edge resource pool. During task allocation, the algorithm matches device capabilities with task requirements based on the service descriptions, achieving precise scheduling. In flexible manufacturing scenarios, this module allows the motion control interfaces of different robot models to be uniformly encapsulated as a "precision assembly service." Upper-layer applications only need to call the service interface and can adapt to new and old equipment without modifying the code.

[0031] The device abstraction layer uses the OpenHarmony driver framework to encapsulate the driver interfaces of different robots into standard APIs. Resource pooling management: A dynamic registration and discovery mechanism enables global visibility of device resources. Upon connection, each device's capability description, real-time status, and communication address are registered to the blockchain ledger, ensuring data immutability. Resource pools are categorized and indexed by device type, geographical location, and idle status, supporting rapid retrieval and combined resource retrieval. For example, in cross-workshop collaborative tasks, the system can jointly utilize idle robotic arms from workshop A and AGVs from workshop B to form virtual equipment groups, increasing resource utilization by over 35%.

[0032] The fault tolerance and self-healing module monitors the system's operating status in real time. Upon detecting a device failure, it triggers task migration and redundant resource scheduling. This implementation's fault tolerance mechanism covers the entire process of fault prediction, rapid diagnosis, and self-healing recovery, ensuring stable system operation even when equipment malfunctions. Specifically, it includes: Redundant equipment pre-deployment strategy unit, used to configure backup machines on the critical path; Fault Detection Unit: This unit enables early warning based on multi-source data fusion and deep learning models. The system collects real-time sensor data on equipment temperature, vibration, and current, combining this data with task load and historical fault records. It then uses a Long Short-Term Memory (LSTM) network to predict equipment health. For example, if the temperature of a robotic arm joint exceeds a threshold for five consecutive minutes, the model predicts a greater than 90% probability of failure within the next 10 minutes, triggering preventative maintenance in advance.

[0033] Fault Diagnosis Unit: Employs a strategy combining rule engine and graph reasoning. The rule engine handles explicit faults (such as communication interruptions or battery depletion), while graph reasoning is used to analyze complex fault chains (such as task sequence errors caused by sensor false alarms). Diagnostic results are categorized by their impact scope: local fault diagnosis results (affecting only a single task), regional fault diagnosis results (affecting a group of devices), and global fault diagnosis results (system-level failures), and a repair priority list is generated.

[0034] Task migration unit: Dynamically selected based on fault type. For hardware failures (such as motor damage), the system activates redundant equipment to take over the task and ensures seamless task status transition through data synchronization; for temporary anomalies (such as network jitter), a task retry and data compensation mechanism is employed; for software errors (such as deadlock), the service is restarted via a watchdog timer. Taking an AGV cluster as an example, when a vehicle gets stuck due to a path planning error, the system migrates its task to a nearby AGV within 3 seconds and updates the global path map, with the overall task latency increasing by only 2%.

[0035] The human-computer interaction module provides a visual interface and operation command interface based on OpenHarmony, supporting real-time intervention and authorization control of collaborative strategies by human operators.

[0036] This implementation provides a multi-layered human-computer interaction interface, achieving seamless human-computer collaboration and refined control through augmented reality (AR), natural language processing (NLP), and multi-level access control technologies. Specifically, it includes: An AR (Augmented Reality) interface based on OpenHarmony: Built on mixed reality devices, it overlays virtual information onto the physical environment. Operators can view equipment status (including AGV battery level and robotic arm task progress), task queues, and environmental heatmaps (such as high-load areas and danger zone markers) in real time via gestures or voice commands. The interface supports remote expert collaboration; when equipment malfunctions are detected, operators can initiate real-time video calls, and remote experts can directly guide troubleshooting steps through AR annotations, shortening troubleshooting time.

[0037] The natural language interaction module integrates speech recognition and semantic parsing technologies, supporting mixed Chinese and English command input. Operators can issue task commands (e.g., "AGV 03 transports materials to workstation B"), query system status, or trigger emergency operations via natural language. The system uses a pre-trained language model to parse contextual semantics, achieving a command recognition accuracy of no less than 95% even in noisy or mispronounced scenarios, with an emergency command response time of less than 0.5 seconds.

[0038] Access control system with hierarchical access control: This system utilizes blockchain technology to achieve fine-grained access control. Operator access levels (administrator, engineer, inspector) and their operational scope (equipment control, log viewing, parameter modification) are defined via smart contracts and encrypted and stored in a distributed ledger. The system monitors operational behavior in real time, immediately terminating commands and triggering alarms when unauthorized operations are detected, ensuring traceability and security.

[0039] Emergency intervention interface: The emergency intervention interface supports one-click pause of tasks or switching to manual control mode.

[0040] In addition, this embodiment also provides a distributed intelligent multi-machine deployment device, including: The edge server cluster is deployed with a distributed intelligent multi-machine collaborative operating system, and adopts a Kubernetes containerized microservice architecture based on OpenHarmony for elastic module expansion; it uses a load balancing algorithm to dynamically allocate computing resources; and it uses hardware acceleration units to improve the computational efficiency of the dynamic task allocation algorithm. A lightweight communication gateway that connects edge server clusters and machine devices, supporting protocol conversion and data caching; The local database stores device configuration information, historical task data, and algorithm model parameters.

[0041] Specific implementation method two: such as Figure 2 As shown, the steps of the distributed intelligent multi-machine cooperative operation method described in this embodiment include: S1: Establish a device communication network; This step is performed based on the communication module. The communication module initializes the distributed intelligent multi-machine collaborative operating system, establishes a communication network between all machines and devices, and connects the robot, sensor and actuator to the distributed intelligent multi-machine collaborative operating system and verifies identity information through the heterogeneous device adaptation module. S2: Device status visualization and task command issuance; This step is performed based on the human-computer interaction module. The operator can view the equipment status, equipment location and environmental 3D map through the human-computer interaction module, issue task instructions, and use a pre-trained language model to parse the context semantics, accurately identify the instructions and send them to the task scheduling module.

[0042] S3: Dynamic task allocation and task scheduling; Dynamic task allocation is the core technology of this invention for achieving efficient collaboration. It solves the problem that traditional static strategies cannot adapt to environmental changes by using a hybrid optimization and machine learning approach. For example... Figure 3 As shown, the task allocation process covers four stages: task modeling, optimization goal definition, algorithm execution, and dynamic adjustment. Task modeling phase: First, task attributes are formally described. Each task is defined as a tuple containing priority, deadline, resource requirements (such as required equipment type, tools, and fixtures), and dependencies. For example, in a mobile phone assembly scenario, the "screen installation" task requires specifying the collaborative robot model, visual inspection camera ID, and completion deadline. Simultaneously, the system collects equipment status data in real time to construct a machine capability vector, including parameters such as current location, remaining battery power, load capacity, and sensor coverage. This dual modeling of tasks and equipment provides the input foundation for the optimization algorithm.

[0043] Optimization of objectives: This phase requires balancing multiple constraints. Core objectives include minimizing task completion time, minimizing total energy consumption, and maximizing equipment utilization, while simultaneously satisfying physical limitations (such as the robotic arm's range of motion), task dependencies (such as task B requiring task A to start after completion), and safety rules (such as AGV obstacle avoidance distance). To quantify these objectives, this invention establishes a multi-objective optimization function, where time cost and energy cost are balanced through a weighted summation. For example, the total cost function can be expressed as: (1); In formula (1), For task completion time, For equipment energy consumption, For resource idle rate, , and This is a weighting coefficient that can be dynamically adjusted based on the scenario.

[0044] Algorithm execution phase: A hybrid strategy of "offline planning + online adjustment" is adopted. Offline planning is based on mixed integer linear programming (MILP) to accurately solve for the optimal solution in small-scale scenarios (e.g., 10 devices); online adjustment relies on a deep reinforcement learning (DRL) model to make adaptive decisions through real-time environmental interaction. The reward function of DRL is designed as a linear combination of task completion rate, path conflict penalty, and energy efficiency score, guiding the agent to balance efficiency and safety in complex scenarios. For example, in a warehouse logistics scenario, the DRL model can respond to changes in shelf location within 0.2 seconds, replan the paths of 50 AGVs, and reduce the conflict rate to below 5%. For ultra-large-scale systems (100+ devices), this invention introduces a distributed auction algorithm, where each device bids for the task based on local information, and a global suboptimal solution is reached through a consensus mechanism, significantly reducing computational complexity.

[0045] Dynamic Adjustment Phase: Continuously monitors task execution status and environmental changes. When equipment failure, order changes, or environmental disturbances (such as personnel entering the work area) are detected, the system triggers strategy replanning. For example, if an AGV pauses due to insufficient power, the algorithm immediately breaks down the unfinished task into sub-tasks, assigns them to other nearby idle AGVs, and ensures data consistency through a two-phase commit protocol. This dynamic adjustment mechanism enables the system to maintain continuous operation in over 95% of unexpected situations, reducing the need for manual intervention by 80%.

[0046] S4: Equipment fault monitoring and self-healing recovery strategy; S401: Equipment malfunction is detected by the fault detection unit; S402: When a device malfunction is detected, assess the scope of the malfunction's impact and classify the affected task levels. S403: Resume task execution by calling the backup equipment configured in the redundant equipment pre-deployment strategy unit or by adopting one of the resource reallocation strategies. The resource reallocation strategy includes the proximity principle, the capability matching principle, and the cost minimization principle. The proximity principle is used to prioritize the physical proximity of the equipment to take over the task; the capability matching principle is used to select the backup equipment with similar performance to the failed equipment; and the cost minimization principle is used to select the optimal solution by comprehensively considering the calculation time, energy consumption, and cost. S404: Record fault logs and update device health score.

[0047] In addition, this embodiment also provides a computer device, such as Figure 4 As shown, it includes a memory and a processor. The memory stores a computer program. The feature is that when the processor executes the computer program, it implements the steps of a distributed intelligent multi-machine collaborative operation method.

[0048] Example 1 Smart factory flexible production line: An electronics manufacturing workshop has deployed 10 collaborative robots, 8 AGVs, and 30 vision sensors for mixed-line production of multiple mobile phone models. Traditional systems are inefficient due to equipment heterogeneity and order fluctuations. The implementation steps after adopting this invention are as follows: System initialization: Device access: The robot accesses the edge server via ROS 2 protocol, the AGV via Modbus TCP protocol, and the vision sensor via MQTT protocol.

[0049] Capability registration: Each device uploads its performance parameters (such as maximum load and motion accuracy) and real-time status to the resource pool.

[0050] Digital twin construction: A virtual image of the production line is generated in the cloud based on the Unity 3D engine, synchronizing the coordinates of physical equipment with task progress.

[0051] Task assignment and execution: Order analysis: The cloud breaks down customer orders into sub-tasks such as welding, assembly, and testing, and distributes them to equipment groups through a reinforcement learning model.

[0052] Dynamic adjustment: When an urgent order is inserted, the edge layer replans the AGV path within 15 milliseconds, and the original task delay only increases by 5%.

[0053] Quality closed loop: After visual inspection detects that the screen installation offset exceeds the limit, the system automatically triggers the rework process, reducing the defect rate by 40%.

[0054] Fault tolerance and recovery: When a robot malfunctions, the system migrates unfinished tasks to backup equipment within 2 seconds and ensures process continuity through data compensation.

[0055] In the event of a network outage, the system switches to local decision-making mode and relies on caching strategies to maintain production line operation for more than 30 minutes.

[0056] Example 2 Automated container terminal at the port: A port has deployed 80 unmanned container trucks, 12 gantry cranes, and 50 rail-mounted gantry cranes. After applying this invention, the following improvements are achieved: Collaborative path planning: Task allocation based on auction algorithms reduced the truck bidding delay from 20 seconds to 1.2 seconds and the path conflict rate from 12% to 0.8%.

[0057] Energy efficiency optimization: After analyzing historical data in the cloud, a new "standby power saving mode" has been added for bridge cranes. When idle for more than 1 minute, the power will be reduced and the energy consumption will be reduced by 18%.

[0058] Cross-device collaboration: The error of loading and unloading by rail-mounted gantry cranes and container trucks is less than ±2 cm, and the loading and unloading efficiency is improved by 25%.

[0059] Furthermore, to further verify the technical effectiveness of the distributed intelligent multi-machine collaborative operation method of the present invention, the system performance was verified through simulation and real-world scenario testing: Experiment 1: Dynamic Task Allocation Efficiency Test In a warehousing scenario with 50 AGVs, the traditional MILP algorithm, deep reinforcement learning (DRL), and the hybrid strategy of this invention were compared: Average task allocation latency: 150ms for this invention (3200ms for traditional MILP, 180ms for DRL). Task completion rate: 98% for this invention (88% for traditional MILP, 94% for DRL); Longest path conflict duration: 1.8 seconds for this invention (12.5 seconds for conventional MILP, 4.3 seconds for DRL).

[0060] Experiment 2: Fault Tolerance and Recovery Capability Test Simulate motor overheating, communication interruption, and software deadlock failures in an automotive welding workshop: Motor overheat recovery time: 4.2 seconds for this invention (300 seconds for the conventional system), reducing task interruption rate from 100% to 0%; Communication interruption recovery time: 8.5 seconds for this invention (traditional systems require device restart), reducing task interruption rate from 100% to 12%; Software deadlock recovery time: 2.1 seconds (traditional systems require manual reset), reducing task interruption rate from 100% to 0%.

[0061] Experiment 3: Heterogeneous Device Compatibility Verification Integration with 6 types of robots from 3 manufacturers, 3 types of AGVs, and 5 types of sensors: ABB robot setup time: 3 minutes for this invention (2 hours for traditional systems); Danfoss AGV protocol conversion time: 5 minutes for this invention (traditional systems require customized middleware); Keyence vision sensor driver adaptation time: 2 minutes for this invention (not supported by traditional systems).

[0062] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent substitutions, and improvements made to the above embodiments without departing from the scope of the present invention, based on the technical essence of the present invention and within the spirit and principles of the present invention, shall still fall within the protection scope of the present invention.

Claims

1. A distributed intelligent multi-machine collaborative operating system, characterized in that, include: The communication module uses the OpenHarmony distributed soft bus method to establish a distributed communication network between machine devices for real-time data transmission and synchronization. The task scheduling module adopts a multi-machine device collaboration strategy based on a dynamic task allocation algorithm and optimizes resource allocation in real time. Heterogeneous device adaptation module, which connects robots, sensors and actuators with different hardware specifications to the distributed intelligent multi-machine collaborative operating system through the OpenHarmony unified interface protocol; The fault tolerance and self-healing module is used to monitor the system's operating status in real time and trigger task migration and redundant resource scheduling when a machine or equipment failure is detected. The human-computer interaction module provides a visual interface and operation command interface based on OpenHarmony, allowing operators to intervene and authorize collaborative strategies in real time.

2. The distributed intelligent multi-machine collaborative operating system according to claim 1, characterized in that, The communication module includes: A real-time communication unit based on OpenHarmony time-sensitive networking is used to ensure that the transmission latency of critical task data is less than 10ms. Distributed blockchain nodes use OpenHarmony-based distributed data management capabilities to store device identity information and operation logs, ensuring that the data is immutable; An adaptive channel switching unit is used to dynamically select one of the communication frequency bands among 5G, Wi-Fi 6, and LoRa based on the intensity of environmental interference.

3. The distributed intelligent multi-machine collaborative operating system according to claim 1, characterized in that, The heterogeneous device adapter module includes: The device abstraction layer uses the OpenHarmony driver framework to encapsulate the driver interfaces of different robots into standard APIs. The resource virtualization unit is used to map the computing power, storage space and sensor data of physical devices into a virtual resource pool; A protocol converter that is compatible with ROS2, OPC UA and Modbus communication protocols; The resource management unit enables global visibility of device resources through a dynamic registration and discovery mechanism, and allows for rapid retrieval and combined invocation of devices.

4. The distributed intelligent multi-machine collaborative operating system according to claim 1, characterized in that, The fault tolerance and self-healing module includes: Redundant equipment pre-deployment strategy unit, used to configure backup machines on the critical path; The fault detection unit determines whether the device is offline or abnormal by using the OpenHarmony distributed heartbeat mechanism and sensor data analysis. The task migration unit is used to reassign unfinished tasks from faulty devices to healthy devices and restore data consistency. The fault diagnosis unit uses a strategy that combines rule engine and graph reasoning to obtain local fault diagnosis results, regional fault diagnosis results and global fault diagnosis results, and generates a repair priority list.

5. The distributed intelligent multi-machine collaborative operating system according to claim 1, characterized in that, The human-computer interaction module includes: An AR augmented reality interface based on OpenHarmony is used to overlay and display the location of machines and equipment, task status, and a 3D map of the environment. The access control unit is used to differentiate the operating permissions of administrators, operators, and visitors; Emergency intervention interface, which supports one-click pause of task or switching to manual control mode; The natural language interaction module integrates speech recognition and semantic parsing to accurately identify user commands.

6. A distributed intelligent multi-machine collaborative operation method, applied to the distributed intelligent multi-machine collaborative operating system described in any one of claims 1-5, characterized in that, include: Step 1: Initialize the distributed intelligent multi-machine collaborative operating system based on the communication module, establish a communication network between all machines and devices, and connect the robot, sensor and actuator to the distributed intelligent multi-machine collaborative operating system and verify their identity information through the heterogeneous device adaptation module; Step 2: The operator views the device status, device location, and 3D map of the environment through the human-computer interaction module, issues task instructions, and uses a pre-trained language model to parse the contextual semantics, accurately identify the instructions, and send them to the task scheduling module. Step 3: The task scheduling module divides the collaborative subgroups according to the task requirements. Each subgroup includes at least one master control node and multiple execution nodes. The master control node generates control instructions based on the dynamic task allocation algorithm and distributes them to the execution nodes through the communication module. The execution nodes feed back status data to the master control node, forming a closed-loop control link. Step 4: The fault tolerance and self-healing module monitors the system's operating status in real time. When a machine or equipment malfunction is detected, the unfinished tasks of the malfunctioning equipment are reassigned to healthy machines or equipment, and data consistency is restored.

7. The distributed intelligent multi-machine collaborative operation method according to claim 6, characterized in that, Step 3 specifically includes: Step 3.1: The task scheduling module receives the task instructions from the human-computer interaction module and parses the task objectives and constraints; Step 3.2: Collect real-time status data of each device using OpenHarmony's distributed capabilities, including location, power level, load capacity, and sensor information; Step 3.3: Divide the collaborative subgroups according to the physical location topology of the devices, task priority and deadline, and the matching degree between device capabilities and task requirements. Each subgroup shall include at least one master node and multiple execution nodes. Step 3.4: The master node calculates the optimal coordination strategy and generates a set of device control instructions by invoking the dynamic task allocation algorithm; the dynamic task allocation algorithm specifically includes: The algorithm is based on one or more of the following: Q-Learning algorithm based on reinforcement learning, Nash equilibrium allocation model based on game theory, and multi-objective optimization model based on genetic algorithm. Specifically, Q-Learning algorithm based on reinforcement learning optimizes long-term task efficiency through reward function; Nash equilibrium allocation model based on game theory is used to balance resource competition among multiple machines; and multi-objective optimization model based on genetic algorithm is used to minimize task completion time and energy consumption. Step 3.5: Using the target device corresponding to each instruction as the execution node, send the device control instructions to the target device through the communication module and monitor the task progress; Step 3.6: Dynamically adjust the allocation of equipment control commands based on environmental changes or equipment status updates.

8. The distributed intelligent multi-machine collaborative operation method according to claim 6, characterized in that, Step 4 specifically includes: Step 4.1: Detect equipment malfunctions using the fault detection unit; Step 4.2: When a device malfunction is detected, assess the scope of the malfunction's impact and classify the affected tasks into levels. Step 4.3: Resume task execution by calling the backup equipment configured in the redundant equipment pre-deployment strategy unit or by adopting one of the resource reallocation strategies. The resource reallocation strategy includes the proximity principle, the capability matching principle, and the cost minimization principle. The proximity principle is used to prioritize the equipment with the closest physical location to take over the task; the capability matching principle is used to select the backup equipment with similar performance to the faulty equipment; and the cost minimization principle is used to select the optimal solution by comprehensively considering the calculation time, energy consumption, and cost. Step 4.4: Record the fault log and update the device health score.

9. A distributed intelligent multi-machine deployment device, applied to the distributed intelligent multi-machine collaborative operating system according to any one of claims 1-5, characterized in that, include: The edge server cluster is deployed with the distributed intelligent multi-machine collaborative operating system as described in any one of claims 1-5, and the edge server cluster adopts a Kubernetes containerized microservice architecture based on OpenHarmony for elastic module expansion; it uses a load balancing algorithm to dynamically allocate computing resources; and it uses hardware acceleration units to improve the computing efficiency of the dynamic task allocation algorithm. A lightweight communication gateway that connects edge server clusters and machine devices, supporting protocol conversion and data caching; The local database stores device configuration information, historical task data, and algorithm model parameters.

10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 6 to 8.

Citation Information

Patent Citations

  • Tracing apparatus and positioning system

    CN110123456A

  • Container and associated methods

    US20210056789A1

Cited By

  • Workshop collaborative carrying system

    CN121340314A

  • Heterogeneous framework fusion device and method for mixed reality or augmented reality

    CN121366269A

  • Integrated dispatching system of automatic guided vehicle

    CN121956925A