Robot control method, system, device and medium based on multi-modal large model

By combining multimodal large models and large language models, the problem of low operating efficiency of autonomous power distribution network live-line working robots has been solved, achieving more efficient and safer operation.

CN120962678BActive Publication Date: 2025-12-12WENZHOU ELECTRIC POWER BUREAU +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511484981.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-12-12
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing autonomous live-line working robots for power distribution networks have low operating efficiency, require operators to manually adjust their positions and confirm via video, rely heavily on human intervention, and are difficult to operate skillfully in high-risk scenarios.

Method used

A multimodal large model is used for data feature extraction and fusion, a large language model is combined for task decomposition, and a human-in-the-loop mechanism is introduced for optimization to generate robot execution instructions.

Benefits of technology

It improves the robot's ability to understand and decompose tasks in complex work scenarios, thereby increasing work efficiency and safety and reducing human intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120962678B_ABST
    Figure CN120962678B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of robot control and discloses a robot control method, system, device and medium based on a multimodal large model, which comprises the following steps: collecting multi-source modal data of a scene where a work task is located, processing the multi-source modal data through a machine learning model, obtaining multimodal features for position coding and Transform fusion processing, obtaining multimodal fusion features, inputting the multimodal fusion features and a constructed work task knowledge base into a large language model to decompose a target work task, introducing a human-in-the-loop mechanism to optimize the decomposition result, and obtaining a subtask sequence; according to a subtask type in the subtask sequence, processing the subtask sequence through a visual language action model or a reinforcement learning model, generating a motion instruction to enable a robot to start an execution process of the target work task; and processing a live work task through a multimodal large model LLM, VLA and the like, so that the work efficiency of an autonomous distribution network live work robot is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of robot control, and in particular to a robot control method and system based on a multi-modal large model, a device and a medium. BACKGROUND

[0002] With the continuous development of technology, robots are often applied to outdoor live working scenes with high risk factors such as harsh weather conditions, high altitude, strong voltage field, etc. However, the existing autonomous distribution live working robots still need to be manually adjusted by operators to adjust the position of the robot, and after confirming that the position is appropriate through the video screen of the mobile control terminal, the specific step program of the corresponding working project is started, and the manual intervention is highly dependent, which further leads to low working efficiency of the robot.

[0003] Therefore, how to solve the problem of low working efficiency of the existing autonomous distribution live working robot has become a technical problem to be solved by those skilled in the art. SUMMARY

[0004] The present application provides a robot control method and system based on a multi-modal large model, which solves the problem of how to improve the working efficiency of the existing autonomous distribution live working robot.

[0005] To solve the above technical problems, the present application provides a robot control method based on a multi-modal large model, comprising:

[0006] Real-time acquisition of multi-source modal data of the scene where the target working task is located, and based on the type of the multi-source modal data, a corresponding machine learning model is used to extract features from the multi-source modal data to obtain multi-modal features;

[0007] The multi-modal features are position encoded and processed by a Transformer fusion to obtain multi-modal fusion features;

[0008] The constructed working task knowledge base and the multi-modal fusion features are input into a large language model to obtain a decomposition result of the target working task, and a human-in-the-loop mechanism is introduced to optimize the decomposition result to obtain a subtask sequence;

[0009] Based on the subtask type in the subtask sequence, the subtask sequence and the multi-source modal data are input into a visual language action model or a reinforcement learning model for processing to generate motion instructions corresponding to the subtask sequence, so that the robot starts the execution process of the target working task.

[0010] Compared with the prior art, the beneficial effects of the present application are as follows:

[0011] The feature extraction can be performed on different types of data by corresponding machine learning models, which can more comprehensively and accurately capture scene information, compared to single modal data processing, and provides a richer feature basis for subsequent task decomposition and instruction generation, and improves the understanding ability of complex work scene; the multi-modal fusion and task decomposition are combined, and the human-in-the-loop mechanism is introduced to optimize the decomposition result, which fully utilizes the language understanding and reasoning ability of the large language model, so that the task decomposition is more in line with the actual demand and logic, and the accuracy and rationality of the task decomposition are improved; the dynamic model selection mechanism can play the advantages of different models according to the characteristics of different sub-tasks, improve the efficiency and accuracy of instruction generation, better control the robot to execute complex work tasks, and further improve the work efficiency of the existing autonomous distribution network live working robot. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the technical solutions of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0013] Figure 1 is a flow chart of a robot control method based on a multi-modal large model provided by an embodiment of the present application;

[0014] Figure 2 is a structural diagram of a robot control system based on a multi-modal large model provided by an embodiment of the present application;

[0015] Figure 3 is a structural diagram of an electronic device provided by an embodiment of the present application;

[0016] Reference signs:

[0017] Among them, 10, feature extraction module; 20, feature fusion module; 30, task decomposition module; 40, instruction execution module; 5000, electronic device; 5001, processor; 5002, bus; 5003, memory; 5004, transceiver. DETAILED DESCRIPTION

[0018] The technical solutions in the embodiments of the present application will be described clearly and completely below in combination with the drawings and embodiments. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. The purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0019] In the description of the present application, the terms "first", "second", "third" and the like are only for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the technical features indicated. Therefore, the features defined with "first", "second", "third" and the like can explicitly or implicitly include one or more of the features. In the description of the present application, unless otherwise stated, the meaning of "multiple" is two or more. The term "and / or" used herein includes any and all combinations of one or more related listed items. The specific meaning of the above terms in the present application can be understood according to the specific circumstances by those skilled in the art.

[0020] In the description of the present application, it should be noted that, unless otherwise defined, all technical and scientific terms used in the present application have the same meaning as generally understood by those skilled in the art. The terms used in the description of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The specific meaning of the above terms in the present application can be understood according to the specific circumstances by those skilled in the art.

[0021] The distribution network live working robot refers to a robot system for carrying out distribution network live working by human-machine cooperation or self-service, which is composed of a robot body, a working tool and an insulating bearing platform. Among them, under the assistance of ground operators, the robot can realize identification and positioning, path planning, end tool replacement, and automatically execute live working tasks, which is called autonomous distribution network live working robot, hereinafter referred to as "robot".

[0022] At present, the robot needs to complete the live working project in the typical working scene. The typical scene is single return horizontal arrangement line, single return triangular arrangement line and double return vertical arrangement line. The working projects that can be carried out are breaking and connecting the current-carrying wire, installing grounding ring, installing fault indicator, pruning branches, supporting insulator, live assisting installation or removal of insulating shield, live replacing lightning arrester, live replacing fuse and live replacing straight pole insulator. However, the operation of the current robot is complex, and manual intervention is dependent on high, so in most cases the robot needs to be manually adjusted by the operator, and the position and working point are confirmed by the mobile control terminal. The robot cannot completely complete the task autonomously, and with the increase of the joint freedom degree of the robot and the diversification of the task, the traditional planning method based on geometry or simplified dynamics model adopted by the robot is difficult to find a feasible solution in high-dimensional and multi-constrained scene, and cannot realize human-like dexterous operation, so the dexterity of the robot is insufficient, which further leads to low working efficiency.

[0023] Based on this, in an embodiment, as Figure 1As shown, the first aspect of the present application provides a robot control method based on a multi-modal large model, comprising:

[0024] S1, real-time acquisition of multi-source modal data of the scene where the target task is located, and based on the type of the multi-source modal data, a corresponding machine learning model is used to extract features from the multi-source modal data to obtain multi-modal features; wherein the multi-source modal data includes force sensation and point cloud data, image data and audio data;

[0025] At present, the actuator of the robot is a pair of mechanical arms carrying dexterous hands, and the robot also carries various types of sensors, including visual sensors, force sensors, tactile sensors, microphones, laser radars and environmental sensors (temperature and humidity sensors, wind speed and direction sensors, etc.), to real-time collect multi-source modal data of the scene where the target task is located, including force sensation and point cloud data, image data, audio data, distance data, real-time sensor data, environmental data, and robot joint torque, speed, temperature and robot end position and attitude, etc.

[0026] For the collected raw data, data preprocessing techniques are used, such as filtering and denoising of images, to ensure the high quality and usability of the data; through timestamp synchronization technology, the consistency of multi-modal data such as laser radar, millimeter wave radar, visual sensor, force sensor, etc. on the time axis is ensured, and the data mismatch problem caused by time difference is reduced; and Kalman filter algorithm is used to perform real-time data fusion and correction on the raw data of laser radar, millimeter wave radar and visual sensor, to improve the accuracy and robustness of the data fusion between different sensors.

[0027] In an embodiment, the use of a corresponding machine learning model to extract features from the multi-source modal data to obtain multi-modal features comprises:

[0028] The force sensation and point cloud data are input into a long short-term memory network model for processing to capture force interaction patterns and obtain time series dynamic features;

[0029] The image data is input into an improved ResNet-50 model for processing to obtain local features of the target object;

[0030] The audio data is input into a Whisper-Base model and a BERT-Base-Chinese model for processing to obtain speech features for parsing execution instructions;

[0031] The time series dynamic features, the local features and the speech features are combined to obtain the multi-modal features.

[0032] Specifically, for the collected visual, force sense and environment multi-source modal data, based on the type, a variety of machine learning models are used to realize feature extraction of multi-modal data, and the specific process includes:

[0033] The force sensor and laser radar collected force sense and point cloud time series data are input into a long short-term memory network (LSTM) for processing to extract dynamic changes and correlations in the time dimension, capture force interaction patterns, and the LSTM gate mechanism retains long-term dependent information and forgets irrelevant information, generating time series features with generalization ability, helping the robot understand the force interaction pattern and the change of the three-dimensional structure of the environment, and outputting time series dynamic features. The processing process of the LSTM on the time series data can refer to the application of the LSTM in the prior art, and will not be described in detail here.

[0034] The RGB image data collected by the vision camera is input into an improved ResNet-50 model for processing, and the local features including the edges, textures, colors and geometric features of the target object (such as the wire and the clamp) are output; wherein the improved ResNet-0 model optimizes the convolutional layer and the residual block based on the original ResNet-50, and enhances the model's ability to extract target edge features. It takes ResNet-50 as a feature extraction backbone network, which takes a 640x480 pixel RGB image as input, normalizes it (pixel scaling to [0, 1]) and data augmentation (random cropping, flipping, brightness adjustment) preprocessing, and then extracts low-level features from the image by a 7x7 convolution kernel (stride 2) in the model, and then extracts high-level semantic features by four residual modules (3, 4, 6, 3 residual blocks) step by step, uses residual connection to alleviate gradient disappearance, and finally compresses the high-dimensional features into a 512-dimensional vector by global average pooling, retains the spatial position and context information of the target, and ensures accurate recognition and positioning in strong light, shadow or occlusion scenes.

[0035] For the 16kHz, 16-bit PCM audio data collected by the microphone, the WebRTC VAD (based on short-time energy and zero-crossing rate) is used to separate the speech and background noise, the obtained speech segments are input into the Whisper-Base model (74M parameters, Chinese optimization), combined with Mel spectrum features and beam search (beam width 5) to convert to high-precision text, then the obtained text is segmented and encoded by BERT-Base-Chinese (12-layer Transformer, 768-dimensional hidden layer) to generate a 768-dimensional semantic vector, capturing deep semantic instructions (such as grasping the clamp and installing it to the wire), mapping to the robot task space, ensuring robust human-computer interaction, and outputting speech features for parsing and executing instructions.

[0036] Finally, the multi-modal features can be obtained by combining the time series dynamic features, local features and speech features extracted by the machine learning model. It should be noted that the above feature extraction process can also be performed using other models, such as using the AlexNet model to extract image features. Of course, speech features and time series dynamic features can also be extracted using other models. The extraction process refers to the implementation process of the corresponding model in the prior art, and will not be described here.

[0037] The present application adopts machine learning models with strong adaptability and strong pertinence for different types of multi-source modal data. This model selection method according to data characteristics fully utilizes the expertise of each model and improves the quality and efficiency of feature extraction. By simultaneously extracting the time series dynamic features of force and point cloud, the geometric features of image, and the speech features of audio and video, comprehensive and rich information is provided for subsequent task processing. The classical model is improved (such as improved ResNet-50), and the excellent models in different fields (such as Whisper-Base in the speech field and BERT-Base-Chinese in the text field) are combined and applied to multi-modal feature extraction, which embodies innovative technical ideas and can better adapt to specific task requirements. Model combination fully utilizes the advantages of different models, realizes functional complementation, and provides a new effective way to solve complex multi-modal data processing problems.

[0038] S2, position encoding and Transformer fusion processing are performed on the multi-modal features to obtain multi-modal fusion features;

[0039] In an embodiment, step S2 includes:

[0040] The time steps are embedded in the multi-modal features by position encoding to obtain multi-modal encoding features;

[0041] The multi-modal encoding features are input into a Transformer encoder to perform feature fusion through self-attention mechanism and multi-head attention mechanism to obtain the multi-modal fusion features;

[0042] Wherein, the position encoding is performed by the following formula:

[0043]

[0044] t ′= t + t

[0045] In the formula, is the position encoding vector for the t-th time step; d is the dimension of the feature vector of each time step in the multi-modal feature; t is the multi-modal encoding feature. t is the multi-modal feature.

[0046] Specifically, the present application adds position information to each time step of the multi-modal feature through position encoding, so that the Transformer can distinguish the sequence order: the multi-modal features such as visual features, force sensation features and language features are combined into a sequence, and the length and dimension of the sequence are determined, an integer index t is assigned to each position of the multi-modal feature, for each position t, a d-dimensional position encoding vector is calculated, for even index (0, 2, 4, , d-2), calculate ; for odd index (1, 3, 5, , d-1), calculate ; finally, the position encoding vector t is added to the feature vector t of the corresponding position, that is, the multi-modal encoding feature t is obtained.

[0047] In the process of Transformer feature fusion, the present application introduces position encoding (Positional Encoding) to provide sequence position information and assist the model in understanding time series and sequence order. Since multi-modal data (especially force / point cloud, etc.) has time series, through additive position encoding, time step information is embedded into the feature vector, and the Transformer can identify input at different time points or spatial order, so as to consider the time series context when fusing. For example, for multi-frame point clouds in the process of executing an action by a mechanical arm, the position encoding can help the Transformer to distinguish early and late frames, so that the fused representation retains the trajectory of the gradual change of the environment.

[0048] Position encoding is very important in feature fusion, because it can preserve the temporal information. Point cloud and force data change over time (e.g. point cloud sequence when the robot is grasping). Position encoding adds a unique time identifier to each frame, and the Transformer distinguishes between early and late frames, preserving dynamic trajectories (e.g. wire clip movement path), while position encoding can distinguish the order of modalities. Visual, force, and language features are arranged in a fixed order (e.g. visual 0-20, force 21-40, language 41-63), ensuring that the Transformer identifies the modality source and correctly associates the instruction "install" with the wire position in the vision. This encoding can significantly enhance context modeling, allowing the Transformer to capture temporal context, such as the association of early visual features (wire clip position) with late force features (grasping feedback), improving the semantic integrity of the fused representation. Ultimately, this can improve the robustness of feature fusion. In contrast to many schemes that ignore timing, this scheme can explicitly model time steps, adapting to scenarios where action sequences are strictly ordered (e.g. grasping before installation), ensuring temporal consistency in the fused representation.

[0049] To fully utilize the information of each modality and obtain unified environmental representation, the scheme adopts a Transformer network structure as the core architecture of multi-modal feature fusion. The Transformer encoder can efficiently model the complex interdependence between different modalities, and process the input multi-modal encoding features to fuse the input features through self-attention and multi-head attention mechanisms to obtain multi-modal fusion features.

[0050] The self-attention mechanism determines the importance of each position's features to other position features by calculating the similarity between query vectors (Query), key vectors (Key), and value vectors (Value). Specifically, for the input multi-modal encoding feature sequence, the query matrix Q, key matrix K, and value matrix V are obtained through linear transformation, then the attention score is calculated, and finally the value matrix is weighted and summed according to the attention score to obtain the output of the self-attention mechanism. In this way, the Transformer can globally associate features of different modalities at each layer: for example, it can focus on the correlation between features of a specific region in the image and force signals at a certain time, or the association between a word in the language instruction and the target object observed in the current vision, thereby fusing these cross-modal associations in the internal hidden state.

[0051] Multi-head attention mechanisms capture feature correlations across different subspaces through multiple parallel attention heads, resulting in a more comprehensive fusion outcome. Specifically, the multi-head attention mechanism performs the self-attention mechanism's computation process multiple times in parallel (number of heads), each time using a different linear transformation matrix to obtain different Q, K, and V values, thus focusing on features from multiple different subspaces. The outputs of each head are concatenated and subjected to a linear transformation to obtain the final output of the multi-head attention mechanism.

[0052] After processing by a multi-head attention mechanism, the output is input into a feedforward neural network for nonlinear transformation. Simultaneously, residual connections are used to add the input to the feedforward neural network output to alleviate the vanishing gradient problem and accelerate model training. After multiple layers of such processing, the Transformer encoder outputs a fused multimodal feature, namely a sequence of context feature vectors. Each vector integrates key information from different sources such as vision, force perception, point clouds, and language, and retains dynamic changes in the working environment over time. It should be noted that the specific fusion process can be found in existing applications of the Transformer encoder, which will not be elaborated upon here. The Transformer's output provides a unified and semantically rich representation of the robot's current environment and task.

[0053] In multimodal feature fusion, the introduction of positional encoding adds time-step information to the multimodal features, enabling the model to better understand the order and correlation of features in the time series. Through positional encoding, the model can more accurately capture these temporal dependencies, thereby improving the quality of fused features. The self-attention mechanism in the Transformer encoder allows the model to automatically focus on important parts of different modal features and uncover the intrinsic connections between features. The multi-head attention mechanism focuses on and fuses features from multiple different subspaces, further enriching the feature representation. This allows the fused multimodal features to more comprehensively and accurately reflect the complex information of the original data, exhibiting higher flexibility and effectiveness compared to traditional feature fusion methods.

[0054] S3. Input the constructed task knowledge base and the multimodal fusion features into the large language model to obtain the decomposition result of the target task, and introduce a human-in-the-loop mechanism to optimize the decomposition result to obtain a sub-task sequence.

[0055] In one embodiment, step S3 includes:

[0056] Acquire live-line work specification data and live-line work expert data, and extract static correlation information of work actions and work environment dependency rule information from them to construct the work task knowledge base;

[0057] transformer architecture, to decompose the target job task in combination with the job task knowledge base, to obtain an initial subtask sequence;

[0058] sequencing the initial subtask sequence according to the job task knowledge base, and parameterizing the sequenced initial subtask sequence to generate a first-optimized subtask sequence;

[0059] optimizing the first-optimized subtask sequence based on the physical motion range of the robot and in combination with the job task knowledge base, to obtain a second-optimized subtask sequence;

[0060] receiving external instructions to introduce the human-in-the-loop mechanism to re-optimize the second-optimized subtask sequence, to obtain the final subtask sequence.

[0061] Specifically, in complex scenarios such as live-line work of distribution network, a high-level job task usually needs to be decomposed into multiple manageable subtasks. For example, for a typical job such as "live-line connection of current lead", the robot cannot complete all steps at once and must follow the process of human-device interaction to execute step by step. The present application utilizes the powerful natural language understanding and reasoning ability of large language model (LLM), combines the experience of power operation experts and industry professional knowledge, and automatically decomposes the complex task semantics into several subtasks. The specific steps include:

[0062] Collect standardized process documents (such as "Live-line Work Safety Regulations") and their operation videos in the power industry as live-line work specification data, and collect experience data of live-line work experts in the live-line work of distribution network as live-line work expert data. Taking "replacement of current lead clamp" as an example, the process provided by experts usually includes steps such as tool inspection, target positioning, clamp installation, and result verification. Then use natural language processing technology and rule engine to extract the sequence of each live-line work step, operation requirements of each step, and other static association information of job actions from the collected data, and extract job environment dependent rule information such as notes for work in different weather and terrain conditions, inspection methods of special equipment from expert data; Finally, based on this information, a job task knowledge base is built, that is, the extracted information is stored in a structured way in a database, and a relational database (such as MySQL) is used for management. For example, create a "replacement of current lead clamp" to store detailed information of each current lead clamp replacement step, including step number, step name, operation content, required parameters, etc.; create a "step association table" to record the sequence and dependency relationship between steps; create an "environment rule table" to store notes for work in different weather and terrain conditions, inspection methods of special equipment, etc.

[0063] A knowledge graph is constructed using a graph database, with nodes representing job actions, environmental factors, etc., and edges representing their relationships; these data are structured into a domain knowledge base, which can be in the form of a knowledge graph or a database, recording information such as task descriptions, step sequences, tool requirements, and safety specifications. For example, the knowledge base specifies that "replacing a drainage clamp" requires insulated gloves and a special wrench, and ensures insulation performance in a high-voltage environment; experts also manually annotate the subtask sequences of typical tasks, define the input (such as tools and environmental conditions), output (such as completion status), and constraints (such as safety requirements and robot movement range) for each subtask, providing high-quality training and reference data for the LLM.

[0064] The target task of live-line work for distribution network is clearly expressed in natural language, for example, "complete a certain live-line work operation in a specific distribution line scenario" is converted into "the robot needs to replace the drainage clamp in a 10kV distribution line" which is easy to understand, providing a basis for subsequent processing. This conversion process can be performed using a trained neural network model that has the ability to extract key information and summarize from the target task, or other existing technologies can be used for conversion, which will not be described in detail here. Then, based on the understanding ability of the pre-trained large language model of the Transformer architecture, the LLM combines the job task knowledge base to analyze the semantic structure and multi-modal fusion features of the input target job task, identifies key actions (such as "replace"), objects (such as "clamp"), and logical sequences to decompose the task and obtain the initial subtask sequence, i.e. the LLM maps the task to the standard process in the knowledge base through template matching, for example, "replace the drainage clamp" is matched to the template "tool preparation → positioning → installation → inspection". For non-standard or complex tasks, the LLM generates an initial subtask sequence using chain reasoning, for example, it deduces that "installing a new clamp" requires "positioning the clamp" to be completed first, and "preparing tools" before positioning. In the decomposition process, the LLM embeds the robot's capabilities and safety constraints to ensure that each subtask is executable. For example, the "position and grasp the target clamp" subtask limits the movement range of the robotic arm to avoid touching the high-voltage line, thereby meeting the safety specifications.

[0065] After generating the subtask sequence, optimization is needed to ensure execution efficiency and logic. The LLM arranges the subtask order according to the job task knowledge base and logical dependencies and priorities derived from expert rules, such as "detect and prepare job tools" must precede "locate and grasp target clamp" to ensure tool availability. Each sorted subtask is parameterized, i.e., specific execution parameters are defined, such as target coordinates and grasping force for "locate and grasp target clamp" and installation angle and torque for "install new clamp". These parameters are determined based on the job task knowledge base and robot hardware specifications (such as force control accuracy of the robotic arm), generating a first-optimized subtask sequence. Then, the feasibility of the subtask sequence is verified through simulation or historical data, such as checking whether "install new clamp" is within the robotic arm's workspace and meets safety distance requirements. If problems are found, the LLM will adjust the subtask definition or order.

[0066] Based on factors such as the robot's physical movement range (such as maximum travel speed, turning radius, mechanical arm operation range, etc.), safety distance (>0.2m) constraints, mechanical structure limitations, and task logic rules in the job task knowledge base, the first-optimized subtask sequence is further optimized. For example, if the robot's arm cannot reach a certain position (does not meet safety distance constraints or exceeds physical movement range), the execution method or order of the subtask is adjusted based on the job task knowledge base, robot hardware specifications, and constraint conditions to ensure that the robot can successfully complete each subtask, resulting in a second-optimized subtask sequence.

[0067] The second-optimized subtask sequence is presented to operators or experts, who receive external instructions and feedback information from them. Operators or experts can make modification suggestions based on their experience and actual situation, such as adjusting the order of subtasks, adding or deleting certain subtasks, modifying subtask parameters, etc. Then, the second-optimized subtask sequence is re-optimized based on the received external instructions to obtain the final subtask sequence, making it not only meet the requirements of robot capabilities and job specifications, but also fully consider human professional knowledge and actual needs, with higher feasibility and effectiveness.

[0068] The application utilizes the powerful language understanding and generation capability of a large language model to intelligently decompose a task according to a knowledge base and multi-modal information, fully utilizes the advantages of multi-modal data and a language model, and improves the accuracy and rationality of task decomposition; a multi-layer optimization mechanism is used to process an initial subtask sequence: sorting and parameterization processing are performed according to a job task knowledge base, so as to ensure that the logical order and parameter settings of the subtasks meet the specifications and actual requirements; optimization is performed in combination with the physical motion range of a robot, so that the subtask sequence can be executed within the capability range of the robot; a human-in-the-loop mechanism is introduced for re-optimization, so as to fully exert the subjective initiative and professional knowledge of a human being, and further improve the quality and feasibility of the subtask sequence; the mechanism can perfect the subtask sequence from different angles, and ensures that the final obtained subtask sequence can efficiently and safely complete a target job task.

[0069] In an embodiment, the target job task is converted into natural language and input into a pre-trained large language model based on a Transformer architecture together with the multi-modal fusion features, to decompose the target job task in combination with a job task knowledge base, and obtain an initial subtask sequence, including:

[0070] The target job task is analyzed to obtain a structured natural language task, which is processed by a feature alignment embedding method together with the multi-modal fusion features, and time sequence position encoding is embedded in the processing result to obtain a cross-modal fusion feature vector containing time sequence information;

[0071] Based on the cross-modal fusion feature vector and the structured natural language task, knowledge is matched from the job task knowledge base to obtain static knowledge matched with the target job task;

[0072] The cross-modal fusion feature vector, the structured natural language task and the static knowledge are input into the pre-trained large language model based on the Transformer architecture for dynamic template matching and chain reasoning, to decompose the target job task through a template self-adaptation mechanism in a first path to generate a preliminary subtask framework, and decompose the target job task through a constraint-guided thinking chain in a second path to generate a subtask logic chain;

[0073] The preliminary subtask framework and the subtask logic chain are subjected to conflict detection and feasibility checking, and after the conflict detection passes and the feasibility checking passes, the initial subtask sequence is generated according to the preliminary subtask framework and the subtask logic chain.

[0074] Specifically, the present application extracts the core information in the target task (oral expression such as "replace the loose drain wire clamp on the 10kV line") by designing a task element extraction template, which includes action type (such as replacement, maintenance, installation, etc.), target object (such as drain wire clamp, insulator, etc., and combined with image recognition results to correct ambiguity, for example, to distinguish between clamp and insulator), and constraint conditions (such as 10kV line, live-line work, etc.), and fills the extracted information into a standardized natural language template to generate a structured natural language task, for example: "In the 10kV live-line distribution line, perform the replacement operation on the loose drain wire clamp, and meet the safety distance ≥ 0.2m, environmental wind speed ≤ 5m / s". The structured natural language task and the multi-modal fusion feature are converted into a 768-dimensional text feature vector using the text encoder of LLM (such as Token Embedding layer of LLaMA-3), and the 1024-dimensional multi-modal fusion feature is reduced to 768-dimensional using a multi-layer perception (MLP) to match the dimension of the text feature vector; the association weight between the text and the multi-modal fusion feature is calculated through the modal attention mechanism (such as the weight value of "clamp position" in the visual feature and "replace the clamp" in the text), and the cross-modal feature vector of 1536 dimensions is obtained by weighted fusion, and the time position encoding is embedded in the feature (the same as the position encoding in the previous text), and the timestamp encoding is added in the cross-modal feature for dynamic work scene (such as wire shaking), so that the LLM can identify the difference between "the current position of the clamp" and "the position 10 seconds ago", and then obtain the cross-modal fusion feature vector containing time sequence information.

[0075] Among them, the modal attention mechanism calculation includes: projecting the text feature and the multi-modal fusion feature into the same dimensional space through linear transformation; generating query vector Q, key vector K and value vector V of the text feature and the multi-modal fusion feature respectively, calculating the similarity between the text query vector and the multi-modal key vector through dot product operation to obtain the original attention score; applying the Softmax function to the original attention score to convert the score into a probability distribution; according to the normalized weight, the multi-modal value vector is weighted and summed to generate the text-guided context feature, and the text feature, the multi-modal feature and its context feature are spliced and averaged or maximized, and the fused feature representation is obtained.

[0076] In addition, a pre-trained feature alignment model can also be used to map the structured natural language task and the multi-modal fusion feature to the same feature space, such as the explicit sequence of disassembling the old clamp and installing the new clamp, adding time position encoding to each feature element, for example, the feature element related to disassembling the old clamp is marked as time sequence 1, and the feature element related to installing the new clamp is marked as time sequence 2, and the cross-modal fusion feature vector containing time sequence information can also be obtained.

[0077] With the cross-modal fusion feature vector and the structured natural language task as the query condition, the similarity between the structured natural language task and the features in the knowledge base is calculated by cosine similarity to find the most matching static knowledge in the knowledge base with the target task; also, through the environmental parameters (such as wind speed 3 m / s) and target attributes (such as clip model) in the cross-modal features, applicable rules are matched from the dynamic rule base as pre-constraints for LLM reasoning.

[0078] Through the target detection model (such as YOLOv8), the target attributes (clip position coordinates, whether blocked, relative distance from the robot arm) are extracted from the visual features in the cross-modal fusion features, and the Fourier transform is performed on the time series features in the cross-modal fusion features to identify abnormal patterns (such as torque fluctuation > 5 N·m and duration > 2 s to determine corrosion jamming), and for standard tasks (that is, tasks with fixed operation procedures, the division of tasks can also be based on the relationship between their complexity and the complexity threshold, such as tasks with complexity exceeding 0.8 are non-standard tasks, and others are standard tasks, and the task complexity is determined based on the completion time of historical tasks or obtained through linear regression model, decision tree / random forest, neural network, etc.), a standard procedure template library is pre-constructed, which is stored in a three-dimensional structure of task type-core step-variable parameter, such as the basic template for "changing drainage clip":

[0079] {

[0080] Task Type: Change Drainage Clip,

[0081] Core Steps: [Tool Preparation, Target Positioning, Old Part Disassembly, New Part Installation, Quality Inspection],

[0082] Variable Parameters: {

[0083] Tool Preparation: {Tool List: [Insulating Wrench, Clamp], Insulation Level: ≥10kV},

[0084] Target Positioning: {Accuracy Requirement: ±5cm, Safety Distance: ≥0.2m},

[0085] Old Part Disassembly: {Default Torque: 10 N·m, Maximum Torque: 15 N·m}

[0086] },

[0087] Trigger Conditions: { / / Feature thresholds for triggering step adjustments

[0088] Robot Arm Extension Limited: {Clip Height: > Maximum Robot Arm Extension Height},

[0089] Old Part Rust: {Torque Fluctuation: > 5 N·m and Duration: > 2 s}

[0090] }

[0091] }。

[0092] By the pre-training large language model based on the Transformer architecture, BERT or its improved model matches the parsed features and these templates, and then decomposes the task based on the matched templates; however, it should be noted that the template is not fixed, but is adapted through a template adaptation mechanism, i.e. LLM dynamically adjusts the step details based on the task knowledge base and the matched static knowledge, such as showing “the clip position is high, and the mechanical arm is limited in extension” in the cross-modal fusion feature vector, then adding a “adjust the robot base position” sub-step in the “positioning” step, such as showing “old clip rust”, then adding a “increase torque to 15N・m” parameter in the “disassembly” step, etc., to decompose the target task based on the template adaptation mechanism in the first path, and generate a preliminary sub-task framework containing dynamically adjusted step details.

[0093] In addition, in the first path, the cosine similarity algorithm can also be used to calculate the matching degree of the parsed multi-modal features and the “trigger condition” in the template, and for the trigger condition with a matching degree ≥0.8, the template adaptation mechanism is started to adjust the steps, and the specific adjustment content can be determined according to the cross-modal features based on the task knowledge base and the matched static knowledge, such as for the mechanical arm limited in extension: inserting the “adjust the robot base position” sub-step before the “positioning target” step, and supplementing the parameters (such as “base elevation 0.5m, ensure that the vertical distance between the mechanical arm and the clip is ≤3.0m”); for “old piece rust”: modify the parameters in the “disassemble old piece” step (such as “initial torque 12N・m, increase by 1N・m every 3s, maximum not exceeding 15N・m”); etc., integrate the adjusted steps and parameters to generate a sub-task sequence with dynamic details.

[0094] For non-standard tasks (other tasks in addition to standard tasks), first, the constraint conditions in the task are quantified, i.e. the safety specifications and robot capabilities are converted into calculable constraint values, such as the maximum extension radius of the mechanical arm 1.5m, the operation range of the insulated bucket ±0.8m, and the prohibition of touching live wires (distance ≥0.2m); then, the cross-modal fusion features, unstructured natural language tasks, and matched static knowledge are input, and a pre-training large language model based on the Transformer architecture, such as Transformer, is used to generate sub-tasks by utilizing its self-attention mechanism through problem decomposition→step-by-step deduction→self-checking:

[0095] First round reasoning (target disassembly): To replace the clip blocked by the tree branch, the tree branch needs to be removed first, but care must be taken not to touch the wire and the clip. Since the tree branch diameter is 8cm < the maximum capacity of the insulated scissors 10cm, it can be directly cut.

[0096] Second round of reasoning (step refinement):

[0097] Step 1: Move the insulated bucket to 1m from the branch (since the insulated bucket has a movement range of ±0.8m, the current distance from the target is 2m, it needs to be moved twice, each time 0.8m and 0.2m, to avoid exceeding the range;

[0098] Step 2: Adjust the angle of the mechanical arm so that the insulated scissors are aligned with the branch (the cutting point is selected at a distance of 30cm from the wire clamp, avoiding the wire clamp and ensuring that there is no residual obstruction after cutting);

[0099] Step 3: Cut the branch (force 50N, as the diameter is 8cm, ensure that it is cut off, and monitor the reaction force to avoid jamming the cutting edge);

[0100] Step 4: Confirm that the branch has fallen (visually check that there is no residual obstruction and that the wire has not been touched) → follow the regular replacement steps.

[0101] Third round of reasoning (parameter verification):

[0102] Verify the mechanical arm extension distance in Step 2: The current distance between the insulated bucket and the branch is 1m, and the mechanical arm needs to be extended by 0.8m to reach it (1m + 0.8m = 1.8m < maximum radius 1.5m? No, calculation error! Recalculate: move the insulated bucket to a distance of 0.7m from the branch, extend the mechanical arm by 0.8m, the total distance is 1.5m = maximum radius, which meets the constraints.

[0103] Then check the dependency between steps through the self-recursive mechanism of the LLM (such as "cutting the branch" must be after "moving the insulated bucket", etc.), and after the logical coherence detection passes, finally generate a sub-task logic chain containing operation details, parameter constraints, and logical order.

[0104] Check the preliminary sub-task framework and sub-task logic chain to determine whether there are any time conflicts (such as multiple sub-tasks not requiring simultaneous execution), resource conflicts (such as the robot not using multiple non-shareable devices simultaneously), and operation conflicts (such as the operation sequence meeting the specifications). If not, consider the physical motion range of the robot to assess whether each sub-task can be executed in the actual work environment, such as whether all sub-tasks are within the reach of the robot. After conflict detection and feasibility verification pass, integrate the preliminary sub-task framework and sub-task logic chain to generate an initial sub-task sequence. If there are conflicts or unfeasible situations, return to adjust the parameters of the template adaptive mechanism or constraint-guided thinking chain and re-decompose the task.

[0105] The application adopts a double-path (template self-adaptive mechanism and constraint-guided thinking chain) mode to decompose the target operation task, which not only ensures the normativity and consistency of decomposition (through template self-adaptation), but also can flexibly reason according to the specific task situation (through constraint-guided thinking chain), thereby improving the accuracy and adaptability of task decomposition.

[0106] Model training is the core of achieving high-quality task decomposition. The training data set is composed of expert-labeled subtask sequences and corresponding operation data (such as visual images, force feedback, and motion trajectories). For example, videos, sensor data, and operator instructions from multiple "replace drainage clamp" operations are collected. Supervised fine-tuning (SFT) is used to enable the LLM to learn the decomposition logic of experts, input task descriptions, and output subtask sequences and their parameters. To further optimize the decomposition strategy, reinforcement learning (RL) can be used with task completion efficiency and safety as the reward function. For example, if the robot fails to perform "install new clamp" (e.g., the clamp is loose), the model will adjust the decomposition strategy to add a "force calibration" subtask. Considering the common tasks of live-line work (such as tool inspection and target positioning), a general model can be pre-trained through transfer learning, and then fine-tuned for specific tasks (such as clamp installation) to improve the model's generalization ability and training efficiency.

[0107] For example, the process of "replacing a drainage clamp" usually includes steps such as tool inspection, positioning operation, installation, and verification, which are provided by power operation experts. The specific implementation process is as follows: input the task description "replace the drainage clamp in the 10kV distribution network line" and the environmental information (such as line location and weather conditions), the LLM extracts key information (target: clamp, action: replace, environment: high-voltage line), and generates a subtask sequence: the first subtask "detect and prepare operation tools" calls the VLA model to verify the integrity of the insulating gloves and special wrench; the second subtask "position and grab the target clamp" processes the visual data through the VLA model to generate a grabbing action; the third subtask "install a new clamp" controls the robot to perform installation with a specified torque; and the fourth subtask "final check" uses the vision system to confirm the installation quality. After each subtask is executed, the VLA model provides feedback (such as success rate of grabbing and stability of installation), and the LLM optimizes subsequent actions based on the feedback. If the operator finds that the clamp model is incorrect, they can input "replace model Y clamp", and the LLM will re-decompose the task and insert a "replace clamp model" subtask.

[0108] In an embodiment, the receiving external instructions introduces the human-in-the-loop mechanism to re-optimize the secondary optimization subtask sequence to obtain a final subtask sequence, including:

[0109] receiving an external instruction and performing standardization processing on the external instruction to obtain a standardized instruction triple including a type, a content, and a timestamp;

[0110] based on a double-threshold check rule constraint, inputting the standardized instruction triple into a job rehearsal model constructed based on a digital twin technology to simulate, to generate an effective instruction queue, and judging the type of the effective instruction queue;

[0111] when the effective instruction queue is a priority intervention instruction, analyzing a logical relationship between subtasks in the secondary optimization subtask sequence by a large language model, and adjusting an execution position of a target subtask corresponding to the effective instruction queue to obtain a first subtask sequence;

[0112] when the effective instruction queue is a task replacement instruction, generating a breakpoint mark based on a target subtask corresponding to the effective instruction queue, and calling the visual language action model to generate a new task according to the breakpoint mark to adjust the secondary optimization subtask sequence to obtain a second subtask sequence;

[0113] when the effective instruction queue is a parameter correction instruction, performing parameter boundary checking on the effective instruction queue, and converting the effective instruction queue into a control signal after passing the checking to correct the secondary optimization subtask sequence to obtain a third subtask sequence.

[0114] Specifically, the application receives an external instruction and performs standardization processing thereon, and converts it into a standardized instruction triple including an instruction type: priority intervention / task replacement / parameter correction, a core content: specific operation description, and a timestamp: a receiving time.

[0115] based on a double-threshold check rule constraint, inputting the standardized instruction triple into a job rehearsal model constructed based on a digital twin technology to simulate, to generate an effective instruction queue, and judging the type of the effective instruction queue;

[0116] When the effective instruction queue is a priority intervention instruction (such as "priority check tool insulation, reposition the clamp"), the instruction and the secondary optimization subtask sequence (original order: positioning → tool check → installation) are taken as inputs, the logical dependencies between subtasks (such as "installation" must be completed after "positioning" and cannot be adjusted; "tool check" has no dependency on "positioning" and can be interchanged, etc.) are analyzed by the large language model LLM, the target subtask (tool check) in the priority intervention instruction is moved to the priority position, and if there are pre-requisite steps in the original sequence (such as "tool check" requires "tool retrieval" first), the pre-requisite steps are automatically completed, a new sequence is generated and output: retrieve tool → tool check → positioning → installation, which is the first subtask sequence after adjustment while preserving logical dependencies.

[0117] When the effective instruction queue is a task replacement instruction (such as "pause installation, clear tree branches first"), the instruction and the state data corresponding to the secondary optimization subtask sequence (such as the installation step has been completed 30% and the current pose of the robot arm) are taken as inputs, the progress parameters of the current subtask (such as installation torque value, robot arm joint angle) are recorded, a breakpoint marker (such as {task ID: install clamp, progress: 30%, recovery condition: tree branch clearing completed}) is generated, and based on the breakpoint marker, the VLA model is called to generate substeps for the new task (clearing tree branches) (positioning tree branches → adjusting the position of the insulation bucket → controlling the insulation cutter to cut), and ensure that the new task path smoothly connects with the current robot arm pose (minimizing movement distance), obtaining a new subtask sequence containing a breakpoint marker.

[0118] When the effective instruction queue is a parameter correction instruction (such as "clamp force reduced from 8N to 6N"), the instruction and the parameter execution range of the secondary optimization subtask sequence (such as "safety range of lead clamp clamping force 5-10N") are taken as inputs, the parameter boundary check of the effective instruction queue is performed first, i.e. confirming that the correction value (6N) is within the safety range (5-10N), if it exceeds the range (such as input 12N), it is automatically clamped to the maximum value (10N) and a prompt "parameter exceeds safety range, adjusted to 10N" is given, the correction parameter in the parameter correction instruction is converted into a robot arm control signal (such as current command), and the actual clamping force is monitored in real time (feedback every 10ms) through a force sensor to ensure that the deviation from the correction value is ≤0.5N, the correction instruction and the robot arm control signal are integrated into the secondary optimization subtask sequence, and the third subtask sequence after updating the parameters is obtained. The subtask sequence optimized by external instructions is the final subtask sequence.

[0119] The application unifies external instructions into a triple including type, content and timestamp, solves the format difference problem of instructions from different sources, provides standardized input for subsequent processing, constructs a job rehearsal model, simulates the influence of instructions on a task sequence, identifies potential conflicts in advance, and avoids errors in actual execution, combines the advantages of digital twin (physical simulation), LLM (logic analysis) and VLA (action generation), forms a "perception-decision-execution" closed loop, and significantly improves the flexibility and accuracy of task optimization.

[0120] S4, based on the subtask type in the subtask sequence, inputting the subtask sequence and the multi-source modal data into a visual language action model or a reinforcement learning model for processing, to generate motion instructions corresponding to the subtask sequence, so as to start an execution process of the target job task by the robot; wherein the visual language action model comprises a visual network layer, a language network layer and a flow matching action network layer; the multi-source modal data further comprises real-time sensor data;

[0121] In an embodiment, the step S4 comprises:

[0122] judging the subtask type in the subtask sequence,

[0123] if a certain subtask in the subtask sequence is determined to be of a first type, inputting the multi-source modal data into the reinforcement learning model for processing, to generate a sub-motion instruction corresponding to the first type subtask;

[0124] if a certain subtask in the subtask sequence is determined to be of a second type, inputting the multi-modal fusion feature and the second type subtask into the visual language action model, so that the visual network layer identifies the multi-modal fusion feature to obtain a target visual feature corresponding to the second type subtask, so that the language network layer analyzes the second type subtask in combination with the job task knowledge base to obtain an ordered subtask sequence, and so that the flow matching action network layer processes the target visual feature, the ordered subtask sequence and the real-time sensor data to generate a sub-motion instruction corresponding to the second type subtask;

[0125] integrating each sub-motion instruction to obtain a motion instruction corresponding to the subtask sequence, so as to start an execution process of the target job task by the robot.

[0126] Specifically, after task decomposition, the application adopts a visual language action model (VLA) or a reinforcement learning model according to the type of subtask to realize action execution, and the specific steps comprise:

[0127] judging the subtask type contained in the finally obtained subtask sequence;

[0128] If a subtask in the subtask sequence is a fixed subtask such as picking up a wire or screwing a screw, it is of the first type, and a pre-trained reinforcement learning model is called to process the multi-source modal data and generate the corresponding split motion instructions (joint speed / torque) for the first type of subtask. Among these reinforcement learning (RL) models, different types are used, but they are all built for a single subtask (such as screwing a screw) to achieve efficient and robust autonomous operation of the robot in a specific task scenario. These skill models are equivalent to expert-level low-level control strategies. When the robot needs to perform a certain specific subtask, the corresponding skill model can be called to complete it. For example, for a subtask such as picking up a wire, which is repeated and has a relatively fixed operation mode, a reinforcement learning model is trained to adjust the robot's pose to stably hold the vibrating wire.

[0129] If a subtask in the subtask sequence is a fixed subtask such as picking up a wire or screwing a screw, it is of the first type, and a pre-trained reinforcement learning model is called to process the multi-source modal data and generate the corresponding split motion instructions (joint speed / torque) for the first type of subtask. Among these reinforcement learning (RL) models, different types are used, but they are all built for a single subtask (such as screwing a screw) to achieve efficient and robust autonomous operation of the robot in a specific task scenario. These skill models are equivalent to expert-level low-level control strategies. When the robot needs to perform a certain specific subtask, the corresponding skill model can be called to complete it. For example, for a subtask such as picking up a wire, which is repeated and has a relatively fixed operation mode, a reinforcement learning model is trained to adjust the robot's pose to stably hold the vibrating wire.

[0130] In the training phase of the skill model, a reinforcement learning algorithm is used to allow the robot agent to repeatedly interact with the environment in a simulated environment, starting from zero experience to explore the optimal strategy. A reasonable reward function is set during the training process to guide learning, such as giving positive rewards for successfully completing subtasks, and negative rewards for failure or timeout, while time consumption and energy consumption can be included in the reward function to balance. Through such trial and error learning, the RL agent gradually adjusts its policy network parameters to maximize the cumulative expected reward, that is, to gradually approach the optimal strategy. However, it may take a long time to converge by random exploration alone. To improve training efficiency and avoid unnecessary collision risks, expert demonstration data is introduced for imitation learning at the beginning of training. Specifically, the robot imitates several demonstration trajectories of human or traditional controller completing subtasks, converts these trajectories into training samples for the policy network, and makes the agent start learning from an initial strategy with higher performance. This can significantly accelerate the convergence of reinforcement learning while ensuring the reasonableness of the strategy. Then gradually transition to pure reinforcement learning to fine-tune and optimize the model based on the existing foundation. The final skill model reaches a level close to or even exceeding that of human experts in its specialized tasks, with high accuracy, fast response, and strong robustness.

[0131] If a subtask in the subtask sequence is a non-fixed subtask, it is a second type of subtask, and the multimodal fusion feature and the second type of subtask are input into the visual language action model VLA for processing to obtain a sub-movement instruction (joint speed / torque instruction corresponding to a collision-free smooth trajectory) corresponding to the second type of subtask.

[0132] The visual language action model processes the input multimodal fusion feature and the second type of subtask through the visual network layer (CLIP-ViT) to identify the target visual feature, i.e., the target position and the safe distance, for output. Specifically, the visual network layer is responsible for processing environment images and sensor data, extracting spatial information and geometric features of obstacles, target objects (such as wires, wire clamps, wire strippers). The CLIP-ViT-L-336px is used to process high-resolution RGB images (from Intel RealSense or similar cameras), depth images (from laser radars or depth sensors), and point cloud data (generated by 3D sensors) after feature extraction and fusion to form multimodal fusion features. For live-line work scenarios, the visual network layer pays special attention to the minimum safe distance of the wire (usually > 0.2m) and the accurate position of the target object (x, y, z coordinates). For example, in the wire stripping task, the visual network can identify the diameter, position of the cable, and clamping point of the wire stripper, ensuring that the action generation meets the insulation and safety constraints.

[0133] The second type of subtask is parsed by the language network layer (LLaMA-3.1) into an action sequence, and the job task knowledge base is combined to generate constraint conditions to obtain an ordered subtask sequence. The language model combines with the expert knowledge base (such as the artificial live working process and safety specifications) to ensure that the subtask sequence meets the power industry standards (such as avoiding touching live components and ensuring strict action order). In addition, the model supports the "man-in-the-loop" mechanism, allowing operators to dynamically adjust tasks through natural language instructions (such as "adjust the wire stripper angle").

[0134] The target visual features, ordered subtask sequence, and real-time sensor data are input into the flow matching action network layer to perform action prediction (the flow matching network based on the Transformer architecture generates action trajectories) and safety verification (checks whether the torque exceeds the knowledge base threshold, etc.) through the flow matching algorithm, generating the corresponding sub-movement instructions for the second type of subtask. Flow matching gradually generates smooth trajectories that meet environmental constraints from random noisy trajectories through iterative denoising. The generation process is conditioned on visual features (obstacle positions, target positions) and sensor data (such as wind speed, conductor sway) to ensure collision-free trajectories and meet safety distance requirements (e.g., avoid conductor by more than 0.2m). The action expert adjusts the motion parameters (such as closing speed, force threshold) based on feedback to ensure gentle grasping and no damage to the target. For example, when the tactile sensor detects that the grasping force reaches the stable threshold (such as 10N), the action expert immediately stops applying greater pressure.

[0135] The outputs of the two models are processed by the TimeSynchronizer node of ROS (Robot Operating System) for time synchronization to align the timestamps of each instruction; then check whether the robot arm has received conflicting instructions such as "move to B point" and "move to A position" at the same time for conflict detection, fuse the instructions that pass the detection in order to generate a global motion instruction sequence, i.e., the motion instructions corresponding to the subtask sequence, to start the execution process of the target job task. It should be noted that the specific process of data processing by some algorithms (flow matching) and models (LLaMA-3.1, etc.) in the above scheme can refer to the application of the algorithm or model in the prior art, and will not be described here.

[0136] In this embodiment, the distributed deep learning framework PyTorch is used, and the GPU cluster is used to accelerate the training. By designing the loss function (including cross-entropy loss, mean square error loss, and inter-modal contrast loss), the performance of the model is effectively evaluated during the training process. In addition, through hyperparameter tuning (including learning rate, dropout, batch size, etc.), the training time is shortened and the model performance is improved. Finally, the precision, response speed and other multi-dimensional indicators are used to comprehensively evaluate the multi-modal large model. Combined with the simulation site test data and the simulation platform verification result, the model structure and parameters are continuously iterated and optimized, and the adaptability and decision accuracy of the robot in the complex dynamic live working environment are continuously improved, so as to ensure the efficiency and safety of the robot in actual operation.

[0137] The application breaks through the limitation of traditional single model processing all tasks, automatically selects reinforcement learning model or visual language action model VLA according to the characteristics of subtasks, significantly improves the adaptability and efficiency of task processing, improves the smoothness and safety of trajectory generation, ensures the stability and repeatability of the robot in the dynamic environment, and meets the insulation and distance requirements of live working; the reinforcement learning model for a single subtask can optimize the strategy through interaction with the environment, introduce expert demonstration data to speed up the learning process, and significantly improve the adaptability and decision accuracy of the robot in the complex dynamic environment.

[0138] In the embodiment of the application, based on the problem of how to improve the working efficiency of the existing autonomous network distribution live working robot, a robot control method based on a multi-modal large model is designed, which collects multi-source modal data of the scene where the target task is located in real time, and uses corresponding machine learning models to extract features for different types of data. This multi-source data fusion processing method can more comprehensively and accurately capture scene information, providing a richer feature basis for subsequent task decomposition and instruction generation compared to single modal data processing, improving the understanding ability of complex working scenes; after position encoding and Transformer fusion processing of multi-modal features to obtain multi-modal fusion features, a task knowledge base is constructed, which is input into a large language model together with the fusion features for task decomposition, and a human-in-the-loop mechanism is introduced to optimize the decomposition result. This combination fully utilizes the language understanding and reasoning ability of the large language model and the professional knowledge and experience of humans, making the task decomposition more in line with actual needs and logic, and improving the accuracy and rationality of task decomposition; based on the subtask types in the subtask sequence, a visual language action model or a reinforcement learning model is flexibly selected for processing to generate motion instructions. This dynamic model selection mechanism can take advantage of different models according to the characteristics of different subtasks, improve the efficiency and accuracy of instruction generation, better control the robot to execute complex working tasks, and thus improve the working efficiency of the existing autonomous network distribution live working robot.

[0139] It should be noted that although the steps in the flowchart above are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order requirement for the execution of these steps, and they can be executed in other orders.

[0140] In another embodiment, such as Figure 2 As shown, a second aspect of the present invention provides a robot control system based on a multimodal large model, comprising:

[0141] The feature extraction module 10 is used to collect multi-source modal data of the scene where the target task is located in real time, and extract features from the multi-source modal data using the corresponding machine learning model based on the type of the multi-source modal data to obtain multi-modal features;

[0142] The feature fusion module 20 is used to perform position encoding and Transformer fusion processing on the multimodal features to obtain multimodal fused features;

[0143] The task decomposition module 30 is used to input the constructed task knowledge base and the multimodal fusion features into the large language model to obtain the decomposition result of the target task, and to introduce a human-in-the-loop mechanism to optimize the decomposition result to obtain a sub-task sequence.

[0144] The instruction execution module 40 is used to input the subtask sequence and the multi-source modal data into a visual language action model or reinforcement learning model for processing based on the subtask type in the subtask sequence, and generate motion instructions corresponding to the subtask sequence so that the robot starts the execution process of the target task.

[0145] It should be noted that the various modules in the aforementioned multimodal large-scale robot control system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module. For specific limitations regarding the multimodal large-scale robot control system, please refer to the limitations of the multimodal large-scale robot control method described above; both have the same function and role, and will not be repeated here.

[0146] A third aspect of the present invention provides an electronic device comprising:

[0147] Processor, memory, and bus;

[0148] The bus is used to connect the processor and the memory;

[0149] The memory is configured to store operation instructions.

[0150] The processor is configured to execute the operation instructions to perform the operation of the robot control method based on the multi-modal large model.

[0151] In an optional embodiment, an electronic device is provided, which comprises Figure 3 As shown in the figure, Figure 3 The electronic device 5000 shown in the figure comprises a processor 5001 and a memory 5003. The processor 5001 and the memory 5003 are connected, for example, through a bus 5002. Optionally, the electronic device 5000 can further comprise a transceiver 5004. It should be noted that in actual applications, the transceiver 5004 is not limited to one, and the structure of the electronic device 5000 does not constitute a limitation on the embodiments of the present application.

[0152] The processor 5001 can be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules and circuits described in combination with the disclosure. The processor 5001 can also be a combination of computing functions, such as one or more microprocessor combinations, DSP and microprocessor combinations, etc.

[0153] The bus 5002 can comprise a channel for transmitting information between the above-mentioned components. The bus 5002 can be a PCI bus or an EISA bus, etc. The bus 5002 can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 3 In the figure, only one thick line is used to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0154] The memory 5003 can be a ROM or other type of static storage device that can store static information and instructions, a RAM or other type of dynamic storage device that can store information and instructions, an EEPROM, a CD-ROM or other optical disk storage, an optical disk storage (including a compact disk, a laser disk, an optical disk, a digital versatile disk, a Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program codes in the form of instructions or data structures and can be accessed by a computer, but not limited thereto.

[0155] The memory 5003 is configured to store application program codes for implementing the solutions of the present application, and the processor 5001 is configured to control the execution of the application program codes stored in the memory 5003. The processor 5001 is configured to execute the application program codes stored in the memory 5003 to implement the content shown in any of the foregoing method embodiments.

[0156] The electronic device includes, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a car terminal (for example, a car navigation terminal), and the like, and a stationary terminal such as a digital TV, a desktop computer, and the like.

[0157] The fourth aspect of the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the program is executed by a processor to implement the robot control method based on a multi-modal large model shown in the first aspect of the present application.

[0158] Another embodiment of the present application provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and when the computer program is executed on a computer, the computer can execute the corresponding content in the foregoing method embodiments.

[0159] In addition, an embodiment of the present application also provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the program is executed by a processor to implement the steps of the above method.

[0160] In summary, the present application relates to the field of robot control, and discloses a robot control method, system, device and medium based on a multi-modal large model. The multi-modal feature is obtained by collecting multi-source modal data of a scene where a task is located and processing the multi-source modal data through a machine learning model, and the multi-modal feature is subjected to position coding and Transformer fusion processing to obtain a multi-modal fusion feature. The multi-modal fusion feature and a constructed task knowledge base are input into a large language model to decompose a target task, and a human-in-the-loop mechanism is introduced to optimize the decomposition result to obtain a sub-task sequence. According to the sub-task type in the sub-task sequence, the sub-task sequence is processed through a visual language action model or a reinforcement learning model to generate a motion instruction to start an execution process of the target task by the robot. The multi-modal large model LLM, VLA, etc. are used to process the live-line task, and the working efficiency of the autonomous distribution network live-line robot is improved.

[0161] Various embodiments are described herein with reference to the drawings, wherein each embodiment is described in a progressive manner, and each embodiment directly or indirectly refers to each other, and each embodiment focuses on the differences from other embodiments. In particular, the system embodiments are described more simply because they are basically similar to the method embodiments, and the relevant parts refer to the description of the method embodiments. It should be noted that the technical features of the above embodiments can be combined in any way, and in order to make the description simple, not all possible combinations of the technical features of the above embodiments are described, but as long as the combination of the technical features does not exist Contradiction, it should be considered within the scope of the description.

[0162] The above-described embodiments only express several preferred embodiments of the present application, which are described in detail and in detail, but should not be construed as limiting the scope of the patent. It should be noted that for ordinary skilled in the art, without departing from the technical principles of the present application, a number of improvements and replacements can be made, which should be considered as the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the protection scope of the claims.

Claims

1. A method for robot control based on a multi-modal large model, characterized in that, The application relates to a method for decomposing a target work task into a subtask sequence. The method comprises the following steps: real-time acquisition of multi-source modal data of a scene where a target work task is located, feature extraction of the multi-source modal data by using a corresponding machine learning model based on the type of the multi-source modal data, and obtaining of multi-modal features; position coding and Transformer fusion processing of the multi-modal features, and obtaining of multi-modal fusion features; inputting of a constructed work task knowledge base and the multi-modal fusion features into a large language model to obtain a decomposition result of the target work task, and optimization of the decomposition result by introducing a human-in-the-loop mechanism to obtain a subtask sequence; 2. The robot control method based on a multi-modal large model according to claim 1, characterized in that, based on the type of a subtask in the subtask sequence, inputting of the subtask sequence and the multi-source modal data into a visual language action model or a reinforcement learning model for processing, generation of motion instructions corresponding to the subtask sequence, and starting of a robot to execute a process of the target work task. The multi-source modal data comprises force sensation and point cloud data, image data, and audio data. The feature extraction of the multi-source modal data by using the corresponding machine learning model comprises the following steps: inputting of the force sensation and point cloud data into a long short-term memory network model for processing to capture a force interaction mode and obtain time-series dynamic features; inputting of the image data into an improved ResNet-50 model for processing to obtain local features of a target object; inputting of the audio data into a Whisper-Base model and a BERT-Base-Chinese model for processing to obtain speech features for parsing execution instructions; 3. The robot control method based on a multi-modal large model according to claim 1, characterized in that, combination of the time-series dynamic features, the local features, and the speech features to obtain the multi-modal features. The position coding and Transformer fusion processing of the multi-modal features comprise the following steps: embedding of a time step into the multi-modal features by position coding to obtain multi-modal coding features; 4. The robot control method based on a multi-modal large model according to claim 1, characterized in that, inputting of the multi-modal coding features into a Transformer encoder to perform feature fusion by a self-attention mechanism and a multi-head attention mechanism to obtain the multi-modal fusion features. The inputting of the constructed work task knowledge base and the multi-modal fusion features into the large language model to obtain the decomposition result of the target work task, and the optimization of the decomposition result by introducing the human-in-the-loop mechanism to obtain the subtask sequence comprise the following steps: acquisition of live work specification data and live work expert data, extraction of work action static association information and work environment dependency rule information from the live work specification data and the live work expert data, and construction of the work task knowledge base; conversion of the target work task into natural language, inputting of the natural language and the multi-modal fusion features into a pre-trained large language model based on a Transformer architecture, decomposition of the target work task by combining the work task knowledge base, and obtaining of an initial subtask sequence; sorting of the initial subtask sequence according to the work task knowledge base, and parameterization processing of the sorted initial subtask sequence to generate a first optimized subtask sequence; optimizing the primary optimization subtask sequence based on a physical motion range of the robot and in combination with the job task knowledge base, to obtain a secondary optimization subtask sequence; receiving an external instruction to introduce the human-in-the-loop mechanism to re-optimize the secondary optimization subtask sequence, to obtain a final subtask sequence.

5. The robot control method based on a multi-modal large model according to claim 4, characterized in that, The target job task is converted into natural language and input into a pre-trained large language model based on a Transformer architecture together with the multi-modal fusion features, to decompose the target job task in combination with the job task knowledge base, to obtain an initial subtask sequence, which includes: The target job task is parsed to obtain a structured natural language task to be processed by a feature alignment embedding method together with the multi-modal fusion features, and a time sequence position code is embedded in the processing result to obtain a cross-modal fusion feature vector containing time sequence information; Based on the cross-modal fusion feature vector and the structured natural language task, knowledge is matched from the job task knowledge base to obtain static knowledge matched with the target job task; The cross-modal fusion feature vector, the structured natural language task, and the static knowledge are input into the pre-trained large language model based on the Transformer architecture for dynamic template matching and chain reasoning, to decompose the target job task through a template self-adaptive mechanism in a first path to generate a preliminary subtask framework, and to decompose the target job task through a constraint-guided thinking chain in a second path to generate a subtask logic chain; The preliminary subtask framework and the subtask logic chain are subjected to conflict detection and feasibility checking, and after passing the conflict detection and the feasibility checking, the initial subtask sequence is generated according to the preliminary subtask framework and the subtask logic chain.

6. The robot control method based on a multi-modal large model according to claim 4, characterized in that, The receiving an external instruction to introduce the human-in-the-loop mechanism to re-optimize the secondary optimization subtask sequence, to obtain a final subtask sequence, includes: Receiving an external instruction and standardizing the external instruction to obtain a standardized instruction triple containing type, content, and timestamp; Based on a double-threshold checking rule constraint, the standardized instruction triple is input into a job rehearsal model constructed based on digital twinning technology for simulation to generate an effective instruction queue, and the type of the effective instruction queue is judged; When the effective instruction queue is a priority intervention instruction, the logical relationship between the subtasks contained in the secondary optimization subtask sequence is analyzed by a large language model, and the execution position of the target subtask corresponding to the effective instruction queue is adjusted to obtain a first subtask sequence; When the effective instruction queue is a task replacement instruction, a breakpoint label is generated based on the target subtask corresponding to the effective instruction queue, and a new task is generated by calling the visual language action model according to the breakpoint label to adjust the secondary optimization subtask sequence to obtain a second subtask sequence; When the effective instruction queue is a parameter modification instruction, a parameter boundary check is performed on the effective instruction queue, and after the check passes, the effective instruction queue is converted into a control signal to modify the secondary optimization subtask sequence to obtain a third subtask sequence.

7. The robot control method based on a multi-modal large model according to claim 1, characterized in that, The visual language action model comprises a visual network layer, a language network layer, and a flow matching action network layer; the multi-source modal data further comprises real-time sensor data; wherein The motion instruction corresponding to the subtask sequence is generated by inputting the subtask sequence and the multi-source modal data into a visual language action model or a reinforcement learning model based on the type of the subtask in the subtask sequence, so as to start the execution process of the target task by the robot, which comprises: judging the type of the subtask in the subtask sequence, if it is determined that a subtask in the subtask sequence is of a first type, inputting the multi-source modal data into the reinforcement learning model for processing to generate a sub-motion instruction corresponding to the first type of subtask; if it is determined that a subtask in the subtask sequence is of a second type, inputting the multi-modal fusion feature and the second type of subtask into the visual language action model, so that the visual network layer identifies the multi-modal fusion feature to obtain a target visual feature corresponding to the second type of subtask, the language network layer combines the task knowledge base to analyze the second type of subtask to obtain an ordered subtask sequence, and the flow matching action network layer processes the target visual feature, the ordered subtask sequence, and the real-time sensor data to generate a sub-motion instruction corresponding to the second type of subtask; The sub-motion instructions are integrated to obtain the motion instruction corresponding to the subtask sequence, so as to start the execution process of the target task by the robot.

8. A multi-modal large model based robot control system, characterized by, It comprises: a feature extraction module for real-time acquisition of multi-source modal data of a scene where a target task is located, and feature extraction of the multi-source modal data based on the type of the multi-source modal data by using a corresponding machine learning model to obtain multi-modal features; a feature fusion module for position coding and Transformer fusion processing of the multi-modal features to obtain multi-modal fusion features; a task decomposition module for inputting a constructed task knowledge base and the multi-modal fusion features into a large language model to obtain a decomposition result of the target task, and introducing a human-in-the-loop mechanism to optimize the decomposition result to obtain a subtask sequence; an instruction execution module for inputting the subtask sequence and the multi-source modal data into a visual language action model or a reinforcement learning model based on the type of the subtask in the subtask sequence for processing to generate a motion instruction corresponding to the subtask sequence, so as to start the execution process of the target task by the robot.

9. An electronic device, comprising: The computer readable storage medium comprises a stored computer program, wherein the device where the computer readable storage medium is located implements the robot control method based on the multi-modal large model according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium comprises a stored computer program, wherein the device where the computer readable storage medium is located implements the robot control method based on the multi-modal large model according to any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Man-machine interaction assembly method and system based on multi-modal large model and reinforcement learning

    CN118744426A

  • Leg-foot robot autonomous behavior control method and system based on multi-modal large model

    CN119077730A