Component industrial robot intelligent process generation method and system based on video language interactive learning

By using a component-based approach to video-based language interaction learning, natural language interaction and cross-platform adaptation of industrial robot systems have been achieved. This solves the problems of reliance on professional programming and low accuracy in complex task execution in existing technologies, and improves the ease of use and adaptability of the system.

CN120996753APending Publication Date: 2025-11-21GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511118931.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing industrial robot control systems rely on specialized programming, are difficult to migrate across tasks, lack cross-modal data utilization, making them difficult for non-professionals to deploy, resulting in low accuracy in complex task execution, and limiting the application of existing technologies in flexible manufacturing scenarios.

Method used

A component-based approach based on video language interaction learning is adopted. Through multimodal data input, video language contrast model training, reinforcement learning rewards and policy optimization, a component-based workflow is generated to achieve natural language interaction and cross-platform adaptation.

Benefits of technology

It lowers the barrier to entry for non-professionals, improves the execution accuracy and adaptability of complex tasks, supports rapid task adaptation and system expansion, and enhances the ease of use and accessibility in the manufacturing industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996753A_ABST
    Figure CN120996753A_ABST
Patent Text Reader

Abstract

The invention discloses a modularized industrial robot intelligent process generation method and system based on video language interactive learning, and aims at solving the technical bottleneck that a traditional industrial robot is complex in programming and difficult in task migration. The method comprises the following steps: analyzing space-time association between a natural language instruction and video demonstration by constructing a video language contrast learning model, and generating cross-task reusable semantic representation; combining a reinforcement learning reward mechanism to fuse sparse task signals and dense video evaluation, and driving adaptive optimization of a strategy model; and dynamically mapping the semantic instruction into an executable workflow of the industrial robot based on a componentized architecture, and verifying logic completeness through an automatic state machine. According to the method, the operation threshold of non-professionals is remarkably reduced, rapid task deployment of cross-entity platforms is realized, the system flexibility and execution precision in a complex manufacturing scene are improved, and a closed-loop self-optimization process control solution is provided for intelligent manufacturing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control technology for industrial robots, and in particular to a component-based intelligent process generation method and system for industrial robots based on video language interaction learning. Background Technology

[0002] As the industrial manufacturing sector accelerates its transformation towards intelligent manufacturing, industrial robots have demonstrated significant value in improving production automation. Traditional industrial robot control systems primarily rely on specialized engineers writing low-level control code to execute tasks, a method with clear application limitations. Faced with increasingly complex manufacturing scenarios and frequently changing production demands, existing technologies reveal three key bottlenecks: First, robot programming heavily relies on domain-specific expertise, making it difficult for non-professionals to participate in system deployment and maintenance, thus limiting the scope of technology application; second, existing solutions generally lack cross-task migration capabilities, requiring the collection of large amounts of demonstration data and customized training whenever a production line needs to process new product models or new processes, resulting in high time and resource costs; finally, most systems employ static image recognition or simple binary reward mechanisms, failing to fully exploit the dynamic operational information contained in video data, making it difficult for robots to meet the adaptability and execution accuracy requirements of complex processes in high-end manufacturing.

[0003] Current mainstream solutions fall into two main technical categories. The imitation learning-based behavior cloning method requires operators to manually control the robot to perform demonstrated actions, training the control model by recording these state actions. While this method performs reasonably well in simple tasks within fixed environments, it suffers from severe generalization flaws: the trained model is highly dependent on the specific robot's mechanical structural parameters, sensor configuration, and controller type; changing equipment or adjusting the task leads to model failure. The vision-language alignment-based pre-training method attempts to map natural language instructions to a visual feature space, using cross-modal similarity to guide robot operations. Although this approach lowers the programming barrier, it faces significant challenges in real-world industrial scenarios: overly coarse reward signal design leads to low training efficiency, data sharing and reuse across robot platforms is difficult, and it cannot handle complex assembly tasks involving multi-step collaboration. These technical shortcomings severely restrict the widespread application of industrial robot systems in flexible manufacturing scenarios.

[0004] The digital transformation of the manufacturing industry urgently requires breaking through the existing technological framework and developing a new generation of control systems with natural interaction capabilities, cross-task generalization capabilities, and continuous optimization capabilities. An ideal solution should achieve three core objectives: lowering the barrier to entry for non-professionals through multimodal interaction, establishing a transferable task representation mechanism to reduce repetitive development costs, and leveraging time-series data analysis to improve the execution accuracy of complex tasks. These needs have spurred component-based system innovation based on video language interaction learning, aiming to build an intelligent control paradigm with deep collaboration between semantic understanding and program execution. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a component-based intelligent process generation method and system for industrial robots based on video language interaction learning.

[0006] To achieve the above-mentioned objectives, the technical solution adopted by the present invention is as follows:

[0007] A component-based intelligent process generation method for industrial robots based on video language interaction learning includes the following steps:

[0008] S1: Receives multimodal input data, including cross-robot entity datasets, video frame sequences and corresponding natural language descriptions, successful and failed example video clips, and low-code platform task flow configuration information;

[0009] S2: Train a video language contrast model, including using a visual encoder to extract video frame features, using a text encoder to extract text embedding representations, and using a temporal aggregator to combine positional embeddings to process frame order information to generate similarity scores between video clips and text descriptions;

[0010] S3: Calculate the contrastive loss function and the sequence ranking loss function, and jointly optimize the model parameters;

[0011] S4: Construct a reinforcement learning reward function, including a dense reward signal generated based on similarity scores and a sparse reward for task completion;

[0012] S5: Employ an entropy regularization strategy to optimize the behavior strategy and maximize cumulative rewards;

[0013] S6: Parse the user's natural language task description, generate action instructions through the optimized strategy, and extract key elements to construct a task attribute set;

[0014] S7: Generate componentized workflows based on task attribute sets, including mapping task steps to executable component instructions and integrating them into a complete workflow;

[0015] S8: Verify the correctness of the workflow logic through an automatic state machine. If the verification fails, an error message is returned and the workflow is regenerated. If the verification passes, the workflow is deployed and executed.

[0016] S9: Collect execution data and iteratively optimize the model in practical applications.

[0017] Furthermore, the similarity score in step S2 is calculated using the following formula:

[0018] S θ (v 1:t ,c)=Transformer([ViT(v1),ViT(v2),…,ViT(v t ),TextEnc(c)])

[0019] Among them, S θ (v 1:t c) represents the similarity score between the video clip and the text description, ViT(v t ) represents the frame features output by the visual encoder, TextEnc(c) represents the text embedding output by the text encoder, Transformer represents the temporal aggregation function, and v 1:t represents the sequence of video clips from frame 1 to frame t, and c represents the natural language task description text.

[0020] Furthermore, the loss function in step S3 is calculated using the following formula:

[0021]

[0022] L total =L xent +αL rank

[0023] Among them, L xent Indicates comparative loss, l rank L represents the sequence ranking loss. total Let |v| represent the total loss function, N represent the number of training samples, and |v| represent the total loss function. i | represents the frame number of the i-th video, and α represents the loss balance coefficient. c represents the similarity between the first t frames of the i-th video and the instruction. i v represents the text description of the i-th training sample. i Let i represent the video data of the i-th training sample.

[0024] Furthermore, the reward function in step S4 is constructed using the following formula:

[0025]

[0026] r dense =r VLC (v t c) = S θ (v1:t ,c)-S θ (v 1:1 c)

[0027] r(s t ,a t ) = r total =r sparse +βr dense

[0028] Where, r dense Indicates a dense reward signal, r total Let r represent the total reward function. sparse S represents the reward for completing a sparse task, β represents the reward adjustment coefficient, and S represents the reward for completing a sparse task. θ (v 1:1 c) represents the similarity benchmark value of the first frame of the video.

[0029] Furthermore, the objective function for strategy optimization in step S5 is:

[0030]

[0031] Where J(π) represents the policy optimization objective function, γ represents the discount factor, and ω represents the exploration intensity coefficient. Represents policy entropy. Indicates the policy in state s t Information entropy.

[0032] Furthermore, the construction of the task attribute set in step S6 is achieved through the following formula:

[0033] Action(k i ,t i ) = π * (D)

[0034] Φ[Action(k i ,t i )]={k1,k2,…,k l}

[0035] T = F({k1,k2,…,k l})={t1,t2,…,t i}

[0036] Where Action(k) i ,t i ) represents the action command output by the strategy, Φ represents the key element parsing function, {k1,k2,…,k l} represents the set of key elements extracted, T represents the set of task attributes, and t i =[desc i ,paratransi [] indicates the subtask description and parameter passing rules.

[0037] Furthermore, the workflow generation in step S7 is achieved through the following formula:

[0038] M i (l i )=Γ(s n )

[0039] W = Ω(M1,…,M) i )

[0040] Among them, M i (l i ) represents the component execution instruction, Γ represents the component matching function, W represents the generated workflow, and Ω represents the workflow integration function.

[0041] Furthermore, the strategy optimization in step S5 includes:

[0042] Critic Network Update:

[0043]

[0044] Actor Network Update:

[0045]

[0046] Among them, Q φ (s,a) is the state value function. Here, φ is the objective value function, and φ is the Critic network parameter. is the target network parameter, β is the entropy regularization coefficient, and γ is the discount factor.

[0047] Furthermore, in step S6, the construction of the task attribute set adopts the chain thinking (CoT) and role setting Prompt method to guide the generation of the large language model step set S = Ψ(T).

[0048] Furthermore, step S8, verifying the correctness of the workflow logic, includes: detecting the compatibility of parameter passing between components, verifying the completeness of the task execution logic, and checking the satisfaction of security constraints.

[0049] This invention discloses a component-based intelligent process generation system for industrial robots based on video language interaction learning. This system can be used to implement the aforementioned component-based intelligent process generation method for industrial robots, specifically including:

[0050] Multimodal data input interface: used to receive datasets across robot entities, video frame sequences and their corresponding natural language task descriptions, example video clips of successful and failed task execution, and task flow parameter information configured by the low-code platform.

[0051] The video language contrastive learning engine includes a visual feature extraction unit, a text semantic encoding unit, and a spatiotemporal information fusion unit. It is used to analyze the semantic relationship between dynamic operation features and language instructions in video sequences and generate a matching degree evaluation between video clips and task descriptions.

[0052] Reinforcement learning reward generator: Based on the matching degree evaluation output of the video language model, a hybrid reward signal containing sparse rewards for task completion state and dense rewards for operation process is constructed to provide fine-grained feedback for policy optimization.

[0053] Policy optimization controller: It adopts an entropy regularization reinforcement learning algorithm to optimize the behavior policy model by maximizing the cumulative reward function, thus balancing task execution efficiency and exploration ability.

[0054] Semantic parsing and task construction module: Includes natural language instruction deconstruction unit and task element mapping unit, which parses the user input task description into a structured task attribute set and clarifies the functional description and parameter passing rules of each subtask.

[0055] Component-based workflow generator: Includes executable component matching unit and process integration unit, maps task attributes to execution instructions of specific industrial components, and assembles them into a complete decentralized workflow.

[0056] Automatic state machine verifier: It uses a formal verification mechanism to detect the compatibility of component parameter passing, the completeness of task logic, and the satisfaction of safety constraints in the workflow, and generates error feedback information for workflows that fail verification.

[0057] Execution feedback loop module: Connects to the industrial robot control platform, collects multimodal data in real time during task execution, drives the iterative optimization of the video language model and strategy controller, and forms a closed-loop learning system.

[0058] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described componentized industrial robot intelligent process generation method.

[0059] The present invention also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described modular industrial robot intelligent process generation method.

[0060] Compared with the prior art, the advantages of the present invention are as follows:

[0061] 1. Through the deep integration of natural language interaction and video demonstration, the system enables non-professionals to directly describe task requirements using everyday language, generating complex industrial processes without the need for specialized programming knowledge. The video language model automatically parses operational intentions and maps them into machine-executable instructions, completely eliminating the reliance on code writing in traditional industrial robots and significantly improving the system's ease of use and widespread adoption in the manufacturing sector.

[0062] 2. An innovative cross-entity component reuse mechanism combined with dynamic parameter mapping technology enables the system to quickly adapt to new tasks based on a small number of samples. The video language model analyzes the operating data of different robot platforms, autonomously extracts common operating patterns, and transforms them into standardized component parameters, solving the industry problem of existing systems needing to repeatedly collect data and retrain during task switching.

[0063] 3. The execution feedback module collects the robot's sensor data and operation trajectory in real time, and drives the model to update the reward function online via video stream. This closed-loop mechanism, which dynamically feeds back the physical execution results to the semantic understanding layer, enables the system to autonomously correct action deviations and optimize operation strategies, significantly improving the success rate and adaptability of tasks under complex working conditions.

[0064] 4. A reinforcement learning framework based on a hybrid reward mechanism organically combines sparse task completion signals with dense video semantic evaluation signals. By analyzing the implicit changes in operation quality in video frame sequences, the system generates more refined training guidance signals than traditional binary rewards, significantly improving data utilization efficiency and model convergence speed.

[0065] 5. The component-based workflow architecture supports plug-and-play feature expansion, allowing for flexible addition and removal of modules based on production line needs. Combined with an automated state machine verification mechanism, it ensures parameter compatibility and logical security between new components and existing systems, providing an evolvable and verifiable technical infrastructure for rapid production line reconfiguration. Attached Figure Description

[0066] Figure 1 This is a framework diagram of the modular industrial robot intelligent process generation method according to an embodiment of the present invention. Detailed Implementation

[0067] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and examples.

[0068] like Figure 1 As shown, this invention provides a component-based intelligent process generation method for industrial robots based on video language interaction learning, comprising the following steps:

[0069] (1) The low-code componentization platform has a rich built-in component library, which includes a variety of commonly used functional components M i

[0070] (2) Configuration information for each step of the task flow created by the low-code platform: L = {l1, l2, ..., l k}

[0071] (3) For example: robotic arm motion component M i (l i ): Controls the robotic arm to perform specific actions, such as grasping and placing. Input parameters include the target position coordinates P. target and Action Type A type The output is the state feedback S of the robotic arm. arm .

[0072] (4) Input cross-entity dataset D from different robot platforms robot ={D1,D2,…,D k}

[0073] (5) Input video-subtitle pairs:

[0074] a) Video frame sequence: The execution process of each task is recorded as a series of video frames v1, v2, ..., v T

[0075] b) Subtitle Description: Each video corresponds to a natural language instruction or description c, explaining the goal and requirements of the task.

[0076] c) Configuration information for each step of the task flow created by the low-code platform: L = {l1, l2, ..., l k}

[0077] (6) Input success and failure examples:

[0078] a) Success Video: Video clips showing the robot successfully completing the task.

[0079] b) Failure Videos: Showing video clips of the robot failing to complete the task correctly. Used to enhance the generalization ability of the model

[0080] (7) Using a pre-trained visual encoder (such as ViT), each frame of video v t Convert to feature vector f t ,in:

[0081] f t =ViT(v t )

[0082] (8) Using a pre-trained text encoder (such as CLIP), convert the caption c into an embedding representation e. c ,in:

[0083] e c =CLIP text (c)

[0084] (9) Define the similarity function Used to measure the degree of matching between video frames and subtitles.

[0085] (10) Define the contrastive loss function L xent :

[0086]

[0087] (11) Use position embedding in Transformer to preserve frame order information

[0088] (12) The temporary aggregator will combine image features f t and subtitle embedding e c Combined, a similarity score is generated for the entire video sequence.

[0089] S θ (v 1:t ,c)=Transformer([ViT(v1),ViT(v2),…,ViT(v t ),TextEnc(c)])

[0090] Here, TextEnc(c) represents text encoding of the text description c, and Transformer is a time aggregation function.

[0091] (13) Define the sequence ranking loss function L rank :

[0092]

[0093] (14) Define the total loss function:

[0094] L total =L xent +αL rank

[0095] Here, α is a hyperparameter that balances the two losses.

[0096] (15) Train the text encoder, view encoder, and temporary aggregator based on the loss function;

[0097] (16) By adjusting the hyperparameter α, the effects of contrast loss and sequence ranking loss are balanced to optimize the overall model performance;

[0098] (17) Define sparse task completion reward r sparse ,For example:

[0099]

[0100] (18) Define dense reward signal r dense :

[0101] r dense =r VLC (v t c) = S θ (v 1:t ,c)-S θ (v 1:1 c)

[0102] (19) Combine sparse task completion rewards r sparse and the learned dense reward signal r dense :

[0103] r(s t ,a t ) = r total =r sparse +βr dense

[0104] Where β is the adjustment coefficient.

[0105] (20) Train the reward function r(s) based on the loss function. t ,a t );

[0106] (21) Soft Actor-Critic (SAC) is used as the policy optimization algorithm to maximize r. VLC (v t c) Reward signal:

[0107] a) Maximize the cumulative entropy regularized reward, objective function:

[0108]

[0109] in Let ω be the policy entropy, and let ω control the exploration intensity.

[0110] b) Minimize the Bellman error, Critic update:

[0111]

[0112] Among them, Q φ(s,a) is the state value function. Here, φ is the objective value function, and φ is the Critic network parameter. β is the target network parameter (delayed update), β is the entropy regularization coefficient (balancing exploration and exploitation), and γ is the discount factor.

[0113] c) Minimize the KL divergence, Actor update:

[0114]

[0115] Both the policy network (Actor) and the value network (Critic) are 3-layer MLPs.

[0116] (22) Optimized behavioral strategies obtained through training π *

[0117] (23) Evaluate the model and workflow performance on the validation set and adjust hyperparameters to optimize the model and workflow effects.

[0118] (24) Deploy the trained model into practical applications.

[0119] (25) The user provides a natural language task description D = {d1, d2, ..., d...} m}, where each d m It represents a piece of information.

[0120] (26) Through the training model strategy π * Obtain the key element k i and task attribute set t i Action output

[0121] Action(k i ,t i ) = π * (D)

[0122] (27) Use the analytic function Φ on Action(k) i ,t i ) Analyze and extract key elements

[0123] Φ[Action(k i ,t i )]={k1,k2,…,k l}

[0124] Where each k i This indicates key elements extracted from the original description, such as specific operational requirements and target product specifications.

[0125] (28) Based on the extracted key element set {k1,k2,…,k lThe system further constructs a task attribute set T = {t1, t2, ..., t}. i The constructed expression is:

[0126] T = F({k1,k2,…,k l})={t1,t2,…,t i}

[0127] Among them, t i =[desc i ,paratrans i ], desc i It is a description of the function of the subtask, while paratrans i These are the parameters required by the subtask and the parameters it outputs to other subtasks.

[0128] Here, F is a mapping function responsible for mapping key elements to specific task attributes.

[0129] (29) Based on the task attribute set T, use Prompt methods such as Chain Thinking (CoT) and role setting to guide LLM in generating a specific set of steps S = Ψ(S) = {s1, s2, ..., s} for a specific industrial decentralized manufacturing scenario. n}

[0130] (30) For each step s n ∈S, the large model generates the corresponding component output M. i (l i )=Γ(s n ), where Γ represents the component matching function, which is responsible for converting steps into specific execution instructions or parameters.

[0131] (31) Output M i It not only serves the current step, but also provides input as a related component for other steps.

[0132] (32) Pass the output set of all steps as input to the process generation module to generate a complete, formalized decentralized workflow.

[0133] W = Ω(M1,…,M) i )

[0134] Ω represents the process of integrating the various step components into a complete workflow.

[0135] (33) The generated workflow W is verified using an automatic state machine verification system to correct any ambiguity in the instructions generated by the LLM and ensure logical correctness and security. If V(W) = true, the workflow has passed the verification.

[0136] (34) If the verification fails, record the error message W. msg The feedback is then sent to the large model to repeat the workflow generation process described above.

[0137] (35) After correction, the effective workflow W* is obtained, imported into the low-code industrial software platform, and tested on actual equipment.

[0138] (36) Continuously collect data in practical applications to further optimize the model and workflow.

[0139] In another embodiment of the present invention, a component-based intelligent process generation system for industrial robots based on video language interaction learning is provided. This system can be used to implement the above-described component-based intelligent process generation method for industrial robots, specifically including:

[0140] Multimodal data input interface: used to receive datasets across robot entities, video frame sequences and their corresponding natural language task descriptions, example video clips of successful and failed task execution, and task flow parameter information configured by the low-code platform.

[0141] The video language contrastive learning engine includes a visual feature extraction unit, a text semantic encoding unit, and a spatiotemporal information fusion unit. It is used to analyze the semantic relationship between dynamic operation features and language instructions in video sequences and generate a matching degree evaluation between video clips and task descriptions.

[0142] Reinforcement learning reward generator: Based on the matching degree evaluation output of the video language model, a hybrid reward signal containing sparse rewards for task completion state and dense rewards for operation process is constructed to provide fine-grained feedback for policy optimization.

[0143] Policy optimization controller: It adopts an entropy regularization reinforcement learning algorithm to optimize the behavior policy model by maximizing the cumulative reward function, thus balancing task execution efficiency and exploration ability.

[0144] Semantic parsing and task construction module: Includes natural language instruction deconstruction unit and task element mapping unit, which parses the user input task description into a structured task attribute set and clarifies the functional description and parameter passing rules of each subtask.

[0145] Component-based workflow generator: Includes executable component matching unit and process integration unit, maps task attributes to execution instructions of specific industrial components, and assembles them into a complete decentralized workflow.

[0146] Automatic state machine verifier: It uses a formal verification mechanism to detect the compatibility of component parameter passing, the completeness of task logic, and the satisfaction of safety constraints in the workflow, and generates error feedback information for workflows that fail verification.

[0147] Execution feedback loop module: Connects to the industrial robot control platform, collects multimodal data in real time during task execution, drives the iterative optimization of the video language model and strategy controller, and forms a closed-loop learning system.

[0148] In another embodiment of the present invention, a terminal device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used for the operation of a modular industrial robot intelligent process generation method.

[0149] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). This computer-readable storage medium is a memory device in a terminal device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and extended storage media supported by the terminal device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device.

[0150] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the modular industrial robot intelligent process generation method in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by a processor.

[0151] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0152] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0153] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0154] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0155] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the implementation methods of the present invention, and should be understood that the scope of protection of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of the present invention.

Claims

1. A component-based intelligent process generation method for industrial robots based on video language interaction learning, characterized in that, Includes the following steps: S1: Receives multimodal input data, including cross-robot entity datasets, video frame sequences and corresponding natural language descriptions, successful and failed example video clips, and low-code platform task flow configuration information; S2: Train a video language contrast model, including using a visual encoder to extract video frame features, using a text encoder to extract text embedding representations, and using a temporal aggregator to combine positional embeddings to process frame order information to generate similarity scores between video clips and text descriptions; S3: Calculate the contrastive loss function and the sequence ranking loss function, and jointly optimize the model parameters; S4: Construct a reinforcement learning reward function, including a dense reward signal generated based on similarity scores and a sparse reward for task completion; S5: Employ an entropy regularization strategy to optimize the behavior strategy and maximize cumulative rewards; S6: Parse the user's natural language task description, generate action instructions through the optimized strategy, and extract key elements to construct a task attribute set; S7: Generate componentized workflows based on task attribute sets, including mapping task steps to executable component instructions and integrating them into a complete workflow; S8: Verify the correctness of the workflow logic through an automatic state machine. If the verification fails, an error message is returned and the workflow is regenerated. If the verification passes, the workflow is deployed and executed. S9: Collect execution data and iteratively optimize the model in practical applications.

2. The method according to claim 1, characterized in that: The similarity score in step S2 is calculated using the following formula: S θ (v 1:t ,c)=Transformer([ViT(v1),ViT(v2),…,ViT(v t ),TextEnc(c)]) Among them, S θ (v 1:t c) represents the similarity score between the video clip and the text description, ViT(v t ) represents the frame features output by the visual encoder, TextEnc(c) represents the text embedding output by the text encoder, Transformer represents the temporal aggregation function, and v 1:t represents the sequence of video clips from frame 1 to frame t, and c represents the natural language task description text.

3. The method according to claim 1, characterized in that: The loss function in step S3 is calculated using the following formula: L total =L xent +αL rank Among them, L xent L represents the comparative loss. rank L represents the sequence ranking loss. total Let |v| represent the total loss function, N represent the number of training samples, and |v| represent the total loss function. i | represents the frame number of the i-th video, and α represents the loss balance coefficient. Let c represent the similarity between the first t frames of the i-th video and the instruction. i v represents the text description of the i-th training sample. i Let i represent the video data of the i-th training sample.

4. The method according to claim 1, characterized in that: The reward function in step S4 is constructed using the following formula: r dense =r VLC (v t ,c)=S θ (v 1:t ,c)-S θ (v 1:1 ,c) r(s t ,a t )=r total =r sparse +βr dense Where, r dense Indicates a dense reward signal, r total Let r represent the total reward function. sparse S represents the reward for completing a sparse task, β represents the reward adjustment coefficient, and S represents the reward for completing a sparse task. θ (v 1:1 c) represents the similarity benchmark value of the first frame of the video.

5. The method according to claim 1, characterized in that: The objective function for strategy optimization in step S5 is: Where J(π) represents the policy optimization objective function, γ represents the discount factor, and ω represents the exploration intensity coefficient. Represents policy entropy. Indicates the policy in state s t Information entropy.

6. The method according to claim 1, characterized in that: The task attribute set construction in step S6 is achieved through the following formula: Action(k i ,t i )=π * (D) Φ[Action(k i ,t i )]={k1,k2,…,k l } T=F({k1,k2,…,k l })={t1,t2,…,t i } Where Action(k) i ,t i ) represents the action command output by the strategy, Φ represents the key element parsing function, {k1,k2,…,k l } represents the set of key elements extracted, T represents the set of task attributes, and t i =[desc i ,paratrans i [] indicates the subtask description and parameter passing rules.

7. The method according to claim 1, characterized in that: The workflow generation in step S7 is achieved through the following formula: M i (l i )=Γ(s n ) W=Ω(M1,…,M i ) Among them, M i (l i ) represents the component execution instruction, Γ represents the component matching function, W represents the generated workflow, and Ω represents the workflow integration function.

8. The method according to claim 1, characterized in that: The strategy optimization in step S5 includes: Critic Network Update: Actor Network Update: Among them, Q φ (s,a) is the state value function. Here, φ is the objective value function, and φ is the Critic network parameter. is the target network parameter, β is the entropy regularization coefficient, and γ is the discount factor.

9. The method according to claim 1, characterized in that: In step S6, the construction of the task attribute set adopts the chain thinking (CoT) and role setting Prompt method to guide the generation of the large language model step set S = Ψ(T).

10. A component-based intelligent process generation system for industrial robots based on video language interaction learning, characterized in that, This system can be used to implement the modular industrial robot intelligent process generation method according to any one of claims 1 to 9, specifically including: Multimodal data input interface: used to receive datasets across robot entities, video frame sequences and their corresponding natural language task descriptions, example video clips of successful and failed task execution, and task flow parameter information configured by the low-code platform; Video language contrastive learning engine: It includes a visual feature extraction unit, a text semantic encoding unit, and a spatiotemporal information fusion unit. It is used to analyze the semantic relationship between dynamic operation features and language instructions in video sequences and generate a matching degree evaluation between video segments and task descriptions. Reinforcement learning reward generator: Based on the matching degree evaluation output by the video language model, a hybrid reward signal containing sparse rewards for task completion state and dense rewards for operation process is constructed to provide fine-grained feedback for policy optimization; Policy optimization controller: It adopts an entropy regularization reinforcement learning algorithm to optimize the behavior policy model by maximizing the cumulative reward function, thus balancing task execution efficiency and exploration ability; Semantic parsing and task construction module: includes natural language instruction deconstruction unit and task element mapping unit, which parses the user input task description into a structured task attribute set, and clarifies the functional description and parameter passing rules of each subtask; Component-based workflow generator: Includes executable component matching unit and process integration unit, maps task attributes to execution instructions of specific industrial components, and assembles them into a complete decentralized workflow; Automatic state machine verifier: It uses a formal verification mechanism to detect the compatibility of component parameter passing, the completeness of task logic, and the satisfaction of safety constraints in the workflow, and generates error feedback information for workflows that fail verification. Execution feedback loop module: Connects to the industrial robot control platform, collects multimodal data in real time during task execution, drives the iterative optimization of the video language model and strategy controller, and forms a closed-loop learning system.