Active man-machine cooperation method driven by process knowledge and interactive semantics

Through a method driven by process knowledge and interactive semantics, and utilizing data sets and deep learning networks, the dynamic and uncertainty problems in the human-machine collaborative assembly of personalized products are solved, operator behavior prediction and adaptive control of robot motion are achieved, and collaborative efficiency and accuracy are improved.

CN120653970AActive Publication Date: 2025-09-16DONGHUA UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510939553.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-09-16
Estimated Expiration
2045-07-08

AI Technical Summary

Technical Problem

In the human-machine collaborative assembly of personalized products, dynamic interactive information is highly dynamic and uncertain. Changes in operator behavior lead to difficulties in information fusion and decision support, and robots face challenges in spatial perception, task planning and reasoning.

Method used

By producing general and domain-specific datasets, using process knowledge and interactive semantics-driven methods, we gradually optimize the generated results, combine semantic similarity evaluation and deep learning networks to extract features, predict operator assembly behavior, and optimize robot motion decisions through reward functions.

Benefits of technology

It realizes diversified prediction of operator assembly behavior and adaptive control of robot motion in a dynamic environment, improves the accuracy and efficiency of human-machine collaboration, and adapts to the assembly needs of personalized products.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653970A_ABST
    Figure CN120653970A_ABST
Patent Text Reader

Abstract

The invention provides an active man-machine cooperation method driven by process knowledge and interactive semantics, and the method comprises the steps: carrying out the step-by-step optimization and updating of a general task description text and a dynamic assembly task interaction description text in the manufacturing field, and obtaining an initial generation result; performing evaluation according to a semantic similarity classification standard and a real label to obtain a final generation result, performing observation sampling on a posture sequence when an operator implements an assembly behavior according to the generation result to obtain observation postures and dynamic scene semantic information, and extracting features to obtain a diffusion matrix; further processing to obtain an inherent matrix and a weight matrix to form point matrix features, mapping the point matrix features to a Gaussian distribution coding set feature map, sampling potential coding factors, mapping the sampling potential coding factors to a posture sequence to obtain a current posture sequence, predicting to obtain an operator assembly behavior prediction result, and combining an assembly environment state and an assembly task state to obtain an operator assembly behavior prediction result. Intelligent agent actions and operator assembly behaviors are evaluated, current state and reward function feedback are provided for the intelligent agents, and the robot is controlled to move.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field, and in particular to an active human-machine collaboration method driven by process knowledge and interactive semantics. Background Art

[0002] In the context of Industry 5.0, with breakthroughs in artificial intelligence technologies such as pre-trained large models, human-machine collaborative assembly for personalized products has attracted significant attention. Active human-machine collaboration has made significant progress in perception and decision-making. As a key research direction in human-centered intelligent manufacturing, human-machine collaboration emphasizes active human-machine collaboration that enhances task execution capabilities through perception, reasoning, and environmental interaction. This initiative not only cultivates intelligence through computation but, more importantly, generates knowledge through dynamic interaction with the environment, thereby shaping active intelligence capabilities.

[0003] Currently, the human-machine collaborative assembly of personalized products faces the following challenges:

[0004] 1. The human-robot collaborative assembly of personalized products involves a large amount of dynamic interactive information, requiring the simultaneous consideration of both process knowledge and interactive semantics. This process is highly dynamic and uncertain, requiring robots to efficiently adapt to rapid changes in operator behavior, component conditions, and the working environment.

[0005] 2. The diversity and complexity of personalized products increase the factors that operators need to consider when choosing assembly paths, which may hinder task progress. This highlights the necessity of effective communication and information exchange between operators and robots.

[0006] 3. The assembly process of personalized products is subject to uncertainty in unstructured environments, and operator behavior will change accordingly. The human-machine collaboration process is dynamic. Compared with structured environments with fixed distributions and rules, the assembly space has unstructured characteristics. Information fusion and decision support are crucial. Robots must quickly process operator instructions and environmental feedback to make accurate decisions in dynamic environments. As a result, robots are facing new challenges in spatial perception, task planning, and reasoning. Summary of the Invention

[0007] The purpose of the technical solution of the present invention is to propose an active human-machine collaboration method driven by process knowledge and interactive semantics.

[0008] The technical solution of the present invention provides an active human-machine collaboration method driven by process knowledge and interactive semantics, comprising the following steps:

[0009] Create a general manufacturing dataset that can enhance the understanding of general manufacturing tasks, including: general manufacturing domain images and corresponding general manufacturing task description text;

[0010] Produce a domain-specific manufacturing dataset that can enhance understanding of dynamic assembly tasks, including: assembly scene images corresponding to dynamic assembly tasks and corresponding dynamic assembly task interaction description text;

[0011] Leveraging process knowledge and interaction semantics, the linear projection layer and low-rank adaptive layer gradually optimize and update the description text of general manufacturing tasks and the interactive description text of dynamic assembly tasks to obtain the initial generation results. GPT-4 is used to evaluate the semantic similarity score between the generated results and the true labels based on the semantic similarity classification standard. When the semantic similarity score falls below the preset threshold, it is determined that the initial generated result does not match the true label. The generated results are then gradually optimized and updated to obtain the final generated results that semantically match the true label.

[0012] The operator's posture sequence when performing assembly behavior according to the generated results is observed and sampled to obtain the observed posture and dynamic scene semantic information. The semantic smoothing enhanced spatiotemporal graph convolutional network with residual connection is used to extract features to obtain the diffusion matrix.

[0013] Obtain the intrinsic matrix from the sampling distribution of the observation posture, sample from the uniform distribution and apply the Gumbel distribution to normalize using the Softmax sampling method to obtain the weight matrix;

[0014] The intrinsic matrix is ​​multiplied by the weight matrix to form a point matrix. MLP learning is used to map the features in the point matrix to a set of Gaussian distribution coding set feature maps with diverse latent codes. Diverse latent coding factors are sampled from the Gaussian distribution coding set feature maps. The latent coding factors are mapped to the posture sequence to obtain the current posture sequence. This realizes the fusion of dynamic scene semantic information and latent coding factors into the observed posture, and performs diversified prediction of the operator's assembly behavior to obtain the operator's assembly behavior prediction results.

[0015] According to the assembly environment status and assembly task status, combined with the task planning operator assembly behavior prediction results, a reward function is used to establish a Markov decision process to evaluate the agent's actions and the operator's assembly behavior, and provide feedback on the current state and reward function to the agent, accumulate assembly operation knowledge, thereby realizing interactive knowledge continuous learning and optimized decision-making, and controlling the robot's motion.

[0016] Preferably, the semantic similarity score formula is:

[0017]

[0018] In the formula, n represents the classification level in the semantic similarity standard, and Numbern represents all generated results that meet the nth level.

[0019] Preferably, in the operator assembly behavior prediction result obtained by the operator assembly behavior diversification prediction, the future posture sequence in the operator assembly behavior prediction result is recursively smoothed and refined, the historical posture sequence remains unchanged in the initial state, and a supervision sequence is generated for feedback adjustment.

[0020] Preferably, the reward function:

[0021]

[0022] Where s or represents the distance between the operator and the robot, s op represents the distance between the operator and the component, s rp Represents the distance between the robot and the component, n t Indicates the total assembly task that needs to be completed, cst indicates whether the current assembly task is completed, n st Indicates the number of completed assembly tasks.

[0023] The technical solution of the present invention proposes an active human-machine collaboration method driven by process knowledge and interactive semantics. The initial generation result is obtained by gradually optimizing and updating the general task description text and the dynamic assembly task interaction description text in the manufacturing field. The final generation result is obtained by evaluation based on the semantic similarity classification standard and the real label. The posture sequence of the operator when performing the assembly behavior according to the generation result is observed and sampled to obtain the observed posture and dynamic scene semantic information, and the features are extracted to obtain the diffusion matrix. The intrinsic matrix and the weight matrix are further processed to form a point matrix feature, which is mapped to the Gaussian distribution coding set feature map. The sampled potential coding factors are mapped to the posture sequence to obtain the current posture sequence, and the operator's assembly behavior prediction result is obtained by prediction. Combined with the assembly environment state and the assembly task state, the intelligent agent action and the operator's assembly behavior are evaluated, and the current state and reward function feedback are provided to the intelligent agent to control the robot movement. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 A schematic diagram of the overall framework of an active human-machine collaboration method driven by process knowledge and interactive semantics provided by an embodiment of the present invention;

[0025] Figure 2 A schematic diagram of a low-rank adaptive fine-tuning process in an active human-machine collaboration method driven by process knowledge and interactive semantics provided in an embodiment of the present invention;

[0026] Figure 3 A schematic diagram of a process flow for predicting operator behavior diversity in a proactive human-machine collaboration method driven by process knowledge and interactive semantics provided by an embodiment of the present invention;

[0027] Figure 4A schematic diagram of the robot motion adaptive decision-making process in an active human-machine collaboration method driven by process knowledge and interactive semantics provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0028] Below in conjunction with specific embodiment, further set forth the present invention.Should be understood that these embodiments are only used to illustrate the present invention and are not used in limiting the scope of the present invention.In addition, should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms fall equally within the scope limited by the appended claims of the application.

[0029] like Figure 1 As shown, an embodiment of the present invention provides an active human-machine collaboration method driven by process knowledge and interactive semantics, including:

[0030] like Figure 2 As shown in the figure, a general manufacturing dataset is created to enhance the understanding of common manufacturing tasks. It includes images in the general manufacturing domain and corresponding textual descriptions of general manufacturing tasks. During fine-tuning, multimodal prompts for manufacturing tasks are designed by combining the embedded features of the manufacturing domain images and the corresponding language instructions for the general manufacturing task descriptions, providing a standardized input format. Specifically, it is represented as: [image] <general manufacturing image> [description] <image description> [prompt instruction] <instruction>. For example, an image description might be "An operator standing next to an uncovered reducer" and the corresponding image might be accompanied by a prompt such as "Please briefly describe this image." By combining visual features and language prompts, the dataset learns the mapping between general manufacturing images and descriptions, providing a foundation for adapting to different specific tasks.

[0031] We created a domain-specific manufacturing dataset that enhances understanding of dynamic assembly tasks. The dataset includes assembly scene images and text describing the corresponding dynamic assembly task interactions. By combining the embedded features of the assembly scene images and the language instructions corresponding to the dynamic assembly task interaction descriptions, we designed multimodal prompts for dynamic assembly tasks during fine-tuning. Specifically, this is represented as: [image] <assembly scene image> [description] <dynamic interaction description> [prompt instruction] <instruction>. For example, the text "The operator is currently holding both ends of the gear shaft with both hands and slightly bent over, standing in front of the reducer. Their position is close to the reducer's third slot. Therefore, it is predicted that the operator will move the gear shaft horizontally from top to bottom and place it into the reducer's third slot" describes the assembly scene in the corresponding image and predicts the dynamic process in the near future. Using prompts such as "Given the given assembly scene image, please describe the current scene as detailed as possible and predict the dynamic process in the near future," we guide the learning of the mapping relationship between the assembly scene image and the dynamic interaction description text through the combination of visual features and language prompts, enabling understanding of the current assembly scene and preliminary planning of future assembly.

[0032] By utilizing process knowledge and interaction semantics, the linear projection layer and low-rank adaptive layer are used to gradually optimize and update the general task description texts in the manufacturing field and the interactive description texts of dynamic assembly tasks. Following the human learning thinking, the system transitions from general tasks to specific tasks from shallow to deep, and obtains the initial generation results. GPT-4 is used according to the semantic similarity classification standard to evaluate the semantic similarity score between the generation results and the true labels. When the semantic similarity score is lower than the preset threshold and it is determined that the initial generation result does not match the true label, the generation result is gradually optimized and updated again to form a closed loop, and finally the final generation result that matches the semantics of the true label is obtained.

[0033] To facilitate GPT-4's more fine-grained assessment of text semantic similarity, a semantic similarity classification standard was designed to enable qualitative assessment of the semantic similarity between generated results and true labels. Furthermore, human experience was introduced to verify GPT-4's evaluation results, ensuring the accuracy of generated results and a precise understanding of human-machine collaboration knowledge, enabling the learning and optimization of human-machine collaborative interaction knowledge in specific domains.

[0034] The semantic similarity score represents a qualitative assessment of semantic similarity, as follows:

[0035]

[0036] In the formula, n represents the classification level in the semantic similarity standard, and Numbern represents all generated results that meet the nth level. This weighted average score can better reflect the overall quality of the generated results in terms of semantic similarity.

[0037] like Figure 3 As shown in the figure, the posture sequence when the operator performs the assembly behavior according to the generated results is observed and sampled to obtain the observed posture X∈RS×C×L and dynamic scene semantic information, and a semantic smoothing enhanced spatiotemporal graph convolutional network with residual connection is used to extract features to obtain the diffusion matrix Where S represents the number of skeleton joints and D_D is the dimension of the feature; the intrinsic matrix is ​​obtained from the sampling distribution of the observation pose X∈RS×C×L Sampling from uniform distribution and applying Gumbel distribution to normalize using Softmax sampling method, we get weight matrix W∈RD_W×M; transform the intrinsic matrix Multiply with the weight matrix W∈RD_W×M to form a point matrix Use MLP to learn the point matrix The features in the image are mapped to a set of Gaussian distribution coding set feature maps with diversified potential codes, and diversified potential coding factors are sampled from the Gaussian distribution coding set feature maps. The potential coding factors are mapped to the posture sequence to obtain the current posture sequence, that is, the dynamic scene semantic information and the potential coding factors are integrated into the observed posture, and the operator's assembly behavior is predicted in a diversified manner to obtain the operator's assembly behavior prediction results, ensuring that the current posture has a basis for diversified prediction and generating a variety of high-precision potential behavior possibilities for dynamic human motion.

[0038] Intrinsic Matrix Each row contains a D_I-dimensional intrinsic vector, the number of all intrinsic vectors is M, and the intrinsic matrix and diffusion matrix The fusion of the two forms the diffusion space, the intrinsic matrix According to the weight matrix W∈RD_W×M, a linear combination is performed in this space to generate multiple points, which are used to generate a diversified distribution. The weight matrix W∈RD_W×M has M weights of D_W dimension.

[0039] In the Gaussian distribution coding set, each row of the point matrix corresponds to a point of D_I dimension and D_W dimension from the diffusion space.

[0040] This step enhances the diversity of the inherent weights derived from the observed posture by combining the weights derived from the diffusion matrix and the Gumbel distribution, ensuring that the motion distribution is diffused as much as possible while maintaining the inherent posture.

[0041] The semantically smoothed and enhanced spatiotemporal graph convolutional prediction network consists of a multi-step structure, each of which includes an encoder-decoder module.

[0042] In the prediction of operator assembly behavior diversity, the future posture sequence SF = (SF1, SF2, ..., SFt) in the operator assembly behavior prediction results is recursively smoothed. The future posture sequence SF is gradually refined, while the historical posture sequence SH remains unchanged in the initial state. The supervisory sequence SHFt-1, SHFt-2, ..., SHF1 is generated to adjust the corresponding historical posture SH = (SH1, SH2, ..., SHt) to build a feedback mechanism. The specific definitions are as follows:

[0043]

[0044] Where i-1 represents the time of the last historical posture sequence.

[0045] like Figure 4As shown in the figure, according to the assembly environment state and assembly task state, combined with the task planning operator assembly behavior prediction results, the reward function is established as a Markov decision process to evaluate the agent action and the operator assembly behavior, and provide feedback on the current state and reward function to the agent, accumulate assembly operation knowledge, thereby realizing interactive knowledge continuous learning and optimized decision-making, controlling the robot movement, and ensuring understanding and generation capabilities.

[0046] The reward function is crucial for the agent to learn to make correct decisions because it provides feedback to the agent to help it make decisions. The reward function ensures the timeliness of the robot's decision-making and the safety of the collaboration. For decision execution, the reward function considers both minimizing the robot's decision time and maximizing human satisfaction. The reward function is as follows:

[0047]

[0048] Where s or represents the distance between the operator and the robot, s op Represents the distance between the operator and the component; s rp Indicates the distance between the robot and the component; n t Indicates the total assembly task that needs to be completed; cst indicates whether the current assembly task is completed, 1 for completed and -1 for uncompleted; n st The reward function can guide the robot to continuously learn decision-making behaviors at different spatial distances, ensuring that the final decision meets the requirements of shorter time and larger safety distance.

[0049] The embodiment of the present invention provides a proactive human-machine collaboration method driven by process knowledge and interactive semantics, which has the following beneficial effects:

[0050] 1. Aiming at the process dynamics and human-machine collaboration issues in human-machine collaborative assembly, an active human-machine collaboration method enhanced with personalized product process knowledge and interactive semantics is proposed from four aspects: spatial perception, process analysis, decision reasoning and task execution.

[0051] 2. Since static domain knowledge cannot enable large models to understand the dynamic interactions in human-machine collaborative assembly, a low-rank adaptive fine-tuning method enhanced with process knowledge and interaction semantics is proposed. By using data representing the dynamic interactions of assembly, only the linear projection layer and low-rank adaptive layer of the large model are updated to form a large model for a specific domain.

[0052] 3. Since the lexical similarity evaluation method cannot meet the open results generated by the large model, a semantic similarity-driven fine-tuning evaluation and optimization method is proposed. GPT-4 and artificial experience are introduced to evaluate the generated results. The inconsistent results will be regenerated to form a closed loop to ensure the large model's understanding of domain knowledge.

[0053] 4. To address the problem that highly random operator assembly behaviors interfere with the robot's embodied reasoning, we introduce human factors into the human-robot collaborative loop and propose a diversified prediction method for operator intentions to improve the ability of large models to predict operator behavior. Through the "human-in-the-loop" approach, we continuously improve the human-robot collaboration to ensure task execution.

[0054] 5. To address the shortcomings of robot motion control and considering the subjectivity of traditional methods such as expert experience, a deep reinforcement learning-driven robot motion adaptive decision-making method is proposed. It continuously learns and optimizes operator intention prediction and scene interaction knowledge, and continuously improves task reasoning and adaptability to dynamic scenes.

Claims

1. A proactive human-machine collaboration method driven by process knowledge and interactive semantics, characterized by: The following steps are involved: Create a general manufacturing dataset that can enhance the understanding of general manufacturing tasks, including: general manufacturing domain images and corresponding general manufacturing task description text; Produce a domain-specific manufacturing dataset that can enhance understanding of dynamic assembly tasks, including: assembly scene images corresponding to dynamic assembly tasks and corresponding dynamic assembly task interaction description text; Leveraging process knowledge and interaction semantics, the linear projection layer and low-rank adaptive layer gradually optimize and update the description text of general manufacturing tasks and the interactive description text of dynamic assembly tasks to obtain the initial generation results. GPT-4 is used to evaluate the semantic similarity score between the generated results and the true labels based on the semantic similarity classification standard. When the semantic similarity score falls below the preset threshold, it is determined that the initial generated result does not match the true label. The generated results are then gradually optimized and updated to obtain the final generated results that semantically match the true label. The operator's posture sequence when performing assembly behavior according to the generated results is observed and sampled to obtain the observed posture and dynamic scene semantic information. The semantic smoothing enhanced spatiotemporal graph convolutional network with residual connection is used to extract features to obtain the diffusion matrix. Obtain the intrinsic matrix from the sampling distribution of the observation posture, sample from the uniform distribution and apply the Gumbel distribution to normalize using the Softmax sampling method to obtain the weight matrix; The intrinsic matrix is ​​multiplied by the weight matrix to form a point matrix. MLP learning is used to map the features in the point matrix to a set of Gaussian distribution coding set feature maps with diverse latent codes. Diverse latent coding factors are sampled from the Gaussian distribution coding set feature maps. The latent coding factors are mapped to the posture sequence to obtain the current posture sequence. This realizes the fusion of dynamic scene semantic information and latent coding factors into the observed posture, and performs diversified prediction of the operator's assembly behavior to obtain the operator's assembly behavior prediction results. According to the assembly environment status and assembly task status, combined with the task planning operator assembly behavior prediction results, a reward function is used to establish a Markov decision process to evaluate the agent's actions and the operator's assembly behavior, and provide feedback on the current state and reward function to the agent, accumulate assembly operation knowledge, thereby realizing interactive knowledge continuous learning and optimized decision-making, and controlling the robot's motion.

2. The active human-machine collaboration method driven by process knowledge and interactive semantics according to claim 1, characterized in that: The semantic similarity score formula is: In the formula, n represents the classification level in the semantic similarity standard, and Numbern represents all generated results that meet the nth level.

3. The active human-machine collaboration method driven by process knowledge and interactive semantics according to claim 1, characterized in that: In the operator assembly behavior prediction result obtained by the operator assembly behavior diversification prediction, the future posture sequence in the operator assembly behavior prediction result is recursively smoothed and refined, the historical posture sequence remains unchanged in the initial state, and a supervision sequence is generated for feedback adjustment.

4. The active human-machine collaboration method driven by process knowledge and interactive semantics according to claim 1, characterized in that: The reward function: Where s or represents the distance between the operator and the robot, s op represents the distance between the operator and the component, s rp Represents the distance between the robot and the component, n t Indicates the total assembly task that needs to be completed, cst indicates whether the current assembly task is completed, n st Indicates the number of completed assembly tasks.

Citation Information

Patent Citations

  • Online reinforcement learning data increasing and expanding method based on diffusion model

    CN119476372A

  • A digital twin simulation method and system for industrial simulation platform

    CN119783407A

  • Surgical robot motion planning method based on diffusion model

    CN120038756A

  • Man-machine cooperation method and system based on robot process automation

    CN120258747A

  • LLM driven multimodal human-robot interaction planning

    EP4570449A1