An operational learning method based on multi-expert knowledge fusion and related equipment

CN122575223APending Publication Date: 2026-08-14LONGCHENG LABORATORY OF INTELLIGENT MANUFACTURING
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-06
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0004]本申请实施例提供一种基于多专家知识融合的操作学习方法及其相关设备,目的在于解决现有的操作设备在焊接操作学习的过程中过度依赖单一专家数据易受个体偏差影响,缺乏多维度数据综合学习焊接特征,导致模型训练偏差,无法确保焊接作业的焊接质量与效率的技术问题

Benefits of technology

[0017]本申请实施例的基于多专家知识融合的操作学习方法及其相关设备,通过采集多位专家操作数据,基于质量、效率、稳定性多维度综合评价筛选构建标准数据集,结合逆强化学习推导奖励函数并开展分层强化训练,有效规避单一专家个体偏差影响,保障训练数据质量与模型学习的效果。并且,无需人工设计奖励函数,精准匹配复杂工业操作工艺要求,通过分层强化训练模式提升了模型对操作任务的适配性与执行精度,可以增强机器人操作技能的稳定性与鲁棒性,大幅提升自主作业质量与效率,适用于钢结构、船舶制造等复杂工业场景的焊接作业或者其他作业。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575223A_ABST
    Figure CN122575223A_ABST
Patent Text Reader

Abstract

This application provides an operation learning method and related equipment based on multi-expert knowledge fusion. This application belongs to the field of intelligent welding manufacturing technology. The method includes: acquiring demonstration industrial operation data from at least two experts performing demonstration industrial operation tasks; evaluating the demonstration industrial operation data according to quality, efficiency, and stability dimensions to obtain a comprehensive evaluation score for each demonstration industrial operation data; based on the comprehensive evaluation scores of each demonstration industrial operation data, selecting demonstration industrial operation data with comprehensive evaluation scores higher than a set threshold to construct a standard operation dataset; determining a reward function based on the standard operation dataset; and performing hierarchical reinforcement training on the initial operation control model based on the standard operation dataset and the reward function to obtain a trained operation control model. This technical solution can enhance the stability of robot operation skills and significantly improve the quality and efficiency of autonomous operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of intelligent welding manufacturing technology, and in particular relates to an operation learning method based on multi-expert knowledge fusion and related equipment. Background Technology

[0002] Welding is a common process in various industries. Against the backdrop of industrial production's transformation towards intelligent manufacturing, automated welding operations using welding robots have become a core direction for improving production efficiency and quality stability, especially suitable for fields relying on complex processes such as steel structures and shipbuilding. Current robot operation learning methods mostly adopt single-expert welding demonstration and imitation or traditional welding programming modes. These methods directly train the robot to perform corresponding industrial tasks by collecting data such as the operation trajectory and process parameters of a single expert.

[0003] However, existing technologies also have significant drawbacks. For example, relying on welding demonstration data from a single expert is susceptible to individual operating habits, fluctuations in condition, or skill deficiencies, which can affect the welding quality of the robot in actual welding operations, making it difficult to achieve ideal levels. Furthermore, the quality, efficiency, and stability of the demonstration data are not comprehensively and quantitatively evaluated, and there is insufficient data support; invalid or low-quality data can easily cause model training bias. Moreover, the reward function design in existing solutions during the learning process relies on human experience, making it difficult to accurately match the process requirements of complex welding tasks. This results in insufficient model robustness, poor welding performance in real-world simulations, or an inability to meet the precision and stability requirements of welding tasks. Therefore, how to enable operating equipment to learn welding operation skills more accurately and improve the quality and efficiency of autonomous welding operations is a technical challenge that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] This application provides an operation learning method and related equipment based on multi-expert knowledge fusion. The purpose is to solve the technical problem that existing operation equipment relies too much on single expert data during the welding operation learning process, which is easily affected by individual biases. It lacks multi-dimensional data to comprehensively learn welding features, resulting in model training bias and failing to ensure the welding quality and efficiency of welding operations.

[0005] In a first aspect, embodiments of this application provide an operation learning method based on multi-expert knowledge fusion, the method comprising: Obtain demonstration industrial operation data performed by at least two experts; A comprehensive evaluation algorithm was used to evaluate the demonstration industrial operation data of each expert according to the dimensions of quality, efficiency and stability, and a comprehensive evaluation score was obtained for each demonstration industrial operation data. Based on the comprehensive evaluation scores of each demonstration industrial operation data, demonstration industrial operation data with comprehensive evaluation scores higher than a set threshold are selected to construct a standard operation dataset. The reward function is determined using an inverse reinforcement learning algorithm based on the aforementioned standard operating dataset. Based on the standard operation dataset and the reward function, the initial operation control model is subjected to hierarchical reinforcement training to obtain the trained operation control model.

[0006] In one feasible embodiment, after obtaining the trained operation control model, the method further includes: The operation control model is deployed to the operating equipment to control the operation of the operating equipment when it performs industrial operation tasks.

[0007] In one feasible embodiment, a comprehensive evaluation algorithm is used to evaluate the demonstration industrial operation data of each expert according to quality, efficiency, and stability dimensions, including: Identify welding quality parameters and assembly performance parameters in each of the aforementioned demonstration industrial operation data; wherein, the welding quality parameters include at least one of weld penetration, weld width, and number of pores; the assembly performance parameters include the gap and / or concentricity parameters of the assembled components; The quality dimension scores for each of the aforementioned demonstration industrial operation data are determined based on the welding quality parameters and assembly performance parameters. Identify at least one of the following parameters in the demonstration industrial operation data: total time to complete the task, path smoothness, and energy consumption parameters, and determine the efficiency dimension score for each demonstration industrial operation data. Identify stability parameters, compliance parameters, and path deviation parameters in the operational data of each of the demonstration industrial projects; wherein, the stability parameters include the stability of operating force, operating torque, and operating speed; wherein, the compliance parameters include whether the operating force and operating torque exceed the specified thresholds; Based on the stability parameter, the compliance parameter, and the path deviation parameter, a stability dimension score is determined for each of the demonstration industrial operation data. The evaluation score for each of the demonstration industrial operation data is determined based on the scores for the quality dimension, the efficiency dimension, and the stability dimension.

[0008] In one feasible embodiment, the evaluation score for each of the exemplary industrial operation data is determined based on the quality dimension score, the efficiency dimension score, and the stability dimension score, including: The evaluation score is calculated using the following formula: ; in, To evaluate the score, Score the quality dimensions. Score the efficiency dimension. Score the stability dimension. , , These are adjustable weighting coefficients.

[0009] In one feasible embodiment, before evaluating the demonstration industrial operation data from each expert according to the quality, efficiency, and stability dimensions using a comprehensive evaluation algorithm, the method further includes: The operational data of each of the demonstration industries are decomposed into tasks to obtain a task set consisting of serial tasks and parallel tasks. Serialize multiple tasks in the task set to obtain the execution order relationship; The optimal order relationship is determined based on the execution order relationship, and the order anomaly filtering is performed on the demonstration industrial operation data to obtain demonstration industrial operation data with the optimal order.

[0010] In one feasible embodiment, before determining the reward function using an inverse reinforcement learning algorithm based on the standard operating dataset, the method further includes: Demonstration industrial operation data is filtered based on a preset screening ratio, and perturbation data is added to obtain an enhanced dataset; wherein, the perturbation data includes one or more of the following: path perturbation, welding torch angle offset perturbation, force control perturbation, welding electrical parameter perturbation, and visual perturbation. The standard operation dataset and the enhanced dataset are merged into an updated standard operation dataset.

[0011] In one feasible embodiment, the standard operating dataset includes: the demonstration industrial operating data and the operating environment data corresponding to the demonstration industrial operating data; Based on the standard operation dataset and the reward function, the initial operation control model is subjected to hierarchical reinforcement training to obtain the trained operation control model, including: Environmental disturbance data is added to the operating environment data; wherein, the environmental disturbance data includes at least one of force noise, visual texture change noise, and illumination change noise; A simulation scenario is constructed based on the aforementioned operating environment data; The simulation robot with the initial model deployed is controlled to perform the demonstration industrial operation task in the simulation scenario to obtain operation simulation data; wherein, the high-level manager network in the initial model is used to determine the expected weld width sub-target and the coordinates of the next welding start point based on the demonstration industrial operation task, and the low-level actuator network in the initial model determines the joint motion parameters and welding torch angle control parameters based on the expected weld width sub-target and the coordinates of the next welding start point; Based on the operation simulation data, the demonstration industrial operation data, and the reward function, the high-level manager network and the low-level actuator network of the initial model are trained hierarchically to obtain the trained operation control model.

[0012] In one feasible embodiment, after deploying the operation control model to the operating device, the method further includes: The welding task is performed using the operating equipment, and the welding results of the operating equipment are quantitatively evaluated through a preset quality assessment subsystem to obtain a welding quality score. During the welding process, the system receives correction operation information provided by the user through a preset interface. The operation results are split into action sequences, and the splitting results are fused with the quality score and user correction operation information to construct a model optimization dataset; Based on the model, the dataset is optimized, and the correction reward function is learned using the maximum entropy inverse reinforcement learning algorithm. In the simulation scenario, the operation control model is subjected to hierarchical reinforcement learning iterative training based on the correction reward function.

[0013] Secondly, embodiments of this application provide an operation learning device based on multi-expert knowledge fusion, the device comprising: The operation data acquisition module is used to acquire demonstration industrial operation data from at least two experts performing demonstration industrial operation tasks. The scoring module is used to evaluate the demonstration industrial operation data of each expert according to the quality dimension, efficiency dimension and stability dimension using a comprehensive evaluation algorithm, and obtain the comprehensive evaluation score of each demonstration industrial operation data. The dataset construction module is used to select demonstration industrial operation data with comprehensive evaluation scores higher than a set threshold based on the comprehensive evaluation scores of each demonstration industrial operation data, and construct a standard operation dataset. The reward function determination module is used to determine the reward function based on the standard operating dataset using an inverse reinforcement learning algorithm. The model training module is used to perform hierarchical reinforcement training on the initial operation control model based on the standard operation dataset and the reward function to obtain the trained operation control model.

[0014] Thirdly, embodiments of this application provide an operation learning device based on multi-expert knowledge fusion, the device comprising: a processor, and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the operation learning method based on multi-expert knowledge fusion as described above.

[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the operation learning method based on multi-expert knowledge fusion as described above.

[0016] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the operation learning method based on multi-expert knowledge fusion as described above.

[0017] This application's embodiment of the operation learning method and related equipment based on multi-expert knowledge fusion collects operation data from multiple experts, selects and constructs a standard dataset based on a comprehensive evaluation of quality, efficiency, and stability, and combines inverse reinforcement learning to derive a reward function and conduct hierarchical reinforcement training. This effectively avoids the influence of individual expert bias and ensures the quality of training data and the effectiveness of model learning. Furthermore, it eliminates the need for manual design of reward functions, accurately matches the requirements of complex industrial operation processes, and improves the model's adaptability and execution accuracy to operation tasks through hierarchical reinforcement training. This enhances the stability and robustness of robot operation skills, significantly improves the quality and efficiency of autonomous operations, and is suitable for welding operations or other operations in complex industrial scenarios such as steel structure and shipbuilding. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating an operation learning method based on multi-expert knowledge fusion provided in an embodiment of this application; Figure 2 This is a schematic diagram of the process based on multi-expert knowledge fusion and adversarial training provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of a robot skill learning system based on multi-expert knowledge fusion and adversarial training provided in an embodiment of this application; Figure 4This is a schematic diagram of the structure of an operation learning device based on multi-expert knowledge fusion provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an operation learning device based on multi-expert knowledge fusion provided in an embodiment of this application. Detailed Implementation

[0020] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.

[0021] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.

[0022] To address the problems of existing technologies, this application provides an operation learning method and related equipment based on multi-expert knowledge fusion. The technical solution provided in this application targets complex industrial scenarios such as steel structure and shipbuilding. It collects operation data from multiple experts through a multimodal module, selects high-quality datasets, derives a reward function using maximum entropy inverse reinforcement learning, and enhances model robustness by combining hierarchical reinforcement learning and adversarial training. Adaptive impedance control enables safe transfer from simulation to reality, and the model can continuously evolve. This solution effectively avoids individual expert bias, eliminates the need for manually designed reward functions, significantly improves robot operation accuracy, stability, and environmental adaptability, enabling the robot to learn the operational skills truly required to complete tasks, and improving the efficiency of skill learning and production.

[0023] The operation learning method based on multi-expert knowledge fusion provided in the embodiments of this application will be introduced first.

[0024] Figure 1This is a flowchart illustrating an operation learning method based on multi-expert knowledge fusion provided in an embodiment of this application. Figure 1 As shown, the method may include the following steps: S101, Obtain demonstration industrial operation data by at least two experts performing demonstration industrial operation tasks; Among them, experts can be senior technicians with industry certification, extensive work experience, high product quality pass rate and peer review, such as senior welders with more than 10 years of experience in ship welding, who have obtained ship industry welding special certification and whose workpiece pass rate in the past three years is not less than 99%.

[0025] Demonstration industrial operation tasks can include ship welding tasks. For example, specific scenarios can be targeted at manual arc welding, submerged arc welding, CO2 gas shielded welding, tungsten inert gas welding, and metallic inert gas welding, such as welding of thin plates or non-ferrous metals.

[0026] Specifically, the demonstration industrial operation tasks can be specific tasks in industrial production that require professional skills to complete, such as the T-shaped corner joint double-sided continuous fillet weld of EH36 steel in shipbuilding, the connection between the hull longitudinal skeleton and the panel, and the welding task with a weld leg of 6mm, which adopts the gas metal arc welding (FCAW) process and must be strictly implemented in accordance with relevant industry standards.

[0027] Demonstration industrial operation data can be multimodal integrated data generated by experts during the execution of demonstration industrial operation tasks, including high-precision visual images, such as real-time weld seam imaging and molten pool status images, motion trajectories, such as the spatial displacement trajectory of the welding torch, six-dimensional force / torque sequences, such as welding torch operating force and torque change data, process parameters, such as wire feed speed, welding current, voltage, welding speed, torch angle, and compressed air pressure, and also includes equipment operating status data and environmental parameter data during task execution.

[0028] This solution can acquire operational data through synchronous acquisition, timestamp alignment, and filtering. For example, by using an integrated acquisition device equipped with a laser sensor, a six-dimensional force sensor, an infrared thermal imager, and a visual guidance device, it can simultaneously acquire data such as the hand movement trajectory, applied force / torque changes, welding current, voltage, wire feed speed, and welding speed of 4-6 experts operating the welding torch. After acquisition, the multi-source data is timestamped, abnormal fluctuation data is removed, and filtering and noise reduction are performed to ensure data consistency.

[0029] S102, The comprehensive evaluation algorithm is used to evaluate the demonstration industrial operation data of each expert according to the quality dimension, efficiency dimension and stability dimension, and the comprehensive evaluation score of each demonstration industrial operation data is obtained. The evaluation score for the quality dimension is determined based on at least one of the following: weld surface smoothness grade, weld penetration depth, weld width, and number of pores; the evaluation score for the efficiency dimension is determined based on at least one of the following: total time to complete the task, path smoothness, and energy consumption parameters; and the stability dimension is determined based on at least one of the following: operating force, operating torque, and operating speed. The comprehensive evaluation algorithm can be a pre-built multi-dimensional quantitative evaluation algorithm model. Specifically, it can be an algorithm that comprehensively evaluates the operational data of demonstration industries by combining industry standard requirements, production cycle requirements, and operational stability standards. Its core includes quantitative scoring formulas and weight coefficient optimization mechanisms.

[0030] The quality dimension can be the core dimension for measuring whether the results of the demonstration industrial operation meet the industry standards and process requirements. Taking welding tasks as an example, the main references are key indicators such as weld surface flatness grade, weld penetration depth, weld width and number of pores, which must meet the mandatory compliance standards of the classification society.

[0031] Efficiency can be a dimension for measuring whether the demonstration industrial operation process meets the production cycle requirements. It mainly includes the total time to complete the task, such as the welding time of a 4m long fillet weld should be ≤8min, path smoothness, such as no unnecessary backtracking in the operation path, energy consumption, such as the power consumption in the welding process, etc., which can directly affect the production cycle.

[0032] The stability dimension can be a dimension that measures the fluctuation of various parameters during the demonstration industrial operation, including the variance of operating force, speed, torque, and deviation from the ideal path. The smaller the fluctuation range, the higher the operational stability. It only affects the consistency of appearance and is a redundancy guarantee dimension.

[0033] This solution can automatically complete the evaluation process according to preset evaluation rules, a quantitative indicator system, and weighting coefficients. For example, the system first extracts key parameters such as weld penetration, welding time, and force control variance from the demonstration industrial operation data. Then, based on the weighting coefficients of each dimension, it calculates the comprehensive score using a quantitative formula. During the evaluation process, it is necessary to combine the process standards in the knowledge graph for rule verification and eliminate operation data that violates the specifications. Finally, the quantitative score obtained by comprehensively scoring the quality, efficiency, and stability dimensions can range from 0 to 100 points. This score is used to intuitively reflect the overall quality of each demonstration industrial operation data. The higher the score, the better the operation data, which can serve as the core basis for selecting high-quality data.

[0034] S103, Based on the comprehensive evaluation score of each demonstration industrial operation data, select the demonstration industrial operation data with a comprehensive evaluation score higher than a set threshold, and construct a standard operation dataset; This solution can automatically complete the screening operation based on the preset quality threshold. Combined with the outlier detection mechanism, it removes low-quality data with a comprehensive evaluation score below the threshold, abnormal data in the operation sequence, and data that violates the process rules, and selects high-quality demonstration industrial operation data that meet the requirements.

[0035] The standard operation dataset can be a dataset composed of high-quality demonstration industrial operation data. It should cover effective operation data of both serial and parallel tasks. Demonstration industrial operation data with comprehensive evaluation scores higher than a set threshold are selected. For example, the top 20% of high-quality welding operation data can be selected. For example, it can include standard operation data of serial tasks such as clamping and positioning, root tack fixing, reverse carbon gouging and root cleaning, and reverse formal welding, as well as operation data of parallel operation of front welding robot start-up and reverse carbon gouging dust removal robot start-up, to provide reliable samples for subsequent model training.

[0036] S104, Determine the reward function using an inverse reinforcement learning algorithm based on the standard operating dataset; Inverse Reinforcement Learning (IRL) algorithms, preferably Maximum Entropy Inverse Reinforcement Learning (Max Ent IRL) algorithms, can deduce the implicit optimal reward rules from high-quality demonstration data without the need for manual design of reward functions, making them suitable for complex industrial operation scenarios.

[0037] Reward functions can be used to evaluate agents, such as the benefits of a welding robot performing a certain action in a specific state. They can quantify the contribution of operational behavior to welding quality, efficiency, and stability. For example, they can reward actions that achieve the required weld penetration, welding time that meets the cycle time requirements, and smooth operation.

[0038] This solution utilizes the maximum entropy inverse reinforcement learning algorithm to perform deep analysis and learning on multimodal data in the standard operation dataset, such as operation trajectories, process parameters, and result quality data. It aims to uncover the mapping relationship between high-quality operational behaviors and reward values, thereby generating a reward function R(s, a). This function must implicitly incorporate industry standard requirements and the optimal operational strategies of the expert group.

[0039] S105, Based on the standard operation dataset and the reward function, perform hierarchical reinforcement training on the initial operation control model to obtain the trained operation control model.

[0040] The initial operation control model can be a hierarchical policy network model built on a multi-layer perceptron (MLP) or a Transformer, which includes two sub-networks: a high-level manager and a low-level worker. It has not been trained or has only undergone preliminary training and does not yet have the ability to accurately execute industrial operation tasks.

[0041] Layered reinforcement training can be a training method in a high-fidelity simulation environment that divides model training into two levels: high-level decision-making training and low-level action execution training. The high-level manager sets sub-goals on a second-level time scale, such as "move from point A to point B while maintaining a weld depth of 2mm", while the low-level worker executes specific actions on a millisecond-level time scale, such as adjusting joint angular velocity and welding torch pressure. Both use the reward function as the optimization objective, and adversarial training is introduced to improve the robustness of the model.

[0042] This approach can be implemented in a high-fidelity simulation environment. First, a virtual welding environment, including hull structure, fixtures, and lighting, is constructed based on the operational environment data from the standard operation dataset. Then, the initial model is deployed to the simulation robot. Using the standard operation dataset as training samples and the reward function R(s, a) as the optimization objective, iterative training is performed using the Hierarchical PPO algorithm (Hierarchical Proximal Policy Optimization Algorithm). During training, adversarial perturbations such as path disturbances, force noise, and lighting changes are introduced to ensure the model can still perform high-quality operations under these perturbations. Dropout (dropout rate = 0.2-0.3) and L2 regularization (λ = 1 × 10⁻⁶) are also employed. -4 Techniques such as […] are used to prevent model overfitting.

[0043] The trained operation control model can be a model with high operational accuracy and stability after multiple rounds of layered reinforcement training and adversarial training. It can autonomously formulate sub-goals according to actual working conditions and accurately execute operation actions. The output control commands, such as welding torch angle of 45° and welding speed of 4.5mm / s, can meet the quality and efficiency requirements of industrial operation tasks.

[0044] The technical solution provided in this embodiment collects multimodal demonstration industrial operation data from multiple experts, further filters it to obtain the highest comprehensive quality evaluation scores, and constructs a standard dataset for learning welding operations during welding task execution. Furthermore, the calculation of the comprehensive quality evaluation score considers the specific needs of the ship welding scenario, such as welding standards for different types of welding and the specific welding requirements of ship structural components. For example, the welding process must ensure flatness to meet subsequent installation requirements, or the welding process must ensure the stability of the operation trajectory to meet the limited welding space requirements, etc. After acquiring demonstration industrial operation data from multiple experts, this solution employs quality, efficiency, and stability dimensions... This approach utilizes reference data across various dimensions to achieve a multi-dimensional, objective, and holistic evaluation of demonstration industrial operation data, ultimately yielding a comprehensive evaluation score. Particularly noteworthy are the consideration of specific reference data across the quality, efficiency, and stability dimensions, including weld surface smoothness level, weld penetration, weld width, number of pores, total task completion time, path smoothness, energy consumption parameters, operating force, operating torque, and operating speed. This provides objective and quantitative evaluation criteria for the assessment of multiple expert demonstration industrial operation data. It offers more accurate data for subsequent reward function determination and operation control model training, leading to more accurate training results and ensuring the stability of welding robots in completing welding tasks. Furthermore, this solution avoids the distortion caused by individual expert operating habits, state fluctuations, or skill deficiencies, ensuring data diversity and quality. This solution automatically determines the reward function based on the maximum entropy inverse reinforcement learning algorithm, which solves the technical bottleneck that the reward function is difficult to define manually in complex industrial operations. By combining hierarchical reinforcement training with adversarial training, the solution improves the model's decision-making ability, execution accuracy and robustness, enabling the trained model to master the optimal operating skills of a group of experts. It is suitable for complex industrial scenarios such as shipbuilding, and improves production quality and efficiency.

[0045] In one feasible embodiment, after obtaining the trained operation control model, the method further includes: The operation control model is deployed to the operating equipment to control the operation of the operating equipment when it performs industrial operation tasks.

[0046] An operation control model is a model that has undergone tiered reinforcement training and possesses industrial operation control capabilities, enabling it to output control commands for operating equipment.

[0047] This solution can be deployed through software installation and parameter configuration. For example, the trained model file can be imported into the control system of the operating device, and relevant parameters can be adjusted.

[0048] Operating equipment can be equipment used to perform industrial operation tasks, such as industrial robots like welding robots and assembly robots.

[0049] Industrial operation tasks can be specific operational tasks in industrial production, such as ship welding tasks and mechanical assembly tasks.

[0050] Operation control can be achieved by the control system of the operating equipment driving the actuators of the equipment to complete corresponding actions according to the control instructions output by the operation control model. For example, a welding robot can automatically complete welding operations according to the control instructions such as the welding torch angle and welding speed output by the model.

[0051] This technical solution deploys the trained operation control model to the operating equipment, realizing the transformation of the model from theoretical training to practical application. This enables the operating equipment to autonomously execute industrial operation tasks. By precisely controlling the operating equipment through the model, the efficiency and quality stability of industrial operation tasks are improved, and the uncertainty caused by manual operation is reduced. It can be applied to large-scale industrial production scenarios.

[0052] In one feasible embodiment, a comprehensive evaluation algorithm is used to evaluate the demonstration industrial operation data of each expert according to quality, efficiency, and stability dimensions, including: Identify welding quality parameters and assembly performance parameters in each of the aforementioned demonstration industrial operation data; wherein, the welding quality parameters include at least one of weld penetration, weld width, and number of pores; the assembly performance parameters include the gap and / or concentricity parameters of the assembled components; The quality dimension scores for each of the aforementioned demonstration industrial operation data are determined based on the welding quality parameters and assembly performance parameters. Identify at least one of the following parameters in the demonstration industrial operation data: total time to complete the task, path smoothness, and energy consumption parameters, and determine the efficiency dimension score for each demonstration industrial operation data. Identify stability parameters, compliance parameters, and path deviation parameters in the operational data of each of the demonstration industrial projects; wherein, the stability parameters include the stability of operating force, operating torque, and operating speed; wherein, the compliance parameters include whether the operating force and operating torque exceed the specified thresholds; Based on the stability parameter, the compliance parameter, and the path deviation parameter, a stability dimension score is determined for each of the demonstration industrial operation data. The evaluation score for each of the demonstration industrial operation data is determined based on the scores for the quality dimension, the efficiency dimension, and the stability dimension.

[0053] Among them, welding quality parameters are parameters that can be used to measure the quality of welding operation results, including at least one of weld penetration depth, weld width, and number of pores. For example, the weld penetration depth of a certain welding operation is 3mm, the weld width is 8mm, and the number of pores is 2.

[0054] Assembly performance parameters are parameters that can be used to measure the quality of assembly operation results, including the gap and / or concentricity parameters of the assembled parts, such as a gap of 0.2 mm and a concentricity of 0.1 mm after assembly.

[0055] This solution can be achieved by analyzing sensor data in the demonstration industrial operation data, such as identifying parameters like weld penetration and weld width through weld image data collected by a vision sensor.

[0056] Quality dimension scoring can be based on quantitative scores given by welding quality parameters and assembly performance parameters. For example, 30 points are set for meeting the weld penetration standard, 30 points for meeting the weld width standard, and 40 points for meeting the porosity standard. The quality dimension score is obtained by comprehensive calculation.

[0057] The total time to complete a task can be the time it takes for an expert to complete a demonstration industrial operation task, such as taking 7 minutes to complete a 4-meter-long fillet weld.

[0058] Path smoothness can be a parameter that measures the smoothness of an expert's operational path, for example, by calculating the curvature changes of the operational path.

[0059] Energy consumption parameters can be the amount of energy consumed during the execution of a demonstration industrial operation task, such as the electrical energy consumed during welding.

[0060] Efficiency scoring can be based on a quantitative score given by at least one of the following parameters: total time to complete the task, path smoothness, and energy consumption. For example, a total time to complete the task within 8 minutes is worth 40 points, path smoothness is worth 30 points, and energy consumption is worth 30 points. The efficiency score is calculated by combining these parameters.

[0061] Stability parameters can be parameters that measure the stability of various operating parameters during operation, including the stability of operating force, operating torque, and operating speed. For example, the variance of operating force is 0.5N, and the fluctuation range of operating speed is ±0.2mm / s.

[0062] Compliance parameters can be parameters that measure whether an operation complies with relevant specifications, including whether the operating force and torque exceed the specified thresholds. For example, if the specification stipulates that the threshold for welding operating force is 10N, if an expert operates with a force of 8N, it complies with the compliance requirements. However, if an expert operates with a force of 13N, it does not comply with the compliance requirements.

[0063] The path deviation parameter can be a parameter that measures the deviation between the actual operating path and the ideal path. For example, the maximum deviation between the actual welding path and the ideal path is 0.3mm.

[0064] The stability dimension score can be a quantitative score given based on stability parameters, compliance parameters, and path deviation parameters. For example, 30 points are set for achieving operational stability, 30 points for achieving compliance, and 40 points for achieving path deviation. The stability dimension score is obtained by comprehensive calculation.

[0065] The final evaluation score can be a combination of scores from the quality dimension, efficiency dimension, and stability dimension. For example, if the quality dimension score is 85, the efficiency dimension score is 80, and the stability dimension score is 90, the evaluation score can be obtained by weighted calculation, such as a final evaluation score of 85.

[0066] This technical solution, through specific evaluation indicators and scoring methods in the dimensions of quality, efficiency, and stability, makes the execution of the comprehensive evaluation algorithm more standardized and operable. By refining various evaluation parameters, it can more comprehensively and accurately assess the quality of demonstration industrial operation data, avoiding the subjectivity of the evaluation. Furthermore, the specific settings of each parameter are all in line with the actual industrial scenario. For example, the selection of welding quality parameters and assembly performance parameters meets the core requirements of industrial production for operation results, ensuring that the selected standard operation dataset has high practicality.

[0067] In one feasible embodiment, the evaluation score for each of the exemplary industrial operation data is determined based on the quality dimension score, the efficiency dimension score, and the stability dimension score, including: The evaluation score is calculated using the following formula: ; in, To evaluate the score, Score the quality dimensions. Score the efficiency dimension. Score the stability dimension. , , These are adjustable weighting coefficients.

[0068] Evaluation score It can be a comprehensive quantitative score calculated by a formula, used to measure the overall quality of the demonstration industrial operation data.

[0069] The quality dimension score can be a quantitative evaluation result of the quality dimension of the demonstration industrial operation data, and the value range can be 0-100 points.

[0070] Efficiency scoring can be a quantitative evaluation of the efficiency of operational data in a demonstration industry, with a value range of 0-100.

[0071] The stability dimension score can be a quantitative evaluation result of the stability dimension of the demonstration industrial operation data, and the value range can be 0-100 points.

[0072] Adjustable weighting coefficients can be set according to the needs of industrial scenarios and process requirements to meet [the requirements of industrial scenarios and processes]. For example, in ship welding scenarios, based on the principles of rigid specifications, shipbuilding cycle time, and stability redundancy, the value range of α can be 0.55-0.65, the value range of β can be 0.20-0.30, and the value range of γ can be 0.10-0.20. The initial default values ​​can be set to α=0.60, β=0.25, and γ=0.15. Later, the values ​​can be iteratively optimized every 100 new data points using a logistic regression algorithm with a learning rate of 0.01 and an L2 regularization of 1×10⁻⁶. -4 .

[0073] This solution can be completed automatically according to the above formula. For example, the scores for the quality dimension, efficiency dimension, and stability dimension, as well as the preset weight coefficients, are substituted into the formula to calculate the evaluation score.

[0074] This technical solution calculates evaluation scores using a weighted formula, making the evaluation process more objective and avoiding the subjectivity of human evaluation. Furthermore, the adjustable weighting coefficients give the method strong flexibility and adaptability, allowing for adjustments to the importance of each dimension according to the needs of different industrial scenarios. For example, in shipbuilding, depending on the type of shipbuilding task, such as high-quality large ships, the quality dimension has the highest weighting coefficient, followed by efficiency, and then stability. Verification through real-world ship cases shows that this formula can effectively filter out high-quality demonstration data, significantly reducing the rework rate in subsequent operations and improving operational quality.

[0075] In one feasible embodiment, before evaluating the demonstration industrial operation data from each expert according to the quality, efficiency, and stability dimensions using a comprehensive evaluation algorithm, the method further includes: The operational data of each of the demonstration industries are decomposed into tasks to obtain a task set consisting of serial tasks and parallel tasks. Serialize multiple tasks in the task set to obtain the execution order relationship; The optimal order relationship is determined based on the execution order relationship, and the order anomaly filtering is performed on the demonstration industrial operation data to obtain demonstration industrial operation data with the optimal order.

[0076] Task breakdown can be automatically completed by the task analysis module according to the process requirements of industrial operation tasks. For example, the task of welding double-sided continuous fillet welds of ship T-joints can be broken down into tasks such as clamping and positioning, root tack welding, reverse carbon gouging and root cleaning, reverse formal welding, interpass temperature monitoring, and flux drying and replenishment.

[0077] A serial task can be a task that must be performed in a strict sequence, such as clamping and positioning, root tack fixing, reverse carbon gouging and root cleaning, and reverse formal welding, first front welding, interpass temperature dropping to ≤150℃, and second front welding.

[0078] Parallel tasks can be tasks that can be started and executed independently within the same time period, such as the start of the front welding robot, the start of the back carbon planing dust removal robot, the synchronous reporting of work by the Manufacturing Execution System (MES), the infrared monitoring of layer temperature, the operation of the clamping robot at adjacent ribs, and the automatic replenishment of flux drying box.

[0079] A task set can be a collection of both serial and parallel tasks, encompassing all subtasks required to complete an industrial operation task.

[0080] Serialization can be performed based on process logic and operational specifications. For example, it involves analyzing the decomposed tasks to determine which tasks must be executed sequentially and which tasks can be executed in parallel, clarifying the temporal sequence relationship of serial tasks and the triggering and termination conditions of parallel tasks.

[0081] The execution order relationship can be the logical relationship between tasks. For example, after the first front welding is completed, the interpass temperature must be lowered to ≤150℃ before the second front welding can be performed.

[0082] The optimal sequence relationship can be the best task execution sequence that can guarantee operational quality and efficiency. For example, the execution process determined according to the ship welding process specification is clamping and positioning, root tack fixing, reverse carbon gouging and root cleaning, and reverse formal welding, or the sequence of carbon gouging and root cleaning, magnetic particle inspection, defect repair and first pass welding on the reverse side.

[0083] Sequence anomaly filtering can be done automatically based on the optimal sequence relationship. For example, it can remove exemplary industrial operation data where the task execution order does not conform to the optimal sequence relationship, such as operation data where reverse welding is performed without carbon gouging and root cleaning.

[0084] Demonstration industrial operation data with optimal order can be demonstration industrial operation data in which the task execution order conforms to the optimal order relationship, such as welding operation data executed in the order of clamping and positioning, root spot fixing, reverse carbon gouging and root cleaning, and reverse formal welding.

[0085] This technical solution clarifies the execution relationships between tasks by decomposing and serializing the demonstration industrial operation data, providing a foundation for subsequent data evaluation and screening. By determining the optimal order relationship and filtering for order anomalies, low-quality data caused by incorrect task execution order can be eliminated, further improving the quality of the demonstration industrial operation data. At the same time, the task decomposition into serial and parallel tasks conforms to the actual process of industrial operation. For example, there are a large number of serial and parallel tasks in the ship welding process. This processing method can adapt to the actual operation scenario and facilitate accurate control of the operating equipment.

[0086] In one feasible embodiment, before determining the reward function using an inverse reinforcement learning algorithm based on the standard operating dataset, the method further includes: Demonstration industrial operation data is filtered based on a preset screening ratio, and perturbation data is added to obtain an enhanced dataset; wherein, the perturbation data includes one or more of the following: path perturbation, welding torch angle offset perturbation, force control perturbation, welding electrical parameter perturbation, and visual perturbation. The standard operation dataset and the enhanced dataset are merged into an updated standard operation dataset.

[0087] The preset screening ratio can be a pre-set ratio for screening high-quality demonstration industrial operation data. For example, the screening ratio can be set to the top 10%, and the demonstration industrial operation data ranked in the top 10% of the comprehensive evaluation scores can be selected.

[0088] This solution can automatically complete the process based on preset filtering ratios, such as filtering high-quality data that meets the ratio requirements from standard operating datasets.

[0089] Perturbation data can be noise added to enhance the robustness of the dataset, including one or more of the following: path perturbation, welding torch angle offset perturbation, force control perturbation, welding electrical parameter perturbation, and visual perturbation.

[0090] Path disturbances can be small disturbances added to the operation path, such as adding Gaussian noise (Δx,Δy,Δz, N(0,5cm)) to the simulation trajectory to simulate hull structure errors or fixture deviations.

[0091] The welding torch angular offset disturbance can be a small disturbance added to the welding torch attitude angle, such as Δ N(0,10°) represents the deviation of the welding torch posture caused by structural interference or human error.

[0092] Force-controlled disturbances can be small disturbances added to force / torque signals, such as adding a small pulse disturbance of ±5N lasting 0.1s to the force / torque signal to simulate unevenness on the hull surface or abrupt changes in welding wire contact.

[0093] Welding electrical parameter disturbances can be small disturbances added to electrical parameters such as welding current and voltage, such as current disturbance ±10A, voltage disturbance ±1V, simulating power supply fluctuations or arc length changes.

[0094] Visual disturbances can be interferences added to visual images, such as adding Gaussian blur, lighting changes, metallic reflection noise to simulate images, or simulating interferences such as strong light, reflections, and smoke inside a ship's cabin.

[0095] Augmented datasets, which can consist of selected high-quality demonstrative industrial operation data and added perturbation data, can be used to improve the robustness of the model.

[0096] This solution can merge the standard operating dataset and the augmented dataset to obtain an updated standard operating dataset, which contains both high-quality demonstration data and perturbation data, providing richer samples for model training.

[0097] This technical solution ensures the basic quality of the augmented dataset by screening high-quality demonstration industrial operation data through a preset screening ratio. At the same time, by adding various perturbation data, it can simulate various complex interference scenarios in industrial production, allowing the model to be exposed to more diverse situations during training. This enables the model to learn the truly important operational features and avoids the model overfitting to the perfect demonstration data. By merging the standard operation dataset and the augmented dataset, the content and diversity of the dataset are enriched, enabling the trained model to maintain good task performance in real complex environments.

[0098] In one feasible embodiment, the standard operating dataset includes: The demonstration industrial operation data and the corresponding operating environment data of the demonstration industrial operation data; Based on the standard operation dataset and the reward function, the initial operation control model is subjected to hierarchical reinforcement training to obtain the trained operation control model, including: Environmental disturbance data is added to the operating environment data; wherein, the environmental disturbance data includes at least one of force noise, visual texture change noise, and illumination change noise; A simulation scenario is constructed based on the operational environment data; a simulation robot with the initial model deployed is controlled to perform the demonstration industrial operation task in the simulation scenario to obtain operation simulation data; wherein, the high-level manager network in the initial model is used to determine the expected weld width sub-target and the coordinates of the next welding start point according to the demonstration industrial operation task, and the low-level actuator network in the initial model determines the joint motion parameters and welding torch angle control parameters based on the expected weld width sub-target and the coordinates of the next welding start point; Based on the operation simulation data, the demonstration industrial operation data, and the reward function, the high-level manager network and the low-level actuator network of the initial model are trained hierarchically to obtain the trained operation control model.

[0099] Among them, the operating environment data can be the execution environment-related data corresponding to the demonstration industrial operating data, such as the hull structure data, lighting condition data, and fixture position data in the welding task.

[0100] Environmental disturbance data can be interference data added to simulate complex real-world environments, including at least one of force noise, visual texture variation noise, and illumination variation noise.

[0101] Force noise can be interference from forces in the simulated environment, such as ±5N Z-axis force disturbances simulating undulations on the surface of a ship's hull.

[0102] Visual texture change noise can be interference from changes in the surface texture of the simulated object, such as changing the texture of a steel plate, or having a certain degree of rust, scratches, and paint, to simulate an old ship hull.

[0103] Lighting change noise can be interference that simulates changes in ambient lighting, such as simulating flickering lights or reflections inside a ship's cabin.

[0104] Simulation scenarios can be virtual operating environments built based on operating environment data, such as virtual welding environments built on the NVIDIA IsaacSim platform, which include virtual robots, workpiece models, and simulated environmental disturbances.

[0105] Add environmental disturbance data to the operating environment data, such as setting different light intensities and fluctuation frequencies in the simulation environment.

[0106] A simulation robot can be a virtual robot deployed in a simulation environment to simulate the performance of industrial operation tasks, such as a virtual welding robot.

[0107] This solution can automatically generate simulation scenarios based on operating environment data and environmental disturbance data. For example, by importing robot and workpiece models and setting environmental parameters and disturbance conditions, a simulation scenario similar to the actual industrial scene can be constructed. The simulation robot is controlled to perform operations based on the output instructions of the initial model; for example, controlling a virtual welding robot to perform welding operations according to the path and parameters output by the model. Relevant data generated by the simulation robot during the execution of the demonstration industrial operation task can be collected, such as welding torch posture, welding current, voltage, and weld formation data during the simulated welding process.

[0108] The high-level manager network can be a high-level decision-making module in the initial model, used to formulate sub-objectives, such as determining the expected weld width sub-objective and the coordinates of the next welding start point based on the demonstration industrial operation task.

[0109] The low-level actuator network can be a low-level execution module in the initial model, used to perform specific actions, such as determining joint motion parameters and welding torch angle control parameters based on the desired weld width sub-target and the coordinates of the next welding start point.

[0110] Joint motion parameters can be parameters that control the movement of robot joints, such as joint angular velocities of [0.1, -0.05, 0.2] rad / s.

[0111] The welding torch angle control parameter can be a parameter that controls the welding torch posture angle, for example, the welding torch angle is kept at 45°±5°.

[0112] This scheme can train the high-level manager network and the low-level actuator network separately. The high-level network is trained with the goal of setting the optimal sub-goal, while the low-level network is trained with the goal of accurately executing the sub-goal. By iteratively training in a simulation scenario and continuously adjusting the model parameters, the deviation between the operational simulation data output by the model and the demonstration industrial operational data is gradually reduced, and finally the trained operation control model is obtained.

[0113] This technical solution adds environmental disturbance data to the operating environment data to construct a simulation environment that closely resembles actual industrial scenarios, resulting in better model training performance. The hierarchical reinforcement training mode conforms to the task logic of complex industrial operations, with the higher level setting reasonable sub-goals and the lower level precisely executing actions, thereby improving the model's decision-making ability and execution accuracy. Furthermore, training the model in a simulated scenario using a simulated robot not only reduces actual training costs and safety risks but also enables rapid iterative optimization of the model. The higher-level manager network and the lower-level actuator network can implement specific functions and training objectives during the training process, enabling the trained model to better adapt to the complex changes in the actual industrial environment.

[0114] In one feasible embodiment, after deploying the operation control model to the operating device, the method further includes: The welding task is performed using the operating equipment, and the welding results of the operating equipment are quantitatively evaluated through a preset quality assessment subsystem to obtain a welding quality score. During the welding process, the system receives correction operation information provided by the user through a preset interface. The operation results are split into action sequences, and the splitting results are fused with the quality score and user correction operation information to construct a model optimization dataset; Based on the model, the dataset is optimized, and the correction reward function is learned using the maximum entropy inverse reinforcement learning algorithm. In the simulation scenario, the operation control model is subjected to hierarchical reinforcement learning iterative training based on the correction reward function.

[0115] Welding tasks can be specific industrial tasks that the operating equipment needs to perform, such as welding a double-sided continuous fillet weld on a ship's T-joint. This task can be completed automatically by the operating equipment under the drive of the control system. For example, a welding robot can automatically perform welding operations based on the control instructions output by the operation control model.

[0116] The quality assessment subsystem can be a pre-defined system for assessing the quality of welding results, such as an assessment system equipped with high-precision visual inspection equipment.

[0117] Quantitative assessment can be achieved by the quality assessment subsystem scanning and detecting the welding results, comparing them with ideal standards, and then giving a quantitative score. For example, parameters such as weld penetration, width, and number of pores can be detected, and the welding quality score can be calculated based on the detection results.

[0118] Welding quality score is a quantitative evaluation score of the quality of welding results, with a value range of 0-100 points, used to intuitively reflect the quality of welding results.

[0119] Preset interfaces can be used to receive user correction operation information, such as physical guidance interfaces, augmented reality / virtual reality (AR / VR) teleoperation interfaces, voice command interfaces, or graphical interface interfaces.

[0120] This solution can receive correction operation information sent by users through preset interfaces in real time, such as welding torch angle adjustment commands sent by users through AR / VR devices. It acquires relevant information when users correct their operation of the equipment, such as optimal path information provided by dragging the robot's end effector, or parameter adjustment information such as increasing force by 10% or slowing down speed given by voice commands.

[0121] The result of the operation can be the outcome after the equipment performs the welding task, such as the weld after welding.

[0122] Action sequence decomposition involves breaking down the operation process corresponding to the operation result into a series of continuous action sequences. For example, the welding process can be broken down into action sequences such as arc initiation, electrode movement, and arc termination. The decomposed action sequences, welding quality scores, and user correction operation information are then integrated to form a model optimization dataset.

[0123] Model optimization datasets can be datasets used to optimize operation control models, containing action sequences, quality evaluations, and correction information. For example, a dataset containing high-quality welding action sequences, corresponding welding quality scores, and user correction instructions.

[0124] Maximum Entropy Inverse Reinforcement Learning (Max Ent IRL) is an algorithm that inversely derives the reward function from data, learning the optimal reward rules implicit in the data.

[0125] The correction reward function can be a reward function learned based on the model optimization dataset. It can reward actions that meet the correction requirements and improve the quality of operation, such as rewarding the user for the welding torch angle and force operation after correction.

[0126] Iterative training allows the operational control model to be trained multiple times in a simulation scenario, with the correction reward function as the optimization objective, thereby continuously improving the model's performance.

[0127] This technical solution uses a quality assessment subsystem to quantitatively evaluate welding results, objectively reflecting the actual application effect of the operation control model. It receives user correction operation information, realizes human-machine collaboration, and enables the model to learn from the user's optimization experience. It constructs a model optimization dataset and learns the correction reward function. In the simulation scenario, it performs hierarchical reinforcement learning iterative training, which can continuously improve the performance of the operation control model, enabling the model to have self-evolution capabilities, constantly adapt to new process requirements and environmental changes, and further improve the operation accuracy and quality stability of the operating equipment.

[0128] To enable those skilled in the art to better understand this technical solution, this application also provides a preferred embodiment.

[0129] This invention provides a robot skill learning system and method that can automatically clean expert data, integrate optimal strategies from a group of experts, and improve robustness through adversarial simulation training, ultimately enabling the robot to master skills that surpass the level of a single expert.

[0130] To achieve the above objectives, this invention proposes a skill training method for embodied intelligent robots based on multi-expert knowledge fusion and adversarial training. Figure 2 This is a schematic diagram of the process based on multi-expert knowledge fusion and adversarial training provided in an embodiment of this application. For example... Figure 2 As shown, the main steps include: Step 1: Collect effective operational data from multiple senior technicians through multimodal data acquisition and preprocessing; First, an expert database was established. The selection criteria included industry-specific welding certification, more than 10 years of ship welding experience, a workpiece qualification rate of no less than 99% in the past three years, and passing peer review. Four to six senior welders who met the standards were selected as demonstration experts.

[0131] The multimodal data acquisition and preprocessing module simultaneously collects multi-source data from experts performing the aforementioned welding tasks. This data includes: high-precision visual images, welding torch spatial motion trajectory, six-dimensional force / torque sequence, and process parameters such as wire feed speed (8 m / min), welding current (220 A), voltage (28 V), welding speed (4.5 mm / s), torch angle (45°), and compressed air pressure (0.6 MPa). Equipment operating status and environmental parameters, such as light intensity and humidity inside the ship's cabin, are also recorded. After acquisition, the multi-source data is timestamped, and abnormal fluctuations caused by equipment vibration and signal interference are removed. Filtering and noise reduction are then applied to ensure data consistency and validity.

[0132] Step 2: Through expert data quality assurance strategies, the data is cleaned, filtered, and scored to form a high-quality demonstration dataset DATA_HQ; First, conduct cross-validation of multi-expert data and task breakdown; For example, the welding task is broken down into a task set consisting of serial tasks and parallel tasks. The serial tasks include, but are not limited to: From clamping and positioning, root tack fixing, reverse carbon gouging and root cleaning to reverse formal welding; From the first front pass weld, the interpass temperature is reduced to ≤150℃ before the second front pass weld; From carbon planing and cleaning, magnetic particle inspection (MT), defect repair to the first pass welding on the reverse side; From completing the seam, performing 3D laser scanning, issuing inspection reports, to releasing the workstation.

[0133] Parallel tasks include, but are not limited to: The front welding robot starts, the back carbon planer dust removal robot starts, and the MES (Manufacturing Execution System) reports work simultaneously; Layer temperature infrared monitoring, adjacent rib clamping robot operation, and automatic replenishment of flux drying box; The magnetic particle inspection robot takes photos, the reverse preheating flame gun is ignited, and the welding wire canister is replaced. Laser scanning, clamping at adjacent workstations, and cloud-based AI-predicted deformation.

[0134] For the above task set, we collected operation data from each expert, and through cross-validation of multi-expert data, we filtered out "outliers" caused by individual experts' operation habits and state fluctuations, and extracted the optimal operation mode shared by the group.

[0135] Then, quantitative indicators are used for scoring and screening; A comprehensive evaluation algorithm is used to conduct a multi-dimensional quantitative assessment of the demonstration data and calculate a comprehensive quality score. ; in, The quality dimension is scored based on welding quality parameters such as weld penetration, width, and number of pores, as well as assembly performance parameters such as component gap and concentricity. The score is quantified after the workpiece is detected by high-precision sensors. Efficiency is scored based on the total time to complete the task, such as ≤8 minutes for a 4m long weld, path smoothness (which can be evaluated by changes in the curvature of the operation path), and energy consumption parameters. The stability dimension is scored based on stability parameters of operating force, torque, and speed, compliance parameters, and path deviation parameters.

[0136] The initial weighting coefficients are set to α=0.60, β=0.25, and γ=0.15, and must satisfy the following conditions: For every 100 new weld data entries thereafter, using "rework / qualified" as 0 / 1 labels, the weight coefficients are iteratively optimized using a logistic regression algorithm with a learning rate of 0.01 and an L2 regularization of 1×10⁻⁶. -4 .

[0137] The comprehensive evaluation score threshold S is set at 80 points, and the top 20% of high-quality data are selected to form a high-quality demonstration dataset DATA_HQ.

[0138] Construct a process knowledge graph that includes safety rules, process standards and physical constraints, and use specifications such as "welding speed must not exceed 4mm / s" and "welding torch angle deviation must be ≤4°" as hard filters to automatically reject demonstration data that violates the rules.

[0139] For disputed data with quality scores falling within the intermediate threshold, such as 70 points ≤ S < 80 points, the data is submitted to the chief expert for adjudication via a manual audit interface. The chief expert's adjudication result, whether "valid" or "invalid," is used as a label to train the anomaly detection model, continuously optimizing the accuracy of the system's automatic judgment.

[0140] We selected the top 10% of expert welding trajectories from DATA_HQ as "good example" data, and introduced various perturbations to form an enhanced dataset. Specific perturbation types included: Path disturbance: Gaussian noise Δx, Δy, Δz, N(0, 0.5mm) is added to the simulated trajectory to simulate hull structure error or fixture deviation; Angular offset disturbance: Δ N(0,2°) represents the deviation of the welding torch posture caused by structural interference or human error. Force-controlled disturbance: A pulse of ±1N lasting 0.1s is added to the force / torque signal to simulate unevenness on the hull surface or sudden changes in welding wire contact. Welding electrical parameter disturbances: current ±10A, voltage ±1V, simulating power supply fluctuations or arc length changes; Visual disturbances: Gaussian blur, lighting changes, and metallic reflection noise are added to the simulated image to simulate interference such as strong light, reflections, and smoke inside the ship's cabin.

[0141] This step merges the standard operating dataset and the augmented dataset to obtain an updated standard operating dataset. At the same time, Dropout regularization is added to the network structure, and weight decay and data augmentation techniques are used to prevent the model from over-memorizing specific patterns from the training data and to improve generalization ability.

[0142] Step 3: Based on the maximum entropy inverse reinforcement learning algorithm, input DATA_HQ into the reward function and derive it in reverse to obtain the reward function R(s, a); The updated standard operation dataset is input into the reward function inverse derivation module. Using the Maximum Entropy Inverse Reinforcement Learning (Max Ent IRL) algorithm, the mapping relationship between high-quality operational behaviors and reward values ​​is mined, automatically deriving the reward function R(s, a) that implicitly reflects industry standard requirements and the optimal operational strategies of the expert group. This function can quantify the contribution of operational behaviors to welding quality, efficiency, and stability. For example, high reward values ​​are assigned to actions that achieve the required weld penetration, meet welding time requirements, and demonstrate smooth operation.

[0143] Step 4: In the simulation environment, R(s,a) is used as the reward signal to drive hierarchical reinforcement learning training, during which adversarial perturbations are introduced. The standard operating dataset includes exemplary industrial operating data and corresponding operating environment data, such as hull structure data, fixture position data, and initial lighting conditions. Environmental disturbance data is added to the operating environment data, including: Force noise: simulates ±0.5NZ force disturbances on the hull surface; Visual texture variation noise: Change the texture of the steel plate, such as rust, scratches and paint stains, to simulate an old ship hull; Lighting variation noise: Simulates the flickering and reflection of lights inside the ship's cabin.

[0144] Based on the above operating environment data and environmental disturbance data, a high-fidelity simulation scene was built on the NVIDIA Isaac Sim platform. Robot and workpiece models were imported, and environmental parameters, such as the range of light intensity fluctuations and metal reflectivity, were set to ensure that the simulation scene closely matches the actual industrial scene.

[0145] The initial operation control model was constructed based on a multilayer perceptron (MLP), which includes a high-level manager network and a low-level actuator network. It was deployed to the simulated robot, using a standard operation dataset as the learning sample and the reward function R(s,a) as the optimization objective. The Hierarchical PPO (Hi PPO) algorithm was used for iterative training.

[0146] High-level manager network: Define sub-goals on a second-level time scale, such as "move from point A to point B while maintaining a weld penetration of 2mm" or "determine the coordinates of the next welding start point and the desired weld width". Low-level actuator network: performs specific actions on a millisecond time scale, and determines joint motion parameters based on high-level sub-objectives, such as joint angular velocity [0.1,-0.05,0.2] rad / s, and welding torch angle control parameters, such as 45°±1°; Adversarial perturbation introduction: During training, perturbations are added to the force feedback, visual input, and operation path of the simulated robot. The layered strategy is required to still meet the task success criteria of weld penetration depth ≥2mm, width error ≤±0.5mm, and no defects such as undercut and porosity under the perturbation, thereby improving the model's cross-environment generalization ability.

[0147] Step 5: Deploy the trained strategy to the physical robot using the Sim-to-Real safe migration strategy; The trained strategy is deployed to the physical welding robot through the Sim-to-Real safe migration module. This module includes an adaptive impedance controller, runs on the RT-Linux real-time system, and is used to compensate for the differences between the simulation and the real dynamics online.

[0148] The adaptive impedance controller takes the desired trajectory, actual force feedback, and position error as inputs, and outputs the corrected joint torque or velocity command. When the solid robot is welding the double-bottom structure of a ship hull, if the actual steel plate thickness deviates from the design value by +2mm, causing a sudden increase in the welding torch contact force, the controller detects a force error >2N and automatically reduces the welding torch's downward speed and adjusts the Z-axis position to avoid torch collisions or poor weld formation, ensuring a smooth and safe migration process.

[0149] Step 6: After the robot runs, it continuously collects new data and continuously optimizes the model through cloud-based online self-learning and expert intervention to correct deviations.

[0150] When a physical robot performs a welding task, a pre-set quality assessment subsystem scans and inspects the welding results, comparing them with an ideal standard to generate a welding quality score Q_actual. For example, the ideal standard is a weld depth of 5 mm and a weld width of 1 cm, while the weld depth in the actual weld is 5.5 mm and the weld width is 1 cm. The welding quality score Q_actual can be calculated based on the deviation between these two values. Specifically, it is determined by the ratio of the weld depth deviation of 0.5 mm to the ideal standard weld depth (0.1), and the ratio of the weld width deviation of 0 mm to the ideal standard weld width (0). If the quality score is... Meanwhile, it receives corrective action information from experts in real time through multiple preset interfaces.

[0151] For example, an expert can manually drag the robot's end effector to demonstrate a better path or force. The system records the correction actions using a high-precision force sensor. With AR / VR (Augmented Reality / Virtual Reality) teleoperation, the expert can plan and correct the robot's movements in a virtual mapping space, with instructions issued in real time. Alternatively, the expert can use voice commands, such as "increase the force by 10%" or "slow down the speed," or adjust the operating parameters using a slider.

[0152] The operation process corresponding to the operation result is decomposed into action sequences. For example, the welding process is broken down into continuous actions such as arc initiation, electrode movement, and arc termination. The decomposed results are then fused with welding quality scores and user correction operation information to construct a model optimization dataset. The Max Ent IRL algorithm is used to learn the correction reward function from this dataset. In the simulation scenario, the operation control model is iteratively trained using hierarchical reinforcement learning based on this function.

[0153] The entire process data of a successful task, such as the state sequence s1, s2, ..., s... n Action sequence a1, a2, ..., a n The model is bound to Q_actual to form positive samples, which are used to fine-tune the reward function and policy network, and the new model version with improved performance is seamlessly deployed to the robot cluster to achieve continuous model evolution.

[0154] On-site verification showed that after adopting the technical solution of this application, the rework rate of 200 automatic welds was reduced from 8% to 1.5%, the welding efficiency was increased by more than 30%, the skills learned by the robot surpassed the level of a single expert, and it can effectively adapt to the process requirements and environmental changes of complex industrial scenarios such as shipbuilding, significantly reduce the dependence on skilled welders, and ensure the continuity and stability of welding operations.

[0155] Based on the above methods, the present invention adopts the following technical solution: a robot skill learning system based on multi-expert knowledge fusion and adversarial training. Figure 3 This is a schematic diagram of the structure of a robot skill learning system based on multi-expert knowledge fusion and adversarial training provided in an embodiment of this application. Figure 3 As shown, it includes: Multimodal expert data acquisition and preprocessing module 310; This is used to simultaneously collect operational data from multiple experts (N≥4), including high-precision visual images, motion trajectories, six-dimensional force / torque sequences, and process parameters; and to perform timestamp alignment and filtering on the multi-source data.

[0156] Multi-expert data quality assurance module 320; If the training data (the operations performed by the skilled technicians) contains systematic biases or errors, the trained robot will learn these errors and even amplify them. Because "garbage in, garbage out" is a typical problem in machine learning, the validity and accuracy of the input data from the skilled technicians directly affect the reliability and ultimate performance ceiling of the entire system. This is a core risk that must be addressed at the system design level.

[0157] Reward function inverse derivation module 330; The maximum entropy inverse reinforcement learning algorithm is used as input to automatically derive the implicit and optimal reward function R(s, a).

[0158] Layered reinforcement learning training module 340; In a high-fidelity simulation environment based on a physics engine, training is performed using the derived reward function R(s, a). This module adopts a hierarchical structure: The high-level manager accepts rewards R(s, a) over a longer timescale and outputs sub-goals.

[0159] Worker: Performs primitive actions to achieve sub-goals in a shorter time scale, and also with R(s,a) as the optimization goal.

[0160] Adversarial training unit: Introduces domain randomization and small perturbations to expert demonstrations (such as force noise and visual texture changes) into the simulation to improve the robustness and generalization ability of the model.

[0161] For an explanation of the timescale of the hierarchical reinforcement learning training module, please refer to Table 1 below; Table 1: Compared with the prior art, the present invention has the following significant advantages: Skill level surpassing individual experts: By integrating the best practices of multiple experts and filtering out bad habits, the robot's learned skills represent the highest level of the group, rather than imitating a single expert who may have flaws.

[0162] Data-driven reward design eliminates the need for manual reward design: By automatically deriving reward functions from data through inverse reinforcement learning, the bottleneck of difficulty in manually defining reward functions for complex skills is solved.

[0163] Strong robustness and high security: The data cleaning mechanism eliminates bad learning operations; adversarial training in simulation greatly improves the model's adaptability and stability in the real world; and the secure migration mechanism ensures the security of the deployment process.

[0164] Self-evolution capability: The online learning module enables the system to continuously optimize and adapt to new process requirements or environmental changes.

[0165] Possessing human-level adaptability and error correction capabilities: Through real-time human-machine collaboration, robots can cope with new situations not anticipated during training, resulting in a qualitative leap in their flexibility and practicality.

[0166] Achieving knowledge transfer and evolution across time and space: Any authorized expert's effective optimization and improvement of a robot can be instantly deployed to all similar robots worldwide via cloud synchronization, enabling the "exponential" accumulation and transfer of expert experience and skills, and allowing the entire robot group's capabilities to continuously evolve.

[0167] Figure 4 This is a schematic diagram of the structure of an operation learning device based on multi-expert knowledge fusion provided in an embodiment of this application. Figure 4 As shown, the device may include: The operation data acquisition unit 410 is used to acquire demonstration industrial operation data from at least two experts performing demonstration industrial operation tasks. Scoring unit 420 is used to evaluate the demonstration industrial operation data of each expert according to the quality dimension, efficiency dimension and stability dimension using a comprehensive evaluation algorithm, and to obtain a comprehensive evaluation score for each demonstration industrial operation data. Data set construction unit 430 is used to select demonstration industrial operation data with comprehensive evaluation scores higher than a set threshold based on the comprehensive evaluation scores of each demonstration industrial operation data, and construct a standard operation dataset. The reward function determination unit 440 is used to determine the reward function based on the standard operating dataset using an inverse reinforcement learning algorithm. The model training unit 450 is used to perform hierarchical reinforcement training on the initial operation control model based on the standard operation dataset and the reward function to obtain the trained operation control model.

[0168] The operation learning device based on multi-expert knowledge fusion provided in this embodiment has the same functional modules and beneficial effects as the operation learning method based on multi-expert knowledge fusion described above. To avoid repetition, it will not be described in detail here.

[0169] Figure 5 This is a schematic diagram of the structure of an operation learning device based on multi-expert knowledge fusion provided in an embodiment of this application. Figure 5 As shown, the operation learning device based on multi-expert knowledge fusion may include a processor 501 and a memory 502 storing computer program instructions.

[0170] Specifically, the processor 501 may include a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.

[0171] Memory 502 may include mass storage for data or instructions. For example, and not limitingly, memory 502 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. In one instance, memory 502 may include removable or non-removable (or fixed) media, or memory 502 may be non-volatile solid-state storage. Memory 502 may be internal or external to the integrated gateway disaster recovery device.

[0172] In one instance, memory 502 may be read-only memory (ROM). In one instance, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these.

[0173] Memory 502 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this disclosure.

[0174] The processor 501 reads and executes computer program instructions stored in the memory 502 to implement the operation learning method based on multi-expert knowledge fusion in the above embodiments.

[0175] In one example, the operation learning device based on multi-expert knowledge fusion may further include a communication interface 503 and a bus 504. For example, Figure 5 As shown, the processor 501, memory 502, and communication interface 503 are connected through bus 504 and complete communication with each other.

[0176] The communication interface 503 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.

[0177] Bus 504 includes hardware, software, or both, that couples components of an operational learning device based on multi-expert knowledge fusion together. For example, and not limitingly, the bus may include an Accelerated GraphicsPort (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 504 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.

[0178] The operation learning device based on multi-expert knowledge fusion can execute the operation learning method based on multi-expert knowledge fusion in the embodiments of this application, thereby realizing the operation learning method based on multi-expert knowledge fusion described in the above embodiments.

[0179] Furthermore, in conjunction with the operation learning device method based on multi-expert knowledge fusion in the above embodiments, this application embodiment can provide a computer storage medium for implementation. The computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the operation learning device methods based on multi-expert knowledge fusion in the above embodiments.

[0180] This application also provides a computer program product, including a computer program that, when executed by a processor, implements any of the operation learning device methods based on multi-expert knowledge fusion in the above embodiments.

[0181] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.

[0182] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, read-only memory (ROM), flash memory, erasable read-only memory (EROM), floppy disks, compact disc read-only memory (CD-ROM), optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.

[0183] It should be noted that the acquisition, storage, use, and processing of data in this application embodiment all comply with the relevant provisions of national laws and regulations.

[0184] In the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, they do not mean that the applicant has used or necessarily used the solution.

[0185] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.

[0186] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.

[0187] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.

Claims

1. An operational learning method based on multi-expert knowledge fusion, characterized in that, The method includes: Obtain demonstration industrial operation data performed by at least two experts; A comprehensive evaluation algorithm was used to evaluate the demonstration industrial operation data of each expert according to the dimensions of quality, efficiency, and stability, and a comprehensive evaluation score was obtained for each demonstration industrial operation data. The evaluation score of the quality dimension was determined based on at least one of the following: weld surface smoothness grade, weld penetration depth, weld width, and number of pores. The evaluation score of the efficiency dimension was determined based on at least one of the following: total time to complete the task, path smoothness, and energy consumption parameters. The stability dimension was determined based on at least one of the following: operating force, operating torque, and operating speed. Based on the comprehensive evaluation score of each demonstration industrial operation data, demonstration industrial operation data with a comprehensive evaluation score higher than a set threshold are selected to construct a standard operation dataset. The reward function is determined using an inverse reinforcement learning algorithm based on the aforementioned standard operating dataset. Based on the standard operation dataset and the reward function, the initial operation control model is subjected to hierarchical reinforcement training to obtain the trained operation control model. The operation control model is deployed to the operating equipment to control the operation of the operating equipment when it performs industrial operation tasks.

2. The operation learning method based on multi-expert knowledge fusion according to claim 1, characterized in that, A comprehensive evaluation algorithm was used to evaluate the demonstration industrial operation data provided by each expert according to three dimensions: quality, efficiency, and stability. Identify welding quality parameters and assembly performance parameters in each of the aforementioned demonstration industrial operation data; wherein, the welding quality parameters include at least one of weld penetration, weld width, and number of pores; the assembly performance parameters include the gap and / or concentricity parameters of the assembled components; The quality dimension scores for each of the aforementioned demonstration industrial operation data are determined based on the welding quality parameters and assembly performance parameters. Identify at least one of the following parameters in the demonstration industrial operation data: total time to complete the task, path smoothness, and energy consumption parameters, and determine the efficiency dimension score for each demonstration industrial operation data. Identify stability parameters, compliance parameters, and path deviation parameters in the operational data of each of the demonstration industrial projects; wherein, the stability parameters include the stability of operating force, operating torque, and operating speed; wherein, the compliance parameters include whether the operating force and operating torque exceed the specified thresholds; Based on the stability parameter, the compliance parameter, and the path deviation parameter, a stability dimension score is determined for each of the demonstration industrial operation data. The evaluation score for each of the demonstration industrial operation data is determined based on the scores for the quality dimension, the efficiency dimension, and the stability dimension.

3. The operation learning method based on multi-expert knowledge fusion according to claim 2, characterized in that, Based on the scores for the quality dimension, the efficiency dimension, and the stability dimension, an evaluation score is determined for each of the demonstration industrial operation data, including: The evaluation score is calculated using the following formula: ; in, To evaluate the score, Score the quality dimensions. Score the efficiency dimension. Score the stability dimension. , , These are adjustable weighting coefficients.

4. The operation learning method based on multi-expert knowledge fusion according to claim 1, characterized in that, Before using a comprehensive evaluation algorithm to evaluate the demonstration industrial operation data from each expert according to the dimensions of quality, efficiency, and stability, the method further includes: The operational data of each of the demonstration industries are decomposed into tasks to obtain a task set consisting of serial tasks and parallel tasks. Serialize multiple tasks in the task set to obtain the execution order relationship; The optimal order relationship is determined based on the execution order relationship, and the order anomaly filtering is performed on the demonstration industrial operation data to obtain demonstration industrial operation data with the optimal order.

5. The operation learning method based on multi-expert knowledge fusion according to claim 1, characterized in that, Before determining the reward function using an inverse reinforcement learning algorithm based on the standard operating dataset, the method further includes: Demonstration industrial operation data is filtered based on a preset screening ratio, and perturbation data is added to obtain an enhanced dataset; wherein, the perturbation data includes one or more of the following: path perturbation, welding torch angle offset perturbation, force control perturbation, welding electrical parameter perturbation, and visual perturbation. The standard operation dataset and the enhanced dataset are merged into an updated standard operation dataset.

6. The operation learning method based on multi-expert knowledge fusion according to claim 5, characterized in that, The standard operating data set includes: the demonstration industrial operating data and the corresponding operating environment data of the demonstration industrial operating data; Based on the standard operation dataset and the reward function, the initial operation control model is subjected to hierarchical reinforcement training to obtain the trained operation control model, including: Environmental disturbance data is added to the operating environment data; wherein, the environmental disturbance data includes at least one of force noise, visual texture change noise, and illumination change noise; A simulation scenario is constructed based on the aforementioned operating environment data; The simulation robot with the initial model is controlled to perform the demonstration industrial operation task in the simulation scenario to obtain operation simulation data; wherein, the high-level manager network in the initial model is used to determine the expected weld width sub-target and the coordinates of the next welding start point based on the demonstration industrial operation task, and the low-level actuator network in the initial model determines the joint motion parameters and welding torch angle control parameters based on the expected weld width sub-target and the coordinates of the next welding start point; Based on the operation simulation data, the demonstration industrial operation data, and the reward function, the high-level manager network and the low-level actuator network of the initial model are trained hierarchically to obtain the trained operation control model.

7. The operation learning method based on multi-expert knowledge fusion according to claim 6, characterized in that, After deploying the operation control model to the operating device, the method further includes: The welding task is performed using the operating equipment, and the welding results of the operating equipment are quantitatively evaluated through a preset quality assessment subsystem to obtain a welding quality score. During the welding process, the system receives correction operation information provided by the user through a preset interface and obtains the operation results. The operation results are split into action sequences, and the splitting results are fused with the quality score and user correction operation information to construct a model optimization dataset; Based on the model, the dataset is optimized, and the correction reward function is learned using the maximum entropy inverse reinforcement learning algorithm. In the simulation scenario, the operation control model is subjected to hierarchical reinforcement learning iterative training based on the correction reward function.

8. An operational learning device based on multi-expert knowledge fusion, characterized in that, The device includes: The operation data acquisition unit is used to acquire demonstration industrial operation data performed by at least two experts in the demonstration industrial operation task; The scoring unit is used to evaluate the demonstration industrial operation data of each expert according to the quality dimension, efficiency dimension, and stability dimension using a comprehensive evaluation algorithm, and obtain a comprehensive evaluation score for each demonstration industrial operation data. The evaluation score of the quality dimension is determined based on at least one of the following: weld surface smoothness grade, weld penetration depth, weld width, and number of pores. The evaluation score of the efficiency dimension is determined based on at least one of the following: total time to complete the task, path smoothness, and energy consumption parameters. The stability dimension is determined based on at least one of the following: operating force, operating torque, and operating speed. The dataset construction unit is used to select demonstration industrial operation data with comprehensive evaluation scores higher than a set threshold based on the comprehensive evaluation scores of each demonstration industrial operation data, and construct a standard operation dataset. The reward function determination unit is used to determine the reward function based on the standard operating dataset using an inverse reinforcement learning algorithm. The model training unit is used to perform hierarchical reinforcement training on the initial operation control model based on the standard operation dataset and the reward function to obtain the trained operation control model. The deployment unit deploys the operation control model to the operating device to perform operation control on the operating device when the operating device performs industrial operation tasks.

9. An operating device, characterized in that, The operating device includes: a robot body and a control unit; the control unit stores an operation control model as described in any one of claims 1-7, or performs operation control on the robot body based on the operation control model as described in any one of claims 1-7.

10. An operational learning device based on multi-expert knowledge fusion, characterized in that, The device includes: a processor and a memory storing computer program instructions; the processor reads and executes the computer program instructions to implement the operation learning method based on multi-expert knowledge fusion as described in any one of claims 1-7.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when executed by a processor, implement the operation learning method based on multi-expert knowledge fusion as described in any one of claims 1-7.

12. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the operation learning method based on multi-expert knowledge fusion as described in any one of claims 1-7.