A control method and device for a flexible assembly robotic arm based on abnormal working conditions

By generating abnormal state data and utilizing weighted training of multiple evaluation models, the problem of stable judgment and action selection of robotic arms under abnormal working conditions in flexible assembly was solved, thereby improving product yield and production stability.

CN122125718APending Publication Date: 2026-06-02XI AN JIAOTONG UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XI AN JIAOTONG UNIV
Filing Date
2026-04-23
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

In existing flexible assembly processes, due to frequent changes in working conditions, robotic arms struggle to maintain stable judgment and make reasonable action selections when faced with unseen disturbances, resulting in low product yield. Existing methods are unstable in taking values ​​under abnormal working conditions, and their action selections are conservative, making it difficult to balance robustness and extrapolation capability.

Method used

By acquiring normal samples, abnormal data is generated by offsetting or interpolating normal data with a preset amplitude. Multiple evaluation models are used to calculate assembly quality evaluation scores, and a target loss value is formed by weighting. The target evaluation model is then trained to determine the optimal action of the robotic arm. By combining the construction of abnormal states at the data level with the conservative aggregation at the evaluation stage, the ability to identify and tolerate unseen states is improved.

Benefits of technology

Without increasing the computational burden during the deployment phase, it improves the ability to identify and recover from abnormal operating conditions, adapts to the disturbance characteristics of flexible assembly, and enhances product yield and production stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122125718A_ABST
    Figure CN122125718A_ABST
Patent Text Reader

Abstract

This invention discloses a control method and device for a flexible assembly robotic arm based on abnormal working conditions, relating to the field of flexible assembly technology. Abnormal data is obtained by offsetting or interpolating normal data with a preset amplitude. Assembly quality assessment scores for performing assembly actions under normal data are calculated using multiple evaluation models. Based on the minimum assembly quality assessment score and the reference quality index corresponding to the normal sample, the evaluation deviation risk value of the normal sample is determined. The input sensitivity vector of each evaluation model is calculated using the abnormal data, and a correlation penalty loss value is calculated based on the input sensitivity vector of each evaluation model. The evaluation deviation risk value and the correlation penalty loss value are weighted to obtain a target loss value. Each evaluation model is trained using the target loss value to obtain a target evaluation model. The target evaluation model is used to evaluate the value of actions in the current state to determine the optimal action of the robotic arm. This method improves product yield.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of flexible assembly technology, and in particular to a control method and device for a flexible assembly robotic arm based on abnormal working conditions. Background Technology

[0002] In general flexible assembly production processes, batch differences in incoming parts, slight deformation of parts, reflection and illumination fluctuations, fixture vibration and occlusion are common phenomena. The robotic arm relies on visual perception, fixture execution and trajectory control to coordinate positioning, grasping and assembly, and the training data mostly comes from historical records and small-batch trial production tests.

[0003] However, due to frequent changes in operating conditions, the assembly process needs to maintain stable judgment and reasonable action selection in the face of unseen disturbances while ensuring production progress and product yield. Existing offline training methods mostly restrict action selection with uniform conservative strategies, which do not adequately consider problems beyond the distribution of working states caused by changes in the environment and perception. This can easily lead to over-constraint on potential feasible solutions, resulting in low product yield. Summary of the Invention

[0004] Therefore, it is necessary to provide a control method and device for a flexible assembly robot arm based on abnormal working conditions to address the above-mentioned technical problems.

[0005] The following technical solution is adopted in this specification: This specification provides a control method for a flexible assembly robot arm based on abnormal working conditions, including: Acquire normal samples in flexible assembly tasks; normal samples include normal data and corresponding assembly actions; normal data represents the state data of parts with dimensions within tolerance range under stable illumination; Without violating equipment tolerances and safety constraints, abnormal data is obtained by offsetting or interpolating normal data by a preset amplitude; abnormal data includes states not covered in the normal samples of flexible assembly tasks. The assembly quality assessment score for performing assembly actions under normal data is calculated using multiple assessment models. Based on the minimum assembly quality assessment score and the reference quality index corresponding to the normal sample, the assessment deviation risk value of the normal sample is determined. Each assessment model focuses on different key factors when scoring. Using out-of-state data, the input sensitivity vector of each evaluation model is calculated, and the correlation penalty loss value is calculated based on the input sensitivity vector of each evaluation model; the input sensitivity vector represents the magnitude of change in the assembly quality evaluation score output by the evaluation model when out-of-state data changes. The target loss value is obtained by weighting the evaluation bias risk value and the correlation penalty loss value, and then training each evaluation model with the target loss value to obtain the target evaluation model. The optimal action for the robotic arm is determined by evaluating the value of the current action using a target evaluation model.

[0006] Optionally, the normal sample includes multiple normal data sets; the normal data sets include workstation images, robotic arm end-effector poses, and attitude data and force-position measurement data of key contact points; abnormal data are obtained by performing a preset amplitude offset or interpolation along the normal data sets, including: For any given normal data, zero-mean noise is added to each type of data in the normal data to obtain the first data; For any given set of normal data, an offset is added to the key features in the normal data that are related to the positioning error to obtain the second set of data; A third set of data is obtained by performing mixed-weight interpolation on multiple normal data sets. The first, second, and third data points are identified as abnormal data.

[0007] Optionally, the method further includes: An interval clamp is applied to the first data to limit the image intensity in the first data to the effective grayscale range and to limit the attitude data and force measurement data to the upper and lower limits allowed by the device tolerance.

[0008] Optionally, multiple normal data points are interpolated to obtain a third set of data, including: Two normal data points are randomly selected from the normal samples of the same workstation and the same assembly task. The two selected normal data points are weighted by a weighting factor to obtain a third data point. The value of the weighting factor is between 0.2 and 0.8.

[0009] Optionally, using anomalous data, the input sensitivity vector for each evaluation model is calculated, including: For any evaluation model, the partial derivative of the assembly quality evaluation score output by the evaluation model with respect to the input abnormal data is determined as the input sensitivity vector.

[0010] Optionally, based on the input sensitivity vector of each evaluation model, a relevance penalty loss value is calculated, including: Calculate the correlation between any two evaluation models based on the input sensitivity vector of each evaluation model; Calculate the correlation penalty loss value based on the correlation between any two evaluation models.

[0011] Optionally, the correlation penalty loss value The calculation formula is: ; in, Indicates the first i The evaluation model and the firstj The correlation between the evaluation models This is the upper limit of relevance. .

[0012] This specification provides a control device for a flexible assembly robot arm based on abnormal working conditions, including: The acquisition module is used to acquire normal samples in flexible assembly tasks; normal samples include normal data and corresponding assembly actions; normal data represents the state data of parts with dimensions within tolerance range under stable illumination. The data generation module is used to obtain abnormal data by offsetting or interpolating normal data by a preset amplitude without violating equipment tolerances and safety constraints; the abnormal data includes states not covered in the normal samples in the flexible assembly task; The first evaluation module is used to calculate the assembly quality evaluation score of the assembly action performed under normal data through multiple evaluation models, and to determine the evaluation deviation risk value of the normal sample based on the minimum assembly quality evaluation score and the reference quality index corresponding to the normal sample; each evaluation model focuses on different key factors when scoring. The second evaluation module is used to calculate the input sensitivity vector of each evaluation model using out-of-state data, and to calculate the correlation penalty loss value based on the input sensitivity vector of each evaluation model; the input sensitivity vector represents the magnitude of change in the assembly quality evaluation score output by the evaluation model when the out-of-state data changes. The weighting module is used to weight the evaluation bias risk value and the correlation penalty loss value to obtain the target loss value, and to train each evaluation model with the target loss value to obtain the target evaluation model; The control module is used to evaluate the value of the current state's actions using a target evaluation model, and to determine the optimal action for the robotic arm.

[0013] This specification provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described control method for a flexible assembly robotic arm based on abnormal working conditions.

[0014] This specification provides a control device for flexible assembly, including a processor and a memory. The memory stores program instructions, and the processor executes the program instructions to implement the above-described control method for a flexible assembly robotic arm based on abnormal working conditions. It is also communicatively connected to a camera, a pose sensor, and an end effector, and outputs control instructions and evaluation results compatible with the trajectory control module.

[0015] Optionally, the control device provides a version management and rollback interface to support parameter fixing, grayscale scaling, and migration across multiple models.

[0016] Optionally, the control device maintains a strategy of aligning the minimum value aggregation with the training objective during the deployment phase to reduce misjudgments caused by uncertainty.

[0017] This specification provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described control method for a flexible assembly robotic arm based on abnormal working conditions.

[0018] The above-mentioned technical solutions adopted in this specification can achieve the following beneficial effects: In the control method for flexible assembly robotic arms based on abnormal working conditions provided in this specification, under the premise of not violating equipment tolerances and safety constraints, a preset amplitude offset or interpolation is performed along the normal data to generate abnormal state data. The abnormal state data is used to train the evaluation model, so that the model has "pre-learned" various boundary scenarios during the training phase. When encountering external disturbances in actual operation, it will not be oversensitive or fail, and will maintain stable judgment. Furthermore, multiple evaluation models focus on different key factors, and conservative minimum value aggregation is implemented on normal samples to suppress the expansion of uncertainty. At the same time, it avoids the excessive rejection of high-return solutions with practical feasibility. Independent target and input sensitivity are introduced to decorrelate abnormal samples, so that the evaluation model forms differentiated focus at the functional level. Combined with minimum value aggregation, a robust target loss value is obtained. Without increasing the computational burden in the deployment phase, the ability to identify and recover from external states is improved, adapting to the disturbance characteristics of flexible assembly, thereby improving product yield. Attached Figure Description

[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0020] Figure 1 This specification provides a schematic flowchart of a control method for a flexible assembly robot arm based on abnormal working conditions. Figure 2 This specification provides a comparison diagram of the state distribution of the learning strategy and offline data in a flexible assembly process. Figure 3 This document provides a comparative diagram of action selection strategies and extrapolation effects under abnormal working conditions in flexible assembly. Figure 4 This specification provides a schematic diagram of the structure and data flow of a flexible assembly offline training and collaborative evaluation system. Figure 5 This is a schematic diagram of a computer device used in this specification to implement a control method for a flexible assembly robotic arm based on abnormal working conditions. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this specification clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments in this specification without creative effort are within the scope of protection of this application.

[0022] In existing technologies, evaluation model cluster schemes generally adopt shared objectives, leading to convergent focus among evaluation models. This makes them prone to consistent misjudgments under complex interference, hindering effective functional division. The introduction of dynamic models or generators to compensate for insufficient data coverage incurs additional modeling costs and integration complexity, impacting parameter controllability and traceability. Some methods also increase computational and communication overhead during deployment, affecting on-site production speed and maintenance convenience. These issues result in unstable strategy values ​​and conservative action selection under abnormal conditions, making it difficult to balance robustness and extrapolation capability.

[0023] Based on this, the present invention provides a control method and device for a flexible assembly robot arm based on abnormal working conditions. By constructing and sampling modelless abnormal states in conjunction with strategies at the data level, performing decorrelation operations on the independent targets and input sensitivities of the evaluation model cluster in the evaluation stage, and adopting conservative minimum value aggregation in the deployment stage, robust values ​​and reasonable extrapolation for unseen disturbances are achieved, while maintaining a simple execution architecture that is easy to connect and maintain.

[0024] Specifically, considering factors such as batch variations, slight part deformation, strong glare and illumination fluctuations, minor end-clamp jitter, and temporary obstructions in flexible assembly lines, frequent occurrences of uncovered training states during operation can lead to value assessment biases and overly conservative action selections. This further results in production rhythm fluctuations, decreased product yield, increased downtime for troubleshooting, and higher debugging costs. This invention constructs an offline training and deployment system for stable production rhythms. For normal ranges covered by historical data, the accuracy and stability of value assessment are emphasized to ensure the assembly process meets predetermined tolerances and quality requirements. For potentially abnormal ranges, controllable disturbance samples are introduced, and a value-taking mechanism based on evaluation model cluster collaboration is adopted to improve the identification and fault tolerance capabilities for unseen state and action combinations. Conservative minimum value aggregation is implemented during the evaluation and backup phases to suppress the expansion of uncertainty while avoiding the excessive rejection of practically feasible high-return solutions. The system maintains a simple execution architecture, facilitating integration with modules such as visual inspection and trajectory control, and supporting rapid deployment and subsequent maintenance.

[0025] Compared to existing methods that often focus on uniform conservative constraints at the action level, neglecting risks arising from external state distributions and suppressing potentially viable high-return solutions, resulting in limited improvements in recovery speed and product yield, this invention integrates both state and action into the training and evaluation objectives. It employs a combined design of constructing abnormal states on the data side and conservative aggregation during the evaluation phase. Without increasing online computational burden or call complexity, it simultaneously achieves robustness and the ability to extrapolate abnormal states, adapting to the perturbation characteristics of flexible assembly.

[0026] To address the issue that under conditions of intertwined illumination reflection, pose displacement, and fixture deformation, a single evaluation model is prone to consistent misjudgments of noise and abnormal features. Furthermore, when the focus of an evaluation model cluster overlaps, errors are amplified, leading to slow response to out-of-distribution states, increased recovery time, and higher debugging frequency, ultimately impacting production rhythm and product yield. This invention constructs a collaborative approach for an evaluation model cluster. For the normal range covered by historical data, independent convergence targets are set for each evaluation model, ensuring sufficient accuracy and consistency in value evaluation under stable conditions. For abnormal ranges, the correlation of input sensitivity between evaluation models is reduced, allowing different evaluation models to perform functional divisions, focusing on key elements such as pose, velocity, and texture reflection, improving the identification and fault tolerance capabilities for unseen states and action combinations. A conservative minimum aggregation method is used in the evaluation and backup phases to suppress overestimation caused by high uncertainty, balancing robustness and executability, and facilitating integration with modules such as visual inspection and trajectory control.

[0027] Compared to existing evaluation model cluster schemes that often employ a unified objective, leading to model homogenization and insufficient extrapolation capabilities, this invention introduces both independent objectives and input sensitivity decorrelation in engineering implementation. This allows the evaluation model to differentiate its focus at the functional level, and, combined with minimum value aggregation, obtains robust values. Without increasing the computational burden during deployment, it improves the ability to identify and recover from out-of-distribution states, adapting to the disturbance characteristics of flexible assembly.

[0028] For common scenarios in flexible assembly lines, such as sudden changes in illumination, partial occlusion, and slight deformation of parts, offline data coverage is insufficient, and the training phase lacks targeted learning for key abnormal states. Under these working conditions, strategies exhibit unstable values ​​and conservative action selection, leading to slow production recovery, limited product yield improvement, and difficulty in developing effective extrapolation capabilities for anomalies. This invention constructs model-free abnormal states based on historical data without introducing an environment model or generator, and integrates this with the current strategy to complete action sampling, forming state-action samples for training. The construction method includes controllable amplitude random noise perturbations, micro-perturbations of key features, and sample hybrid interpolation to cover various possible out-of-distribution states; the generated samples are directly supplied to the collaborative device of the evaluation model cluster for functional division training.

[0029] Compared to traditional methods that rely heavily on dynamic models or generators for distributed operation, resulting in higher engineering costs and integration complexity, as well as insufficient traceability and parameter controllability, this invention employs a purely data-level construction approach linked to strategy sampling. This approach offers shorter paths, stronger controllability, and facilitates rapid iteration on the assembly site.

[0030] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.

[0031] Figure 1 This is a flowchart illustrating a control method for a flexible assembly robot arm based on abnormal working conditions, as described in this specification. The method specifically includes the following steps: S101, Obtain normal samples in flexible assembly tasks; normal samples include normal data and corresponding assembly actions; normal data represents the state data under stable illumination and when the part dimensions are within tolerance range.

[0032] The routine samples can include samples selected from historical assembly records and small-batch trial production data based on assembly tolerance and first-time assembly success rate thresholds. Routine samples include routine data and corresponding assembly actions. Routine data includes traceable physical quantities such as station images captured by overhead and lateral cameras, the pose of the end effector in the workpiece coordinate system, and the end effector clamping force and contact state. The combination of these physical quantities within the same cycle is recorded as the assembly state, i.e., routine data. s The assembly action is the combination of control commands such as speed, displacement step size, insertion force, and alignment fine-tuning executed by the robotic arm within the corresponding cycle. a The standard sample also includes the assembly result corresponding to the assembly status (one piece is qualified, reworked, or scrapped).

[0033] Based on the correspondence between "assembly status - assembly action - assembly result" in the historical records, reference quality indicators that can reflect the performance of the assembly are calculated, such as the normalized scores of successful assembly, rework, and scrap, or statistical indicators related to the first-time success rate and defect rate.

[0034] S102, without violating equipment tolerances and safety constraints, perform a preset amplitude offset or interpolation along the normal data to obtain abnormal data; abnormal data includes states not covered in the normal samples in the flexible assembly task.

[0035] To address situations in flexible assembly lines where sudden changes in illumination, partial occlusion, and slight deformation of parts are insufficiently covered by normal samples, this approach constructs traceable and controllable abnormal states at the data level and integrates them with the current assembly strategy to complete motion sampling. This provides training samples specifically covering "out-of-distribution" scenarios for evaluating the model cluster. This method does not introduce additional dynamic models or generators; it only transforms existing images, attitudes, and force-position records, facilitating rapid implementation and iteration in production settings.

[0036] Optionally, routine data includes workstation images, robotic arm end-effector pose, and attitude data and force-position measurement data of key contact points; routine samples include multiple routine data sets, i.e., multiple assembly records; for ease of explanation, the routine data of a single assembly will be used. Represented as: (1); in, The image of the workstation or its feature vector reflects the appearance and reflectivity of the part; The pose data for the robotic arm's end effector and key contact points reflects the insertion direction and depth; The force measurement data, such as clamping force and insertion force, reflects the resistance and interference during the contact process. The goal of constructing abnormal states is to obtain states that may occur on-site but are not yet fully covered by the data, by making small offsets or interpolations along reasonable directions of the above physical quantities, without violating equipment tolerances and safety constraints.

[0037] In one embodiment, abnormal data is obtained by performing a preset amplitude offset or interpolation on normal data, including the following steps: S201, For any normal data, zero-mean noise is superimposed on each type of data in the normal data to obtain the first data.

[0038] First, random noise disturbances of sensing error are simulated for normal data. Specifically, normal data... By adding zero-mean noise, we get: (2); in, As the first data, Each component corresponds to a small jitter in image brightness, attitude angle and force position signal, and its standard deviation is determined by on-site calibration: for example, the standard deviation of the image component is selected within a small segment of the pixel intensity quantization range, the standard deviation of the attitude component does not exceed the angle and displacement measurement accuracy, and the standard deviation of the force position component does not exceed the normal working jitter. It follows a mean of 0 and a covariance of normal distribution To avoid generating states beyond the device's capabilities, the first data... Perform interval clamping to limit the image intensity in the first data to the effective grayscale range and to limit the attitude data and force measurement data to the upper and lower limits allowed by the device tolerance.

[0039] S202, for any normal data, add offset to the key features related to positioning error in the normal data to obtain the second data.

[0040] Secondly, we simulate the perturbation of key features related to positioning errors. Key features related to positioning errors in normal data include a subset of positioning-related dimensions in the feature space. For example, normal data may include workstation images, the pose of the robotic arm's end effector and the attitude of key contact points, force and displacement measurements, etc. Among these, those directly corresponding to positioning errors are typically: the relative position and attitude of the end effector and contact points, the insertion direction angle, and the position of key points in the image and coordinate system. These dimensions are defined as a set. M For each Generate offset ,get:

[0041] (3); in, For the second data, The first in normal data m The components in each key dimension This is the offset. The offset threshold, jointly determined by detection accuracy and attitude tolerance, can cover sub-pixel level translation, minute rotation, or slight scale changes, while remaining unchanged for non-critical dimensions. This type of sample mainly simulates state offsets caused by sensor calibration errors, visual positioning errors, or assembly alignment errors.

[0042] S203, perform mixed weighted interpolation on multiple normal data to obtain the third data.

[0043] Finally, sample hybrid interpolation is constructed. In one embodiment, multiple normal data are hybridized to obtain third data, including: randomly selecting two normal data points from the normal samples of the same workstation and the same assembly task, and weighting the two selected normal data points by a weighting factor to obtain the third data; the weighting factor ranges from 0.2 to 0.8. The two randomly selected normal data points can be representative of two work points or two batches of incoming parts, and are weighted according to the weighting factor after aligning the coordinate system. Interpolation yields:

[0044] (4); in, This is the third data point. For the first A normal data point, For the first A normal data point, This represents the minimum value of the action. This represents the maximum value of the action. and The interpolation strength is set according to the differences between tasks and the desired interpolation intensity, for example, only taking values ​​in the range of 0.2–0.8 to avoid degenerate to the original sample. After interpolation, the attitude and force components are physically verified to ensure that no state that significantly violates the mechanism's motion range or safety boundaries is generated. Optionally, multiple mixed interpolation operations can be performed.

[0045] S204, the first data, the second data, and the third data are identified as abnormal data.

[0046] The first, second, and third data obtained from the above three types of construction are collectively referred to as abnormal data. .

[0047] Optionally, for each abnormal data Call the current version of the assembly strategy module Output the assembly action corresponding to this state. : (5); in, For assembly actions corresponding to abnormal data, For plug-in speed commands, For single-step displacement or feed step size, For the target insertion force or force threshold, This refers to the allowed number of fine-tuning steps. If the strategy module itself provides the action as continuous parameters, amplitude and velocity constraints can be applied after the output. This is confined within the safe limits allowed by the robotic arm and gripper. The final result is anomalous data – assembly action pairs – used for training. This is to support subsequent traceability and quality verification.

[0048] It should be noted that the parameter range of the abnormal data construction corresponds to the on-site sensing accuracy and attitude tolerance, and all abnormal data construction and linkage sampling record the source, parameters and time information for audit traceability.

[0049] The distinction between normal and abnormal data is based on inspection records and assembly tolerances, and can be updated in conjunction with production rhythm and first-time assembly success rate thresholds.

[0050] S103 calculates the assembly quality assessment score for performing assembly actions under normal data through multiple assessment models, and determines the assessment deviation risk value of normal samples based on the minimum assembly quality assessment score and the reference quality index corresponding to the normal samples; each assessment model focuses on different key factors when scoring.

[0051] In flexible assembly environments, factors such as changes in illumination, reflections, slight positional and orientation shifts, and fixture deformation all affect camera images and pose sensor signals, causing the same physical disturbance to exhibit highly correlated changes across different sensing channels. If only a single evaluation model is used to score assembly quality, and this model primarily relies on a type of susceptible signal, it is prone to inconsistent misjudgments of the entire batch of assembly records under abnormal operating conditions. To address this, this invention introduces multiple evaluation models, allowing each model to have a physical division of labor, focusing on key factors such as attitude shifts, velocity disturbances, abnormal insertion forces, or changes in texture reflection, ensuring that each evaluation model has an independent convergence objective.

[0052] With the first k Taking one evaluation model as an example, in the offline phase, normal data and corresponding assembly actions are input into the first... k The evaluation model yields the... k Each evaluation model is used for each normal data-assembly action pair. Provide a predicted score , That is, the first k An evaluation model for normal data Next, perform assembly actions. Assembly quality assessment score.

[0053] To ensure robust values ​​upon deployment, this embodiment employs conservative minimum value aggregation during both the training and deployment phases. This is used as the final score for that state-action. The sample set of data determined to be normal based on inspection records and assembly tolerances. The risk minimization objective is adopted as shown in the following formula:

[0054] (6); in, This represents the risk value of assessment bias for normal samples; To evaluate the conservative score of the model cluster for state-action pairs, i.e. the minimum assembly quality assessment score, This is a reference quality index derived from the actual assembly results. It is achieved by minimizing... The conservative score given by the evaluation model under normal operating conditions, such as stable lighting and part dimensions within tolerance range. Try to match the actual assembly effect as closely as possible. This ensures that, under normal production conditions, neither excessive conservatism nor overestimation of risk is achieved.

[0055] The above formula means that under normal working conditions where the lighting is stable and the size and orientation of the parts are within tolerance, each evaluation model is required to give a score that is highly consistent with the actual assembly result, so as to form a reliable normal baseline for multiple paths and avoid sacrificing basic accuracy for the sake of division of labor.

[0056] Minimum aggregation is used in both the training objective and deployment inference phases to maintain consistency between the objective and deployment and to avoid overestimation caused by high uncertainty.

[0057] S104. Calculate the input sensitivity vector for each evaluation model using out-of-state data, and calculate the correlation penalty loss value based on the input sensitivity vector of each evaluation model. The input sensitivity vector represents the magnitude of change in the assembly quality evaluation score output by the evaluation model when out-of-state data changes.

[0058] This embodiment uses anomalous data as the data source for anomalous intervals, feeding it into the anomalous training branch of a multi-assessment model. This encourages different assessment models to form complementary points of interest under conditions such as changes in illumination, partial occlusion, and slight deformation, achieving stable coverage of out-of-distribution conditions. For samples classified as anomalous risk intervals, this embodiment does not directly use assembly quality indicators as optimization targets. Instead, it follows the principle of ensuring that the scores approximate the true quality in the normal interval, while forming dispersed and conservative values ​​through multi-assessment collaboration in the anomalous interval.

[0059] This embodiment uses multiple evaluation models and applies decorrelation constraints to their input sensitivity to reduce the correlation between the input sensitivity vectors of each evaluation model. This allows different evaluation models to form a functional division in a physical sense, focusing on key factors such as attitude deviation, abnormal insertion force, or changes in texture reflection. Then, a final score (correlation penalty loss value) is formed through conservative aggregation, thereby improving the ability to identify and tolerate abnormal working conditions without changing the deployment link.

[0060] In one embodiment, the input sensitivity vector for each evaluation model is calculated using out-of-state data, including: for any evaluation model, determining the partial derivative of the assembly quality evaluation score output by the evaluation model with respect to the input out-of-state data as the input sensitivity vector.

[0061] To highlight the different roles of each evaluation model under abnormal operating conditions, this invention introduces input sensitivity decorrelation constraints in the abnormal risk interval. Let the joint input vector be... This includes pre-processed workstation images, end-effector pose parameters, and force-position measurements during the insertion process. k The input sensitivity vector of an evaluation model is defined as the partial derivative of the assembly quality evaluation score with respect to the abnormal input data, i.e.:

[0062] (7); in, For the first k The input sensitivity vector of an evaluation model. Each component characterizes the magnitude of change in the evaluation model's score when a certain type of physical quantity undergoes a small change. For example, the component corresponding to image brightness reflects the device's sensitivity to reflections and illuminance fluctuations, the component corresponding to attitude angle reflects its sensitivity to insertion direction deviations, and the component corresponding to contact force reflects its sensitivity to excessive insertion force or interference collisions.

[0063] In one embodiment, calculating the correlation penalty loss value based on the input sensitivity vector of each evaluation model includes: calculating the correlation between any two evaluation models based on the input sensitivity vector of each evaluation model; and calculating the correlation penalty loss value based on the correlation between any two evaluation models.

[0064] Specifically, the correlation between the input sensitivity vectors of different evaluation models is constrained using the Pearson correlation coefficient, denoted as... For the first i , No. j The correlation between the input sensitivity vectors of the evaluation model is assessed, and an upper limit on the allowed correlation is preset. Correlation penalty loss value The calculation formula is: (8); in, Indicates the first i The evaluation model and the first j The correlation between the evaluation models This is the upper limit of relevance. .

[0065] This embodiment achieves this by constraining the correlation between the input sensitivity vectors of the joint input to not exceed a preset correlation upper limit, so as to avoid all evaluation models relying on the same type of easily disturbed physical quantity.

[0066] S105, the evaluation bias risk value and the correlation penalty loss value are weighted to obtain the target loss value, and each evaluation model is trained using the target loss value to obtain the target evaluation model.

[0067] The evaluation bias risk value and the correlation penalty loss value are weighted and merged into the overall training objective. When the correlation between a pair of evaluation models does not exceed a certain threshold... When the correlation is zero, it does not affect training; when the correlation exceeds... As the threshold increases, the corresponding loss term grows larger, driving the training process to automatically adjust the sensitivity of the two evaluation models to different physical quantities, causing them to converge near or below the threshold. The engineering implication is that, while maintaining the normal accuracy of each evaluation model, it deliberately avoids all evaluation models being "highly sensitive" to the same type of perturbation, such as all heavily relying on the brightness of the connected region or a certain pose component. Instead, it encourages some evaluation models to primarily focus on pose and trajectory deviations, while others focus on force-potential anomalies or changes in texture reflection. In this way, when the overall lighting is dim or a certain type of sensor drifts briefly, at least some evaluation models can still maintain reliable judgments through other physical quantities, reducing the probability of consistent misjudgments from a cluster perspective.

[0068] The training and evaluation of this invention cover three types of operations: grasping, positioning, and insertion, and use production rhythm fluctuations, downtime troubleshooting time, and first-time assembly success rate as acceptance indicators. Furthermore, the training employs a single backpropagation to jointly optimize strategy parameters and evaluation model parameters, with regularization constraints only taking effect during the training period, and the original inference chain of the camera, encoder, and control system retained on the deployment side.

[0069] This embodiment combines accurate evaluation of normal samples with conservative processing of abnormal data, enabling the robotic arm to have stable extrapolation and recovery capabilities under abnormal working conditions such as light fluctuations, batch differences, and slight deformation.

[0070] S106: The target evaluation model is used to evaluate the value of the current state's actions and determine the optimal action of the robotic arm.

[0071] First, obtain the current assembly state. The candidate motion space (including workstation images, pose and force data) is used to determine the current assembly state. With candidate actions The inputs are fed into the evaluation model that has processed multiple targets. China (among them) (As the index of the target evaluation model), the assembly quality evaluation score corresponding to each target evaluation model is calculated; subsequently, based on the consistency strategy maintained during the training and deployment phases of this invention, the conservative value score of each candidate action is calculated. By taking the minimum value of each model's output, the potential overestimation of the score under abnormal operating conditions is suppressed, ensuring the robustness of the value assessment. Then, the action that maximizes the conservative value score is searched from the candidate action space, and this action is determined as the optimal action. (in (This represents the selected optimal action); finally, considering the preset equipment tolerances and physical safety constraints, the optimal action is... The system performs interval clamping to limit the insertion speed, displacement step size, and force threshold within the legal and safe range allowed by the robotic arm and gripper, and finally outputs instructions to control the robotic arm's execution.

[0072] Figure 2 This is a schematic diagram illustrating the state distribution of offline data and online operating conditions. Figure 2 Figure (a) shows the state distribution for a moderately disturbed scenario. Figure 2 Figure (b) shows the state distribution of the channel transport and obstacle avoidance scenario; Figure 3 The differences between the traditional unified pressure reduction strategy and the proposed method in terms of action selection and extrapolation effect under abnormal operating conditions were compared. Figure 4 This is a schematic diagram of the structure and data flow of a flexible assembly offline training and collaborative evaluation system.

[0073] Figure 2 This visually demonstrates the misalignment between the learning strategy and the state distribution of offline data in flexible assembly conditions. Figure 2 Figure (a) in the figure corresponds to the overlap and offset of the distribution of normal data and actual operating conditions under a moderate disturbance scenario, while Figure 2 Figure (b) shows the severe out-of-distribution phenomenon that occurs in complex environments such as occlusion and extreme size deviation, proving that the model trained on normal data alone is at risk of failure when faced with abnormal online conditions, thus demonstrating the necessity of constructing abnormal data in this invention. Figure 3 The extrapolation differences between traditional strategies and the method of this invention are compared using value estimation curves. The horizontal axis in the figure represents the space formed by states and actions. The further to the right, the greater the difference between the state-space pair and common normal data, indicating anomalous data. Traditional reinforcement learning methods show value assessment values ​​that conform to value trends in the in-distribution data (normal data) region. However, once entering the out-of-distribution data (abnormal data) region, the assessment value will rise uncontrollably. This is because the assessment model will make unreliable extrapolations for unfamiliar regions, giving overly optimistic high valuations. In order to solve the disastrous consequences of traditional reinforcement learning methods, traditional techniques have proposed a traditional conservative reinforcement learning method, which applies a uniform penalty to all out-of-distribution data, assuming that all out-of-distribution data is bad. However, in practical applications, some out-of-distribution data may indeed be high-reward (just not collected in historical data). This over-penalty can easily lead to over-penalizing potentially feasible high-reward schemes. In contrast, this method uses conservative aggregation of multiple assessment models to improve the identification and extrapolation ability of effective actions in out-of-distribution data while suppressing the risk of overestimation. Figure 4 This details the overall architecture and data flow of the method of the present invention, by obtaining normal samples from the offline dataset. ( This represents normal data. For assembly actions, As a reward for feedback, For the next normal data, perform error minimization training on known data, and simultaneously process the constructed outlier samples. ( For constructed anomalous data, For the corresponding actions of linkage sampling, multiple evaluation models are input respectively. Input sensitivity decorrelation technology is used to induce each model to form a functional division. Finally, the robust value evaluation score is output through minimum value aggregation to guide control decisions.

[0074] In one embodiment, the present invention also provides a control method for a flexible assembly robotic arm based on abnormal working conditions, the embodiment including: S301: Sample routine samples from historical data and complete quality verification. Collect historical assembly data and inspection records, combine assembly tolerances and first-time assembly success rate thresholds to determine routine samples, and label the result of each assembly (qualified, reworked, or scrapped).

[0075] Routine samples include routine data and corresponding assembly actions, extracted from historical assembly data and small-batch trial production data that have undergone quality verification. s Organized uniformly into image features Pose parameters With force measurement The state vector is composed of the workstation number, task type and batch information for each sample.

[0076] S302 uses random noise, feature perturbation, or sample mixing interpolation to construct abnormal samples.

[0077] Abnormal samples include anomalous data and corresponding assembly actions. Without introducing an environmental model, anomalous data construction is performed on normal data using three methods: random noise perturbation, key feature perturbation, and sample mixture interpolation. First, zero-mean Gaussian noise is superimposed on selected normal data and interval clamping is applied to obtain… Secondly, in the set of key dimensions related to positioning and alignment. M Apply a controlled offset to obtain For two normal data points of the same task, weight factors are used to determine their approximations. Interpolate and perform physical consistency checks to obtain The ranges of various construction parameters correspond one-to-one with the field sensing accuracy and attitude tolerance, and the construction method and parameters are written into the metadata.

[0078] Abnormal data is sampled using the current strategy to obtain the corresponding assembly actions.

[0079] S303 deploys a parallel evaluation model cluster structure to complete interface and verification.

[0080] Within the overall architecture of the flexible assembly robotic arm's abnormal operating condition response system, the number of evaluation models K is selected, and a system is established. K A parallel evaluation model is established, with an independent target cache and parameter storage structure configured for each model; an information cache is also established to record assembly state-action pairs. Reference quality indicators And labels for normal and abnormal operating conditions.

[0081] S304 employs a training process that combines collaborative evaluation and conservative aggregation to create a reproducible model version. Each model converges to its respective objective on normal samples, stabilizing baseline performance.

[0082] In normal samples Each evaluation model is trained independently according to the convergence objective, so that each evaluation model gives a score under normal operating conditions where illumination, size, and pose all meet the tolerances. Compared with reference quality indicators The deviation does not exceed the preset range, forming a consistent baseline with basic accuracy for multiple paths.

[0083] Specifically, minimum value aggregation is uniformly adopted in both the training objective and the deployment inference phase. As an assembly quality score, it ensures that the training objectives are consistent with the online value rules, and suppresses the amplified impact of high uncertainty conditions on the score. Under simulation and small-batch conditions, the response time, score fluctuation and first-time assembly success rate were verified for three scenarios: illumination change, partial occlusion and slight deformation. After meeting the standards, the application scope was gradually expanded.

[0084] For anomalous data, only the sample proportion and working condition type are recorded, serving as an interface for subsequent anomalous sample construction and collaborative evaluation training.

[0085] S305 performs input sensitivity decorrelation training in abnormal regions to form functional division of labor.

[0086] For each constructed anomalous data The current assembly strategy module is invoked to generate the corresponding assembly action. Safety constraints are imposed on the amplitude and speed of the movements to form state-action sample pairs. The sample pairs and their metadata are written into the abnormal risk interval sample library. On the abnormal risk interval sample consisting of abnormal state data and corresponding assembly actions, the evaluation model is calculated for the joint input. Input sensitivity vector ,Will The formal relevance metric is embedded in the overall training objective, which constrains the sensitivity relevance, making some evaluation models more sensitive to pose shifts, some more sensitive to insertion force or velocity perturbations, and some more sensitive to texture reflection changes, thereby forming a functional division of labor and reducing the risk of the cluster making consistent misjudgments of a single perturbation pattern.

[0087] During the evaluation and backup phase, minimum value aggregation is implemented to output stable value values ​​and action suggestions, complete the value closure loop, and conduct simulation and small-batch verification. The production rhythm, product yield and abnormal response are evaluated, parameters are solidified and version management is completed, gray-scale rollout and monitoring strategies are implemented, and a backtracking and continuous improvement mechanism is established.

[0088] The simulation and small-batch verification process includes at least three scenarios: contrast change, partial occlusion, and slight deformation. The response time, variance, and first-time success rate are evaluated, and grayscale scaling is completed accordingly.

[0089] The method provided by this invention can complete the integration of training and evaluation without changing the online inference link, ensuring that there is no increase in the computation and communication overhead during the deployment phase.

[0090] This invention addresses disturbances present in flexible assembly lines, such as batch variations, light reflection, slight shifts in work posture, and fixture deformation. It establishes an engineering-implementable offline training and evaluation method to ensure accurate strategy values ​​under normal conditions and extrapolation capabilities under abnormal conditions. To this end, a model-free abnormal state construction is proposed and sampled in conjunction with the current strategy to form training samples covering key abnormal disturbances. During the evaluation phase, independent convergence targets are set for the evaluation model cluster, and functional division is achieved through decorrelation with input sensitivity. In the inference and backup phases, conservative minimum aggregation is employed to suppress the expansion of uncertainty. Through this design, stable judgment and reasonable action selection for unseen disturbances are achieved, balancing production rhythm and product yield, and facilitating integration and maintenance with vision inspection and trajectory control modules.

[0091] Compared with the prior art, the present invention has the following advantages: Compared with existing technologies, this invention has the following advantages in productization and task optimization. Addressing the frequent occurrences of illumination changes, reflections, slight deformations, and pose deviations in flexible assembly, this invention integrates states and actions into the training and evaluation objectives, forming a stable set of values ​​and reasonable extrapolation capabilities for abnormal situations. For the production side, this translates to reduced production rhythm fluctuations, shorter downtime for troubleshooting, and improved first-pass yield and first-time assembly success rates.

[0092] The data layer of this invention employs model-free anomaly state construction and strategy-linked sampling, eliminating the need for additional dynamic models or generators. The path is simple, parameters are controllable, and traceability is excellent, facilitating rapid reuse and portability in multi-category, small-batch, and frequently switching production organizations. For operations and maintenance personnel, the parameter space and parameter adjustment steps are clear, and the causal relationship between training data and results is easy to audit.

[0093] This invention introduces independent target and input sensitivity decorrelation into the evaluation model cluster on the evaluation side, enabling the evaluation models to functionally specialize and focus on key elements such as pose, velocity, and texture reflection. Furthermore, it suppresses the expansion of uncertainty through conservative minimum value aggregation. This combination offers controllable computational overhead during deployment, simple integration with existing vision detection and trajectory control modules, and is suitable for stable operation under industrial computing resource conditions, balancing performance and maintainability.

[0094] The subject executing the method of this invention can be a server, which can be a server set up on a business platform, or a device such as a desktop computer or laptop computer that can execute the solution of this specification.

[0095] When applying the control method for the flexible assembly robot arm based on abnormal working conditions provided in this manual, it is not necessary to consider... Figure 1 The steps shown are executed in sequence. The specific execution order of each step can be determined as needed, and this manual does not impose any restrictions on it.

[0096] The above describes a control method for a flexible assembly robot arm based on abnormal working conditions, provided by one or more embodiments of this specification. Based on the same concept, this specification also provides a corresponding control device for a flexible assembly robot arm based on abnormal working conditions, which includes: The acquisition module is used to acquire normal samples in flexible assembly tasks; normal samples include normal data and corresponding assembly actions; normal data represents the state data of parts with dimensions within tolerance range under stable illumination. The data generation module is used to obtain abnormal data by offsetting or interpolating normal data by a preset amplitude without violating equipment tolerances and safety constraints; the abnormal data includes states not covered in the normal samples in the flexible assembly task; The first evaluation module is used to calculate the assembly quality evaluation score of the assembly action performed under normal data through multiple evaluation models, and to determine the evaluation deviation risk value of the normal sample based on the minimum assembly quality evaluation score and the reference quality index corresponding to the normal sample; each evaluation model focuses on different key factors when scoring. The second evaluation module is used to calculate the input sensitivity vector of each evaluation model using out-of-state data, and to calculate the correlation penalty loss value based on the input sensitivity vector of each evaluation model; the input sensitivity vector represents the magnitude of change in the assembly quality evaluation score output by the evaluation model when the out-of-state data changes. The weighting module is used to weight the evaluation bias risk value and the correlation penalty loss value to obtain the target loss value, and to train each evaluation model with the target loss value to obtain the target evaluation model; The control module is used to evaluate the value of the current state's actions using a target evaluation model, and to determine the optimal action for the robotic arm.

[0097] Specific limitations regarding the control device for the flexible assembly robot arm based on abnormal working conditions can be found in the limitations of the control method for the flexible assembly robot arm based on abnormal working conditions mentioned above, and will not be repeated here. Each module in the control device for the flexible assembly robot arm based on abnormal working conditions described above can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of the processor, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0098] This specification also provides a computer-readable storage medium storing a computer program that can be used to execute the above-described... Figure 1 A control method for a flexible assembly robot arm based on abnormal working conditions is provided.

[0099] This specification also provides a control device for flexible assembly, including a processor and a memory, wherein the memory stores program instructions, and the processor executes the program instructions to achieve the above. Figure 1 The method provided is a control method for a flexible assembly robot arm based on abnormal working conditions; and it communicates with a camera, posture sensor and end effector to output control commands and evaluation results compatible with the trajectory control module.

[0100] Optionally, the control device provides a version management and rollback interface to support parameter fixing, grayscale scaling, and migration across multiple models.

[0101] Optionally, the control device maintains a strategy of aligning the minimum value aggregation with the training objective during the deployment phase to reduce misjudgments caused by uncertainty.

[0102] This instruction manual also provides Figure 5 The schematic diagram of the computer device shown is as follows: Figure 5At the hardware level, the computer device includes a processor, internal bus, network interface, memory, and non-volatile memory, and may also include other hardware required for business operations. The processor reads the corresponding computer program from the non-volatile memory into memory and then runs it to achieve the above-mentioned functions. Figure 1 A control method for a flexible assembly robot arm based on abnormal working conditions is provided.

[0103] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0104] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A control method for a flexible assembly robotic arm based on abnormal working conditions, characterized in that, The method includes: Acquire normal samples in flexible assembly tasks; normal samples include normal data and corresponding assembly actions; normal data represents the state data of parts with dimensions within tolerance range under stable illumination; Without violating equipment tolerances and safety constraints, abnormal data is obtained by offsetting or interpolating normal data by a preset amplitude; abnormal data includes states not covered in the normal samples of flexible assembly tasks. The assembly quality assessment score for performing assembly actions under normal data is calculated using multiple assessment models. Based on the minimum assembly quality assessment score and the reference quality index corresponding to the normal sample, the assessment deviation risk value of the normal sample is determined. Each assessment model focuses on different key factors when scoring. Using out-of-state data, the input sensitivity vector of each evaluation model is calculated, and the correlation penalty loss value is calculated based on the input sensitivity vector of each evaluation model; the input sensitivity vector represents the magnitude of change in the assembly quality evaluation score output by the evaluation model when out-of-state data changes. The target loss value is obtained by weighting the evaluation bias risk value and the correlation penalty loss value, and then training each evaluation model with the target loss value to obtain the target evaluation model. The optimal action for the robotic arm is determined by evaluating the value of the current action using a target evaluation model.

2. The method according to claim 1, characterized in that, Normal samples include multiple normal data sets; normal data includes workstation images, robotic arm end-effector poses, and attitude data and force-position measurement data of key contact points; abnormal data is obtained by performing preset amplitude offsets or interpolation along the normal data, including: For any given normal data, zero-mean noise is added to each type of data in the normal data to obtain the first data; For any given set of normal data, an offset is added to the key features in the normal data that are related to the positioning error to obtain the second set of data; A third set of data is obtained by performing mixed-weight interpolation on multiple normal data sets. The first, second, and third data points are identified as abnormal data.

3. The method according to claim 2, characterized in that, The method further includes: An interval clamp is applied to the first data to limit the image intensity in the first data to the effective grayscale range and to limit the attitude data and force measurement data to the upper and lower limits allowed by the device tolerance.

4. The method according to claim 2, characterized in that, A third set of data is obtained by interpolating multiple normal data sets, including: Two normal data points are randomly selected from the normal samples of the same workstation and the same assembly task. The two selected normal data points are weighted by a weighting factor to obtain a third data point. The value of the weighting factor is between 0.2 and 0.

8.

5. The method according to claim 1, characterized in that, Using anomalous data, calculate the input sensitivity vector for each evaluation model, including: For any evaluation model, the partial derivative of the assembly quality evaluation score output by the evaluation model with respect to the input abnormal data is determined as the input sensitivity vector.

6. The method according to claim 1, characterized in that, Based on the input sensitivity vector of each evaluation model, calculate the relevance penalty loss value, including: Calculate the correlation between any two evaluation models based on the input sensitivity vector of each evaluation model; Calculate the correlation penalty loss value based on the correlation between any two evaluation models.

7. The method according to claim 6, characterized in that, Correlation penalty loss value The calculation formula is: ; in, Indicates the first i The evaluation model and the first j The correlation between the evaluation models This is the upper limit of relevance. .

8. A control device for a flexible assembly robotic arm based on abnormal working conditions, characterized in that, The device includes: The acquisition module is used to acquire normal samples in flexible assembly tasks; normal samples include normal data and corresponding assembly actions; normal data represents the state data of parts with dimensions within tolerance range under stable illumination. The data generation module is used to obtain abnormal data by offsetting or interpolating normal data by a preset amplitude without violating equipment tolerances and safety constraints; the abnormal data includes states not covered in the normal samples in the flexible assembly task; The first evaluation module is used to calculate the assembly quality evaluation score of the assembly action performed under normal data through multiple evaluation models, and to determine the evaluation deviation risk value of the normal sample based on the minimum assembly quality evaluation score and the reference quality index corresponding to the normal sample; each evaluation model focuses on different key factors when scoring. The second evaluation module is used to calculate the input sensitivity vector of each evaluation model using out-of-state data, and to calculate the correlation penalty loss value based on the input sensitivity vector of each evaluation model; the input sensitivity vector represents the magnitude of change in the assembly quality evaluation score output by the evaluation model when the out-of-state data changes. The weighting module is used to weight the evaluation bias risk value and the correlation penalty loss value to obtain the target loss value, and to train each evaluation model with the target loss value to obtain the target evaluation model; The control module is used to evaluate the value of the current state's actions using a target evaluation model, and to determine the optimal action for the robotic arm.