Weight dynamic detection method and system based on robot carrying

By constructing an intelligent adaptive sampling model and a parallel computing framework, the robot's dynamic weight detection method is optimized, the impact of instrument sensitivity changes on the detection results is solved, and efficient and accurate dynamic weight detection is achieved to adapt to different detection environments and task requirements.

CN120593876APending Publication Date: 2025-09-05CHANGZHOU ACCURATE WEIGHT +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510827142.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing dynamic detection technology of weights based on robot handling ignores the dynamic measurement uncertainty caused by changes in instrument sensitivity, which challenges the accuracy of the detection results. In addition, the existing technology mainly focuses on the robot's motion control and weight handling path planning, but ignores the impact of changes in instrument sensitivity.

Method used

An intelligent adaptive sampling model is constructed, and the sampling strategy is optimized through Markov decision process and reinforcement learning algorithm. The estimation accuracy is predicted by combining Gaussian process regression model, the sampling probability distribution is dynamically adjusted, the sensitivity change characteristics of the instrument are evaluated using parallel computing framework, and the optimal motion parameters are selected for dynamic detection of weights.

Benefits of technology

It achieves the goal of minimizing the number of sampling samples while ensuring detection accuracy, improving detection efficiency and accuracy, being able to adapt to different detection environments and task requirements, optimizing robot handling operation parameters, and improving the accuracy and reliability of robot-assisted measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120593876A_ABST
    Figure CN120593876A_ABST
Patent Text Reader

Abstract

The invention provides a weight dynamic detection method and system based on robot carrying, and relates to the technical field of detection, and the method comprises the steps: generating a motion parameter sample of a robot carrying weight through a constructed intelligent adaptive sampling model, and controlling the robot to carry a to-be-detected weight according to the motion parameter sample, meanwhile, an indicating value sample of the instrument to the weight is collected, and a sensitivity coefficient interval of instrument sensitivity relative to robot motion parameters is obtained; according to the sensitivity coefficient interval, the uncertainty of the robot motion parameters is transmitted to an instrument indicating value through interval arithmetic, the confidence interval of the instrument indicating value is obtained, the robot motion parameter combination with the minimum difference value between the upper limit and the lower limit of the confidence interval is selected as the optimal motion parameter, and the robot is controlled to carry weights according to the optimal motion parameter for dynamic detection. And taking the confidence interval of the instrument indicating value as the uncertainty interval of the weight magnitude caused by considering the instrument sensitivity, and completing the dynamic detection and uncertainty evaluation of the weight.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of detection technology, and in particular to a method and system for dynamic detection of weights based on robot handling. Background Art

[0002] Dynamic weight detection is a key technology in metrology and verification, widely used in industrial production, scientific research, trade settlement, and other fields. Traditional dynamic weight detection relies primarily on manual operation, which suffers from low detection efficiency, high workload, and unstable test results. To improve the automation level and detection accuracy of dynamic weight detection, researchers have begun to introduce robotics technology to achieve automated handling and dynamic detection of weights.

[0003] However, existing dynamic weight detection technologies based on robot handling still have some shortcomings. For one thing, existing technologies primarily focus on robot motion control and weight handling path planning, while neglecting the analysis of dynamic measurement uncertainty caused by variations in instrument sensitivity during the detection process. Furthermore, fluctuations in robot motion parameters and shifts in the weight's center of mass lead to dynamic variations in instrument sensitivity, posing challenges to the accuracy of detection results. Summary of the Invention

[0004] The embodiments of the present invention provide a method and system for dynamic detection of weights based on robot handling, which can at least solve some of the problems existing in the prior art.

[0005] A first aspect of an embodiment of the present invention provides a method for dynamic detection of weights handled by a robot, comprising:

[0006] An intelligent adaptive sampling model is constructed, modeling the sampling process as a Markov decision process. The number of samples required to achieve a predetermined accuracy is used as the optimization objective, and the sampling probability distribution parameters are used as optimization variables. A reinforcement learning algorithm is used to adaptively optimize the sampling strategy. Based on the statistical characteristics of the samples already drawn, Gaussian process regression is used to learn the relationship between the weight mass, robot motion parameters, and instrument indications from the samples. This generates an approximate estimate of the sampling probability distribution, dynamically adjusting the sampling probability distribution to minimize the number of samples required to achieve the predetermined accuracy requirements of the sensitivity analysis.

[0007] The constructed intelligent adaptive sampling model is used to generate motion parameter samples of a robot handling weights. The robot is then controlled to carry the weights to be measured according to the motion parameter samples. Simultaneously, samples of the instrument's indications of the weights are collected. The collected motion parameter samples and indication samples are input into a parallel computing framework. The mathematical expectation and variance of the instrument's indications under the motion parameter samples are calculated in parallel on a high-performance computing platform. The variation characteristics of the instrument's sensitivity under different motion parameter combinations are evaluated, and the sensitivity coefficient range of the instrument's sensitivity relative to the robot's motion parameters is obtained.

[0008] According to the sensitivity coefficient interval, interval arithmetic is used to transfer the uncertainty of the robot motion parameters to the instrument indication to obtain the confidence interval of the instrument indication. The robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval is selected as the optimal motion parameter. The robot is controlled to carry the weight according to the optimal motion parameters for dynamic detection. The confidence interval of the instrument indication is used as the uncertainty interval of the weight value caused by the instrument sensitivity to complete the dynamic detection and uncertainty assessment of the weight.

[0009] In an optional embodiment,

[0010] An intelligent adaptive sampling model is constructed, and the sampling process is modeled as a Markov decision process. The number of samples required to achieve a predetermined accuracy is used as the optimization target, and the sampling probability distribution parameters are used as optimization variables. The adaptive optimization of the sampling strategy is achieved through a reinforcement learning algorithm, including:

[0011] The sampling process is modeled as a Markov decision process, defining the state space, action space, state transition probability, reward function, discount factor, and policy function. The state space is constructed based on the statistical properties of the samples already drawn and the parameters of the current sampling probability distribution. The adjustment strategy of the sampling probability distribution parameters is used as the action space. A reward function is designed that comprehensively considers the number of samples and estimation accuracy. The discount factor is used to control the weight of future rewards. The policy function represents the probability distribution of selecting an action in a given state.

[0012] A reinforcement learning algorithm based on policy gradient is used. The gradient of the expected value of the cumulative reward with respect to the policy function parameters is calculated through Monte Carlo estimation. The policy function parameters are iteratively updated using the gradient ascent method. Through multiple iterations, the policy function gradually converges to the optimal policy.

[0013] At the same time, the Gaussian process regression model is integrated, which takes the statistical characteristics of the extracted samples and the sampling probability distribution parameters as input and the estimation accuracy as output. The nonlinear mapping relationship between sample characteristics and estimation accuracy is learned to predict the estimation accuracy under different sampling probability distributions.

[0014] In an optional embodiment,

[0015] Adopting a reinforcement learning algorithm based on policy gradient, Monte Carlo estimation is used to calculate the gradient of the expected value of the cumulative reward with respect to the policy function parameters. The policy function parameters are iteratively updated using the gradient ascent method. Through multiple iterations, the policy function gradually converges to the optimal policy, including:

[0016] Construct a cumulative reward function to evaluate the long-term benefits of the state-action trajectory. The cumulative reward function consists of two parts: an immediate reward and a discount factor. The immediate reward is constructed based on the robot's dynamic detection performance indicators, including a weighted combination of multiple optimization objectives such as detection accuracy, detection efficiency, and energy consumption. The discount factor is used to balance the immediate reward and long-term reward and control the reward decay rate.

[0017] The Monte Carlo estimation algorithm is used to approximate the policy gradient. By executing the current policy function, multiple state-action trajectory samples are generated. The cumulative reward of each trajectory sample is calculated, and the policy gradient is estimated using the cumulative reward. The policy gradient is calculated as the cumulative reward multiplied by the gradient of the state-action log probability with respect to the policy function parameters. The policy gradients estimated from multiple trajectories are averaged.

[0018] The Adam optimization algorithm is used to adaptively iteratively update the policy function parameters. According to the estimated policy gradient, the policy function parameters are optimized through multiple iterations with an adaptively adjusted learning rate as the step size, so that it can accelerate convergence to the optimal strategy while satisfying the policy improvement theory. The trust region is introduced to control the step size of the policy update to improve the sample efficiency and training stability of the algorithm, and finally the optimal strategy that can autonomously adapt to different detection environments and task requirements is obtained.

[0019] In an optional embodiment,

[0020] The constructed intelligent adaptive sampling model is used to generate motion parameter samples of the robot carrying the weight. The robot is controlled to carry the weight to be tested according to the motion parameter samples. At the same time, the instrument's indication samples of the weight are collected. The collected motion parameter samples and indication samples are input into the parallel computing framework. The mathematical expectation and variance of the instrument indication under the motion parameter samples are calculated in parallel on the high-performance computing platform. The variation characteristics of the instrument sensitivity under different motion parameter combinations are evaluated. The sensitivity coefficient range of the instrument sensitivity relative to the robot motion parameters is obtained, including:

[0021] A parallel computing framework is constructed to process the collected motion parameter samples and instrument indication samples. The parallel computing framework includes a data distribution module, a parallel computing module, and a result summary module. The data distribution module distributes the motion parameter samples and indication samples to different computing nodes for parallel processing. The parallel computing module calculates the mathematical expectation and variance of all indication samples corresponding to the j-th motion parameter sample on each computing node. The result summary module summarizes the calculation results of each computing node to obtain the complete statistical characteristics of the instrument indication.

[0022] Based on the statistical characteristics of the instrument indications obtained by parallel computing, the finite difference method is used to numerically estimate the instrument sensitivity coefficient. For the j-th motion parameter sample, by applying positive and negative disturbances to its i-th motion parameter, the difference between the mathematical expectations of the instrument indications under positive and negative disturbances is calculated and divided by twice the disturbance amount as the sensitivity coefficient estimate corresponding to its i-th motion parameter.

[0023] Based on the sensitivity coefficient estimation of each motion parameter sample, the sensitivity coefficient interval of each motion parameter is constructed. The upper limit of the sensitivity coefficient interval takes the maximum value of the sensitivity coefficient estimation corresponding to all motion parameter samples, and the lower limit takes the minimum value, thereby quantitatively evaluating the variation range of the instrument sensitivity relative to each motion parameter.

[0024] In an optional embodiment,

[0025] Determining the sensitivity coefficient interval includes:

[0026] ;

[0027] ;

[0028] in, represents the estimated value of the sensitivity coefficient of the i-th motion parameter corresponding to the j-th motion parameter sample, Indicates that a positive perturbation is applied to the i-th component of the j-th motion parameter sample After that, the mathematical expectation of the instrument indication is Indicates applying a negative perturbation to the i-th component of the j-th motion parameter sample - After that, the mathematical expectation of the instrument indication is represents the perturbation amount on the i-th motion parameter.

[0029] In an optional embodiment,

[0030] Based on the sensitivity coefficient interval, interval arithmetic is used to transfer the uncertainty of the robot motion parameters to the instrument indication to obtain the confidence interval of the instrument indication. The robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval is selected as the optimal motion parameter. The robot is controlled to carry the weight according to the optimal motion parameter for dynamic detection, including:

[0031] A velocity planning method is used to generate the velocity curve of the robot's motion process. By setting the duration of the acceleration ramp, constant speed, and deceleration stages, as well as the acceleration rate of change in the acceleration ramp and deceleration stages, the generated velocity curve is made smooth and continuous in both time and velocity dimensions. By adjusting the duration ratio of the three stages and the acceleration rate of change, the smoothness and volatility of the velocity curve are optimized, thereby reducing the uncertainty introduced by velocity fluctuations.

[0032] A trajectory planning method is used to generate the displacement curve of the robot's motion process. By using six boundary conditions (starting position, velocity, acceleration, and end position, velocity, and acceleration) as constraints, the six coefficients of trajectory planning are solved so that the generated displacement curve meets the continuity requirements of position, velocity, and acceleration in both time and displacement dimensions, resulting in a motion trajectory with continuous acceleration, smooth velocity, and no sudden displacement changes.

[0033] By constructing a robot kinematic and dynamic model, using the generated velocity curve and displacement curve as input, and using numerical integration and iterative optimization algorithms for simulation analysis, the changes in velocity, acceleration, and torque in the joint space and Cartesian space during the robot's motion are obtained. By setting the optimization objective function, the duration of the three stages of velocity planning and the acceleration change rate parameters as well as the polynomial coefficients of trajectory planning are optimized, so that the velocity, acceleration fluctuations, and joint torque changes of the robot's actual motion process meet the set performance index requirements;

[0034] The robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval is selected as the optimal motion parameter, and the robot is controlled to carry the weight for dynamic detection according to the optimal motion parameter.

[0035] A second aspect of an embodiment of the present invention provides a dynamic detection system for weights handled by a robot, comprising:

[0036] The first unit is used to build an intelligent adaptive sampling model. The sampling process is modeled as a Markov decision process. The number of samples required to achieve a predetermined accuracy is the optimization objective, and the sampling probability distribution parameters are the optimization variables. The sampling strategy is adaptively optimized through a reinforcement learning algorithm. Based on the statistical characteristics of the samples, Gaussian process regression is used to learn the relationship between the weight mass, robot motion parameters, and instrument indications from the samples. An approximate estimate of the sampling probability distribution is generated, and the sampling probability distribution is dynamically adjusted to minimize the number of samples required to achieve the predetermined accuracy requirements of the sensitivity analysis.

[0037] The second unit is used to generate motion parameter samples for a robot handling weights using the constructed intelligent adaptive sampling model. The robot is then controlled to handle the weights to be tested according to the motion parameter samples. Simultaneously, samples of the instrument's indications of the weights are collected. The collected motion parameter samples and indication samples are input into a parallel computing framework. The mathematical expectation and variance of the instrument's indications under the motion parameter samples are then calculated in parallel on a high-performance computing platform. The variation characteristics of the instrument's sensitivity under different motion parameter combinations are evaluated, and the sensitivity coefficient range of the instrument's sensitivity relative to the robot's motion parameters is obtained.

[0038] The third unit is used to transfer the uncertainty of the robot motion parameters to the instrument indication using interval arithmetic based on the sensitivity coefficient interval, obtain the confidence interval of the instrument indication, select the robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval as the optimal motion parameters, control the robot to carry the weight according to the optimal motion parameters for dynamic detection, and use the confidence interval of the instrument indication as the uncertainty interval of the weight value caused by the instrument sensitivity to complete the dynamic detection and uncertainty assessment of the weight.

[0039] According to a third aspect of the embodiments of the present invention,

[0040] An electronic device is provided, comprising:

[0041] processor;

[0042] a memory for storing processor-executable instructions;

[0043] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0044] According to a fourth aspect of the embodiments of the present invention,

[0045] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0046] The intelligent adaptive sampling model proposed in this paper uses a Markov decision process to model the sampling process, optimizes the sampling strategy function using a policy gradient reinforcement learning algorithm, and integrates a Gaussian process regression model to predict estimation accuracy, achieving adaptive optimization of the sampling strategy. Compared with traditional fixed sampling methods, this method can dynamically adjust the sampling probability distribution based on sample characteristics and estimation requirements, minimizing the number of samples while maintaining estimation accuracy, thereby improving sampling efficiency.

[0047] This paper proposes a method for analyzing the impact of robotic handling motion on instrument sensitivity, utilizing an intelligent adaptive sampling model and a parallel computing framework. This method efficiently evaluates the changing characteristics of instrument sensitivity under different motion parameter combinations and quantitatively determines the sensitivity of instrument sensitivity to the robot's motion parameters. This method can guide the optimization of robotic handling parameters, select motion parameter combinations that are insensitive to instrument sensitivity, and improve the accuracy and reliability of robot-assisted metrology. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 Schematic diagram of the flow of a method for dynamic detection of weights based on robot handling according to an embodiment of the present invention;

[0049] Figure 2 Schematic diagram of the structure of a dynamic detection system for weights handled by a robot according to an embodiment of the present invention. DETAILED DESCRIPTION

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0051] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0052] Figure 1 FIG. 1 is a flow chart of a method for dynamic detection of weights handled by a robot according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0053] S101. Build an intelligent adaptive sampling model, modeling the sampling process as a Markov decision process. The optimization objective is the number of samples required to achieve a predetermined accuracy, and the sampling probability distribution parameters are used as optimization variables. Adaptive optimization of the sampling strategy is achieved through a reinforcement learning algorithm. Based on the statistical characteristics of the samples, Gaussian process regression is used to learn the relationship between the weight mass, robot motion parameters, and instrument readings from the samples. This generates an approximate estimate of the sampling probability distribution, dynamically adjusting the sampling probability distribution to minimize the number of samples required to achieve the predetermined accuracy requirements for sensitivity analysis.

[0054] S102. Utilize the constructed intelligent adaptive sampling model to generate motion parameter samples for a robot handling a weight. Control the robot to handle the weight to be tested according to the motion parameter samples. Simultaneously, collect instrument readings of the weight. Input the collected motion parameter and readings into a parallel computing framework. On a high-performance computing platform, parallel calculations are performed on the mathematical expectation and variance of the instrument readings for the motion parameter samples. The variation in instrument sensitivity for different motion parameter combinations is evaluated, resulting in a sensitivity coefficient range for the instrument sensitivity relative to the robot's motion parameters.

[0055] S103. Based on the sensitivity coefficient interval, use interval arithmetic to transfer the uncertainty of the robot's motion parameters to the instrument indication, and obtain the confidence interval of the instrument indication. Select the robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval as the optimal motion parameters. Control the robot to carry the weight according to the optimal motion parameters for dynamic testing. Use the confidence interval of the instrument indication as the uncertainty interval of the weight value caused by the instrument sensitivity to complete the dynamic testing and uncertainty assessment of the weight.

[0056] In an optional embodiment,

[0057] An intelligent adaptive sampling model is constructed, and the sampling process is modeled as a Markov decision process. The number of samples required to achieve a predetermined accuracy is used as the optimization target, and the sampling probability distribution parameters are used as optimization variables. The adaptive optimization of the sampling strategy is achieved through a reinforcement learning algorithm, including:

[0058] The sampling process is modeled as a Markov decision process, defining the state space, action space, state transition probability, reward function, discount factor, and policy function. The state space is constructed based on the statistical properties of the samples already drawn and the parameters of the current sampling probability distribution. The adjustment strategy of the sampling probability distribution parameters is used as the action space. A reward function is designed that comprehensively considers the number of samples and estimation accuracy. The discount factor is used to control the weight of future rewards. The policy function represents the probability distribution of selecting an action in a given state.

[0059] A reinforcement learning algorithm based on policy gradient is used. The gradient of the expected value of the cumulative reward with respect to the policy function parameters is calculated through Monte Carlo estimation. The policy function parameters are iteratively updated using the gradient ascent method. Through multiple iterations, the policy function gradually converges to the optimal policy.

[0060] At the same time, the Gaussian process regression model is integrated, which takes the statistical characteristics of the extracted samples and the sampling probability distribution parameters as input and the estimation accuracy as output. The nonlinear mapping relationship between sample characteristics and estimation accuracy is learned to predict the estimation accuracy under different sampling probability distributions.

[0061] For example, this paper proposes a method for constructing an intelligent adaptive sampling model. The sampling process is modeled as a Markov decision process. The number of samples required to achieve a predetermined accuracy is used as the optimization target, and the sampling probability distribution parameter is used as the optimization variable. The adaptive optimization of the sampling strategy is achieved through a reinforcement learning algorithm. The specific technical implementation steps are as follows:

[0062] Markov decision process modeling. The sampling process is abstracted into a Markov decision process, defining the state space, action space, state transition probabilities, reward function, discount factor, and policy function. The state space includes the statistical properties of the samples already drawn (such as sample mean, variance, skewness, etc.) and the parameters of the current sampling probability distribution (such as the mean and standard deviation of the normal distribution). For example, for a two-dimensional sampling problem, the state can be expressed as (sample mean 1, sample mean 2, sample variance 1, sample variance 2, sampling distribution mean 1, sampling distribution mean 2, sampling distribution standard deviation 1, sampling distribution standard deviation 2).

[0063] The action space is a strategy for adjusting the parameters of the sampled probability distribution. For example, for a normal distribution, an action can be to increase or decrease the mean or standard deviation. The state transition probability represents the probability of transitioning to the next state after taking a certain action in the current state. The reward function comprehensively considers the number of samples and estimation accuracy. For example, the reward can be designed as the ratio of the improvement in estimation accuracy to the number of new samples, encouraging the maximum possible improvement in accuracy with the fewest possible samples. The discount factor controls the weight of future rewards and takes a value between 0 and 1. The policy function represents the probability distribution of selecting an action in a given state. For example, the softmax function can convert action values ​​into a probability distribution.

[0064] For example, suppose the current state is (sample mean 1 = 10, sample mean 2 = 20, sample variance 1 = 5, sample variance 2 = 8, sampling distribution mean 1 = 12, sampling distribution mean 2 = 18, sampling distribution standard deviation 1 = 2, sampling distribution standard deviation 2 = 3). The available actions include increasing / decreasing the mean and standard deviation of the sampling distribution in two dimensions, for a total of eight actions. If you choose the "Increase sampling distribution mean 1" action, the state shifts to (sample mean 1 = 10, sample mean 2 = 20, sample variance 1 = 5, sample variance 2 = 8, sampling distribution mean 1 = 13, sampling distribution mean 2 = 18, sampling distribution standard deviation 1 = 2, sampling distribution standard deviation 2 = 3). Simultaneously, the sample statistics are updated based on the newly extracted samples. If the estimation accuracy improves by 2% and the number of new samples is 10, the reward is 2% / 10 = 0.002.

[0065] Policy Gradient Reinforcement Learning. This algorithm uses a policy gradient-based reinforcement learning algorithm to optimize a sampled policy function. Using the Monte Carlo method, the algorithm samples multiple state-action trajectories, calculates the cumulative reward for each trajectory, and uses this as a sample to estimate the gradient of the expected value of the cumulative reward with respect to the policy function parameters. The gradient estimation formula is the cumulative reward multiplied by the derivative of the state-action log probability with respect to the policy function parameters, and then averaged over all trajectory samples. The policy function parameters are iteratively updated using the gradient ascent method, and the policy function gradually converges to the optimal policy over multiple iterations.

[0066] For example, suppose 10 trajectories are sampled, each taking 20 steps from the initial state to the final state. The state, action, and reward at step t of the i-th trajectory are (sit, ait, rit), respectively, and the cumulative reward of the trajectory is Ri. The policy function parameter is θ, and the derivative of the state-action log probability with respect to the parameter θ is ▽logπ(ait|sit, θ). The gradient estimate is then (1 / 10)∑iRi ∑t ▽logπ(ait|sit, θ), and the parameter θ is updated based on this gradient estimate. Repeat multiple iterations to ultimately obtain the optimal sampling policy function.

[0067] Gaussian process regression accuracy prediction. To efficiently evaluate estimation accuracy under different sampling probability distributions during policy optimization, a Gaussian process regression model is integrated to learn the nonlinear mapping relationship between sample features and estimation accuracy. The statistical characteristics of the extracted samples and the sampling probability distribution parameters are used as the input of the Gaussian process regression model, and the corresponding estimation accuracy is used as the output. A kernel function (such as the Gaussian kernel function) is used to describe the similarity between input features, and the mapping relationship is learned from the training samples. For new inputs, the Gaussian process regression model can predict the corresponding estimation accuracy mean and confidence interval.

[0068] For example, consider 100 sets of historical data, each containing sample statistical characteristics (sample mean, variance), sampling probability distribution parameters (sample mean, standard deviation), and corresponding estimated accuracy. 80 of these sets are used as training sets, and 20 as test sets. A Gaussian process regression model is trained using the training set samples, and model hyperparameters (such as kernel function parameters) are optimized to minimize the negative log-likelihood function. Model performance is evaluated on the test set. If the difference between the predicted estimated accuracy and the true estimated accuracy is small, the model can be used for accuracy prediction in policy optimization.

[0069] In summary, the intelligent adaptive sampling model proposed in this paper uses a Markov decision process to model the sampling process, optimizes the sampling policy function using a policy gradient reinforcement learning algorithm, and integrates a Gaussian process regression model to predict estimation accuracy, achieving adaptive optimization of the sampling strategy. Compared with traditional fixed sampling methods, this method can dynamically adjust the sampling probability distribution based on sample characteristics and estimation requirements, minimizing the number of samples while maintaining estimation accuracy, thereby improving sampling efficiency.

[0070] In an optional embodiment,

[0071] Adopting a reinforcement learning algorithm based on policy gradient, Monte Carlo estimation is used to calculate the gradient of the expected value of the cumulative reward with respect to the policy function parameters. The policy function parameters are iteratively updated using the gradient ascent method. Through multiple iterations, the policy function gradually converges to the optimal policy, including:

[0072] Construct a cumulative reward function to evaluate the long-term benefits of the state-action trajectory. The cumulative reward function consists of two parts: an immediate reward and a discount factor. The immediate reward is constructed based on the robot's dynamic detection performance indicators, including a weighted combination of multiple optimization objectives such as detection accuracy, detection efficiency, and energy consumption. The discount factor is used to balance the immediate reward and long-term reward and control the reward decay rate.

[0073] The Monte Carlo estimation algorithm is used to approximate the policy gradient. By executing the current policy function, multiple state-action trajectory samples are generated. The cumulative reward of each trajectory sample is calculated, and the policy gradient is estimated using the cumulative reward. The policy gradient is calculated as the cumulative reward multiplied by the gradient of the state-action log probability with respect to the policy function parameters. The policy gradients estimated from multiple trajectories are averaged.

[0074] The Adam optimization algorithm is used to adaptively iteratively update the policy function parameters. According to the estimated policy gradient, the policy function parameters are optimized through multiple iterations with an adaptively adjusted learning rate as the step size, so that it can accelerate convergence to the optimal strategy while satisfying the policy improvement theory. The trust region is introduced to control the step size of the policy update to improve the sample efficiency and training stability of the algorithm, and finally the optimal strategy that can autonomously adapt to different detection environments and task requirements is obtained.

[0075] For example, this paper uses a policy gradient-based reinforcement learning algorithm. Monte Carlo estimation is used to calculate the gradient of the expected value of the cumulative reward with respect to the policy function parameters. The policy function parameters are iteratively updated using the gradient ascent method. Through multiple iterations, the policy function gradually converges to the optimal policy. The specific technical implementation steps are as follows:

[0076] Construct a cumulative reward function. The cumulative reward function is used to evaluate the long-term benefits of a complete state-action trajectory and consists of two parts: an immediate reward and a discount factor. The immediate reward is constructed based on the robot's dynamic detection performance indicators. It comprehensively considers multiple optimization objectives such as detection accuracy, detection efficiency, and energy consumption, and obtains a scalar reward value through weighted summation. For example, assuming that the detection accuracy weight is 0.5, the detection efficiency weight is 0.3, and the energy consumption weight is 0.2, and the current detection accuracy is 95%, the detection efficiency is 100 items / hour, and the energy consumption is 1 kW·h, then the immediate reward = 0.5 × 95% + 0.3 × 100 + 0.2 × (-1) = 47.3.

[0077] The discount factor is a value between 0 and 1 that balances immediate and long-term rewards and controls the rate at which future rewards decay. A larger discount factor emphasizes long-term gains, while a smaller discount factor emphasizes short-term gains. For example, if the discount factor is 0.9, the cumulative reward = 47.3 + 0.9 × (reward at the next moment) + 0.9^2 × (reward at the next moment) + ... As time steps increase, the weight of future rewards decreases.

[0078] Monte Carlo estimation of policy gradients. By executing the current policy function, multiple complete trajectory samples are generated from the initial state to the final state. For each trajectory sample, its cumulative reward is calculated as an estimate of the long-term return under the current policy. The policy gradient is estimated using the cumulative reward. The policy gradient is calculated as the cumulative reward of the trajectory multiplied by the gradient of the state-action log probability with respect to the policy function parameters, and then the policy gradient estimated for all trajectory samples is averaged.

[0079] Specifically, assume that 10 trajectory samples are generated, each lasting 20 time steps. The state at step t of the i-th trajectory is sit, the action is ait, and the immediate reward is rit. The gradient of the state-action log probability with respect to the policy function parameter θ is ▽log π(ait|sit,θ). Then the cumulative reward of the i-th trajectory is Ri=ri1+γ·ri2+γ^2·ri3+...+γ^19·ri20, where γ is the discount factor. Let the gradient estimate gi=Ri·(∑t=120 ▽log π(ait|sit,θ)), then the policy gradient ĝ=(1 / 10)·∑i=110gi.

[0080] Adam optimizes policy function parameters. The Adam optimization algorithm adaptively and iteratively updates the policy function parameters. Based on the estimated policy gradient, the Adam optimization algorithm uses an adaptively adjusted learning rate as the step size. Through multiple iterations, the policy function parameters are optimized to accelerate convergence to the optimal policy while satisfying policy improvement theory. To improve the algorithm's sample efficiency and training stability, a trust region is introduced to control the step size of the policy update.

[0081] The Adam optimization algorithm contains four hyperparameters: the learning rate α, the exponential decay rate β1 of the first-order moment estimate, the exponential decay rate β2 of the second-order moment estimate, and the numerical stability constant ϵ. For example, we set α = 0.001, β1 = 0.9, β2 = 0.999, and ϵ = 10-8. At the kth iteration, the first-order moment estimate mk and the second-order moment estimate vk are first calculated. Then, they are bias-corrected to obtain m̂k and v̂k, and the parameters are finally updated according to the policy gradient.

[0082] To introduce trust region control, the parameter update step size is constrained to meet the trust region radius δ. The iterative optimization process is repeated until the policy function parameters converge or the preset maximum number of iterations is reached, and the optimal policy function is finally obtained.

[0083] The aforementioned policy gradient-based reinforcement learning algorithm enables adaptive optimization of the robot's dynamic detection strategy, yielding an optimal strategy that can autonomously adapt to different detection environments and task requirements. In practical applications, it's necessary to consider the specific robot system and detection task, rationally set factors such as the state space, action space, and reward function, and fine-tune the algorithm's hyperparameters to achieve optimal performance.

[0084] In an optional embodiment,

[0085] The constructed intelligent adaptive sampling model is used to generate motion parameter samples of the robot carrying the weight. The robot is controlled to carry the weight to be tested according to the motion parameter samples. At the same time, the instrument's indication samples of the weight are collected. The collected motion parameter samples and indication samples are input into the parallel computing framework. The mathematical expectation and variance of the instrument indication under the motion parameter samples are calculated in parallel on the high-performance computing platform. The variation characteristics of the instrument sensitivity under different motion parameter combinations are evaluated. The sensitivity coefficient range of the instrument sensitivity relative to the robot motion parameters is obtained, including:

[0086] A parallel computing framework is constructed to process the collected motion parameter samples and instrument indication samples. The parallel computing framework includes a data distribution module, a parallel computing module, and a result summary module. The data distribution module distributes the motion parameter samples and indication samples to different computing nodes for parallel processing. The parallel computing module calculates the mathematical expectation and variance of all indication samples corresponding to the j-th motion parameter sample on each computing node. The result summary module summarizes the calculation results of each computing node to obtain the complete statistical characteristics of the instrument indication.

[0087] Based on the statistical characteristics of the instrument indications obtained by parallel computing, the finite difference method is used to numerically estimate the instrument sensitivity coefficient. For the j-th motion parameter sample, by applying positive and negative disturbances to its i-th motion parameter, the difference between the mathematical expectations of the instrument indications under positive and negative disturbances is calculated and divided by twice the disturbance amount as the sensitivity coefficient estimate corresponding to its i-th motion parameter.

[0088] Based on the sensitivity coefficient estimation of each motion parameter sample, the sensitivity coefficient interval of each motion parameter is constructed. The upper limit of the sensitivity coefficient interval takes the maximum value of the sensitivity coefficient estimation corresponding to all motion parameter samples, and the lower limit takes the minimum value, thereby quantitatively evaluating the variation range of the instrument sensitivity relative to each motion parameter.

[0089] For example, this paper uses the constructed intelligent adaptive sampling model to generate motion parameter samples of a robot carrying a weight. The robot is controlled to carry the weight to be tested according to the motion parameter samples. At the same time, the instrument's indication samples of the weight are collected. The collected motion parameter samples and indication samples are input into a parallel computing framework. The mathematical expectation and variance of the instrument's indication under the motion parameter samples are calculated in parallel on a high-performance computing platform. The variation characteristics of the instrument's sensitivity under different motion parameter combinations are evaluated, and the sensitivity coefficient range of the instrument's sensitivity relative to the robot's motion parameters is obtained. The specific technical implementation steps are as follows:

[0090] Build a parallel computing framework. The parallel computing framework includes a data distribution module, a parallel computing module, and a result aggregation module. The data distribution module is responsible for distributing the collected motion parameter samples and instrument value samples to different computing nodes to achieve parallel processing. For example, suppose 100 sets of motion parameter samples and corresponding instrument value samples are collected, each set containing 10 motion parameters and 100 instrument value samples. These 100 sets of samples are evenly distributed to 5 computing nodes, with each node processing 20 sets of samples.

[0091] The parallel computing module calculates the mathematical expectation and variance of all the indication samples corresponding to the jth group of motion parameter samples on each computing node. For example, for the first group of motion parameter samples processed by the first computing node, the mathematical expectation of the corresponding 100 indication samples is calculated as the arithmetic mean of these indications, and the variance is calculated as the arithmetic mean of the squares of the differences between each indication and the mathematical expectation.

[0092] The result summary module aggregates the calculation results of each calculation node to obtain the mathematical expectation and variance of the instrument indication corresponding to all motion parameter samples, forming a complete instrument indication statistical characteristic. For example, the mathematical expectation and variance of the indication of 100 sets of motion parameter samples obtained by 5 calculation nodes are sequentially spliced ​​in the order of sample numbers to obtain a vector containing 100 mathematical expectations and 100 variances, which is used as the instrument indication statistical characteristic.

[0093] Numerical estimation of sensitivity coefficients. Based on the statistical characteristics of the instrument indications obtained through parallel computing, the finite difference method is used to numerically estimate the instrument sensitivity coefficients. For the jth group of motion parameter samples, perturbation analysis is performed on each of its motion parameters in turn. For example, for the i-th motion parameter of the jth group of samples, a small positive perturbation and a small negative perturbation (e.g., ±1%) are applied to its original value to obtain two new motion parameter samples.

[0094] Using the new motion parameter samples, the robot is controlled to carry weights and collect indication samples. The difference between the mathematical expectations of the instrument indications under positive and negative perturbations is calculated. This difference, divided by twice the perturbation amount, is used to estimate the sensitivity coefficient of the i-th motion parameter in the j-th group of samples. For example, if the original value of the i-th motion parameter in the j-th group of samples is 10, the positive perturbation is 10.1, and the negative perturbation is 9.9, the mathematical expectation of the indication under positive perturbation is 100.5, and the mathematical expectation of the indication under negative perturbation is 99.5. The sensitivity coefficient is estimated to be (100.5-99.5) / (10.1-9.9)=5.

[0095] Sensitivity coefficient interval construction. Based on the sensitivity coefficient estimates for each motion parameter sample, a sensitivity coefficient variation interval is constructed for each motion parameter. For the i-th motion parameter, the upper limit of the sensitivity coefficient interval is the maximum sensitivity coefficient estimate corresponding to the i-th motion parameter across all motion parameter samples, and the lower limit is the minimum sensitivity coefficient estimate. This quantitatively evaluates the range of instrument sensitivity relative to each motion parameter.

[0096] For example, consider the three motion parameters of a robot carrying a weight: translational velocity, lifting height, and acceleration. Assuming a sample size of 100, numerical estimation yields 100 translational velocity sensitivity coefficients, with the maximum value being 8.5 and the minimum being -3.2. The translational velocity sensitivity coefficient range is [-3.2, 8.5], indicating that for a 1-unit change in translational velocity, the mathematically expected change in the instrument reading is between -3.2 and 8.5 units.

[0097] Similarly, the sensitivity coefficient range for lift height is [-1.8, 4.6], and the sensitivity coefficient range for acceleration is [-5.3, 2.9]. Using the sensitivity coefficient range, we can quantitatively compare the impact of different motion parameters on instrument sensitivity. If the sensitivity coefficient range for a particular motion parameter is wide and includes zero, it indicates that the instrument sensitivity is highly sensitive to its changes and requires careful control during actual handling operations.

[0098] In summary, the proposed method for analyzing the impact of robotic handling motion on instrument sensitivity using an intelligent adaptive sampling model and a parallel computing framework can efficiently evaluate the changing characteristics of instrument sensitivity under different motion parameter combinations and quantitatively determine the sensitivity of instrument sensitivity to the robot's motion parameters. This method can guide the optimization of robotic handling operation parameters, select motion parameter combinations that are insensitive to instrument sensitivity, and improve the accuracy and reliability of robot-assisted metrology.

[0099] In an optional embodiment,

[0100] Determine the sensitivity coefficient interval including:

[0101] ;

[0102] ;

[0103] in, represents the estimated value of the sensitivity coefficient of the i-th motion parameter corresponding to the j-th motion parameter sample, Indicates that a positive perturbation is applied to the i-th component of the j-th motion parameter sample After that, the mathematical expectation of the instrument indication is Indicates applying a negative perturbation to the i-th component of the j-th motion parameter sample - After that, the mathematical expectation of the instrument indication is represents the perturbation amount on the i-th motion parameter.

[0104] In an optional embodiment,

[0105] Based on the sensitivity coefficient interval, interval arithmetic is used to transfer the uncertainty of the robot motion parameters to the instrument indication to obtain the confidence interval of the instrument indication. The robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval is selected as the optimal motion parameter. The robot is controlled to carry the weight according to the optimal motion parameter for dynamic detection, including:

[0106] A velocity planning method is used to generate the velocity curve of the robot's motion process. By setting the duration of the acceleration ramp, constant speed, and deceleration stages, as well as the acceleration rate of change in the acceleration ramp and deceleration stages, the generated velocity curve is made smooth and continuous in both time and velocity dimensions. By adjusting the duration ratio of the three stages and the acceleration rate of change, the smoothness and volatility of the velocity curve are optimized, thereby reducing the uncertainty introduced by velocity fluctuations.

[0107] A trajectory planning method is used to generate the displacement curve of the robot's motion process. By using six boundary conditions (starting position, velocity, acceleration, and end position, velocity, and acceleration) as constraints, the six coefficients of trajectory planning are solved so that the generated displacement curve meets the continuity requirements of position, velocity, and acceleration in both time and displacement dimensions, resulting in a motion trajectory with continuous acceleration, smooth velocity, and no sudden displacement changes.

[0108] By constructing a robot kinematic and dynamic model, using the generated velocity curve and displacement curve as input, and using numerical integration and iterative optimization algorithms for simulation analysis, the changes in velocity, acceleration, and torque in the joint space and Cartesian space during the robot's motion are obtained. By setting the optimization objective function, the duration of the three stages of velocity planning and the acceleration change rate parameters as well as the polynomial coefficients of trajectory planning are optimized, so that the velocity, acceleration fluctuations, and joint torque changes of the robot's actual motion process meet the set performance index requirements;

[0109] The robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval is selected as the optimal motion parameter, and the robot is controlled to carry the weight for dynamic detection according to the optimal motion parameter.

[0110] For example, based on the sensitivity coefficient interval, interval arithmetic is used to transfer the uncertainty of the robot's motion parameters to the instrument indication, obtaining a confidence interval for the instrument indication. The robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval is selected as the optimal motion parameter. The robot is then controlled to carry the weight for dynamic testing according to the optimal motion parameter. The specific technical implementation steps are as follows:

[0111] Speed ​​planning. A speed planning method is used to generate a velocity curve for the robot's motion. By setting the duration of the acceleration ramp, constant speed, and deceleration stages, as well as the jerk of the acceleration ramp and deceleration stages, the generated velocity curve is made smooth and continuous in both time and speed dimensions.

[0112] For example, let's assume the robot's total motion time is 5 seconds, with an acceleration ramp lasting 1 second, a constant speed ramp lasting 3 seconds, and a deceleration ramp lasting 1 second. The acceleration jerk of the acceleration ramp is set to 2 m / s³, and the acceleration jerk of the deceleration ramp is set to -2 m / s³. Based on an initial velocity of 0 and a constant speed of 1 m / s, the final velocity of the acceleration ramp is calculated to be 2 m / s, and the initial velocity of the deceleration ramp is 1 m / s. This creates a velocity curve where the velocity increases from 0 to 2 m / s between 0 and 1 second, remains at 1 m / s between 1 and 4 seconds, and then decreases from 1 m / s to 0 between 4 and 5 seconds.

[0113] By adjusting the duration ratio and acceleration rate of the three phases, the smoothness and volatility of the velocity curve can be optimized. For example, increasing the duration of the acceleration ramp and deceleration phases and reducing the acceleration rate can make the acceleration of the velocity curve smoother, thereby reducing the uncertainty introduced by velocity fluctuations.

[0114] Trajectory planning. A trajectory planning method is used to generate the displacement curve of the robot's motion process. Using six boundary conditions (starting position, velocity, acceleration, and end position, velocity, and acceleration) as constraints, the six coefficients of trajectory planning are solved so that the generated displacement curve meets the continuity requirements of position, velocity, and acceleration in both time and displacement dimensions.

[0115] For example, suppose the robot needs to move from position 0 to position 10m in 5 seconds, and the initial and final speeds and accelerations are both 0. The relationship between time t and displacement s is expressed as a 5th-degree polynomial, that is, s=a5t 5 +a4t 4+a3t³+a2t²+a1t+a0. Substituting the six boundary conditions yields six equations for the polynomial coefficients a5, a4, a3, a2, a1, and a0. Solving these six equations determines the values ​​of the polynomial coefficients, and thus yields the robot's displacement curve from 0 to 5 seconds. This displacement curve ensures continuous acceleration, smooth velocity, and no sudden changes in displacement.

[0116] Kinematic and dynamic analysis: By constructing the robot's kinematic and dynamic models, using the velocity curve generated in step 1 and the displacement curve generated in step 2 as input, numerical integration and iterative optimization algorithms are used for simulation analysis to obtain the changes in velocity, acceleration, and torque in the joint space and Cartesian space during the robot's motion.

[0117] Assuming the robot's dynamic model is known, numerical integration is used to calculate the angle, angular velocity, and angular acceleration of each robot joint at each moment between 0 and 5 seconds. This is then used to determine the position, velocity, and acceleration of the robot's end effector in Cartesian space based on the forward solution of the robot's kinematics. The driving torque of each joint is then calculated based on the inverse solution of the robot's dynamics.

[0118] By setting optimization objective functions, such as minimizing the sum of squared terminal velocity and acceleration and the change in joint torque, and using an iterative optimization algorithm to search for the duration of the three velocity planning stages and the acceleration rate of change parameters, as well as the polynomial coefficients of trajectory planning, the robot's actual motion process speed, acceleration fluctuations, and joint torque changes meet the set performance requirements. For example, increasing the duration of the acceleration rise segment of the velocity curve by 20%, increasing the duration of the deceleration segment by 20%, and reducing the acceleration rate of change by 15% can reduce the maximum terminal velocity from 2m / s to 1.8m / s, the maximum acceleration from 4m / s² to 3.5m / s², and the maximum rate of change of joint torque from 100N·m / s to 80N·m / s, meeting the requirements for motion smoothness.

[0119] Optimal motion parameter selection. Based on the preceding analysis, confidence intervals for the robot's motion parameters (such as velocity, acceleration, and trajectory position) are obtained. For each parameter, its true value is assumed to be uniformly distributed within the confidence interval. Interval arithmetic is used to calculate the confidence intervals and obtain the confidence intervals for the instrument's indications. The robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval is selected as the optimal motion parameter combination. The robot is then controlled to carry the weight for dynamic testing using the optimal motion parameters.

[0120] For example, assuming the velocity confidence interval is [0.8, 1.2] m / s, the acceleration confidence interval is [1.8, 2.2] m / s², and the trajectory endpoint position confidence interval is [9.9, 10.1] m, the instrument indication confidence interval is [99.5, 100.5] through interval arithmetic. If another set of motion parameters has a velocity confidence interval of [0.9, 1.1] m / s, an acceleration confidence interval of [1.9, 2.1] m / s², and a trajectory endpoint position confidence interval of [9.95, 10.05] m, and the corresponding instrument indication confidence interval is [99.8, 100.2], then the second set of motion parameters is selected as the optimal parameters because its indication confidence interval width (0.4) is smaller than the indication confidence interval width (1.0) of the first set of parameters, indicating that it introduces a smaller instrument indication uncertainty.

[0121] In summary, the proposed robot handling motion optimization method generates a smooth and continuous motion curve through velocity and trajectory planning. It then optimizes the curve parameters using kinematic and dynamic analysis. Sensitivity analysis results are then used to select the optimal motion parameters. This method improves the stability and efficiency of robot handling while maintaining the accuracy of instrument dynamic detection. This method can provide an important reference for motion control in robot-assisted metrology and has application value in improving the level of metrology automation.

[0122] Figure 2 FIG. 1 is a structural diagram of a dynamic detection system for weights handled by a robot according to an embodiment of the present invention. Figure 2 As shown, the system includes:

[0123] The first unit is used to build an intelligent adaptive sampling model. The sampling process is modeled as a Markov decision process. The number of samples required to achieve a predetermined accuracy is the optimization objective, and the sampling probability distribution parameters are the optimization variables. The sampling strategy is adaptively optimized through a reinforcement learning algorithm. Based on the statistical characteristics of the samples, Gaussian process regression is used to learn the relationship between the weight mass, robot motion parameters, and instrument indications from the samples. An approximate estimate of the sampling probability distribution is generated, and the sampling probability distribution is dynamically adjusted to minimize the number of samples required to achieve the predetermined accuracy requirements of the sensitivity analysis.

[0124] The second unit is used to generate motion parameter samples for a robot handling weights using the constructed intelligent adaptive sampling model. The robot is then controlled to handle the weights to be tested according to the motion parameter samples. Simultaneously, samples of the instrument's indications of the weights are collected. The collected motion parameter samples and indication samples are input into a parallel computing framework. The mathematical expectation and variance of the instrument's indications under the motion parameter samples are then calculated in parallel on a high-performance computing platform. The variation characteristics of the instrument's sensitivity under different motion parameter combinations are evaluated, and the sensitivity coefficient range of the instrument's sensitivity relative to the robot's motion parameters is obtained.

[0125] The third unit is used to transfer the uncertainty of the robot motion parameters to the instrument indication using interval arithmetic based on the sensitivity coefficient interval, obtain the confidence interval of the instrument indication, select the robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval as the optimal motion parameters, control the robot to carry the weight according to the optimal motion parameters for dynamic detection, and use the confidence interval of the instrument indication as the uncertainty interval of the weight value caused by the instrument sensitivity to complete the dynamic detection and uncertainty assessment of the weight.

[0126] According to a third aspect of the embodiments of the present invention,

[0127] An electronic device is provided, comprising:

[0128] processor;

[0129] a memory for storing processor-executable instructions;

[0130] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0131] According to a fourth aspect of the embodiments of the present invention,

[0132] A computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0133] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A dynamic detection method for weights based on robot handling, characterized in that: include: An intelligent adaptive sampling model is constructed, modeling the sampling process as a Markov decision process. The number of samples required to achieve a predetermined accuracy is used as the optimization objective, and the sampling probability distribution parameters are used as optimization variables. A reinforcement learning algorithm is used to adaptively optimize the sampling strategy. Based on the statistical characteristics of the samples already drawn, Gaussian process regression is used to learn the relationship between the weight mass, robot motion parameters, and instrument indications from the samples. This generates an approximate estimate of the sampling probability distribution, dynamically adjusting the sampling probability distribution to minimize the number of samples required to achieve the predetermined accuracy requirements of the sensitivity analysis. The constructed intelligent adaptive sampling model is used to generate motion parameter samples of a robot handling weights. The robot is then controlled to carry the weights to be measured according to the motion parameter samples. Simultaneously, samples of the instrument's indications of the weights are collected. The collected motion parameter samples and indication samples are input into a parallel computing framework. The mathematical expectation and variance of the instrument's indications under the motion parameter samples are calculated in parallel on a high-performance computing platform. The variation characteristics of the instrument's sensitivity under different motion parameter combinations are evaluated, and the sensitivity coefficient range of the instrument's sensitivity relative to the robot's motion parameters is obtained. According to the sensitivity coefficient interval, interval arithmetic is used to transfer the uncertainty of the robot motion parameters to the instrument indication to obtain the confidence interval of the instrument indication. The robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval is selected as the optimal motion parameter. The robot is controlled to carry the weight according to the optimal motion parameters for dynamic detection. The confidence interval of the instrument indication is used as the uncertainty interval of the weight value caused by the instrument sensitivity to complete the dynamic detection and uncertainty assessment of the weight.

2. The method according to claim 1, characterized in that An intelligent adaptive sampling model is constructed, and the sampling process is modeled as a Markov decision process. The number of samples required to achieve a predetermined accuracy is used as the optimization target, and the sampling probability distribution parameters are used as optimization variables. The adaptive optimization of the sampling strategy is achieved through a reinforcement learning algorithm, including: The sampling process is modeled as a Markov decision process, defining the state space, action space, state transition probability, reward function, discount factor, and policy function. The state space is constructed based on the statistical properties of the samples already drawn and the parameters of the current sampling probability distribution. The adjustment strategy of the sampling probability distribution parameters is used as the action space. A reward function is designed that comprehensively considers the number of samples and estimation accuracy. The discount factor is used to control the weight of future rewards. The policy function represents the probability distribution of selecting an action in a given state. A reinforcement learning algorithm based on policy gradient is used. The gradient of the expected value of the cumulative reward with respect to the policy function parameters is calculated through Monte Carlo estimation. The policy function parameters are iteratively updated using the gradient ascent method. Through multiple iterations, the policy function gradually converges to the optimal policy. At the same time, the Gaussian process regression model is integrated, which takes the statistical characteristics of the extracted samples and the sampling probability distribution parameters as input and the estimation accuracy as output. The nonlinear mapping relationship between sample characteristics and estimation accuracy is learned to predict the estimation accuracy under different sampling probability distributions.

3. The method according to claim 2, characterized in that Adopting a reinforcement learning algorithm based on policy gradient, Monte Carlo estimation is used to calculate the gradient of the expected value of the cumulative reward with respect to the policy function parameters. The policy function parameters are iteratively updated using the gradient ascent method. Through multiple iterations, the policy function gradually converges to the optimal policy, including: Construct a cumulative reward function to evaluate the long-term benefits of the state-action trajectory. The cumulative reward function consists of two parts: an immediate reward and a discount factor. The immediate reward is constructed based on the robot's dynamic detection performance indicators, including a weighted combination of multiple optimization objectives such as detection accuracy, detection efficiency, and energy consumption. The discount factor is used to balance the immediate reward and long-term reward and control the reward decay rate. The Monte Carlo estimation algorithm is used to approximate the policy gradient. By executing the current policy function, multiple state-action trajectory samples are generated. The cumulative reward of each trajectory sample is calculated, and the policy gradient is estimated using the cumulative reward. The policy gradient is calculated as the cumulative reward multiplied by the gradient of the state-action log probability with respect to the policy function parameters. The policy gradients estimated from multiple trajectories are averaged. The Adam optimization algorithm is used to adaptively iteratively update the policy function parameters. According to the estimated policy gradient, the policy function parameters are optimized through multiple iterations with an adaptively adjusted learning rate as the step size, so that it can accelerate convergence to the optimal strategy while satisfying the policy improvement theory. The trust region is introduced to control the step size of the policy update to improve the sample efficiency and training stability of the algorithm, and finally the optimal strategy that can autonomously adapt to different detection environments and task requirements is obtained.

4. The method according to claim 1, wherein The constructed intelligent adaptive sampling model is used to generate motion parameter samples of the robot carrying the weight. The robot is controlled to carry the weight to be tested according to the motion parameter samples. At the same time, the instrument's indication samples of the weight are collected. The collected motion parameter samples and indication samples are input into the parallel computing framework. The mathematical expectation and variance of the instrument indication under the motion parameter samples are calculated in parallel on the high-performance computing platform. The variation characteristics of the instrument sensitivity under different motion parameter combinations are evaluated. The sensitivity coefficient range of the instrument sensitivity relative to the robot motion parameters is obtained, including: A parallel computing framework is constructed to process the collected motion parameter samples and instrument indication samples. The parallel computing framework includes a data distribution module, a parallel computing module, and a result summary module. The data distribution module distributes the motion parameter samples and indication samples to different computing nodes for parallel processing. The parallel computing module calculates the mathematical expectation and variance of all indication samples corresponding to the j-th motion parameter sample on each computing node. The result summary module summarizes the calculation results of each computing node to obtain the complete statistical characteristics of the instrument indication. Based on the statistical characteristics of the instrument indications obtained by parallel computing, the finite difference method is used to numerically estimate the instrument sensitivity coefficient. For the j-th motion parameter sample, by applying positive and negative disturbances to its i-th motion parameter, the difference between the mathematical expectations of the instrument indications under positive and negative disturbances is calculated and divided by twice the disturbance amount as the sensitivity coefficient estimate corresponding to its i-th motion parameter. Based on the sensitivity coefficient estimation of each motion parameter sample, the sensitivity coefficient interval of each motion parameter is constructed. The upper limit of the sensitivity coefficient interval takes the maximum value of the sensitivity coefficient estimation corresponding to all motion parameter samples, and the lower limit takes the minimum value, thereby quantitatively evaluating the variation range of the instrument sensitivity relative to each motion parameter.

5. The method according to claim 4, characterized in that Determining the sensitivity coefficient interval includes: ; ; in, represents the estimated value of the sensitivity coefficient of the i-th motion parameter corresponding to the j-th motion parameter sample, Indicates that a positive perturbation is applied to the i-th component of the j-th motion parameter sample After that, the mathematical expectation of the instrument indication is Indicates applying a negative perturbation to the i-th component of the j-th motion parameter sample - After that, the mathematical expectation of the instrument indication is represents the perturbation amount on the i-th motion parameter.

6. The method according to claim 1, characterized in that Based on the sensitivity coefficient interval, interval arithmetic is used to transfer the uncertainty of the robot motion parameters to the instrument indication to obtain the confidence interval of the instrument indication. The robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval is selected as the optimal motion parameter. The robot is controlled to carry the weight according to the optimal motion parameter for dynamic detection, including: A velocity planning method is used to generate the velocity curve of the robot's motion process. By setting the duration of the acceleration ramp, constant speed, and deceleration stages, as well as the acceleration rate of change in the acceleration ramp and deceleration stages, the generated velocity curve is made smooth and continuous in both time and velocity dimensions. By adjusting the duration ratio of the three stages and the acceleration rate of change, the smoothness and volatility of the velocity curve are optimized, thereby reducing the uncertainty introduced by velocity fluctuations. A trajectory planning method is used to generate the displacement curve of the robot's motion process. By using six boundary conditions (starting position, velocity, acceleration, and end position, velocity, and acceleration) as constraints, the six coefficients of trajectory planning are solved so that the generated displacement curve meets the continuity requirements of position, velocity, and acceleration in both time and displacement dimensions, resulting in a motion trajectory with continuous acceleration, smooth velocity, and no sudden displacement changes. By constructing a robot kinematic and dynamic model, using the generated velocity curve and displacement curve as input, and using numerical integration and iterative optimization algorithms for simulation analysis, the changes in velocity, acceleration, and torque in the joint space and Cartesian space during the robot's motion are obtained. By setting the optimization objective function, the duration of the three stages of velocity planning and the acceleration change rate parameters as well as the polynomial coefficients of trajectory planning are optimized, so that the velocity, acceleration fluctuations, and joint torque changes of the robot's actual motion process meet the set performance index requirements; The robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval is selected as the optimal motion parameter, and the robot is controlled to carry the weight for dynamic detection according to the optimal motion parameter.

7. A dynamic detection system for weights based on robot handling, used to implement the method according to any one of claims 1 to 6, characterized in that: include: The first unit is used to build an intelligent adaptive sampling model. The sampling process is modeled as a Markov decision process. The number of samples required to achieve a predetermined accuracy is the optimization objective, and the sampling probability distribution parameters are the optimization variables. The sampling strategy is adaptively optimized through a reinforcement learning algorithm. Based on the statistical characteristics of the samples, Gaussian process regression is used to learn the relationship between the weight mass, robot motion parameters, and instrument indications from the samples. An approximate estimate of the sampling probability distribution is generated, and the sampling probability distribution is dynamically adjusted to minimize the number of samples required to achieve the predetermined accuracy requirements of the sensitivity analysis. The second unit is used to generate motion parameter samples for a robot handling weights using the constructed intelligent adaptive sampling model. The robot is then controlled to handle the weights to be tested according to the motion parameter samples. Simultaneously, samples of the instrument's indications of the weights are collected. The collected motion parameter samples and indication samples are input into a parallel computing framework. The mathematical expectation and variance of the instrument's indications under the motion parameter samples are then calculated in parallel on a high-performance computing platform. The variation characteristics of the instrument's sensitivity under different motion parameter combinations are evaluated, and the sensitivity coefficient range of the instrument's sensitivity relative to the robot's motion parameters is obtained. The third unit is used to transfer the uncertainty of the robot motion parameters to the instrument indication using interval arithmetic based on the sensitivity coefficient interval, obtain the confidence interval of the instrument indication, select the robot motion parameter combination with the smallest difference between the upper and lower limits of the confidence interval as the optimal motion parameters, control the robot to carry the weight according to the optimal motion parameters for dynamic detection, and use the confidence interval of the instrument indication as the uncertainty interval of the weight value caused by the instrument sensitivity to complete the dynamic detection and uncertainty assessment of the weight.

8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.