A data acquisition and analysis system for a six-dimensional sensor

Through the data acquisition and analysis system of the six-dimensional sensor, diversified data of the robotic arm is collected and processed, and the neural network model is monitored and asynchronously trained in real time. Data simulation is used to simulate dynamic changes, which solves the problem of insufficient data acquisition of six-dimensional sensors on the automated production line, improving the generalization ability of the model and the operation efficiency of the robotic arm.

CN119871454BActive Publication Date: 2025-06-10HANGZHOU KELIN ELECTRIC CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510352933.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-10
Estimated Expiration
2045-03-25

AI Technical Summary

Technical Problem

The existing six-dimensional sensors are difficult to collect diversified force/moment data in high-intensity and high-frequency repeated operations on automated production lines that are running continuously all day, resulting in insufficient training data required for reinforcement learning or too single data mode, resulting in deviations or instability in the model when working within the new action or new torque range.

Method used

Provide a data acquisition and analysis system for six-dimensional sensors, including a data acquisition module, a data processing module, anomaly detection module and a data simulation module. By collecting standard output data and auxiliary data of the robotic arm, the state vector is constructed in conjunction and preprocessed and annotated. The anomaly detection module is used to monitor the prediction error of the neural network model in real time, and use mixed working condition data for asynchronous training to improve the generalization ability of the model. The data simulation module optimizes and trains the neural network model by generating dynamically changing simulation data.

Benefits of technology

It effectively solves the problems of insufficient data acquisition and single mode, significantly improves the generalization ability of neural network models, allowing them to accurately predict force or torque in a brand new environment, and ensures efficient and accurate operation of the robotic arm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119871454B_ABST
    Figure CN119871454B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of sensor detection, and specifically relates to a data acquisition and analysis system for a six-dimensional sensor, including: a data acquisition module, which is used to collect the standard output data and auxiliary data of the six-dimensional sensor, and fuse and construct them into a state vector; then preprocess and label the state vector; it also includes: a data processing module, which is used to receive the state vector within a fixed period of time as the input of the neural network model and output the compensated and decoupled result of the state vector. By generating dynamically changing simulation data to simulate the uncertainties in the real environment, the present invention provides rich training samples for the neural network model; combined with the reinforcement learning model, the system can adaptively adjust the behavior of the robotic arm in the virtual environment and optimize its operations under different temperature and external force conditions; this dynamic simulation and optimization training mechanism not only enhances the adaptability of the robotic arm in complex environments, but also improves the robustness and stability of its control system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of sensor detection, and particularly to a data acquisition and analysis system for a six-dimensional sensor. Background Art

[0002] With the rapid development of industrial automation, robotics, and intelligent equipment, the demand for multi-dimensional real-time monitoring and precise control is becoming increasingly urgent. A six-dimensional sensor combines three-axis acceleration and three-axis angular velocity (or other combinations of multi-dimensional sensors), and can capture key information such as the spatial position, attitude, velocity, and acceleration of the object to be measured in a real-time dynamic environment.

[0003] Upon retrieval, Chinese Patent No. CN202311042421.6 discloses a decoupling calibration system for a six-dimensional force / torque sensor based on reinforcement learning, which is a six-dimensional sensor calibration system for an application scenario, including an intelligent agent, an external environment, and an experience pool. Among them: the intelligent agent includes a neural network structure of a Critic part and an Actor part, and a performance optimization training module; the above solution reduces the dependence on special calibration equipment and improves the deployment flexibility by using an existing robotic arm and sensor to perform multi-dimensional loading and calibration in a specific operation task.

[0004] However, on an automated production line that operates continuously throughout the day, the robotic arm usually needs to perform high-intensity and high-frequency repetitive operations; during the online calibration process, the algorithm model requires a certain amount of time for data acquisition, update, and learning; if the production beat is strict, it is difficult to spare extra time to collect diverse force / torque data, resulting in insufficient training data required for reinforcement learning or overly single data patterns; when the algorithm model works in a new action or new torque range, large deviations or instability may occur. In addition, due to the lack of sufficient diverse data, existing neural network models often cannot correctly predict forces and torques when facing new working environments (such as different postures, temperatures, or vibration conditions). This lack of generalization ability makes the robotic arm perform poorly in practical applications and unable to meet the requirements of efficient and precise operations.

[0005] Therefore, a data acquisition and analysis system for a six-dimensional sensor is proposed to solve the above-mentioned problems. Summary of the Invention

[0006] Technical Problems to be Solved:

[0007] Aiming at the above-mentioned shortcomings of the prior art, the present invention provides a data acquisition and analysis system for a six-dimensional sensor, which can effectively solve the problem that the algorithm model in the prior art cannot spare extra time to collect diverse force / torque data in a continuously working robotic arm, resulting in insufficient training data required for reinforcement learning or overly single data patterns.

[0008] Technical solution:

[0009] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0010] The present invention provides a data acquisition and analysis system for a six-dimensional sensor. The technical solution adopted by the present invention is as follows: It includes: a data acquisition module, which is used to collect the standard output data and auxiliary data of the six-dimensional sensor, and fuse and construct them into a state vector; then preprocess and label the state vector; it also includes:

[0011] A data processing module, which is used to receive the state vector within a fixed period of time as the input of the neural network model, and output the compensation and decoupling result of the state vector; calculate the prediction error vector based on the compensation and decoupling result;

[0012] An anomaly detection module, which compares the prediction error vector with a preset threshold. If the prediction error vector is greater than the error threshold and lasts for N sampling periods, it indicates that there is a hardware failure in the sensor or the robotic arm itself, or the training data of the neural network model for this working condition is insufficient; where N is the continuous length threshold for judging anomalies; after manually eliminating the hardware failure, the model parameters of the neural network model are improved; the improvement method is: mark the error with abnormal data as abnormal working condition data, then collect the normal working condition data within a historical fixed time period, merge its abnormal working condition data and normal working condition data to construct mixed working condition data, and use the mixed working condition data to perform asynchronous training on the neural network model to obtain the optimal model parameters;

[0013] A data simulation module, which generates a certain number of random pose data based on the physical model of the robotic arm; applies dynamic interference data and load data to each pose data, and fuses and constructs them into simulation data; under the optimal model parameters of the neural network model, uses the simulation data to train the neural network model for unknown data.

[0014] Among them, the output formula of the compensation and decoupling result is:

[0015] ; in the formula, represents the neural network model in the case of the initial strategy; is the predicted value of the force or torque of the six-dimensional sensor at time t; is the state vector; is the model parameter of the neural network;

[0016] The calculation formula of the prediction error vector is: ; in the formula, is the actual measured value of the six-dimensional sensor at time t; is the error between the predicted value and the actual measured value.

[0017] Among them, the asynchronous training method is as follows:

[0018] Divide the mixed working condition data into S individuals; and randomly initialize the model parameters in each individual;

[0019] Define the loss function of the neural network model:

[0020] ; where is the loss function value; is the actual measurement value of the i-th individual; represents the output predicted by the neural network model for the i-th individual ;

[0021] Divide all the individuals in the mixed working condition data into several batches; completely traverse all the individuals in each batch, perform forward propagation on the model parameters of each individual, and calculate the average loss of each individual. The calculation formula is:

[0022] ; where B represents the index set of the currently selected batch; is the average loss of the individual; then calculate the gradient g by backpropagation, and update the model parameters using the optimization algorithm; repeat the training for different batches until the preset number of iterations is reached to obtain new model parameters; perform evolutionary algorithm operations on the individuals in the trained batches to obtain the next generation population; repeat the above steps starting from gradient update; until the loss drops to the preset threshold or the number of iteration rounds exceeds the set upper limit, stop the asynchronous training; and output the optimal model parameters of the optimal individual with the highest fitness value in the next generation population.

[0023] Among them, the evolutionary algorithm operation method is as follows:

[0024] Select the updated individuals from the batch and calculate the loss function value, and use this value as the fitness of the individual; randomly select several individuals from all the individuals in the mixed working condition data as candidates, compare the fitness of each individual, and mark the individual with the smallest fitness as the promoted individual;

[0025] Repeat the operation multiple times until a sufficient number of promoted individuals are selected as parent individuals; combine all the parent individuals in pairs or in multiple pairs to perform gene crossover operations to obtain the final offspring individuals;

[0026] Repeat the above operations for multiple pairs of parental individuals to obtain several new batches of final offspring individuals; combine the new batches of final offspring individuals and parental individuals into a new population, sort them in descending order according to the fitness value, and select the top several individuals from the new population as the reserved population according to the pre-set retention ratio; then randomly select from the remaining individuals to fill the remaining positions of the reserved population until the reserved population is full, thus forming the next generation population;

[0027] Repeat the operations on the next generation population.

[0028] Among them, in the selection process of the promoted individuals, based on the simulated annealing algorithm, one inferior individual is retained in each selection process of the promoted individuals, and the retention formula is: ; in the formula, is the selection probability of the inferior individual; is the fitness difference between the inferior individual and the promoted individual; T is the simulated annealing coefficient; is the exponential function.

[0029] Among them, the way of the gene crossover operation is:

[0030] Finally, generate temporary offspring individuals through segmented crossover; and after completing the segmented cutting, calculate the attention score of each independent segment in each temporary offspring individual based on the cumulative value of the absolute value of the gradient of each individual, and the calculation formula is:

[0031] ; in the formula, is the attention score of the i-th segment F; M is the number of iterations; is the gradient value calculated at the t-th iteration; compare the attention score of each segment with the pre-set attention threshold, and select the segment with the attention score greater than the attention threshold as the key segment ;

[0032] Define a small autoencoder, and input the key segments and in the parental individuals into the small autoencoder respectively to obtain the latent vectors and the latent vector ; perform a crossover operation on and in the latent vector space to form a latent fusion vector ; among them, is a fixed constant; decode the latent fusion vector back to the key segment space to obtain the decrypted segment ; then use the decrypted segment to replace the corresponding key segment in the temporary offspring individual, thus forming the final offspring individual.

[0033] Among them, the dynamic interference data includes dynamic temperature changes and dynamic external force changes, and the application method of the dynamic interference data is as follows:

[0034] Generate dynamic temperature changes based on a random fluctuation function, and the generation formula is:

[0035] ; In the formula, represents the temperature value at time t; is the average temperature; A represents the amplitude of the temperature change; w is the frequency of the temperature change; represents the phase;

[0036] Generate dynamic external force changes based on a random number generator, and the generation formula is:

[0037] ; In the formula, represents the external force value at time t; is the maximum external force value; represents the generated random number; represents the direction of the external force.

[0038] Among them, the data simulation module constructs a dynamic change simulation model based on a reinforcement learning model, and optimizes and trains the generation of simulation data. The optimization and training method is as follows:

[0039] Define a state vector and an action vector, and set a reward mechanism;

[0040] Define a deep Q-network model as the basic structure of the dynamic change simulation model. The basic structure includes an input layer, a hidden layer, and an output layer; the input layer is used to receive the state vector as input, and the output layer is used to receive the output of the hidden layer and calculate the Q value of each action vector;

[0041] Set up a virtual scenario for trying different actions; the virtual scenario returns a new state vector and a reward based on the action vector of the robotic arm; initialize the parameters of the deep Q-network model, and create an experience replay pool for storing the action attempt records of the robotic arm;

[0042] Obtain the current state vector of the robotic arm; select and execute the action vector with a random probability P to obtain a new state vector and a reward; store the current state vector, action vector, new state vector, and reward in the experience replay pool; then randomly sample data from the experience replay pool and calculate the update formula of the Q value to train the deep Q-network model; stop training until the set training steps are reached or the adjustment effect meets the requirements, that is, complete the optimization training of the dynamic change simulation model.

[0043] Among them, the optimization formula for selecting the action vector with a random probability P is:

[0044] ; wherein, is the selection probability of the action vector a; represents the Q value of selecting the action vector a under the state vector s; c is a control constant for adjusting the exploration degree; R is the total number of selections, is the number of times the action vector a is selected; is an adjustment coefficient; is a possible new action vector is the number of times it is selected.

[0045] Among them, the joint optimization of the adjustment coefficient and the control constant is as follows:

[0046] Set the initial adjustment coefficient and the initial control constant based on historical data, and record them as the initial policy parameters ; Generate multiple candidate policy parameters using the initial policy parameters; The virtual scenario uses each candidate policy parameter for learning. At each time step, evaluate the effectiveness of the current candidate policy parameter according to the reward of each candidate policy parameter, and mark the candidate policy parameter with effectiveness as the effective policy parameter; Then update the initial policy parameters using the effective policy parameters to obtain the evolutionary policy parameters; The update formula is:

[0047] ; wherein, is the evolutionary policy parameter; is the standard deviation of the noise; H is the total number of candidate policy parameters; is the reward of the i-th candidate policy parameter, is the average reward of the i-th candidate policy parameter; is the random noise of the i-th candidate policy parameter;

[0048] Repeat the above steps until the preset learning goal or convergence condition is reached, then the joint optimization of the adjustment coefficient and the control constant is completed.

[0049] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:

[0050] 1. In the present invention, by collecting the standard output data of the robotic arm during the working process and combining with auxiliary data fusion to construct a state vector, it provides a high-quality data basis for the subsequent training of the neural network model; and by combining auxiliary data, it ensures that the system can quickly and accurately obtain diversified data in the robotic arm working continuously, avoiding the problems of insufficient six-dimensional sensor data acquisition or single pattern analysis results caused by strict production beats in traditional methods.

[0051] 2. In the present invention, the anomaly detection module monitors the prediction error of the neural network model in real time, and when an anomaly is detected, the model is asynchronously trained using hybrid operating condition data. This training method combines an evolutionary algorithm and a simulated annealing strategy, enabling the model to continuously iterate and optimize under complex operating conditions, significantly enhancing the generalization ability of the model. Even in a completely new posture, temperature, or vibration environment, the model can still accurately predict force or torque, ensuring the efficient and precise operation of the robotic arm.

[0052] 3. In the present invention, the data simulation module provides rich training samples for the neural network model by generating dynamically changing simulation data to simulate the uncertainties in the real environment. Combined with the reinforcement learning model, the system can adaptively adjust the behavior of the robotic arm in the virtual environment and optimize its operation under different temperature and external force conditions. This dynamic simulation and optimization training mechanism not only enhances the adaptability of the robotic arm in complex environments but also improves the robustness and stability of its control system. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 is a schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0055] Embodiment: Referring to Figure 1 , a data acquisition and analysis system for a six-dimensional sensor is proposed in this case, including a data acquisition module, a data processing module, an anomaly detection module, and a data simulation module.

[0056] Among them, the data acquisition module is used to acquire the standard output data of the six-dimensional sensor of the robotic arm in a preset experimental environment. The standard output data includes force or torque information in six degrees of freedom directions, that is, the linear force components along three orthogonal coordinate axes and the torque components around three orthogonal coordinate axes.

[0057] The data processing module is used to receive the standard output data as the training input of the neural network model, and perform basic calibration under the set experimental environment (the temperature and humidity are relatively constant; there is no unnecessary external force interference around the robotic arm area; the vibration sources are relatively few or controllable) and known loads, output the decoupling result of the standard output data (i.e., the estimated value after force or torque correction), and collect a series of samples (the input standard output data and the output decoupling result) to form an initial data set; and input the initial data set into the neural network model, perform several rounds of iterative training, and output the model parameters that meet the requirements of the default working conditions (i.e., the initial weights and biases of the neural network model), denoted as the initial strategy.

[0058] The data acquisition module is also used to collect auxiliary data of the six-dimensional sensor itself and the key components of the robotic arm (such as joints, motors or reducers, etc.), including temperature changes, vibrations, and the joint angles of the robotic arm; it is obtained by installing auxiliary sensors around the six-dimensional sensor and the key components of the robotic arm. When the robotic arm operates for a long time, if the temperature of the key components is too high, it may affect the measurement accuracy of the force or torque sensor; by measuring the temperature of the key components of the robotic arm, the overall thermal environment of the machine can be evaluated, and data thermal drift compensation can be achieved together with the standard output of the six-dimensional sensor. Collecting the vibration of the six-dimensional sensor can directly reflect problems such as impacts, jitters, or loose installations on the sensor body, especially when the robotic arm moves at high speed or encounters collisions; once abnormal vibration readings occur, it may mean that the six-dimensional sensor is loose, the force or torque of the robotic arm is unstable, etc., and timely adjustment or maintenance is required; while the vibration of the key components of the robotic arm can be used to monitor the movement smoothness of the robotic arm itself, as well as whether the joint reducer is worn and whether the harmonic reducer generates backlash, etc.; if abnormally high vibration is detected, it will not only affect the decoupling accuracy of force or torque, but also indicate that the health of the mechanical body is worrying. The force or torque output is usually closely related to the current posture of the robotic arm, load distribution, etc. When performing decoupling calibration or online reinforcement learning, it is necessary to know the current joint angles and the end posture to determine the relationship of force or torque under different joint configurations.

[0059] Then, according to a unified timestamp or synchronous trigger event (for example, each time the robotic arm controller sends joint angles in a loop, it triggers a six-dimensional sensor data reading, and aligns the standard output data and auxiliary data at this time point), the collected standard output data and auxiliary data are merged into a state vector, and then the state vector is normalized and labeled (the label information is the current load and the current action mode) to facilitate subsequent machine model learning; and as the robotic arm continues to operate, the state vector will form a series of time-series sample data.

[0060] The data processing module then receives the state vector within a fixed period of time (i.e., a series of time-series sample data) as the input of the neural network model and outputs the compensated decoupling result of the state vector; its output formula is:

[0061] ; where represents the neural network model under the initial policy; is the predicted value of the force or torque of the six-dimensional sensor by the neural network model at time t, which is an output in vector form and includes the force or torque information in six degrees of freedom directions; is the input of the neural network model, that is, the state vector; are the model parameters of the neural network; define the prediction error vector of the neural network model, that is, the difference between the output result of the neural network model and the actual output result of the six-dimensional sensor. The calculation formula of the prediction error vector is: ; where is the actual measured value of the six-dimensional sensor at time t; is the error between the predicted value and the actual measured value, and it is also an output in vector form.

[0062] The anomaly detection module is used to compare the error with the preset error threshold . If and lasts for N sampling periods, it means that there is a fault or data deviation in the sensor or the robotic arm itself (for example, the sensor has hardware damage, looseness, poor wiring, or zero-point deviation due to aging; the components of the robotic arm are worn, the joints are loose, or a certain connecting rod is deformed, etc., resulting in a large change in the actual force or torque situation), or the neural network model lacks sufficient training data for this working condition, making the neural network model unable to generalize to complex scenarios (for example, a brand-new attitude, temperature, or vibration environment, resulting in the model being unable to correctly predict the force or torque); where represents performing a norm operation on the vector contained in the error, which can be the second norm or the maximum absolute value, etc.; N is the continuous length threshold for judging anomalies. When the error persists in being too large for N sampling periods, manual inspection is carried out to eliminate hardware faults and drifts. If it is determined that there is no hardware fault after manual inspection, the model parameters of the neural network model are improved; the improvement method is:

[0063] And mark the error with abnormal data as abnormal working condition data, then collect the normal working condition data within the historical fixed time period, merge its abnormal working condition data and normal working condition data to construct mixed working condition data, and use the mixed working condition data to perform asynchronous training on the neural network model; the asynchronous training methods include:

[0064] Divide the mixed working condition data into S individuals; and randomly initialize the model parameters in each individual;

[0065] Define the loss function of the neural network model:

[0066] ; where is the loss function value, that is, the objective function to be optimized in asynchronous training; S is the total number of individuals; is the actual measurement value of the i-th individual; represents the output predicted by the neural network model for the i-th individual All individuals in the mixed working condition data are divided into several batches (for example, each batch includes 32 individuals or 64 individuals); completely traverse all individuals in each batch, perform forward propagation on the model parameters of each individual, and calculate the average loss of the individual. The calculation formula is:

[0067] ; where B represents the index set of the currently selected batch. For example, if the size of the batch is 32, then B contains the indices of 32 individuals; is the average loss of the individual;

[0068] Then, backpropagate to calculate the gradient g, and use an optimization algorithm (such as gradient descent, Adam, etc.) to update the model parameters; repeat the training for different batches until the pre-set number of iterations is reached to obtain new model parameters.

[0069] Perform operations of the evolutionary algorithm on the trained batches. Select the updated individuals from the batches and calculate the loss function value, and use this value as the fitness of the individual (the smaller the value, the better the performance). Randomly select several individuals from all individuals in the mixed working condition data as candidates, compare the fitness of each individual, and mark the individual with the smallest fitness as the promoted individual; in addition, during the selection of the promoted individual, based on the simulated annealing algorithm, retain a disadvantaged individual (that is, an individual with a larger fitness value) in each selection of the promoted individual. The retention formula is: ; where is the selection probability of the disadvantaged individual; is the fitness difference between the disadvantaged individual and the promoted individual, The larger the value of, the more difficult it is for the disadvantaged individual to be accepted; T is the simulated annealing coefficient, which is used to control the looseness of accepting the disadvantaged individual. At the initial stage of screening, usually set a higher T, so that even if the fitness difference is large, the corresponding disadvantaged individual has a certain probability of being accepted. As the iteration progresses, T decreases continuously, and the probability of accepting the disadvantaged individual decreases accordingly, so that the selection process of the promoted individual gradually focuses on refining around the superior individuals; It is an exponential function. During the screening process of the tournament, the operation of the simulated annealing algorithm is combined, increasing the possibility of a small number of selections for inferior individuals, and trying more seemingly suboptimal directions in the early stage, so as to further improve the chance of jumping out of the local optimum and achieve a more flexible and robust optimization process.

[0070] Repeat the operation multiple times until a sufficient number of promoted individuals are selected as parent individuals; combine all parent individuals in pairs or in multiple pairs for gene crossover operations (that is, regard the model parameters of each individual as a vector arranged in index order, and determine the position index according to a certain segmentation method, such as single-point, double-point or multi-point crossover. For example, cut at indices p and q, divide the vector into multiple segments and cut, and then splice them again), and finally generate temporary offspring individuals through segmented crossover. After completing the segmented cutting, calculate the attention score of each independent segment in each temporary offspring individual based on the cumulative value of the absolute value of the gradient of each individual. The calculation formula is: ; In the formula, is the attention score of the i-th segment F; M is the size of the sliding window, that is, the number of iterations; is the gradient value calculated at the t-th iteration; compare the attention score of each segment with a pre-set attention threshold, and select the segment with an attention score greater than the attention threshold as the key segment ;

[0071] Define a small autoencoder, including: , ; Among them, E is the encoder, which is used to map the input high-dimensional data (that is, model parameters) to a more compact latent vector z, that is, the compression of the original information; D is the decoder, which is used to restore the latent vector to the same space as the model parameters and reconstruct a result as close as possible to the original model parameters; is the reconstructed model parameter after decoding; input the key segments and in the parent individuals into the small autoencoder respectively to obtain and ; Then perform a crossover operation on and in the latent vector space to form a latent fusion vector ; Among them, is a fixed constant; finally, decode and restore the latent fusion vector to the key segment space to obtain the decrypted segment ; Then use the decrypted segment to replace the corresponding key segment in the temporary offspring individual, thereby forming the final offspring individual.

[0072] A small amount of random disturbance is added to the final offspring individuals, that is, a mutation intensity parameter is set for each final offspring individual, and random noise is sampled from a normal distribution; the random noise is added to different segments of the final offspring individuals to generate new candidate solutions, so that the final offspring individuals will not only stay at the local position obtained by mixing the parent individuals, but also have the opportunity to expand to a larger range of solution space; and in the process of adding noise, an attention coefficient with a value less than 1 is set for the decrypted fragment, so that the random variation amplitude of the decrypted fragment will be smaller, thereby retaining the high-quality features formed after the fusion of the small autoencoder as much as possible, forming an evolutionary model that takes into account both robustness and innovation.

[0073] Repeat the above operation for multiple pairs of parent individuals to obtain several new batches of final offspring individuals; merge the new batch of final offspring individuals and parent individuals into a new population, and sort them in descending order according to the fitness value, and select the first few individuals from the new population as the reserved population according to the pre-set retention ratio; then randomly select from the remaining individuals to fill the remaining positions of the reserved population until the reserved population is filled, thereby forming the next generation population, including multiple optimal individuals.

[0074] Enter the next cycle, start from the gradient update, and repeat the above steps; until the loss drops to the preset threshold, or the iteration round exceeds the set upper limit, stop asynchronous training, and output the optimal result (that is, the optimal individual with the highest fitness value in the most generation population), and obtain the optimal optimal model parameters. Through asynchronous training of the neural network model, the collected abnormal data can be processed centrally, allowing the model to repeatedly iterate and learn more complex or extreme working conditions, thereby significantly improving the ability to identify and adapt to abnormal conditions, and improving the final accuracy and generalization of the model; and once new abnormal or new scene data is added in production, it can be strengthened in the next training, so that the neural network model can be continuously iterated and upgraded, and the intelligence and safety level of the production line can be continuously improved.

[0075] The data simulation module generates a certain number of random pose data based on the existing physical model of the robotic arm (the random number generator generates several sets of joint angles within each joint); at the same time, dynamic interference data and load data are applied to each pose data; the generated data is screened to remove unreasonable or extreme poses (such as joint angles exceeding physical limits), and finally fused to construct simulation data; under the optimal model parameters of the neural network model, the simulation data is used to train the neural network model for unknown data. By generating simulation data in a dynamically changing manner, the uncertainty in the real environment can be effectively simulated in the virtual environment. This method can not only quickly explore a wide state space, but also provide rich training samples for the neural network model, enhancing its adaptability to complex working conditions. By reasonably setting the dynamic temperature and external force, the generated data will be closer to the actual application scenario, providing a solid data foundation for subsequent online calibration and reinforcement learning.

[0076] Among them, the dynamic interference data includes dynamic temperature changes and dynamic external force changes; among them, the dynamic temperature changes are simulated and generated based on a random fluctuation function, and the generation formula is: ; In the formula, represents the temperature value at time t; is the average temperature; A represents the amplitude of the temperature change, that is, the maximum range of temperature fluctuation above and below the average value; w is the frequency of the temperature change; represents the phase, which determines the starting point of the temperature fluctuation and affects the time delay of the temperature fluctuation; the dynamic external force change generates the magnitude and direction of the external force based on a random number generator, and the generation formula is: ; In the formula, represents the external force value at time t; is the maximum external force value; represents the generated random number, ranging from [0,1], which is used to simulate the randomness of the dynamic external force, so that the external force has different values at each time point, thus increasing the uncertainty of the simulation; represents the direction of the external force, usually a unit vector, indicating the direction in which the external force is applied.

[0077] During the generation of simulation data, random changes in temperature and external forces may cause the control system of the robotic arm to face problems of uncertainty and insufficient robustness. Such uncertainty may lead to the following situations: when the external force or temperature changes exceed the expected range, the robotic arm may not respond correctly, resulting in control failure or performance degradation; or the dynamically changing environment may introduce noise, affecting the accuracy of sensors, thereby leading to incorrect state estimation. The data simulation module constructs a dynamic change simulation model based on the reinforcement learning model, optimizes and trains the generation of simulation data, enabling the generation of simulation data to adaptively adjust its behavior in a dynamic environment (i.e., the applied dynamic interference data and load data). Through interaction with the environment, the robotic arm will learn how to optimize its operations under different temperature and external force conditions. The way of optimization training is as follows:

[0078] Define the state vector, including the current temperature, external force, and joint angles of the robotic arm;

[0079] Define the action vector, including the joint angle changes and applied loads of the robotic arm;

[0080] By defining the state vector and action vector, the state and action patterns of the robotic arm that can be observed at each time step can be clarified, providing an effective data basis for subsequent learning.

[0081] Set up a reward mechanism. Give a positive reward when the robotic arm successfully completes a task (such as grasping an object) under the current simulation data situation; give a negative reward when the robotic arm control fails or there is a collision; and give an additional reward according to the operation efficiency of the robotic arm (such as the shorter the time required to complete the task, the higher the reward).

[0082] Define the deep Q-network model as the basic structure of the dynamic change simulation model. The basic structure includes an input layer, hidden layers, and an output layer. The input layer is used to receive the state vector as input. The hidden layers are used to receive the output of the input layer. There are L hidden layers, each hidden layer has m nodes, and the calculation formula for each node is: ; where is the output of the hidden layer, that is, the output of the j-th node in the l-th layer; is the activation function (such as ReLU); is the output of the j-th node in the l-th layer; represents the number of nodes in the l - 1-th layer; is the weight between the i-th node and the j-th node in the l - 1-th layer; is the output of the i-th node in the l - 1-th layer; is the bias of the j-th node in the l-th layer. The output layer is used to receive the output of the hidden layer and calculate the Q value of each action vector. The calculation formula is: ; where is the Q-value of the j-th action vector; is a linear activation function; represents the number of nodes in the L-th layer; is the weight between the i-th node in the L-th layer and the j-th node in the output layer; is the output of the i-th node in the L-th layer; is the bias of the j-th node in the L-th layer.

[0083] Set up a virtual scenario where the robotic arm can try different actions; the virtual scenario returns a new state vector and a reward based on the action vector of the robotic arm; initialize the parameters of the deep Q-network model and create an experience replay pool to store the action attempt records of the robotic arm;

[0084] Obtain the current state vector of the robotic arm; select and execute an action vector with a random probability P to obtain a new state vector and a reward; store the current state vector, action vector, new state vector, and reward in the experience replay pool; then randomly sample data from the experience replay pool and calculate the update formula of the Q-value to train the deep Q-network model; stop training until the set number of training steps is reached or the adjustment effect meets the requirements, that is, complete the optimization training of the dynamic change simulation model.

[0085] When selecting an action vector with a random probability P, optimize the selection of the action vector by combining the UCB strategy. The optimization formula is: ; where is the selection probability of the action vector a; represents the Q-value of selecting the action vector a under the state vector s; is the exploration term, which aims to encourage the robotic arm to select those action vectors with fewer selection times; c is a control constant that adjusts the exploration degree and controls the influence of the exploration term on the final selection probability; R is the total number of selections, is the number of times the action vector a is selected; is the adjustment coefficient. A high adjustment coefficient value will make the selection probability more uniform and increase the randomness of exploration; while a low adjustment coefficient value makes the selection probability more concentrated on the action vectors with higher Q-values, thereby reducing exploration and increasing exploitation; is a possible new action vector is the number of times it is selected.

[0086] By combining the UCB strategy and the adjustment coefficient, the robotic arm can achieve a dynamic balance between exploring new action vectors and exploiting the known best action vectors. A higher adjustment coefficient makes the selection of random probabilities more random, encouraging the exploration of different action vectors, thereby obtaining more simulation data in the simulation environment; while a lower temperature value makes the selection more inclined to action vectors with high Q-values, increasing exploitation and ensuring that effective decisions can be made quickly in practical applications. This balance mechanism enables the robotic arm to flexibly adjust its strategy according to environmental changes and its own learning progress, adapting to different situations.

[0087] During the optimization training process of the dynamically changing simulation model, the optimal strategy for randomly selecting action vectors may change over time and with the environment, and fixed adjustment coefficients and control constants cannot flexibly adapt to these changes, which may lead to the inability to continue improving its strategy during the optimization training process, resulting in poor performance in practical applications and inability to meet task requirements. By combining evolutionary strategies to jointly optimize the adjustment coefficient and the control constant, the random selection of action vectors can converge to the optimal strategy faster, reducing unnecessary exploration. The way of joint optimization is as follows:

[0088] Set the initial adjustment coefficient and the initial control constant based on historical data and record them as the initial policy parameters ; Generate multiple candidate policy parameters using the initial policy parameters. The generation method is as follows: by adding Gaussian noise to the initial policy parameters. The virtual scenario uses each candidate policy parameter for learning. At each time step, evaluate the effectiveness of the current candidate policy parameter according to the reward of each candidate policy parameter, and mark the candidate policy parameter with effectiveness as the effective policy parameter.

[0089] Then update the initial policy parameters using the effective policy parameters to obtain the evolutionary policy parameters. The update formula is:

[0090] ; In the formula, is the evolutionary policy parameter; is the standard deviation of the noise, which controls the amplitude of the noise and affects the amplitude of the update of the initial policy parameters; H is the total number of candidate policy parameters; is the reward of the i-th candidate policy parameter, is the average reward of the i-th candidate policy parameter; is the random noise of the i-th candidate policy parameter.

[0091] Repeat the above steps until the preset learning goal or convergence condition is reached, then the joint optimization of the adjustment coefficient and the control constant is completed. The joint optimization effectively enhances the learning ability and adaptability of the manipulator in complex environments. The dynamically adjusted adjustment coefficient and control constant enable the manipulator to achieve a better balance between exploration and exploitation, thereby improving the learning efficiency and decision-making quality. Through the real-time update and feedback mechanism, the manipulator can flexibly respond to environmental changes and find better solutions.

[0092] The way to evaluate the reward of the candidate policy parameters is as follows:

[0093] At each time step, select an action training to execute according to the candidate policy parameters. After execution, a new state vector and the corresponding reward will be returned according to the action vector. Accumulate the rewards within a fixed period of time, calculate the average reward according to the time steps within a fixed period of time, and compare it with the preset reward threshold. When the average reward is greater than the reward threshold, it indicates that the current candidate policy parameters perform well and are recorded as effective policy parameters.

[0094] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A data acquisition and analysis system for a six-dimensional sensor, characterized in that: include: The data acquisition module collects the standard output data and auxiliary data of the six-dimensional sensor and constructs them into a state vector; Label the state vector; also include: The data processing module receives the state vector within a fixed period of time as the input of the neural network model, and outputs the compensation decoupling result of the state vector; and calculates the prediction error vector based on the compensation decoupling result; The anomaly detection module compares the predicted error vector with a preset threshold. If the predicted error vector is greater than the error threshold and lasts for N sampling cycles, it indicates that the sensor, robotic arm or neural network model is abnormal; where N is the duration threshold for judging the abnormality; if the neural network model is abnormal, it is improved; the improvement method is: the error with abnormal data is marked as abnormal working condition data, and then the normal working condition data within a fixed historical time period is collected, and the abnormal working condition data and the normal working condition data are combined to construct mixed working condition data, and the mixed working condition data is used to asynchronously train the neural network model to obtain the optimal model parameters; The data simulation module generates a number of random posture data based on the physical model of the robotic arm; applies dynamic interference data and load data to each posture data, and fuses them into simulation data; and uses the simulation data to train the neural network model with unknown data under the optimal model parameters of the neural network model.

2. A six-dimensional sensor data acquisition and analysis system as claimed in claim 1, characterized in that: The output formula of the compensation decoupling result is: ; In the formula, Represents the neural network model under the initial policy situation; is the predicted value of the force or torque of the six-dimensional sensor at time t; is the state vector; are the model parameters of the neural network; The calculation formula of the prediction error vector is: ; In the formula, is the actual measurement value of the six-dimensional sensor at time t; is the error between the predicted value and the actual measured value.

3. A six-dimensional sensor data acquisition and analysis system as claimed in claim 2, characterized in that: The asynchronous training method is: Divide the mixed working condition data into S individuals; and randomly initialize the model parameters in each individual; Define the loss function of the neural network model: ; In the formula, is the loss function value; is the actual measurement value of the i-th individual; Represents the neural network model for the i-th individual The predicted output; Divide all individuals in the mixed working condition data into several batches; completely traverse all individuals in each batch, perform forward propagation on the model parameters of each individual, and calculate the average loss of the individual. The calculation formula is: ;Wherein, B represents the index set of the currently selected batch; is the average loss of the individuals; then back-propagate to calculate the gradient g, and use the optimization algorithm to update the model parameters; repeat the training of different batches until the preset number of iterations is reached to obtain the new model parameters; perform the evolutionary algorithm operation on the individuals in the trained batches to obtain the next generation population; repeat the above steps starting from the gradient update; until the loss drops to the preset threshold, or the number of iterations exceeds the set upper limit, stop the asynchronous training; And output the optimal model parameters of the best individual with the highest fitness value in the next generation population.

4. A six-dimensional sensor data acquisition and analysis system as claimed in claim 3, characterized in that: The evolutionary algorithm operates in the following way: Select the updated individual from the batch and calculate the loss function value, and use the value as the fitness of the individual; randomly select several individuals from all individuals in the mixed working condition data as candidates, compare the fitness of each individual, and mark the individual with the smallest fitness as the promoted individual; Repeat the operation several times until a sufficient number of individuals are selected as parent individuals; combine all parent individuals in pairs or multiple pairs, perform gene crossover operation to obtain the final offspring individuals; Repeat the above operation for multiple pairs of parent individuals to obtain several new batches of final offspring individuals; merge the new batch of final offspring individuals and parent individuals into a new population, and sort them in descending order according to the fitness value, and select the first few individuals from the new population as the reserved population according to the pre-set retention ratio; then randomly select from the remaining individuals to fill the remaining positions of the reserved population until the reserved population is filled, thereby forming the next generation population; Repeat the operation for the next generation of population.

5. A six-dimensional sensor data acquisition and analysis system as claimed in claim 4, characterized in that: In the selection process of the promoted individuals, based on the simulated annealing algorithm, one disadvantaged individual is retained in each selection process of the promoted individuals, and the retention formula is: ; In the formula, is the selection probability of the disadvantaged individual; is the fitness difference between the disadvantaged individuals and the promoted individuals; T is the simulated annealing coefficient; is an exponential function.

6. A six-dimensional sensor data acquisition and analysis system as claimed in claim 4, characterized in that: The method of the gene crossover operation is: The temporary offspring individuals are generated by segmented crossover. After the segmented cutting is completed, the attention score of the independent segment in each temporary offspring individual is calculated based on the cumulative value of the absolute value of the gradient of each individual. The calculation formula is: ; In the formula, is the attention score of the i-th fragment F; M is the number of iterations; is the gradient value calculated at the tth iteration; the attention score of each segment is compared with the preset attention threshold, and the segment with an attention score greater than the attention threshold is selected as the key segment ; Define a small autoencoder that converts the key fragments in the parent individual and Input them into the small autoencoder to get the latent vector and the latent vector ; In the latent vector space and Perform cross operation to form a potential fusion vector ;in, is a fixed constant; the potential fusion vector is decoded and restored to the key segment space to obtain the decrypted segment ; Reuse the decrypted fragment Replace the corresponding key fragments in the temporary offspring individuals to form the final offspring individuals.

7. A six-dimensional sensor data acquisition and analysis system as claimed in claim 1, characterized in that: The dynamic interference data includes dynamic temperature change and dynamic external force change, and the dynamic interference data is applied in the following manner: Based on the random fluctuation function simulation, dynamic temperature changes are generated, and the generation formula is: ; In the formula, represents the temperature value at time t; is the average temperature; A represents the amplitude of temperature change; w is the frequency of temperature change; Indicates phase; Generate dynamic external force changes based on a random number generator, and the generation formula is: ; In the formula, represents the external force value at time t; is the maximum external force value; Represents a generated random number; Indicates the direction of external force.

8. A six-dimensional sensor data acquisition and analysis system as claimed in claim 7, characterized in that: The data simulation module builds a dynamic change simulation model based on the reinforcement learning model, and optimizes the generation of simulation data. The optimization training method is: Define the state vector and action vector, and set the reward mechanism; Define the deep Q network model as the basic structure of the dynamic change simulation model, the basic structure includes an input layer, a hidden layer and an output layer; The input layer is used to receive the state vector as input, and the output layer is used to receive the output of the hidden layer to calculate the Q value of each action vector; Set up a virtual scene to try different actions; the virtual scene returns a new state vector and reward based on the action vector of the robot arm; initialize the parameters of the deep Q network model and create an experience replay pool to store the action attempt records of the robot arm; Get the current state vector of the robot arm; select an action vector with random probability P and execute it to obtain a new state vector and reward; store the current state vector, action vector, new state vector and reward in the experience replay pool; Then, data is randomly sampled from the experience replay pool, and the update formula of the Q value is calculated to train the deep Q network model; the training is stopped until the set number of training steps is reached or the adjustment effect meets the requirements, and the optimization training of the dynamically changing simulation model is completed.

9. A six-dimensional sensor data acquisition and analysis system as claimed in claim 8, characterized in that: The optimization formula for selecting the action vector with random probability P is: ; In the formula, is the selection probability of action vector a; represents the Q value of selecting action vector a under state vector s; c is the control constant for adjusting the degree of exploration; R is the total number of selections, is the number of times the action vector a is selected; is the adjustment coefficient; For possible new action vectors The number of times selected.

10. A six-dimensional sensor data acquisition and analysis system as claimed in claim 9, characterized in that: The joint optimization of the adjustment coefficient and the control constant is: Set the initial adjustment coefficient and initial control constant based on historical data and record them as initial strategy parameters ; The initial strategy parameters are used to generate multiple candidate strategy parameters. The virtual scene uses each candidate strategy parameter for learning. At each time step, the effectiveness of the current candidate strategy parameter is evaluated according to the reward of each candidate strategy parameter, and the candidate strategy parameters with effectiveness are marked as valid strategy parameters. The initial strategy parameters are then updated with the valid strategy parameters to obtain the evolution strategy parameters. The update formula is: ; In the formula, is the evolution strategy parameter; is the standard deviation of the noise; H is the total number of candidate strategy parameters; is the reward of the i-th candidate strategy parameter, is the average reward of the i-th candidate strategy parameter; is the random noise of the i-th candidate strategy parameter; Repeat the above steps until the preset learning goal or convergence condition is reached, and the joint optimization of the adjustment coefficient and the control constant is completed.

Citation Information

Patent Citations

  • Six-dimensional force / torque sensor decoupling calibration system based on reinforcement learning

    CN117109803A

  • Robot motion control method and device, computing equipment and storage medium

    CN116766191A

  • Multi-dimensional force sensor decoupling method based on transfer learning

    CN118565684A