Clamping auxiliary device performance evaluation method and system based on reinforcement learning

By constructing a working condition parameter space and optimizing the performance testing model through a reinforcement learning-based performance evaluation method for clamping auxiliary devices, the problems of low accuracy and efficiency in traditional evaluation methods are solved, and efficient performance testing is achieved.

CN120558539BActive Publication Date: 2025-12-12JINXIN PRECISION COMPONENTS KUNSHAN CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510473640.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-12-12
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

Traditional methods for evaluating the performance of clamping auxiliary devices suffer from low accuracy and efficiency in performance testing, leading to a large number of unnecessary blind tests and wasted resources.

Method used

A reinforcement learning-based approach is adopted, which involves constructing a working condition parameter space, performing similarity analysis and deviation evaluation, and optimizing the performance detection model by combining a deep Q-network algorithm, and then gradually conducting performance testing.

Benefits of technology

It improves the accuracy and efficiency of performance testing, reduces blind testing and repetitive steps, and saves testing time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120558539B_ABST
    Figure CN120558539B_ABST
Patent Text Reader

Abstract

The application provides a clamping auxiliary device performance evaluation method and system based on reinforcement learning, and relates to the field of intelligent evaluation. The method comprises the following steps: taking standard working condition parameters as a benchmark, evaluating the deviation of a plurality of working condition parameters in the working condition parameter space, and constructing a working condition parameter sequence; constructing a performance detection model based on reinforcement learning; and performing performance detection under a plurality of working condition parameters in sequence by using the performance detection model according to the working condition parameter sequence, and outputting a performance evaluation result. The technical problem of low performance test accuracy and efficiency in the traditional clamping auxiliary device performance evaluation method is solved. By evaluating the difficulty of different working conditions and combining the reinforcement learning algorithm, performance testing is gradually performed from easy to difficult, which can reduce blind testing and unnecessary repeated steps, effectively improve the testing precision, save the testing time, and significantly improve the accuracy and efficiency of performance testing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent evaluation, and in particular to a clamping auxiliary device performance evaluation method and system based on reinforcement learning. BACKGROUND

[0002] As an important device for ensuring the stable positioning and fixing of workpieces during machining, clamping auxiliary devices are widely used in fields such as numerical control machine tools, automated assembly lines, and robot operations. The performance of clamping auxiliary devices directly affects the precision, efficiency, and stability of workpiece machining, so it is particularly important to accurately evaluate their performance.

[0003] Currently, traditional performance evaluation is usually based on experience and fixed test procedures. Even if certain working condition parameters do not significantly affect clamping accuracy and stability, they will still be included in the test process, resulting in a large number of unnecessary blind tests. This inefficient testing method not only wastes time but also increases the cost of experiments. SUMMARY

[0004] The purpose of the present application is to provide a clamping auxiliary device performance evaluation method and system based on reinforcement learning to solve the technical problem of low performance test accuracy and efficiency in traditional clamping auxiliary device performance evaluation methods, which includes:

[0005] In a first aspect, the present application provides a clamping auxiliary device performance evaluation method based on reinforcement learning, which includes: randomly combining parameters according to the operating condition indicators of the clamping auxiliary device to construct a working condition parameter space; taking the standard working condition parameters of the clamping auxiliary device as a reference, respectively evaluating the deviation of several working condition parameters in the working condition parameter space to construct a working condition parameter sequence; constructing a performance detection model based on reinforcement learning, wherein the performance detection model is a clamping mechanics relationship model between the clamping auxiliary device and the workpiece, which can optimize the clamping operation parameters; according to the working condition parameter sequence, using the performance detection model to perform performance detection under several working condition parameters in turn, and outputting the performance evaluation result.

[0006] Preferably, the clamping auxiliary device performance evaluation method based on reinforcement learning further includes configuring the operating condition indicators of the clamping auxiliary device, wherein the operating condition indicators at least include workpiece material, workpiece shape, processing type, processing parameters, and environmental parameters.

[0007] Preferably, the clamping auxiliary device performance evaluation method based on reinforcement learning further comprises: dividing the machining parameters and the environmental parameters according to a predetermined step size, obtaining a machining parameter interval set and an environmental parameter interval set, wherein the machining parameters are set according to the machining type, and the environmental parameters at least include temperature, humidity and vibration; performing parameter random combination according to the workpiece material, the workpiece shape, the machining type, the machining parameter interval set and the environmental parameter interval set, and generating a plurality of working condition parameters to construct a working condition parameter space.

[0008] Preferably, the clamping auxiliary device performance evaluation method based on reinforcement learning further comprises: obtaining standard working condition parameters of the clamping auxiliary device, wherein the standard working condition parameters include standard workpiece material, standard workpiece shape, standard machining type, standard machining parameter and standard environmental parameter; taking the standard workpiece material, the standard workpiece shape, the standard machining type, the standard machining parameter and the standard environmental parameter as a benchmark, respectively analyzing the similarity of the plurality of working condition parameters, and outputting a plurality of similarities; and subtracting the plurality of similarities respectively by 1 to obtain a plurality of deviation degrees.

[0009] Preferably, the clamping auxiliary device performance evaluation method based on reinforcement learning further comprises: randomly selecting any working condition parameter in the plurality of working condition parameters as a first working condition parameter; configuring a similarity comparison operator, wherein the similarity comparison operator at least includes cosine similarity and Euclidean distance; using the cosine similarity and the Euclidean distance, taking the standard workpiece material, the standard workpiece shape, the standard machining type, the standard machining parameter and the standard environmental parameter as a benchmark, analyzing the similarity of the first working condition parameter, and obtaining a first similarity after mean calculation, and sequentially analyzing to obtain a plurality of similarities of a plurality of working condition parameters.

[0010] Preferably, the clamping auxiliary device performance evaluation method based on reinforcement learning further comprises: based on the plurality of deviation degrees, arranging the plurality of working condition parameters in order from small to large according to the deviation degrees to construct a working condition parameter sequence.

[0011] Preferably, the clamping auxiliary device performance evaluation method based on reinforcement learning further comprises: collecting modeling parameters related to the clamping auxiliary device, establishing a mechanical relationship model between the clamping auxiliary device and the workpiece through finite element analysis; defining a state space and an action space, and designing a reward function; based on reinforcement learning, combining a deep Q network algorithm, the state space, the action space and the reward function to train the mechanical relationship model until a convergence condition is met, and generating a performance detection model.

[0012] Preferably, the performance evaluation method for a clamping auxiliary device based on reinforcement learning further includes: a state space including clamping force, clamping speed, and fixture stiffness; an action space including adjusting clamping force, adjusting clamping speed, and adjusting fixture stiffness; and evaluation indicators for the reward function including clamping accuracy, fixture stability, workpiece deformation control, and production efficiency.

[0013] Preferably, the method for evaluating the performance of a clamping auxiliary device based on reinforcement learning further includes: performing performance tests under several operating parameters sequentially using the performance testing model according to the operating parameter sequence, and outputting several performance test data; performing a comprehensive performance evaluation of the clamping auxiliary device based on the several operating parameters and the several performance test data, and generating a performance test report as the performance evaluation result.

[0014] Secondly, the present invention also provides a performance evaluation system for clamping auxiliary devices based on reinforcement learning, used to execute a performance evaluation method for clamping auxiliary devices based on reinforcement learning as described in the first aspect, comprising: a working condition parameter space construction module, used to construct a working condition parameter space by randomly combining parameters according to the working condition indicators of the clamping auxiliary device; a working condition parameter sequence construction module, used to construct a working condition parameter sequence by performing deviation evaluation on several working condition parameters in the working condition parameter space based on the standard working condition parameters of the clamping auxiliary device; a performance detection model construction module, used to construct a performance detection model based on reinforcement learning, wherein the performance detection model is a clamping mechanical relationship model between the clamping auxiliary device and the workpiece, which can optimize the clamping operation parameters; and a performance evaluation result output module, used to sequentially perform performance detection under several working condition parameters according to the working condition parameter sequence and the performance detection model, and output the performance evaluation result.

[0015] The embodiments of the present invention have the following advantages:

[0016] A working condition parameter space is constructed by randomly combining parameters according to the operating condition indicators of the clamping auxiliary device. Then, using the standard working condition parameters of the clamping auxiliary device as a benchmark, deviation evaluations are performed on several working condition parameters within the working condition parameter space to construct a working condition parameter sequence. Furthermore, a performance testing model is constructed based on reinforcement learning. This performance testing model is a model of the clamping mechanics relationship between the clamping auxiliary device and the workpiece, which can optimize the clamping operation parameters. Finally, according to the working condition parameter sequence, performance testing is performed sequentially under several working condition parameters using the performance testing model, and the performance evaluation results are output. In other words, by evaluating the difficulty of different working conditions and combining them with reinforcement learning algorithms, performance testing is performed step-by-step from easy to difficult. This reduces blind testing and unnecessary repetitive steps, effectively improving testing accuracy and saving testing time, thereby significantly improving the accuracy and efficiency of performance testing. Attached Figure Description

[0017] Figure 1 A flow chart of steps of a clamping auxiliary device performance evaluation method based on reinforcement learning of the present application;

[0018] Figure 2 A structural schematic diagram of a clamping auxiliary device performance evaluation system based on reinforcement learning of the present application.

[0019] Explanation of reference signs:

[0020] The working condition parameter space construction module 11, the working condition parameter sequence construction module 12, the performance detection model construction module 13, and the performance evaluation result output module 14. DETAILED DESCRIPTION

[0021] The present application provides a clamping auxiliary device performance evaluation method and system based on reinforcement learning, which solves the technical problem of low performance test accuracy and efficiency in the traditional clamping auxiliary device performance evaluation method. By evaluating the difficulty of different working conditions and combining reinforcement learning algorithm, the performance test is gradually performed from easy to difficult, which can reduce blind test and unnecessary repeated steps, effectively improve test precision and save test time, thereby significantly improving the accuracy and efficiency of performance test.

[0022] The technical solutions in the present application will be described clearly and completely below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application. In addition, it should be noted that, for convenience of description, only parts related to the present application are shown in the drawings, rather than all parts.

[0023] Embodiment one, please refer to the accompanying Figure 1 The present application provides a clamping auxiliary device performance evaluation method based on reinforcement learning, which is applied to a clamping auxiliary device performance evaluation system based on reinforcement learning, and specifically includes the following steps:

[0024] S10: Randomly combining parameters according to the operation working condition indexes of the clamping auxiliary device to construct a working condition parameter space.

[0025] Further, the step S10 of the present application further includes:

[0026] S11: Configuring the operation working condition indexes of the clamping auxiliary device, wherein the operation working condition indexes at least include workpiece material, workpiece shape, processing type, processing parameter and environmental parameter.

[0027] Specifically, when evaluating the performance of the clamping auxiliary device, the configuration of the working condition parameters is crucial. These working condition indicators will affect the efficiency, precision, and stability of the clamping auxiliary device. In order to comprehensively evaluate the clamping performance, the working condition indicators need to cover multiple aspects. The following is the configuration of the working condition indicators of the clamping auxiliary device based on the enhanced learning performance evaluation method, including workpiece material, workpiece shape, machining type, machining parameters, and environmental parameters, etc. Among them, the material of the workpiece directly affects the selection of clamping force, the rigidity requirement of the clamp, and the friction characteristics during clamping. Common workpiece material indicators include metal material, non-metal material, composite material, etc. The shape of the workpiece affects the contact method, clamping method, and clamp design of the clamping device, mainly including geometric shape, size, etc. Different machining types require different clamping methods and operating parameters, such as cutting machining, drilling machining, etc. At the same time, multiple parameters involved in the machining process will affect the performance of the clamping auxiliary device, such as cutting depth, cutting force. Environmental factors may have a significant impact on the performance of the clamping auxiliary device during machining, such as temperature, humidity, vibration, etc. By configuring the above multi-dimensional working condition indicators, the working performance of the clamping auxiliary device under different machining conditions can be simulated comprehensively.

[0028] S12: dividing the machining parameters and environmental parameters according to a predetermined step size, obtaining a machining parameter interval set and an environmental parameter interval set, wherein the machining parameters are set according to the machining type, and the environmental parameters at least include temperature, humidity, and vibration; S13: randomly combining parameters according to the workpiece material, the workpiece shape, the machining type, the machining parameter interval set, and the environmental parameter interval set, to generate a plurality of working condition parameters, and constructing a working condition parameter space.

[0029] Specifically, in this step, the main purpose is to finely divide the machining parameters and environmental parameters in order to evaluate and optimize different working conditions in the subsequent steps. During the division process, the parameter intervals should be set according to different machining types and environmental factors according to a predetermined step size, so as to generate adaptive machining parameter interval set and environmental parameter interval set. First, the machining parameters are divided according to a predetermined step size, for example, the interval range of the machining parameters will be set according to the actual machining conditions for each machining type. For milling machining, the cutting speed may be between 50 and 200 meters per minute. Then, according to the predetermined step size (such as 5% or 10% interval division), the parameters are segmented within these parameter intervals to ensure that various working conditions are covered for comprehensive evaluation, and the machining parameter interval set is obtained. On the other hand, the temperature, humidity, vibration, and other environmental parameters are divided according to a predetermined step size to obtain the environmental parameter interval set.

[0030] Next, according to the workpiece material, workpiece shape, processing type, processing parameter interval set and environmental parameter interval set, parameter random combination is carried out, each working condition combination can be regarded as an independent working condition, and all related processing and environmental parameters are included; then all possible working condition parameter combinations form a working condition parameter space, and each working condition parameter combination corresponds to a possible processing environment, which is used for simulating and testing the performance of the clamping auxiliary device. Representative working conditions are generated by random combination to ensure comprehensive performance evaluation and optimization of the clamping auxiliary device under various actual production conditions.

[0031] S20: Taking the standard working condition parameters of the clamping auxiliary device as a reference, the deviation of the working condition parameters in the working condition parameter space is evaluated, and a working condition parameter sequence is constructed.

[0032] Further, the step S20 of the present application further comprises:

[0033] S21: Obtain the standard working condition parameters of the clamping auxiliary device, wherein the standard working condition parameters include standard workpiece material, standard workpiece shape, standard processing type, standard processing parameter and standard environmental parameter.

[0034] Specifically, the standard working condition parameters of the clamping auxiliary device are obtained, wherein the standard working condition parameters are important basis for comparison and optimization as a reference point, and the standard working condition parameters represent the running parameters of the clamping auxiliary device under normal or ideal working conditions, including standard workpiece material, standard workpiece shape, standard processing type, standard processing parameter and standard environmental parameter.

[0035] S22: Taking the standard workpiece material, standard workpiece shape, standard processing type, standard processing parameter and standard environmental parameter as a reference, the similarity of the working condition parameters is analyzed, and the similarity is output.

[0036] Further, the step S22 of the present application further comprises:

[0037] S221: Randomly select any working condition parameter in the working condition parameters as a first working condition parameter; S222: Configure a similarity comparison operator, wherein the similarity comparison operator at least includes cosine similarity and Euclidean distance; S223: Using cosine similarity and Euclidean distance, taking the standard workpiece material, standard workpiece shape, standard processing type, standard processing parameter and standard environmental parameter as a reference, the similarity of the first working condition parameter is analyzed, and the first similarity is obtained after mean calculation, and the similarity of the working condition parameters is obtained in turn.

[0038] Specifically, first, any one of the working condition parameters is randomly selected as the first working condition parameter; then, a similarity comparison operator is configured, wherein the similarity comparison operator at least includes a cosine similarity and a Euclidean distance. The cosine similarity is used to measure the similarity between two vectors (in this scheme, the working condition parameter vectors). The value of the cosine similarity is between -1 and 1. The closer the value is to 1, the more similar the two working condition parameters are. The closer the value is to -1, the less similar they are. The Euclidean distance is used to calculate the straight-line distance between two points, which represents the spatial distance between the two points. The smaller the Euclidean distance, the smaller the difference between the two working condition parameters, and the more similar they are. The larger the distance, the greater the difference.

[0039] Next, using the cosine similarity and the Euclidean distance, the first working condition parameter is analyzed based on the standard workpiece material, the standard workpiece shape, the standard processing type, the standard processing parameter, and the standard environment parameter. That is, the cosine similarity between the vector of the first working condition parameter and the vector of the standard working condition parameter (such as the standard workpiece material, the standard workpiece shape, etc.) is calculated to obtain a cosine similarity value. The Euclidean distance between the vector of the first working condition parameter and the vector of the standard working condition parameter is calculated to obtain a distance value. Then, for the multiple similarity values (cosine similarity and Euclidean distance) between the first working condition parameter and the standard working condition parameter, the average value is calculated to obtain a comprehensive similarity measure. Finally, using the same method, similar similarity analysis is performed on other generated working condition parameters to obtain several similarities of several working condition parameters.

[0040] S23: Subtract 1 from each of the similarities to obtain several deviation degrees; S24: Based on the several deviation degrees, arrange the several working condition parameters in order of increasing deviation degree to construct a working condition parameter sequence.

[0041] Specifically, 1 is subtracted from each of the similarities to obtain several deviation degrees. The higher the deviation degree, the greater the difference between the working condition parameter and the standard working condition parameter. Conversely, the lower the deviation degree, the higher the similarity between the working condition parameter and the standard working condition parameter. Then, based on the several deviation degrees, the several working condition parameters are arranged in order of increasing deviation degree to construct a working condition parameter sequence. The working condition parameter sequence will be used as the input parameter sequence in subsequent performance evaluation, ensuring that the working condition closest to the standard working condition is tested first, and the working condition that deviates more is tested later. Through this sequential testing, those working conditions that are most likely to work stably can be tested first, while avoiding excessive testing of working conditions that deviate greatly in the early stage, thereby improving efficiency. This method reduces unnecessary blind testing and optimizes the testing process, ensuring that the performance of the clamping auxiliary device under different working conditions can be evaluated most effectively.

[0042] S30: constructing a performance detection model based on reinforcement learning, wherein the performance detection model is a clamping mechanical relationship model between the clamping auxiliary device and the workpiece, and the clamping operation parameters can be optimized.

[0043] Further, the step S30 of the present application further comprises:

[0044] S31: collecting modeling parameters related to the clamping auxiliary device, and establishing a mechanical relationship model between the clamping auxiliary device and the workpiece through finite element analysis.

[0045] Specifically, first, the parameters related to the clamping auxiliary device and the workpiece are collected, and common modeling parameters include clamping auxiliary device parameters (clamping fixture material properties, clamping fixture geometric shape, clamping force and mechanical state, etc.), workpiece parameters (workpiece material properties, workpiece geometric shape and size, etc.); Then, a mechanical relationship model between the clamping auxiliary device and the workpiece is established through finite element analysis method (FEA), which is a numerical method for simulating the behavior of a structure by dividing the physical structure into discrete small elements. First, a three-dimensional model is established according to the geometric parameters of the clamping auxiliary device and the workpiece. This model needs to include various parts of the clamp, contact surfaces, clamping force application points, etc.; Then, assign material properties (such as elastic modulus, yield strength, Poisson's ratio, etc.) to each part in the model, which is crucial to the mechanical response of the model; And apply boundary conditions and loads, boundary conditions refer to the fixing conditions of the clamping auxiliary device, which simulates how the clamping device is connected with the workpiece and its workbench; The load is to apply appropriate clamping force and external load in the model to simulate the actual machining process. Through the finite element solver, the stress and strain distribution when the clamping device and the workpiece are in contact are calculated, the deformation of the workpiece and the clamp under the action of the clamping force is analyzed, and it is judged whether the clamping is uniform, whether there is excessive clamping or unstable clamping phenomenon.

[0046] S32: defining state space and action space, and designing reward function.

[0047] Further, the step S32 of the present application further comprises:

[0048] S321: the state space includes clamping force, clamping speed and clamp stiffness, the action space includes adjusting clamping force, adjusting clamping speed and adjusting clamp stiffness, and the evaluation index of the reward function includes clamping accuracy, clamp stability, workpiece deformation control and production efficiency.

[0049] Specifically, then, the state space and action space are defined, and the reward function is designed, wherein the state space includes clamping force, clamping speed, clamp stiffness, the action space includes adjusting clamping force, adjusting clamping speed and adjusting clamp stiffness; the reward function is used to evaluate the result obtained by the agent after executing a specific action, and the evaluation index of the reward function includes clamping accuracy, clamp stability, workpiece deformation control and production efficiency. By reasonably designing the state space, action space and reward function, feedback can be provided for the reinforcement learning algorithm to help the agent gradually optimize the parameters in the clamping process to achieve the goal of improving clamping accuracy, stability and production efficiency.

[0050] S33: Based on reinforcement learning, the mechanical relationship model is trained in combination with a deep Q network algorithm, a state space, an action space and a reward function until a convergence condition is met, and a performance detection model is generated.

[0051] Specifically, reinforcement learning is a machine learning method that involves an agent interacting with an environment, learning how to take actions in different situations to maximize cumulative rewards through a trial-and-error process. Then, based on reinforcement learning, the mechanical relationship model is trained in combination with the deep Q-network algorithm, state space, action space, and reward function. Deep Q-network (DQN) is a reinforcement learning algorithm based on Q-learning, which uses a deep neural network to approximate the Q-value function, effectively handling large-scale state space and action space. Q-learning is a model-free reinforcement learning algorithm in which the Q-value represents the expected return (reward) obtained by taking a certain action in a particular state. DQN approximates the Q-value using a deep neural network, thereby avoiding the limitations of tabular storage of Q-values and enabling the handling of more complex and high-dimensional state spaces and action spaces. The training process is as follows: first, a deep neural network is initialized to approximate the Q-value function. The network accepts the current state (clamping force, clamping speed, clamp stiffness) as input and outputs the Q-value for each action. Then, the experience replay mechanism is used to store the interaction between the agent and the environment (i.e., state, action, reward, next state) in the experience replay pool. During training, experiences are randomly drawn from the replay pool for learning, avoiding overfitting and improving training stability. Then, in each round of training, the agent obtains the current state from the environment, selects an action (according to the epsilon-greedy strategy, i.e., most of the time selecting the action with the highest current Q-value, but occasionally performing random selection to explore new strategies); after executing the action, the environment provides a reward and enters the next state, and the data in the experience replay pool is used to update the Q-value in the neural network, minimizing the error of the Bellman equation (i.e., training the network by minimizing the difference between the current Q-value and the predicted Q-value); training will continue until the output of the Q-network is stable and the training error no longer changes significantly, or the predetermined maximum number of training rounds and convergence criteria are met. After the training process is completed and the convergence conditions are met, the deep Q-network generates a performance detection model. This model can be used to predict the performance of the clamping auxiliary device under different working condition parameters and optimize the operating parameters based on the strategies learned by the agent.

[0052] S40: According to the sequence of working condition parameters, the performance detection model is used to perform performance detection under several working condition parameters in sequence, and performance evaluation results are output.

[0053] Further, the step S40 of the present application further comprises:

[0054] S41: According to the sequence of working condition parameters, the performance detection model is used to perform performance detection under several working condition parameters in sequence, and several performance detection data are output; S42: Based on the several working condition parameters and the several performance detection data, the performance of the clamping auxiliary device is comprehensively evaluated, and a performance detection report is generated as the performance evaluation result.

[0055] Specifically, according to the sequence of working condition parameters, the performance detection model is used to perform performance detection under several working condition parameters in sequence, and in this process, the performance detection model is used to perform clamping auxiliary device performance detection under each working condition parameter, and the result of each detection output is a performance data set, which can reflect the performance change of the clamping auxiliary device under different working conditions, and several performance detection data are output. Then, the clamping auxiliary device performance is comprehensively evaluated according to the several working condition parameters and the several performance detection data, the performance of different working conditions is compared and analyzed, and a detailed performance detection report is generated. The report will be used as the final evaluation result of the performance of the clamping auxiliary device, and will provide data support for subsequent optimization or improvement.

[0056] In summary, the clamping auxiliary device performance evaluation method based on reinforcement learning provided by the present application has the following technical effects:

[0057] By combining parameters according to the working condition indicators of the clamping auxiliary device, a working condition parameter space is constructed. Then, taking the standard working condition parameters of the clamping auxiliary device as the benchmark, the deviation of several working condition parameters in the working condition parameter space is evaluated, and a sequence of working condition parameters is constructed. Further, a performance detection model is constructed based on reinforcement learning, wherein the performance detection model is a clamping mechanics relationship model between the clamping auxiliary device and the workpiece, which can optimize the clamping operation parameters. Finally, according to the sequence of working condition parameters, the performance detection model is used to perform performance detection under several working condition parameters in sequence, and the performance evaluation result is output. That is, by evaluating the difficulty of different working conditions and combining the reinforcement learning algorithm, the performance test is gradually performed from easy to difficult, which can reduce blind test and unnecessary repeated steps, effectively improve the test precision and save the test time, thereby significantly improving the accuracy and efficiency of performance test.

[0058] Embodiment two, based on the same inventive concept as the clamping auxiliary device performance evaluation method based on reinforcement learning in the foregoing embodiments, the present application also provides a clamping auxiliary device performance evaluation system based on reinforcement learning. Please refer to the accompanying Figure 2, comprising: a working condition parameter space construction module 11, configured to combine parameters randomly according to working condition indicators of the clamping auxiliary device to construct a working condition parameter space; a working condition parameter sequence construction module 12, configured to evaluate the deviation of each working condition parameter in the working condition parameter space based on a standard working condition parameter of the clamping auxiliary device to construct a working condition parameter sequence; a performance detection model construction module 13, configured to construct a performance detection model based on reinforcement learning, wherein the performance detection model is a clamping mechanical relationship model between the clamping auxiliary device and the workpiece, and can optimize clamping operation parameters; and a performance evaluation result output module 14, configured to perform performance detection under each working condition parameter in sequence by using the performance detection model according to the working condition parameter sequence, and output a performance evaluation result.

[0059] Further, the clamping auxiliary device performance evaluation system based on reinforcement learning is further configured to: configure working condition indicators of the clamping auxiliary device, wherein the working condition indicators at least include workpiece material, workpiece shape, processing type, processing parameter and environmental parameter.

[0060] Further, the clamping auxiliary device performance evaluation system based on reinforcement learning is further configured to: divide the processing parameter and the environmental parameter according to a predetermined step length to obtain a processing parameter interval set and an environmental parameter interval set, wherein the processing parameter is set according to the processing type, and the environmental parameter at least includes temperature, humidity and vibration; and combine parameters randomly according to the workpiece material, the workpiece shape, the processing type, the processing parameter interval set and the environmental parameter interval set to generate a plurality of working condition parameters to construct a working condition parameter space.

[0061] Further, the clamping auxiliary device performance evaluation system based on reinforcement learning is further configured to: obtain a standard working condition parameter of the clamping auxiliary device, wherein the standard working condition parameter includes standard workpiece material, standard workpiece shape, standard processing type, standard processing parameter and standard environmental parameter; and analyze the similarity of each working condition parameter based on the standard workpiece material, the standard workpiece shape, the standard processing type, the standard processing parameter and the standard environmental parameter to output a plurality of similarities; and subtract the plurality of similarities from 1 respectively to obtain a plurality of deviation degrees.

[0062] Further, the clamping auxiliary device performance evaluation system based on reinforcement learning is further configured to: randomly select any working condition parameter in the plurality of working condition parameters as a first working condition parameter; configure a similarity comparison operator, wherein the similarity comparison operator at least includes cosine similarity and Euclidean distance; and analyze the similarity of the first working condition parameter based on the standard workpiece material, the standard workpiece shape, the standard processing type, the standard processing parameter and the standard environmental parameter by using the cosine similarity and the Euclidean distance, and obtain a first similarity after mean calculation, and obtain a plurality of similarities of a plurality of working condition parameters in sequence.

[0063] Further, the clamping auxiliary device performance evaluation system based on reinforcement learning is further used for: arranging the plurality of working condition parameters in ascending order of the plurality of deviation degrees, and constructing a working condition parameter sequence.

[0064] Further, the clamping auxiliary device performance evaluation system based on reinforcement learning is further used for: collecting modeling parameters related to the clamping auxiliary device, establishing a mechanical relationship model between the clamping auxiliary device and the workpiece through finite element analysis; defining a state space and an action space, and designing a reward function; training the mechanical relationship model based on reinforcement learning, combining a deep Q network algorithm, the state space, the action space and the reward function, until a convergence condition is met, and generating a performance detection model.

[0065] Further, the clamping auxiliary device performance evaluation system based on reinforcement learning is further used for: the state space includes clamping force, clamping speed and jig rigidity, the action space includes adjusting clamping force, adjusting clamping speed and adjusting jig rigidity, and the evaluation indexes of the reward function include clamping accuracy, jig stability, workpiece deformation control and production efficiency.

[0066] Further, the clamping auxiliary device performance evaluation system based on reinforcement learning is further used for: performing performance detection under the plurality of working condition parameters in sequence by using the performance detection model according to the working condition parameter sequence, and outputting a plurality of performance detection data; performing comprehensive performance evaluation of the clamping auxiliary device according to the plurality of working condition parameters and the plurality of performance detection data, generating a performance detection report as a performance evaluation result.

[0067] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The clamping auxiliary device performance evaluation method based on reinforcement learning in the first embodiment is also applicable to the clamping auxiliary device performance evaluation system based on reinforcement learning in the present embodiment, and the clamping auxiliary device performance evaluation system based on reinforcement learning in the present embodiment can be clearly understood by the person skilled in the art through the foregoing detailed description of the clamping auxiliary device performance evaluation method based on reinforcement learning. Therefore, in order to make the specification concise, it will not be described in detail here. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant part is described in the method part.

[0068] The foregoing description of the disclosed embodiments enables a person skilled in the art to make or use the application. Numerous modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without the use of the inventive faculty. Therefore, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0069] It will be readily apparent to one skilled in the art that varying substitutions and modifications can be made to the application without departing from the scope and spirit of the application. Accordingly, it is intended that all such alterations and modifications be considered as within the scope of the application.

Claims

1. A clamping assist device performance evaluation method based on reinforcement learning, characterized by, The method comprises: According to the working condition index of the clamping auxiliary device, parameters are randomly combined to construct a working condition parameter space; Taking the standard working condition parameters of the clamping auxiliary device as a reference, the deviation of a plurality of working condition parameters in the working condition parameter space is evaluated respectively to construct a working condition parameter sequence; Based on reinforcement learning, a performance detection model is constructed, wherein the performance detection model is a clamping mechanical relationship model between the clamping auxiliary device and the workpiece, which can optimize the clamping operation parameters; According to the working condition parameter sequence, the performance detection model is used to perform performance detection under a plurality of working condition parameters in sequence, and a performance evaluation result is output; Taking the standard working condition parameters of the clamping auxiliary device as a reference, the deviation of a plurality of working condition parameters in the working condition parameter space is evaluated respectively, comprising: Obtaining the standard working condition parameters of the clamping auxiliary device, wherein the standard working condition parameters include standard workpiece material, standard workpiece shape, standard processing type, standard processing parameters and standard environmental parameters; Taking the standard workpiece material, standard workpiece shape, standard processing type, standard processing parameters and standard environmental parameters as a reference, the similarity of a plurality of working condition parameters is analyzed respectively, and a plurality of similarities are output; The plurality of similarities are subtracted by 1 respectively to obtain a plurality of deviation degrees; Taking the standard workpiece material, standard workpiece shape, standard processing type, standard processing parameters and standard environmental parameters as a reference, the similarity of a plurality of working condition parameters is analyzed respectively, comprising: Randomly selecting any working condition parameter in the plurality of working condition parameters as a first working condition parameter; Configuring a similarity comparison operator, wherein the similarity comparison operator at least includes cosine similarity and Euclidean distance; Using cosine similarity and Euclidean distance, taking the standard workpiece material, standard workpiece shape, standard processing type, standard processing parameters and standard environmental parameters as a reference, the similarity of the first working condition parameter is analyzed, and a plurality of similarities of a plurality of working condition parameters are obtained after mean calculation, and a plurality of similarities of a plurality of working condition parameters are obtained in sequence; Based on the plurality of deviation degrees, a plurality of working condition parameters are arranged according to the deviation degrees from small to large to construct a working condition parameter sequence.

2. The method of claim 1, wherein, The working condition index of the clamping auxiliary device is configured, wherein the working condition index at least includes workpiece material, workpiece shape, processing type, processing parameters and environmental parameters.

3. The method of claim 2, wherein, According to the working condition index, parameters are randomly combined to construct a working condition parameter space, comprising: According to the predetermined step, the processing parameters and the environmental parameters are divided to obtain a processing parameter interval set and an environmental parameter interval set, wherein the processing parameters are set according to the processing type, and the environmental parameters at least include temperature, humidity and vibration; According to the workpiece material, the workpiece shape, the processing type, the processing parameter interval set and the environmental parameter interval set, a plurality of working condition parameters are randomly combined to construct a working condition parameter space.

4. The method of claim 1, wherein, Based on reinforcement learning, a performance detection model is constructed, comprising: Collecting modeling parameters related to the clamping auxiliary device, and establishing a mechanical relationship model between the clamping auxiliary device and the workpiece through finite element analysis; Defining state space and action space, and designing reward function; Based on reinforcement learning, the mechanical relationship model is trained by combining deep Q network algorithm, state space, action space and reward function until the convergence condition is met, and a performance detection model is generated.

5. The method of claim 4, wherein, The state space includes clamping force, clamping speed and clamp stiffness, the action space includes adjusting clamping force, adjusting clamping speed and adjusting clamp stiffness, and the evaluation indexes of the reward function include clamping accuracy, clamp stability, workpiece deformation control and production efficiency.

6. The method of claim 1, wherein, The performance evaluation result includes: According to the sequence of working condition parameters, the performance detection model is used to perform performance detection under several working condition parameters in turn, and several performance detection data are outputted; According to several working condition parameters and several performance detection data, the performance of the clamping auxiliary device is comprehensively evaluated, and a performance detection report is generated as the performance evaluation result.

7. A clamping assist device performance evaluation system based on reinforcement learning, characterized by, The steps for implementing the clamping auxiliary device performance evaluation method based on reinforcement learning in any one of claims 1 to 6 include: A working condition parameter space construction module is configured to combine parameters randomly according to the working condition indexes of the clamping auxiliary device to construct a working condition parameter space; A working condition parameter sequence construction module is configured to evaluate the deviation of several working condition parameters in the working condition parameter space based on the standard working condition parameters of the clamping auxiliary device to construct a working condition parameter sequence; A performance detection model construction module is configured to construct a performance detection model based on reinforcement learning, wherein the performance detection model is a clamping mechanical relationship model between the clamping auxiliary device and the workpiece, which can optimize the clamping operation parameters; A performance evaluation result output module is configured to use the performance detection model to perform performance detection under several working condition parameters in turn according to the sequence of working condition parameters, and output the performance evaluation result.

Citation Information

Patent Citations

  • Clamping force prediction model training method, prediction method, device, equipment and medium

    CN118485184A

  • Cutter clamping state evaluation method based on self-encoding reconstruction characteristics

    CN119046875A