A magnetic grating fixed-length adjustment system based on wireless transmission control
By adopting wireless transmission control and layered reinforcement learning technology in the magnetic gate measurement system, the control parameters are optimized and adaptive adjustment is realized, the problems of low optimization efficiency and migration difficulties in existing systems are solved, and the measurement accuracy and system intelligence level are improved.
Patent Information
- Application Number
- CN202510432220.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The existing magnetic grid measurement system relies on manual setting of control parameters, has low optimization efficiency, weak environmental adaptability, and is difficult to migrate across devices and across scenarios, resulting in limited measurement accuracy and system intelligence level.
The magnetic gate fixed-length adjustment system based on wireless transmission control is adopted, including a digital twin simulation module, a layered reinforcement learning control module, a cross-domain migration evolution module and a biological stress response module. Adaptive adjustment and cross-domain migration are achieved through physical modeling and reinforcement learning.
Adaptive control parameter optimization is realized, measurement accuracy is improved, human intervention is reduced, system intelligence is improved, and the deployment process of different devices and scenarios is simplified through cross-domain migration evolution method.
Smart Images

Figure CN119937295B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of magnetic grating measurement optimization, and particularly relates to a magnetic grating fixed-length adjustment system based on wireless transmission control. Background Art
[0002] Magnetic grating measurement systems are widely used in the fields of precision positioning and length measurement, such as industrial automation, numerical control machine tools, and precision measurement equipment. Currently, magnetic grating measurement systems usually rely on manual setting of control parameters to optimize measurement accuracy. However, this method has the following problems: 1. Parameter adjustment depends on experience, and the optimization efficiency is low: under different working conditions, it is difficult for manual adjustment to find the global optimal control parameters, resulting in poor system adaptability; 2. Weak environmental adaptability: factors such as temperature changes, mechanical vibrations, and electromagnetic interference may affect measurement accuracy, and existing systems lack an adaptive adjustment mechanism; 3. Difficult to migrate across devices and scenarios: under different devices and application scenarios, control strategies need to be re-adjusted, increasing the deployment cost. Summary of the Invention
[0003] The present invention provides a magnetic grating fixed-length adjustment system based on wireless transmission control, which solves the technical problems in related technologies that rely on manual setting of control parameters, have low optimization efficiency, weak environmental adaptability, and are difficult to migrate across devices and scenarios, resulting in limited measurement accuracy and system intelligence level.
[0004] The present invention provides a magnetic grating fixed-length adjustment system based on wireless transmission control, including:
[0005] A digital twin simulation module, used to construct a 3D simulation model of the magnetic grating system through physical modeling and synchronize multi-source sensor data, where the multi-source sensor data includes: temperature, vibration, magnetic field strength, and stress, and the 3D simulation model includes: electromagnetic field simulation, mechanical vibration simulation, and thermal deformation simulation;
[0006] A hierarchical reinforcement learning control module, used to execute a macro-layer control strategy and a micro-layer control strategy at the macro layer and the micro layer respectively. Among them, the macro-layer control strategy uses a multi-objective hierarchical evolutionary strategy to optimize the first control parameter; the micro-layer control strategy optimizes the second control parameter by constructing a reinforcement learning model; the first control parameter includes: PID gain parameter, filtering coefficient, and wireless communication power, and the second control parameter includes: temperature compensation parameter, hysteresis compensation parameter, and adaptive learning rate;
[0007] A cross-domain migration and evolution module, used to extract common features from the macro-layer control strategy and the micro-layer control strategy trained under different devices and working conditions, construct a parameter gene library, and obtain a control strategy suitable for new devices according to the parameter gene library;
[0008] A biological stress response module is used to perform real-time analysis on multi-source sensor data according to the 3D simulation model of the magnetic grating system, and trigger the first preset control strategy when abnormal behavior is detected.
[0009] Furthermore, a 3D simulation model of the magnetic grating system is constructed through physical modeling, and multi-source sensor data is synchronized in real time. The specific steps include:
[0010] S201, Use CAD software to establish the 3D geometric structure of the magnetic grating system and perform finite element mesh division;
[0011] S202, By simulating the electromagnetic-mechanical and thermal coupling behaviors of the magnetic grating system, perform electromagnetic field simulation, mechanical vibration simulation and thermal deformation simulation on the magnetic grating system respectively;
[0012] S203, Obtain the multi-source sensor data running in real time in the real environment, and input the multi-source sensor data into the 3D simulation model to achieve simulation-physical synchronization.
[0013] Furthermore, the specific steps of the multi-objective hierarchical evolution strategy include:
[0014] S301, Initialize the population using the Latin hypercube sampling method. The population includes P individuals, and each individual is represented by a first control parameter vector composed of first control parameters;
[0015] S302, Construct a multi-objective optimization function;
[0016] S303, Use the non-dominated sorting genetic algorithm to optimize the individuals.
[0017] Furthermore, the multi-objective optimization function includes:
[0018] ;
[0019] ;
[0020] ;
[0021] ;
[0022] Among them, F represents the comprehensive fitness of the individual, that is, the total loss value of the optimization goal, a represents the index of the goal, represents the weight coefficient of the a-th goal, represents the loss value of the a-th goal, , and respectively represent the first fitness, the second fitness and the third fitness, that is, the tracking error loss value, the energy consumption loss value and the noise suppression loss value, t represents the index of the time step, and T represents the number of time steps. represents the preset target value expected by the magnetic grating system at the t-th time step, represents the actual measurement result of the magnetic grating system at the t-th time step, represents the wireless communication power of the magnetic grating system at the t-th time step, represents the filtered sensor measurement value at the t-th time step, represents the unfiltered sensor measurement value at the t-th time step.
[0023] Further, the specific steps of step S303 include:
[0024] S501, calculate the first fitness, second fitness, third fitness, and comprehensive fitness of the individual according to the multi-objective optimization function;
[0025] S502, divide the population into three groups according to the comprehensive fitness of the individual: high-fitness individuals, medium-fitness individuals, and low-fitness individuals;
[0026] S503, retain the high-fitness individuals, perform crossover and mutation operations on the medium-fitness individuals, and perform high-probability mutation operations on the low-fitness individuals;
[0027] S504, according to the first fitness, second fitness, and third fitness of the individual, perform screening using non-dominated sorting, and select 3 Pareto optimal individuals. Among them, the individuals with the highest first fitness, second fitness, and third fitness respectively are used as Pareto optimal individuals;
[0028] S505, when the number of iterations reaches the preset maximum number of iterations, output the current 3 Pareto optimal individuals, and select the individual that best suits the current application scenario according to actual needs, otherwise return to S502.
[0029] Further, the crossover operation uses the simulated binary crossover method to generate new individuals. The calculation formula of the simulated binary crossover method is:
[0030] ;
[0031] where n represents the element index of the first control parameter vector of the individual, represents the value of the n-th element of the first control parameter vector of the newly generated offspring individual, represents a random number between 0 and 1, represents the crossover distribution index that controls the search range, and represent the values of the n-th elements of the first control parameter vectors of two parent individuals, c represents the current iteration number, and C represents the maximum iteration number, represents a Gaussian random variable with a mean of 0 and a standard deviation of 1;
[0032] The mutation operation uses the polynomial mutation method to update individuals. The calculation formula for polynomial mutation is:
[0033] ;
[0034] where represents the value of the nth element of the first control parameter vector after the mutation of the offspring individual, represents the value of the nth element of the first control parameter vector of the offspring individual, represents the mutation distribution index.
[0035] Furthermore, the micro-layer control strategy optimizes the second control parameter by constructing a reinforcement learning model. The specific steps include:
[0036] S601, construct a state vector containing the second control parameter according to the operating state of the magnetic grating system;
[0037] S602, use the TinyML model to construct a policy network. This policy network takes the state vector as input and outputs the adjustment amount of the second control parameter. The calculation formula for the adjustment amount is: , represents the adjustment amount of the second control parameter, represents the policy network, represents the state vector at the t-th time step, represents the network parameters;
[0038] S603, construct an action space, and the action space represents the geometric range of the adjustment amount of the second control parameter;
[0039] S604, construct the reward function of the reinforcement learning model;
[0040] S605, use the policy gradient method to update the policy network.
[0041] Furthermore, the macro-layer control strategy and the micro-layer control strategy are respectively transformed into a macro feature vector and a micro feature vector, and then spliced to obtain a control strategy vector. For the control strategy vectors of different devices and different working conditions, calculate the standard deviation of each parameter in the control strategy vector in turn. When the standard deviation is less than the first preset threshold, it is determined that the parameter changes little among different devices and working conditions, and it is regarded as a common feature. Combine the extracted common features to obtain a parameter gene pool, and the parameter gene pool provides a general control strategy framework for new devices.
[0042] Further, the bio - like stress response module performs real - time analysis on multi - source sensor data based on the 3D simulation model of the magnetic grating system. When the multi - source sensor data exceeds the preset safety threshold, it determines that the magnetic grating system is operating abnormally and triggers the first preset control strategy. A preset time period is set, and the number of times the same abnormality is triggered within this time period is counted. If the same abnormality occurs multiple times within the preset time period and the magnetic grating system still operates stably, the triggering condition of this abnormality is adjusted.
[0043] The beneficial effects of the present invention are as follows: The present invention adopts a hierarchical reinforcement learning control strategy. By optimizing the PID gain, filtering coefficient, and wireless communication power at the macro level, and optimizing temperature compensation, hysteresis compensation, and adaptive learning rate at the micro level, it realizes the optimization of adaptive control parameters. Compared with the traditional method of manually setting parameters, the present invention can automatically learn and adjust the optimal control parameters, improve the measurement accuracy, reduce human intervention, and improve the intelligent level of the system.
[0044] The present invention uses a cross - domain transfer and evolution method to extract common features from control strategies trained under different devices and working conditions, construct a parameter gene pool, and generate a control strategy applicable to new devices accordingly. Compared with the traditional magnetic grating measurement system that requires re - adjusting control parameters, the present invention can quickly transfer the optimal control strategy under different devices and different scenarios, reduce the cost of repeated debugging, and improve the versatility and deployment efficiency of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic diagram of the modules of a magnetic grating fixed - length adjustment system based on wireless transmission control of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] Now, the subject matter described herein will be discussed with reference to exemplary embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described relative to some examples can also be combined in other examples.
[0047] As Figure 1 shown, a magnetic grating fixed - length adjustment system based on wireless transmission control includes:
[0048] A digital twin simulation module 101, which is used to construct a 3D simulation model of the magnetic grating system through physical modeling and synchronize multi - source sensor data in real time to simulate the dynamic behavior of the magnetic grating control system, realize online testing and optimization. The multi - source sensor data includes: temperature, vibration, magnetic field strength, and stress. The 3D simulation model includes: electromagnetic field simulation, mechanical vibration simulation, and thermal deformation simulation.
[0049] Electromagnetic field simulation is used to simulate the propagation of magnetic grating signals and the influence of electromagnetic interference. Mechanical vibration simulation is used to evaluate the rigidity and dynamic stability of the system. Thermal deformation simulation is used to analyze the influence of temperature change on the measurement accuracy of the magnetic grating;
[0050] The hierarchical reinforcement learning control module 102 is used to execute the macro-layer control strategy and the micro-layer control strategy at the macro layer and the micro layer respectively. Among them, the macro-layer control strategy uses a multi-objective hierarchical evolutionary strategy to optimize the first control parameter; the micro-layer control strategy optimizes the second control parameter by constructing a reinforcement learning model; the first control parameter includes: PID gain parameter, filtering coefficient and wireless communication power, and the second control parameter includes: temperature compensation parameter, hysteresis compensation parameter and adaptive learning rate;
[0051] The cross-domain transfer and evolution module 103 is used to extract common features from the macro-layer control strategy and the micro-layer control strategy trained under different devices and different working conditions, construct a parameter gene pool, and obtain a control strategy suitable for the new device according to the parameter gene pool;
[0052] The bio-inspired stress response module 104 is used to trigger a first preset control strategy when an abnormal behavior is detected according to the real-time data of the sensors of the 3D simulation model of the magnetic grating system.
[0053] In an embodiment of the present invention, the vibration of the multi-source sensor data is represented by vibration acceleration, and a MEMS sensor is used to obtain the vibration acceleration; a thermocouple sensor is used to collect the temperature; a Hall sensor is used to collect the magnetic field intensity; a piezoresistive strain gauge is used to collect the stress.
[0054] In an embodiment of the present invention, a 3D simulation model of the magnetic grating system is constructed through physical modeling, and the multi-source sensor data is synchronized in real time. The specific steps include:
[0055] S201, use CAD software to establish the 3D geometric structure of the magnetic grating system and perform finite element mesh division. Specifically, use AutoCAD software to create 3D models of components such as the magnetic grating ruler, guide rail, and actuator in the magnetic grating system, and use HyperMesh for finite element mesh division;
[0056] S202. By simulating the electromagnetic-mechanical and thermal coupling behaviors of the magnetic grating system, perform electromagnetic field simulation, mechanical vibration simulation, and thermal deformation simulation on the magnetic grating system respectively. Specifically, use ANSYS Maxwell to perform finite element analysis of the electromagnetic field, set boundary conditions: define the magnetic induction intensity, set the magnetization direction, magnetic permeability, remanence, and coercive force of the magnetic grating, and calculate the magnetic field distribution and hysteresis error; use ANSYS Mechanical to perform vibration mode analysis and transient analysis, set boundary conditions: constrain the fixed boundary of the magnetic grating guide rail, apply a mechanical vibration source, and calculate the vibration frequency and stress distribution; use ANSYS Fluent to perform thermal deformation simulation, set thermal boundary conditions: apply environmental temperature changes, set the thermal conductivity and specific heat capacity of the material, and calculate the temperature distribution and measurement error caused by thermal expansion.
[0057] S203. Obtain multi-source sensor data of real-time operation in the real environment, and input the multi-source sensor data into the 3D simulation model to achieve simulation-physical synchronization. Specifically, use Kalman filtering to process the multi-source sensor data, reduce the influence of noise, and input it into the 3D simulation model to optimize the magnetic grating measurement accuracy.
[0058] In an embodiment of the present invention, the specific steps of the multi-objective hierarchical evolution strategy include:
[0059] S301. Use the Latin hypercube sampling method to initialize the population. The population includes P individuals, and each individual is represented by a first control parameter vector composed of first control parameters.
[0060] The encoding format of the individual is: , where Q represents the individual represented by the first control parameter vector, , and respectively represent the proportional gain, integral gain, and derivative gain of the PID gain, represents the filtering coefficient, represents the wireless communication power;
[0061] S302. Construct a multi-objective optimization function;
[0062] S303. Use the non-dominated sorting genetic algorithm to optimize the individuals.
[0063] In an embodiment of the present invention, the specific steps of using the Latin hypercube sampling method to initialize the population include:
[0064] S401. Determine the range of each parameter in the first control parameters;
[0065] S402. Uniformly divide the range of each parameter into P intervals;
[0066] S403. Randomly select a value within each of the P intervals of each parameter as the value of the individual parameter, and finally obtain a population consisting of P individuals.
[0067] In an embodiment of the present invention, the multi-objective optimization function includes:
[0068] ;
[0069] ;
[0070] ;
[0071] ;
[0072] where F represents the comprehensive fitness of the individual, that is, the total loss value of the optimization objective, a represents the index of the objective, represents the weight coefficient of the a-th objective, represents the loss value of the a-th objective, , and respectively represent the first fitness, the second fitness, and the third fitness, that is, the tracking error loss value, the energy consumption loss value, and the noise suppression loss value. t represents the index of the time step, and T represents the number of time steps. represents the preset target value expected by the magnetic grating system at the t-th time step, such as the expected position, the expected speed, and the expected state value. represents the actual measurement result of the magnetic grating system at the t-th time step. represents the wireless communication power of the magnetic grating system at the t-th time step. represents the filtered sensor measurement value at the t-th time step. represents the unfiltered sensor measurement value at the t-th time step.
[0073] In an embodiment of the present invention, the specific steps of step S303 include:
[0074] S501. Calculate the first fitness, the second fitness, the third fitness, and the comprehensive fitness of the individual according to the multi-objective optimization function. Among them, the weight coefficient in the multi-objective optimization function adopts an adaptive adjustment method, and the calculation formula of the weight coefficient is: , represents the weight coefficient of the a-th objective at the (b + 1)-th iteration. represents the weight coefficient of the a-th objective at the b-th iteration. represents the learning rate that controls the weight adjustment speed. represents the gradient change of the loss value of the a-th objective with respect to the individual. represents the individual represented by the first control parameter vector;
[0075] S502, divide the population into three groups according to the comprehensive fitness of individuals: high-fitness individuals, medium-fitness individuals, and low-fitness individuals;
[0076] S503, retain the high-fitness individuals, perform crossover and mutation operations on the medium-fitness individuals, and perform high-probability mutation operations on the low-fitness individuals;
[0077] S504, according to the first fitness, second fitness, and third fitness of individuals, perform non-dominated sorting for screening, and select 3 Pareto optimal individuals. Among them, the individuals with the highest first fitness, second fitness, and third fitness respectively are used as Pareto optimal individuals;
[0078] S505, when the number of iterations reaches the preset maximum number of iterations, output the current 3 Pareto optimal individuals, and select the individual that best meets the current application scenario according to actual requirements. Otherwise, return to S502, where the actual requirements include but are not limited to: reducing the tracking error loss value, reducing the energy consumption loss value, or reducing the noise suppression loss value.
[0079] In an embodiment of the present invention, the crossover operation uses the simulated binary crossover method to generate new individuals. The calculation formula of the simulated binary crossover method is:
[0080] ;
[0081] where n represents the element index of the first control parameter vector of the individual, represents the value of the nth element of the first control parameter vector of the newly generated offspring individual, represents a random number between 0 and 1, represents the crossover distribution index for controlling the search range, The larger it is, the closer the offspring is to the parent, The smaller it is, the more dispersed the offspring is, and represent the values of the nth element of the first control parameter vectors of two parent individuals, c represents the current number of iterations, C represents the maximum number of iterations, represents a Gaussian random variable with a mean of 0 and a standard deviation of 1;
[0082] The mutation operation uses the polynomial mutation method to update the individual. The calculation formula of the polynomial mutation is:
[0083] ;
[0084] where, represents the value of the nth element of the first control parameter vector of the offspring individual after mutation, represents the value of the nth element of the first control parameter vector of the offspring individual, represents the mutation distribution index, The larger it is, the smaller the change in individual update, The smaller it is, the greater the change in individual update.
[0085] The high-probability mutation operation is based on the mutation operation and dynamically adjusts the mutation probability. The calculation formula of the high-probability mutation operation is:
[0086] , represents the mutation probability of the high-probability mutation operation, represents the mutation probability of the mutation operation, represents the mutation increment coefficient, and its value range is from 0 to 1, represents the comprehensive fitness of the individual, represents the maximum comprehensive fitness of the individuals in the population.
[0087] In an embodiment of the present invention, at the microscopic level, the magnetic grating system needs to finely adjust the second control parameter in real time to compensate for the control deviation caused by temperature change, hysteresis effect or other disturbances. For this purpose, the present invention uses a lightweight reinforcement learning model to run on an embedded controller to achieve adaptive adjustment of the second control parameter, enabling it to quickly respond to environmental changes and optimize system performance; the microscopic layer control strategy optimizes the second control parameter by constructing a reinforcement learning model. The specific steps include:
[0088] S601. According to the operating state of the magnetic grating system, construct a state vector containing the second control parameter. The state vector contains factual information related to the adjustment of the second control parameter. The expression of the state vector is: , where, represents the state vector at the tth time step, represents the temperature, represents the error between the preset target value expected by the magnetic grating system at the tth time step and the actual measurement result, represents the vibration acceleration at the tth time step, represents the magnetic field strength at the tth time step;
[0089] S602. Use the TinyML model to construct a policy network. This policy network takes the state vector as input and outputs the adjustment amount of the second control parameter. The calculation formula of the adjustment amount is: , represents the adjustment amount of the second control parameter, represents the policy network, represents the network parameters;
[0090] S603. Construct an action space. The action space represents the geometric range of the adjustment amount of the second control parameter. An action represents the adjustment amount of the second control parameter. Update the second control parameter through the adjustment amount. The update formula is: , represents the updated adjustment amount, represents the value of the current second control parameter;
[0091] S604. Construct the reward function of the reinforcement learning model. The reward function is:
[0092] , where represents the value of the reward function at the t-th time step, and respectively represent the first weight coefficient and the second weight coefficient, which are used to balance the relationship between reducing errors and avoiding drastic adjustments;
[0093] S605. Update the policy network using the policy gradient method. The calculation formula of the policy gradient method is:
[0094] , represents the gradient calculation of the network parameter , that is, adjust the direction of the policy network to make its output a better adjustment amount, represents the performance index of the policy network, represents at the state , when the adjustment amount is the probability distribution,
[0095] represents the expected value of, that is, statistical estimation is carried out under different state and action combinations; Through the policy gradient method, the policy network continuously optimizes its own parameters to make it able to output a better adjustment amount, and finally realizes the adaptive optimization of the magnetic grating system.
[0096] The macro-layer control strategy is used for long-term and global optimization. The multi-objective hierarchical evolutionary strategy is used to optimize the first control parameter and adjust the overall operating state of the device to make it reach the best performance in different working modes. The time scale of the policy adjustment is long and it will not be modified frequently; The micro-layer control strategy is mainly responsible for adjusting the second control parameter in real time and adaptively.
[0097] In an embodiment of the present invention, the macro-layer control strategy and the micro-layer control strategy are respectively transformed into a macro feature vector and a micro feature vector, and are spliced to obtain a control strategy vector; For example, the control strategy vector can be represented by
[0098] , where represents the temperature compensation parameter, represents the hysteresis compensation parameter, represents the adaptive learning rate; for the control strategy vectors of different devices and operating conditions, calculate the standard deviation of each parameter in the control strategy vector in turn. When the standard deviation is less than the first preset threshold, it is determined that the parameter changes little among different devices and operating conditions, and it is regarded as a common feature. Combine the extracted common features to obtain a parameter gene pool, and the parameter gene pool provides a general control strategy framework for new devices; for example, the parameter gene pool contains: the proportional gain is set to 1.5, and the derivative gain is set to 0.8. When generating a control strategy for a new device, first extract the control strategy framework from the parameter gene pool, and then fine-tune it according to the specific requirements of the new device. For example, if the temperature of the new device is relatively high, adjust the filtering coefficient to optimize the system performance.
[0099] In an embodiment of the present invention, the bio-inspired stress response module performs real-time analysis on multi-source sensor data according to the 3D simulation model of the magnetic grating system. When the multi-source sensor data exceeds the preset safety threshold, it is determined that the magnetic grating system is operating abnormally, and a first preset control strategy is triggered, where the safety threshold is preset according to historical operation data and device operating conditions; set a preset time period, and count the number of times the same abnormality is triggered within this time period. If the same abnormality occurs multiple times within the preset time period and the magnetic grating system still operates stably, adjust the trigger condition for this abnormality; for example, when the temperature of the magnetic grating system exceeds the preset temperature threshold, adjust the temperature compensation parameter of the magnetic grating system and reduce the wireless communication power to reduce system heat generation; when the situation that the temperature of the magnetic grating system exceeds the preset temperature threshold is triggered more than 5 times within 10 minutes and the system still operates stably, increase the preset temperature threshold to the highest temperature among the triggered abnormalities.
[0100] The above describes the embodiments of this embodiment, but this embodiment is not limited to the above specific implementation manners. The above specific implementation manners are only illustrative and not restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make many forms, all of which fall within the protection scope of this embodiment.
Claims
1. A magnetic grid fixed length adjustment system based on wireless transmission control, characterized in that: include: A digital twin simulation module is used to construct a 3D simulation model of the magnetic grid system through physical modeling and synchronize multi-source sensor data, wherein the multi-source sensor data includes temperature, vibration, magnetic field strength and stress, and the 3D simulation model includes electromagnetic field simulation, mechanical vibration simulation and thermal deformation simulation; A hierarchical reinforcement learning control module is used to execute a macro-layer control strategy and a micro-layer control strategy at the macro-layer and the micro-layer, respectively, wherein the macro-layer control strategy uses a multi-objective hierarchical evolutionary strategy to optimize a first control parameter; the micro-layer control strategy optimizes a second control parameter by constructing a reinforcement learning model; the first control parameter includes: a PID gain parameter, a filter coefficient, and a wireless communication power; the second control parameter includes: a temperature compensation parameter, a hysteresis compensation parameter, and an adaptive learning rate; Among them, the specific steps of the multi-objective hierarchical evolution strategy include: S301, using a Latin hypercube sampling method to initialize a population, the population including P individuals, each individual being represented by a first control parameter vector composed of first control parameters; S302, constructing a multi-objective optimization function, wherein the multi-objective optimization function includes: ; ; ; ; Among them, F represents the comprehensive fitness of the individual, that is, the total loss value of the optimization target, a represents the index of the target, represents the weight coefficient of the ath target, represents the loss value of the a-th target, , and They represent the first fitness, the second fitness and the third fitness, namely the tracking error loss value, the energy consumption loss value and the noise suppression loss value, t represents the index of the time step, T represents the number of time steps, represents the preset target value expected by the magnetic grid system at the tth time step, represents the actual measurement result of the magnetic grid system at the tth time step, represents the wireless communication power of the magnetic grid system at the tth time step, represents the filtered sensor measurement value at the tth time step, represents the sensor measurement value without filtering at the tth time step; S303, optimizing individuals using a non-dominated sorting genetic algorithm; The specific steps of the micro-level control strategy include: S601, constructing a state vector including a second control parameter according to the operating state of the magnetic grid system, wherein the expression of the state vector is: ,in, represents the state vector at the tth time step, Indicates temperature, It represents the error between the preset target value expected by the magnetic grid system at the tth time step and the actual measurement result. represents the vibration acceleration at the tth time step, represents the magnetic field intensity at the tth time step; S602, using the TinyML model to build a policy network, the policy network takes the state vector as input and outputs the adjustment amount of the second control parameter, and the calculation formula of the adjustment amount is: , represents the adjustment amount of the second control parameter, represents the policy network, Represents network parameters; S603, constructing an action space, where the action space represents a geometric range of an adjustment amount of the second control parameter; S604, constructing a reward function of a reinforcement learning model; S605, updating the policy network using a policy gradient method; The cross-domain migration evolution module is used to extract common features from macro-level control strategies and micro-level control strategies trained under different equipment and different working conditions, build a parameter gene library, and obtain control strategies suitable for new equipment based on the parameter gene library; The biological stress response module is used to perform real-time analysis of multi-source sensor data based on the 3D simulation model of the magnetic grid system, and trigger the first preset control strategy when abnormal behavior is detected.
2. According to the magnetic grid fixed length adjustment system based on wireless transmission control according to claim 1, it is characterized in that: The 3D simulation model of the magnetic grid system is constructed through physical modeling, and multi-source sensor data is synchronized in real time. The specific steps include: S201, use CAD software to establish the 3D geometric structure of the magnetic grid system and perform finite element meshing; S202, by simulating the electromagnetic mechanical and thermal coupling behaviors of the magnetic grid system, performing electromagnetic field simulation, mechanical vibration simulation and thermal deformation simulation on the magnetic grid system respectively; S203, acquiring multi-source sensor data of real-time operation in a real environment, and inputting the multi-source sensor data into a 3D simulation model to achieve simulation-physical synchronization.
3. According to the magnetic grid fixed length adjustment system based on wireless transmission control according to claim 1, it is characterized in that: The specific steps of step S303 include: S501, calculating the first fitness, the second fitness, the third fitness and the comprehensive fitness of the individual according to the multi-objective optimization function; S502, dividing the population into three groups according to the comprehensive fitness of individuals: high fitness individuals, medium fitness individuals and low fitness individuals; S503, retaining high fitness individuals, performing crossover and mutation operations on medium fitness individuals, and performing high probability mutation operations on low fitness individuals; S504, selecting three Pareto optimal individuals by using non-dominated sorting according to the first fitness, the second fitness and the third fitness of the individuals, wherein the individual with the highest first fitness, the highest second fitness and the highest third fitness is selected as the Pareto optimal individual; S505, when the number of iterations reaches the preset maximum number of iterations, output the current three Pareto optimal individuals, and select the individual that best meets the current application scenario according to actual needs, otherwise return to S502.
4. According to the magnetic grid fixed length adjustment system based on wireless transmission control according to claim 3, it is characterized in that: The crossover operation uses the simulated binary crossover method to generate new individuals. The calculation formula of the simulated binary crossover method is: ; Where n represents the element index of the first control parameter vector of the individual, Represents the nth element value of the first control parameter vector of the newly generated offspring individual, Represents a random number between 0 and 1. represents the cross-distribution index that controls the search range, and represents the nth element value of the first control parameter vector of the two parent individuals, c represents the current number of iterations, and C represents the maximum number of iterations. It represents a Gaussian random variable with a mean of 0 and a standard deviation of 1. The mutation operation uses the polynomial mutation method to update individuals. The calculation formula of polynomial mutation is: ; in, Represents the nth element value of the first control parameter vector after the offspring individual mutation, represents the nth element value of the first control parameter vector of the offspring individual, Represents the variation distribution index.
5. According to the magnetic grid fixed length adjustment system based on wireless transmission control according to claim 1, it is characterized in that: The macro-level control strategy and the micro-level control strategy are respectively converted into a macro-feature vector and a micro-feature vector, and are concatenated to obtain a control strategy vector; for the control strategy vectors of different devices and different working conditions, the standard deviation of each parameter in the control strategy vector is calculated in turn. When the standard deviation is less than a first preset threshold, it is determined that the parameter varies little between different devices and working conditions, and is regarded as a common feature. The extracted common features are combined to obtain a parameter gene library, which provides a general control strategy framework for new devices.
6. The magnetic grid fixed length adjustment system based on wireless transmission control according to claim 1 is characterized in that: The biological stress response module performs real-time analysis on multi-source sensor data based on the 3D simulation model of the magnetic grid system. When the multi-source sensor data exceeds a preset safety threshold, the magnetic grid system is judged to be operating abnormally and the first preset control strategy is triggered. Set a preset time period and count the number of times the same abnormality is triggered within the time period. If the same abnormality occurs multiple times within the preset time period and the magnetic grid system is still running stably, adjust the triggering condition of the abnormality.
Citation Information
Patent Citations
Improved PSO and SVM based encoder fault diagnosis system and method
CN108828944A
Magnetic grid fixed-length structure controlled by wireless transmission
CN113551586A