Production decision adaptive optimization method and system

Through real-time monitoring of the entropy value of the strategy and the reconstruction of the digital twin simulation model, the dynamic generation strategy explores priority, solving the path dependence problem of industrial production optimization systems, and achieving adaptive optimization and continuous adaptability improvement of production strategies.

CN120406165AActive Publication Date: 2025-08-01AUTOMOTIVE ENGINEERING CORPORATION +1

Patent Information

Application Number
CN202510898870.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-08-01
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

Existing industrial production optimization systems are susceptible to historical policy data, resulting in strong path dependence of optimization solutions, difficulty in adapting to dynamically changing production demands and external disturbances, and limited adaptability.

Method used

By monitoring the policy entropy value in real time, the digital twin simulation model is triggered to be reconstructed, the policy parameter space is reset, the digital twin simulation model after the policy space is reset for simulation verification, filter the subset of feasible policies, and dynamically generate the strategy based on the KL divergence difference to explore priority, and finally deploy the policy with the highest priority to the actual production system.

Benefits of technology

It improves the adaptability and sustainability of production strategies, ensures that the optimization system fully releases the strategy iteration potential while ensuring stability, adapts to external disturbances and internal parameter drifts, and realizes the intelligent evolution of the optimization process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406165A_ABST
    Figure CN120406165A_ABST
Patent Text Reader

Abstract

The invention discloses a production decision adaptive optimization method and system, particularly relates to the technical field of industrial production process optimization, and is used for solving the problems of strategy stiffness and limited adaptive ability caused by historical strategy data dependence of an existing optimization system. Model reconstruction is dynamically triggered by monitoring strategy entropy in real time, candidate strategies are generated in combination with parameter space resetting and multi-physics field simulation verification, strategy exploration priorities are dynamically allocated based on KL divergence difference, and a strategy set is continuously optimized through a closed-loop feedback mechanism. According to the method, path dependence can be actively broken when strategy diversity is attenuated, a dynamic priority distribution balance strategy is utilized to inherit and explore, a data acquisition window is adjusted in combination with deviation rate feedback, a self-adaptive optimization closed loop is formed, the adaptability and decision elasticity of a production strategy are effectively improved, and the efficiency stability in long-term operation is guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of industrial production process optimization, and more specifically, to a production decision adaptive optimization method and system. Background Art

[0002] In the field of industrial production process optimization, digital twin technology and simulation verification methods have been widely used in the generation and verification of production strategies; existing technologies usually evaluate candidate strategies in multiple dimensions by constructing a virtual simulation environment, and screen out strategy solutions that meet preset goals based on historical optimization results. This method can significantly reduce the trial-and-error cost in actual production and improve the resource allocation efficiency within a certain period.

[0003] However, the long-term running optimization system is vulnerable to the influence of historical strategy data, resulting in the subsequent generated optimization solutions showing path dependence. Specifically, the iterative evolution direction of the optimization strategy gradually solidifies, making it difficult to break through the existing decision-making mode, and thus unable to adapt to the dynamic production requirements and external disturbances. This strategy rigidity phenomenon limits the adaptive ability of the optimization system and ultimately affects the sustainability of production efficiency improvement. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide a production decision adaptive optimization method and system to solve the problems proposed in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A production decision adaptive optimization method, comprising the following steps:

[0007] S1. Real-time monitor the strategy entropy value of the current optimization strategy set, and the strategy entropy value is obtained by calculating the distribution dispersion degree of each strategy in historical iterations;

[0008] S2. When the strategy entropy value is lower than the preset entropy value threshold, trigger the reconstruction instruction of the digital twin simulation model;

[0009] S3. Reset the strategy space of the digital twin simulation model according to the reconstruction instruction to generate a candidate strategy set including a new strategy parameter space;

[0010] S4. Use the digital twin simulation model after strategy space reset to simulate and verify the candidate strategy set, and screen out a feasible strategy subset that meets the preset constraint conditions;

[0011] S5. Dynamically generate the strategy exploration priority based on the KL divergence difference between each strategy in the feasible strategy subset and the historical strategy;

[0012] S6. Deploy the strategy with the highest exploration priority to the actual production system and update the current set of optimization strategies.

[0013] In a preferred embodiment, the policy entropy value of the current set of optimization strategies is monitored in real time. The policy entropy value is obtained by calculating the distribution dispersion of each policy in historical iterations, including:

[0014] Obtain the policy parameter vector of each policy in the current set of optimization strategies in historical iterations;

[0015] Extract the feature dimensions of all policy parameter vectors within a preset time window;

[0016] Calculate the variance value of each dimension parameter of the policy parameter vector based on the numerical fluctuation range of the feature dimensions;

[0017] Generate the policy entropy value by weighted summation of the variance values of each dimension parameter.

[0018] In a preferred embodiment, when the policy entropy value is lower than the preset entropy threshold, a reconstruction instruction for the digital twin simulation model is triggered, including:

[0019] Obtain the dynamic adjustment rule of the preset entropy threshold, compare the currently monitored policy entropy value with the preset entropy threshold, and generate a threshold deviation evaluation result;

[0020] Generate a reconstruction instruction for the digital twin simulation model including the reconstruction intensity level according to the judgment that the threshold deviation evaluation result exceeds the preset deviation threshold;

[0021] Transmit the reconstruction instruction to the control interface of the digital twin simulation model, and the control interface performs corresponding parameter initialization operations according to the reconstruction intensity level.

[0022] In a preferred embodiment, according to the reconstruction instruction, the policy parameter space of the digital twin simulation model is reset to generate a candidate policy set including a new policy parameter space, including:

[0023] Analyze the reconstruction intensity level in the reconstruction instruction to generate a parameter dependency constraint table corresponding to the reconstruction intensity level. The parameter dependency constraint table records the physical coupling rules and priority weights between process parameters;

[0024] According to the reconstruction intensity level and the parameter dependency constraint table, dynamically decouple the conflict parameter dimensions and lock the independently adjustable dimensions;

[0025] Based on the device real-time sensor data and the parameter dependency constraint table, perform multi-level boundary constraint iterative correction on the decoupled independently adjustable dimensions;

[0026] Generate a candidate policy set within the modified policy parameter space through dynamic adaptive pseudo-random sampling.

[0027] In a preferred embodiment, the locking condition for the independently adjustable dimension is that the parameter adjustment range does not violate the physical limits of the device and meets the process standards of the production order.

[0028] The correction rule for multi-level boundary constraint iteration correction is to preferentially relax the parameter dimension constraints with high-priority weights and synchronously compress the parameter dimension constraints with low-priority weights.

[0029] The sampling density of dynamic adaptive pseudo-random sampling is dynamically adjusted according to the priority weights of the parameter dimensions.

[0030] In a preferred embodiment, use the digital twin simulation model after resetting the policy space to perform simulation verification on the candidate policy set, and screen out the feasible policy subset that meets the preset constraint conditions, including:

[0031] Generate a time series of dynamic parameters for each policy in the candidate policy set. The dynamic parameters include vibration amplitude, temperature gradient, and electromagnetic interference intensity.

[0032] Generate nodes of the topological network based on the local maximum and minimum points of the time series of dynamic parameters, and the absolute value of the Pearson correlation coefficient of the time series of dynamic parameters in adjacent time windows constitutes the edges of the topological network.

[0033] Extract the persistent homology features of the topological network, calculate the Betti numbers and persistence interval lengths of each dimension, and generate a set of topological invariants.

[0034] Calculate the dynamic constraint satisfaction parameter according to the evolution trend of the set of topological invariants. The evolution trend is jointly characterized by the change rate of the persistence interval length and the stability of the Betti numbers.

[0035] Screen out the policies with dynamic constraint satisfaction parameters higher than the dynamic threshold to generate a feasible policy subset. The dynamic threshold is dynamically adjusted according to the real-time health index of the device.

[0036] In a preferred embodiment, based on the KL divergence difference between each policy in the feasible policy subset and the historical policies, dynamically generate the policy exploration priority, including:

[0037] Generate a policy diversity evaluation parameter based on the KL divergence difference between the parameter distributions of each policy in the feasible policy subset and the historical policy set. The parameter distribution is constructed by the kernel density estimation method.

[0038] Divide the priority interval according to the real-time health index of the device.

[0039] Dynamically sort the KL divergence difference and the historical strategy similarity. When the KL divergence difference is greater than the historical strategy similarity, it is classified as an exploration - priority strategy; otherwise, it is an inheritance - priority strategy.

[0040] Generate a strategy exploration priority based on the priority interval and the strategy classification result. For exploration - priority strategies, they are sorted in descending order of the KL divergence difference within the interval, and for inheritance - priority strategies, they are sorted in ascending order of the historical strategy similarity. The number of priority levels is positively correlated with the number of strategies in the current strategy set.

[0041] In a preferred embodiment, when dividing the priority interval according to the real - time health index of the device, when the real - time health index of the device is lower than the first threshold, the priority interval is limited to high - diversity strategies, and when it is higher than the second threshold, the priority interval is extended to high - similarity strategies.

[0042] In a preferred embodiment, deploy the strategy with the highest strategy exploration priority to the actual production system and update the current optimization strategy set, including:

[0043] Extract the strategy with the highest priority from the strategy exploration priority sorting result and verify that it meets the real - time health index of the device and the process standards of the production order.

[0044] Deploy the verified strategy to the actual production system and update the current optimization strategy set according to the real - time health index of the device: when the real - time health index of the device is lower than the first threshold, replace the strategy with the smallest KL divergence difference, and when it is higher than the second threshold, replace the strategy with the highest historical strategy similarity.

[0045] Record the execution result in the historical strategy set, dynamically adjust the moving average window length according to the deviation rate between the actual and the simulation, and generate a strategy deployment execution log.

[0046] On the other hand, the present invention provides a production decision - making adaptive optimization system, including the following modules:

[0047] An entropy value monitoring module for real - time monitoring of the strategy entropy value of the current optimization strategy set. The strategy entropy value is obtained by calculating the distribution dispersion of each strategy in historical iterations.

[0048] An instruction triggering module for triggering a reconstruction instruction of the digital twin simulation model when the strategy entropy value is lower than a preset entropy value threshold.

[0049] A parameter resetting module for resetting the strategy parameter space of the digital twin simulation model according to the reconstruction instruction to generate a candidate strategy set containing a new strategy parameter space.

[0050] A simulation verification module for simulating and verifying the candidate strategy set using the digital twin simulation model after the strategy space reset, and screening out a feasible strategy subset that meets the preset constraint conditions.

[0051] A priority generation module, which is used to dynamically generate policy exploration priorities based on the KL divergence differences between each policy in the feasible policy subset and the historical policy;

[0052] A policy deployment module, which is used to deploy the policy with the highest policy exploration priority to the actual production system and update the current optimized policy set.

[0053] Compared with the prior art, the present invention has the following beneficial effects:

[0054] 1. By constructing a dynamic closed-loop optimization system, the adaptability and sustainability of production policies are effectively improved; by real-time monitoring the policy entropy value and triggering the model reconstruction mechanism, the attenuation trend of policy diversity can be timely sensed, and the parameter space is actively reset before the policy evolution solidifies. During the parameter space reset process, a multi-physical field coupling simulation and a priority-driven exploration mechanism are integrated to ensure that the newly generated candidate policies not only meet the equipment operation constraints but also cover a wider optimization direction, avoiding the local optimal trap; in the simulation verification link, dynamic threshold adjustment and cross-cycle feedback are introduced, so that the selected feasible policies can accurately match the real-time production status, and the potential of policy iteration is fully released on the premise of ensuring stability.

[0055] 2. Through the policy priority classification and closed-loop update mechanism, the intelligent evolution of the optimization process is realized; based on the dynamic priority sorting of KL divergence differences, the equipment health status is deeply bound to the policy exploration direction. The verified policies are preferentially inherited in the low-risk state, and innovative policy exploration is emphasized in the high-fault-tolerance interval, forming a decision-making mode that takes into account both efficiency and robustness; the deviation rate feedback after deployment and execution directly acts on the dynamic adjustment of the data acquisition window to ensure that the system always optimizes the policy generation logic based on the latest production status; enables the optimization system to continuously adapt to external disturbances and internal parameter drifts, and maintain a high level of decision-making flexibility and production efficiency during long-term operation. Description of the Drawings

[0056] Figure 1 It is a flowchart of the production decision adaptive optimization method of the present invention;

[0057] Figure 2 It is a structural schematic diagram of the production decision adaptive optimization system of the present invention. Detailed Embodiments

[0058] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0059] Example 1: Figure 1 The self - adaptive optimization method for production decision - making of the present invention is given, including the following steps:

[0060] S1. Monitor the policy entropy value of the current optimization policy set in real - time. The policy entropy value is obtained by calculating the distribution dispersion degree of each policy in historical iterations;

[0061] S2. When the policy entropy value is lower than the preset entropy threshold, trigger the reconstruction instruction of the digital twin simulation model;

[0062] S3. Reset the policy space of the digital twin simulation model according to the reconstruction instruction to generate a candidate policy set containing the new policy parameter space;

[0063] S4. Use the digital twin simulation model after policy space reset to simulate and verify the candidate policy set, and screen out a feasible policy subset that meets the preset constraint conditions;

[0064] S5. Dynamically generate the policy exploration priority based on the KL - divergence difference between each policy in the feasible policy subset and the historical policies;

[0065] S6. Deploy the policy with the highest policy exploration priority to the actual production system and update the current optimization policy set.

[0066] S1. Monitor the policy entropy value of the current optimization policy set in real - time. The policy entropy value is obtained by calculating the distribution dispersion degree of each policy in historical iterations, and it can be specifically implemented as follows:

[0067] Obtain the policy parameter vectors recorded by each policy included in the current optimization policy set in the historical iteration process from the historical database of the production system. The policy parameter vector is composed of the process parameters corresponding to the policy during historical execution. The process parameters include the working temperature of production equipment, processing pressure, material conveying rate, and energy consumption coefficient. Each process parameter corresponds to a dimension of the policy parameter vector. When obtaining, screen the historical data within a preset time period according to the policy execution timestamp. The preset time period is set according to the process update cycle of the production system. The process update cycle is the adjustment interval of the main process parameters defined in the production line equipment technical manual. If the equipment technical manual is not clear, by default, select the historical data within the most recent 12 months. For the parameter vectors of the same policy executed multiple times in iteration, after arranging them in chronological order, take the sliding average of the parameters of each dimension as the representative parameter vector of the policy. The calculation window length of the sliding average is consistent with the data acquisition frequency of the production system. The data acquisition frequency is determined by the sensor configuration of the production line. For example, when data is collected every 5 minutes, the sliding average window is set to 1 hour.

[0068] Set the time window length to the preset number of iterations, which is dynamically adjusted according to the policy update frequency of the production system. The policy update frequency is the policy optimization period defined in the production plan. If the production plan is not specified, the last 30 policy iteration periods are selected by default. Extract the complete data set of all policy parameter vectors within this time window from the historical database. The feature dimension of the policy parameter vector refers to the type of process parameters it contains. When extracting, perform data validity verification on each policy parameter vector. The data validity verification includes integrity verification and rationality verification: The integrity verification requires that the number of missing dimension parameters in the policy parameter vector does not exceed 20% of the total dimensions. The rationality verification requires that each dimension parameter value is within the safe operating range specified in the equipment technical manual. For example, the working temperature parameter needs to be between 50°C and 200°C, and the processing pressure parameter needs to be between 0.1 MPa and 5.0 MPa. If a policy parameter vector fails the verification, it is marked as invalid data and excluded. The vacant position after exclusion is filled by the valid parameter vector at the adjacent time point of the same policy. If there is no valid filling data, the vacancy is retained.

[0069] For each extracted feature dimension, calculate the numerical fluctuation range within all policy parameter vectors in the time window, specifically by calculating the variance value of each dimension parameter. The calculation process of the variance value is as follows: For a feature dimension, obtain the parameter value sequence of this feature dimension in all policy parameter vectors within the time window, calculate the arithmetic mean of this sequence, then calculate the square of the difference between each parameter value and the arithmetic mean, sum all the squared values and divide by the total number of parameter values to obtain the variance value of this dimension parameter. When calculating, the dimension of the parameter value is consistent with the unit of the original data collected by the equipment sensor. For example, the working temperature is in degrees Celsius, the processing pressure is in megapascals, the material conveying rate is in kilograms per minute, and the energy consumption coefficient is in kilowatt-hours per piece. If there are multiple groups of parameter value sequences for the same dimension parameter, calculate the variance value for each sequence separately and take the maximum value as the final variance value.

[0070] The strategy entropy value is generated by weighted summation of the variance values of each dimension parameter. The weight coefficients are determined by the analytic hierarchy process. The implementation steps of the analytic hierarchy process are as follows: establish a hierarchical structure model of the impact of process parameters on production efficiency, and the production efficiency indicators include equipment utilization rate, product yield, and energy consumption efficiency; design an expert questionnaire, invite at least three domain experts to conduct pairwise comparison and scoring of the importance of each process parameter, and the scoring scale is from 1 to 9 points. 1 point means that the two parameters are equally important, and 9 points means that the former is extremely important compared to the latter; after collecting the expert scores, construct a judgment matrix, calculate the maximum eigenvalue of the judgment matrix and its corresponding eigenvector, and normalize the eigenvector to obtain the initial weight coefficients of each dimension parameter; conduct a consistency test on the initial weight coefficients. When the consistency ratio is less than 0.1, the weights are confirmed to be valid, otherwise, readjust the scores until the test is passed. The final weight coefficients satisfy that the sum of the weights of all dimensions is 1. When performing weighted summation, multiply the variance value of each dimension parameter by the corresponding weight coefficient and then accumulate. The accumulated result is reserved to two decimal places as the strategy entropy value. If the variance value of a certain dimension parameter cannot be calculated due to abnormal data, its weight coefficient is temporarily evenly distributed to other dimensions, and recalculated after the data returns to normal.

[0071] Data preprocessing includes dynamic adjustment of the time window and standardization of the parameter range: The dynamic adjustment rule of the time window is that when the strategy update frequency of the production system accelerates to more than 1.5 times the original frequency, the time window length is shortened to 2 / 3 of the original length; when the strategy update frequency decreases to less than 2 / 3 of the original frequency, the time window length is extended to 1.5 times the original length. The standardization method of the parameter range is to perform min-max normalization on each dimension parameter and linearly map it to the [0,1] interval. The normalization formula is (current value - minimum value) / (maximum value - minimum value), and the minimum value and maximum value are taken from the safe operating range specified in the equipment technical manual.

[0072] When the strategy entropy value is lower than the preset threshold, a digital twin simulation model reconstruction instruction is triggered. The preset threshold is set according to the distribution of the strategy entropy value in historical data. Specifically, it is the lower limit value of the normal fluctuation range of the strategy entropy value in the past year. The normal fluctuation range is determined by calculating the mean and standard deviation of the strategy entropy value, and the lower limit value is the mean minus twice the standard deviation. The reconstructed digital twin model generates a new strategy parameter space through simulation verification. The boundary of the new strategy parameter space is dynamically adjusted according to the current health state of the equipment. The health state of the equipment is evaluated by a preset predictive maintenance model. The input of the predictive maintenance model is the equipment vibration spectrum, temperature time series data, and lubricant detection results.

[0073] The exception handling mechanism includes empty dataset warning and calculation overflow protection: If there is no valid policy parameter vector within the time window, a warning signal is sent to the production management system and the policy optimization process is suspended until manual intervention to confirm the data source status; If numerical overflow occurs during the variance value calculation, it automatically switches to the logarithmic transformation calculation mode. The calculation method of logarithmic transformation is to take the natural logarithm of the parameter value first and then perform variance calculation, and then restore the final result through exponential operation after the calculation is completed. The boundary condition is defined as that the dimension number of the policy parameter vector does not exceed 100. If it exceeds 100 dimensions, the first 100 dimensions are selected according to the weight coefficient sorting to participate in the calculation.

[0074] Hardware dependencies include the sensor network and data acquisition system of production equipment. The sensor network consists of temperature sensors, pressure sensors, flow meters, and electricity meters. The data acquisition system is implemented based on the industrial Internet of Things platform, and the software version of the Internet of Things platform is not lower than 4.2.1. Software dependencies include the numerical calculation library NumPy version 1.21.0 and above, which is used for variance calculation and weighted summation, and the Analytic Hierarchy Process toolkit AHPy 0.9.3, which is used for weight coefficient calculation. The experimental results are verified by recording the correlation between the change curve of the policy entropy value and production efficiency. For example, in a certain automotive parts production line, when the policy entropy value is lower than the threshold, the model reconstruction is triggered. After the reconstruction, the policy diversity is increased by 25%, and the equipment utilization rate is increased from 82% to 87%.

[0075] It should be noted that the current optimization policy set refers to the set composed of candidate policies to be evaluated within the real-time optimization cycle of the production system. Each policy is composed of a set of predefined process parameter combinations, and the process parameter combinations include but are not limited to equipment operating temperature, processing pressure threshold, upper limit of material conveying rate, and energy consumption efficiency target value. The set is generated by screening historical iterative data. Specifically: Extract the policies that meet the recent production efficiency from the policy library of the digital twin simulation model. The efficiency indicators include a good product rate ≥ 95% and equipment utilization rate ≥ 85%, and eliminate the policies that conflict with the current equipment health status (such as parameter combinations that exceed the equipment vibration tolerance range). The update frequency of the set is synchronized with the production plan. After each policy iteration, the policies are added or deleted according to the simulation verification results to ensure that the set capacity is maintained within the range of 50 - 200 policies to balance the calculation efficiency and policy diversity requirements.

[0076] S2. When the policy entropy value is lower than the preset entropy value threshold, trigger the reconstruction instruction of the digital twin simulation model, which can be specifically implemented as:

[0077] Obtain the dynamic adjustment rule of the preset entropy value threshold, which is calculated based on the mean and standard deviation of the historical policy entropy values. The historical policy entropy values are extracted from the historical database of the production system, and the extraction range is the policy entropy value data recorded in all policy optimization cycles in the most recent year. The mean is calculated by taking the arithmetic mean of the historical policy entropy value sequence, and the standard deviation is calculated by taking the square root of the average of the squares of the differences between each data point in the historical policy entropy value sequence and the mean. The preset entropy value threshold is set to the mean minus twice the standard deviation. If the calculation result is lower than the minimum policy diversity threshold specified in the equipment technical manual, the minimum policy diversity threshold is used as the preset entropy value threshold. The update cycle of the dynamic adjustment rule is synchronized with the policy optimization cycle of the production system, and the policy optimization cycle is set according to the production plan. For example, the dynamic adjustment rule is updated once a month.

[0078] Compare the currently monitored current policy entropy value with the preset entropy value threshold to generate a threshold deviation degree evaluation result. When comparing, calculate the absolute value of the difference between the current policy entropy value and the preset entropy value threshold, and divide it by the preset entropy value threshold to obtain the relative deviation degree. For example, if the preset entropy value threshold is 40 and the current policy entropy value is 35, the relative deviation degree is (40 - 35) / 40 = 0.125. The generation rule of the threshold deviation degree evaluation result is: it is determined as a significant deviation when the relative deviation degree is greater than or equal to 0.1, it is determined as a general deviation when the relative deviation degree is between 0.05 and 0.1, and it is determined as a slight deviation when the relative deviation degree is less than 0.05. The determination threshold is set according to the statistical results of policy failure events in historical data. For example, when the relative deviation degree exceeds 0.1, the historical policy failure probability increases to 30%.

[0079] Generate a digital twin simulation model reconstruction instruction including the reconstruction intensity level based on the judgment that the threshold deviation degree evaluation result exceeds the preset deviation threshold. The preset deviation threshold is set to 0.1, and the reconstruction instruction is triggered when the threshold deviation degree evaluation result is a significant deviation or a general deviation. The reconstruction intensity level is divided into three levels: level 1 reconstruction corresponds to a significant deviation, and the reconstruction intensity is a full parameter space reset; level 2 reconstruction corresponds to a general deviation, and the reconstruction intensity is a core parameter space reset; level 3 reconstruction corresponds to a slight deviation, and the reconstruction intensity is a local parameter fine-tuning. The mapping rule of the reconstruction intensity level is determined by an expert experience table, and the expert experience table records the corresponding relationship between different deviation degree intervals and parameter adjustment ranges. For example, level 1 reconstruction requires resetting more than 80% of the policy parameters, level 2 reconstruction resets 50% - 80% of the parameters, and level 3 reconstruction adjusts 20% - 50% of the parameters.

[0080] The reconstruction instructions are passed to the control interface of the digital twin simulation model, and the control interface performs corresponding parameter initialization operations according to the reconstruction intensity level. The control interface is a standard API interface provided for the digital twin simulation model, which receives JSON-format instructions containing the reconstruction intensity level. The specific logic of the parameter initialization operation is as follows: for the first-level reconstruction, all historical constraint conditions in the current policy parameter space are cleared, and the parameter boundaries are redefined based on the latest device status data; for the second-level reconstruction, the basic process parameter constraints are retained, and the optimization target weight coefficients are reset; for the third-level reconstruction, only the three parameter dimensions with the highest correlation with the current production requirements are adjusted. The data sources for parameter initialization include device real-time sensor data, material property databases, and process standard documents.

[0081] The protection mechanism for the dynamic adjustment rule is as follows: if the amount of historical policy entropy value data is less than 100, a fixed preset entropy value threshold is adopted, and the fixed value is the larger of the minimum policy diversity threshold recommended in the device technical manual and the industry standard value. The optimization method for deviation calculation is: perform truncation processing on extreme values. For example, when the current policy entropy value is lower than the minimum policy diversity threshold, it is directly determined as a significant deviation, and the relative deviation does not need to be calculated. The fault tolerance mechanism for the reconstruction intensity level is: when the intensity level in the reconstruction instruction conflicts with the current device status, for example, the device is in the maintenance period and cannot perform the first-level reconstruction, it is automatically downgraded to the next intensity level and an operation log is generated.

[0082] For example, when applying this embodiment in an electronic product assembly line, when the policy entropy value drops to 38 and the preset entropy value threshold is 40, the system triggers a second-level reconstruction instruction. After the core parameter space is reset, the policy diversity is improved and the device failure rate is reduced. The matching accuracy between the reconstruction intensity level and the deviation is relatively high, and the false trigger rate is reduced. For example, the verification data can be sourced from the real-time monitoring system of the production line, and the reconstruction operation is performed within the data collection cycle.

[0083] The data missing processing rule is: if all historical policy entropy value data are invalid values, for example, data anomalies caused by sensor failures, the benchmark entropy value threshold of the same type of production line in the same industry is temporarily adopted until the local data resumes to be valid. The instruction transmission interruption processing is: when the reconstruction instruction transmission fails, it is automatically retried three times and the backup communication channel is enabled. If it still fails, the optimization process is paused and an artificial intervention alarm is triggered. The parameter initialization conflict processing is: if the redefined parameter boundary exceeds the physical limit of the device, for example, the reset value of the processing pressure exceeds 5.0 MPa, it is automatically truncated to the safe range and marked as a parameter to be reviewed.

[0084] Hardware dependencies include an industrial sensor network that supports real-time data acquisition, such as temperature sensor model PT100 and pressure sensor model MPX5700, as well as edge computing devices with at least a 4-core CPU and 16GB of memory. Software dependencies include the digital twin simulation platform ANSYS Twin Builder version 2022R2 and above, which supports a parameter dynamic reset interface, and the JSON instruction parsing library uses Newtonsoft.Json version 13.0.1 and above.

[0085] In the calculation of the mean and standard deviation, the dimension of the historical policy entropy value data is a dimensionless index, and the unit is uniformly the entropy value unit (Entropy Unit, EU). The calculation result of the relative deviation is reserved to two decimal places. For example, 0.125 is recorded as 0.13. The adjustment basis of the weight coefficient in the parameter initialization operation is the contribution degree of the process parameters to the production efficiency, and the contribution degree is obtained through regression analysis of historical data. For example, the contribution degree coefficient of the working temperature parameter is 0.3, and the processing pressure is 0.2.

[0086] S3. Reset the policy parameter space of the digital twin simulation model according to the reconstruction instruction to generate a candidate policy set containing the new policy parameter space. Specifically, it can be implemented as follows:

[0087] Parse the reconstruction intensity level in the reconstruction instruction to generate a parameter dependency relationship constraint table corresponding to the reconstruction intensity level. The parameter dependency relationship constraint table is constructed by extracting physical coupling rules from the process parameter correlation clauses in the equipment technical manual. The process parameter correlation clauses clearly record the linkage relationship between parameters in tabular form. For example, the changes in the processing pressure parameter and the working temperature parameter need to satisfy the thermodynamic equilibrium formula. Specifically, for every 1MPa increase in pressure, the temperature decreases by 2°C.

[0088] The priority weight is determined according to the process standard document of the production order. The process standard document defines the influence coefficient of each parameter on the core quality index, and the influence coefficient is calculated through the orthogonal experiment method. The implementation steps of the orthogonal experiment method are to design a multi-factor and multi-level experimental matrix and analyze the contribution degree of each parameter. The design of the experimental matrix follows the L9(3^4) standard orthogonal table, and the contribution degree is calculated using the analysis of variance method. The update mechanism of the parameter dependency relationship constraint table is bound to the equipment maintenance cycle, and the equipment maintenance cycle is specified by the preventive maintenance plan in the equipment technical manual. The maintenance plan clearly states that the coupling rule library is verified and revised every quarter, and the revised content needs to be signed and confirmed by the equipment engineer before it takes effect.

[0089] Based on the reconstruction intensity level and the parameter dependency constraint table, conflicting parameter dimensions are dynamically decoupled and independently adjustable dimensions are locked. The operational process for dynamically decoupling conflicting parameter dimensions is to load the coupling rules in the parameter dependency constraint table. When the reconstruction intensity level is a full parameter space reset, the coupling constraints between all parameters are released. When the reconstruction intensity level is a core parameter space reset, only the coupling rules directly related to the key quality indicators of the current production order are retained. The key quality indicators are determined by the Class A inspection items in the order process standard document, such as dimensional accuracy and surface roughness. When the reconstruction intensity level is local parameter fine-tuning, only the coupling relationship related to abnormal parameters monitored by the equipment's real-time sensors is released. Abnormal parameters are defined as parameters that exceed the safety range specified in the equipment technical manual for three consecutive sampling cycles. The locking conditions of independently adjustable dimensions are verified by checking whether the parameter adjustment range is within the physical limits specified in the equipment technical manual. The physical limits include a maximum operating temperature of 300°C, a maximum processing pressure of 50 MPa, and a maximum spindle speed of 8000 rpm. At the same time, the parameter range is checked to see whether it meets the tolerance requirements of the production order process standards. The tolerance requirements include dimensional accuracy of ±0.05 mm, surface roughness Ra ≤ 0.8 μm, and other indicators. If the verification fails, the parameter range redefinition process is triggered.

[0090] Based on real-time device sensor data and parameter dependency constraint tables, we iteratively correct multi-level boundary constraints for decoupled, independently adjustable dimensions. This correction rule prioritizes relaxing constraints on high-priority parameter dimensions while simultaneously compressing constraints on low-priority parameter dimensions. Priority weights are assigned based on the degree of impact of the parameter on production order delivery quality, as determined through regression analysis of historical production data. The correlation coefficient between quality defect events and parameter fluctuations in this regression analysis serves as the basis for weight assignment. The correlation coefficient is calculated using the Pearson product-moment correlation coefficient formula.

[0091] The specific correction operation is to expand the adjustment range of high-priority parameters with a weight ≥ 0.3 to 90%-95% of the physical limit of the equipment. For example, when the physical limit of the operating temperature parameter is 300°C, the adjustment range is expanded from the original 250-280°C to 240-290°C. For medium-priority parameters with a weight between 0.1 and 0.3, the adjustment range remains within the safe operating range recommended by the equipment technical manual, such as maintaining the processing pressure parameter at 40-45MPa. For low-priority parameters with a weight < 0.1, the adjustment range is compressed to 50%-70% of the original range. For example, when the original range of the coolant flow parameter is 5-10L / min, it is compressed to 6-8L / min. Coupling conflict detection is required after each round of correction. If it is detected that the parameter combination violates the coupling rules, the system will roll back to the previous round of correction and reduce the adjustment range by 10%. The iteration is terminated when the adjustment range decays to the minimum value, such as the temperature adjustment step size ≤ 1°C.

[0092] Generate a set of candidate policies within the modified policy parameter space through dynamic adaptive pseudo-random sampling. The implementation of dynamic adaptive pseudo-random sampling is to allocate sampling density according to the priority weights of parameter dimensions. The sampling density is dynamically adjusted according to the priority weights of parameter dimensions. The priority weights are calculated by the orthogonal experiment method and normalized to the range of 0-1. Dimensions with weights ≥ 0.3 are divided into the high-priority group, and 10 sampling points are generated for each dimension; dimensions with weights of 0.1-0.3 are the medium-priority group, and 5 sampling points are generated for each dimension; dimensions with weights < 0.1 are the low-priority group, and 3 sampling points are generated for each dimension. The sampling density of the high-priority group is 3.33 times that of the low-priority group (10 / 3). This ratio is set according to the difference in the contribution of parameters to production efficiency to ensure that key parameters receive denser optimization exploration. The sampling point allocation rule is written into the configuration file of the digital twin platform, and spatial uniform coverage is achieved through Latin hypercube design.

[0093] The sampling method adopts Latin hypercube design. Latin hypercube design evenly divides each parameter dimension into several intervals and randomly selects points within each interval to ensure uniform coverage of the parameter space. The calculation formula for the number of sampling points is the total number of points = Σ (the number of dimensions in each weight group × the corresponding sampling density). For example, if a policy parameter space contains 6 high-weight dimensions, 4 medium-weight dimensions, and 2 low-weight dimensions, then the total number of sampling points is 6×10 + 4×5 + 2×3 = 86. Each sampling point corresponds to a parameter combination of a candidate policy. The numerical values of the parameter combination retain the unit precision consistent with the data collected by the device sensors. For example, the temperature parameter retains one decimal place (unit: °C), the pressure parameter retains an integer (unit: MPa), and the cutting speed parameter retains two decimal places (unit: m / min).

[0094] When constructing the parameter dependency relationship constraint table, the physical coupling rule is obtained by parsing the process relevance clauses in the device technical manual. The mathematical relationship or empirical rule of parameter linkage is clearly defined in the clauses. For example, for every 10°C increase in the working temperature, the processing pressure decreases by 5%. The priority weights are calculated by the orthogonal experiment method. The orthogonal experiment method uses the L9(3^4) experimental matrix to analyze the variance contribution rate of each parameter to the quality index. Parameters with a variance contribution rate exceeding 15% are defined as high-priority.

[0095] When dynamically decoupling conflicting parameter dimensions, the correlation coefficient in the coupling rule is calculated from historical production data. The screening condition for historical data is the data during the period when the device health index ≥ 80 and the production yield rate ≥ 95%. The Pearson product-moment correlation coefficient formula is used for the calculation of the correlation coefficient. Parameter groups with a correlation coefficient ≥ 0.7 are determined to be strongly coupled. In the iterative correction of multi-level boundary constraints, the attenuation mechanism of the adjustment amplitude is that each time a conflict is detected, the adjustment amplitude is multiplied by the attenuation coefficient 0.9 until the conflict is eliminated or the minimum adjustment step size is reached. For example, the temperature adjustment step size ≤ 1°C.

[0096] If there is a version conflict in the parameter dependency constraint table, such as inconsistent rules between the old and new manuals, the version of the device technical manual with the latest timestamp is preferred, and a manual review process is triggered. The review process needs to be completed within 24 hours. When the weight calculation of a certain dimension is abnormal, such as a division-by-zero error, its weight is automatically reset to 0.1 and the abnormal event code is recorded. The code format is ERR_WEIGHT_parameter name_timestamp, and the timestamp is accurate to milliseconds, for example, ERR_WEIGHT_TEMP_20231005143035999. If there is a system failure during the correction process, it is restored from the correction state stored in the most recent persistent storage. The persistent storage interval is set to once every 5 minutes, and the storage medium is an SSD hard drive with a read / write speed ≥ 500MB / s.

[0097] Hardware dependencies include the Kistler 9232A dynamic force sensor with an accuracy of ±0.5% and a range of 0 - 20kN, and the NI PXIe-6368 data acquisition card with a sampling rate of 500kHz and a 16-bit resolution. Software dependencies include Matlab 2022a for orthogonal experiment analysis and the pyDOE library in Python 3.9 to implement the Latin hypercube design, with a version number of 0.3.8. When calculating the parameter priority weights, weight normalization is processed to ensure that the sum is 1. In the Latin hypercube sampling, the number of interval divisions for each dimension is consistent with the sampling density. For example, for a high-weight dimension, 10 intervals are divided, and 1 point is randomly sampled within each interval. The interval boundary values are distributed in equal proportion. Dimensions and units strictly follow the sensor configuration. For example, the temperature unit is degrees Celsius (°C), and the pressure unit is megapascals (MPa), and the accuracy is consistent with the device technical manual.

[0098] S4. Use the digital twin simulation model after the policy space reset to perform simulation verification on the candidate policy set, and screen out a feasible policy subset that meets the preset constraint conditions. Specifically, it can be implemented as follows:

[0099] Generate a time series for the dynamic parameters of each policy in the candidate policy set. The dynamic parameters include vibration amplitude, temperature gradient, and electromagnetic interference intensity. The dynamic parameters are extracted from the real-time simulation results of the digital twin simulation model, and the real-time simulation results are obtained by running according to the reset policy parameter space in step S3. The length of the time series is consistent with the number of iterations of the preset time window in step S2. For example, when the time window is set to the most recent 30 policy iterations, each dynamic parameter generates a time series containing 30 data points.

[0100] The sampling frequency of the time series is determined by the sensor configuration of the production equipment. For example, the sampling frequency of the vibration sensor is 1000 times per second, and the sampling frequency of the temperature sensor is 10 times per second. For non-real-time parameters, such as the electromagnetic interference intensity, data is output by the calculation module of the digital twin model according to the simulation step size, and the simulation step size is synchronized with the device control period. For example, data is output once every 0.1 seconds. The storage format of the time series is a two-dimensional array. The first dimension is the timestamp, and the second dimension is the parameter value. The unit of the parameter value is consistent with the data collected by the sensor. The unit of the vibration amplitude is meters per second squared (m / s²), the unit of the temperature gradient is degrees Celsius per minute (℃ / min), and the unit of the electromagnetic interference intensity is millitesla (mT).

[0101] Nodes of the topological network are generated based on the local maximum and minimum points of the time series of dynamic parameters. A local maximum point is defined as a parameter value at a certain time point being greater than the parameter values at the two adjacent time points before and after it, and a local minimum point is defined as a parameter value at a certain time point being less than the parameter values at the two adjacent time points before and after it. After the nodes are generated, calculate the absolute value of the Pearson correlation coefficient of the time series of dynamic parameters within adjacent time windows, and use it as the edge weight of the topological network.

[0102] The division rule of adjacent time windows is to divide the time series according to a fixed window length, which is the same as the sliding average window in step S1. For example, when the sliding average window is 1 hour, the time window length is set to the number of data points within 1 hour. The calculation range of the edge weight is different dynamic parameter sequences within the same time window. For example, calculate the absolute value of the Pearson correlation coefficient between the vibration amplitude and the temperature gradient within the same time window. The calculation formula is the covariance divided by the product of the standard deviations of the two sequences, and the absolute value of the calculation result is taken to eliminate the directional influence.

[0103] Extract the persistent homology features of the topological network, calculate the Betti numbers and the lengths of the persistence intervals in each dimension, and generate a set of topological invariants. The persistent homology features are realized through the persistent homology calculation of the open-source topological data analysis library GUDHI. The input is the topological network data generated in step S4, and the output is the Betti numbers in each dimension and the corresponding persistence interval lengths. The calculation dimensions of the Betti numbers include dimension 0 and dimension 1. The Betti number in dimension 0 represents the number of connected components in the topological network, and the Betti number in dimension 1 represents the number of circular structures. The length of the persistence interval represents the survival time span of the corresponding topological feature during the change of the parameter scale. For example, a circular structure appears starting from the parameter scale of 0.5 and disappears at the scale of 1.2, and its persistence interval length is 0.7. The storage format of the set of topological invariants is a multi-dimensional array, and each element contains the Betti number value, the start and end times of the persistence interval, and the associated dynamic parameter type, such as the Betti number corresponding to the combined parameter of the vibration amplitude and the temperature gradient.

[0104] Calculate the dynamic constraint satisfaction parameter according to the evolution trend of the set of topological invariants. The evolution trend is jointly characterized by the change rate of the duration interval length and the stability of the Betti number. The change rate of the duration interval length is calculated by arranging the duration interval lengths of the same topological feature in chronological order and calculating the sum of the absolute values of the length differences in adjacent time windows. For example, if the duration interval lengths of a certain ring structure are 0.7, 0.8, and 0.6 in three consecutive time windows, the change rate is |0.8 - 0.7| + |0.6 - 0.8| = 0.3.

[0105] The stability of the Betti number is realized through variance calculation, which statistically analyzes the fluctuation degree of the same Betti number in consecutive time windows. For example, if the 0-dimensional Betti number is 2, 3, 2, 3, 2 in five time windows, its variance is 0.5. The calculation formula of the dynamic constraint satisfaction parameter is the normalized value of the change rate multiplied by the weight coefficient plus the normalized value of the Betti number stability multiplied by the complementary weight coefficient. The weight coefficient is dynamically adjusted according to the device type. For example, when the vibration sensitivity of a numerically controlled machine tool is relatively high, the change rate weight is set to 0.7 and the stability weight is set to 0.3. The normalization method is to divide the original value by the historical maximum value. For example, if the historical maximum change rate is 1.0 and the current change rate is 0.3, the normalized value is 0.3.

[0106] Screen the strategies with dynamic constraint satisfaction parameters higher than the dynamic threshold to generate a subset of feasible strategies. The dynamic threshold is dynamically adjusted according to the real-time health index of the device. The real-time health index of the device is obtained by linear regression prediction of the vibration spectrum characteristics and the temperature drift amount. The vibration spectrum characteristics are extracted by calculating the energy proportion of the main frequency component through fast Fourier transform. For example, after converting the vibration signal from the time domain to the frequency domain, calculate the percentage of the energy in the 5 - 100 Hz frequency band in the total energy. The temperature drift amount is calculated as the absolute value of the difference between the current temperature and the set value. For example, if the set temperature is 200 °C and the current temperature is 210 °C, the drift amount is 10 °C.

[0107] The training data of the linear regression model comes from historical device maintenance records. The inputs are the main frequency energy of the vibration spectrum (unit: decibel) and the temperature drift amount (unit: degree Celsius), and the output is the health index (range 0 - 100). The mapping rule between the dynamic threshold and the health index is that when the health index is greater than or equal to 80, the dynamic threshold is set to 0.8; when the health index is between 60 and 80, the dynamic threshold is 0.7; when the health index is less than 60, the dynamic threshold is 0.6. During the screening process, the strategies with dynamic constraint satisfaction parameters lower than the dynamic threshold are marked as strategies to be verified. After manual review of the strategies to be verified, if they conform to the actual production experience, they are added to the subset of feasible strategies, otherwise they are excluded.

[0108] Data preprocessing includes dynamic parameter alignment, fault tolerance in topological network construction, and optimization of persistent homology calculation. Dynamic parameter alignment is achieved through linear interpolation, which aligns parameters with different sampling frequencies to the same timestamp. For example, temperature data sampled 10 times per second is interpolated into a time series sampled 1000 times per second. The fault tolerance mechanism for topological network construction is that if there are no local extreme points within a certain time window, the window length is automatically extended to adjacent windows and recalculated. The optimization of persistent homology calculation enables an approximate calculation mode for large-scale topological networks. By using the sparse matrix compression function of the GUDHI library, the calculation error is controlled within 5%. For example, when the number of nodes exceeds 1000, a compression algorithm is enabled to reduce memory occupancy.

[0109] Boundary conditions include data missing handling, calculation overflow handling, and threshold mapping fault tolerance. The rule for data missing handling is that if the time series data of a certain dynamic parameter is missing by more than 50%, the topological network construction of this strategy is skipped and marked as an invalid strategy. The rule for calculation overflow handling is that when the continuous interval length exceeds the total simulation duration, it is automatically truncated to the total simulation duration and an abnormal event is recorded. The rule for threshold mapping fault tolerance is that when the health index calculation is abnormal, for example, when the input data is all zero due to sensor failure, the dynamic threshold is default set to 0.5 and a device status warning is triggered.

[0110] Hardware dependencies include a data acquisition card that supports multi-channel synchronous acquisition, such as NI PXIe-6368, which is used for high-frequency acquisition of vibration and temperature data. Its sampling rate is not less than 1000 times per second and the resolution is not less than 16 bits. Software dependencies include integrating a topological analysis library of GUDHI version 3.4.0 and above in the digital twin simulation platform. The LinearRegression class of the Scikit-learn 1.0.2 library is used for linear regression model training, and the simulation step size control needs to be compatible with Python 3.9 and above.

[0111] In this step S4, by constructing a topological network of the dynamic parameter time series and extracting persistent homology features, it solves the problem of policy rigidity caused by traditional simulation verification methods relying on static thresholds or simple variance analysis. The topological network captures the dynamic correlation between parameters (such as the co-fluctuation of vibration and temperature), and the persistent homology features (Betti numbers, persistence intervals) quantify the system stability, overcoming the defect of existing technologies that only focus on single parameters or linear relationships. The dynamic threshold is adaptively adjusted based on the device health index, ensuring that the screening strategy not only meets the preset constraints but also adapts to the real-time state of the device. Compared with traditional methods, through dynamic network modeling and topological invariant analysis, it improves the policy verification accuracy and system robustness, providing interpretable and adaptive optimization decision support for complex industrial scenarios.

[0112] S5. Based on the KL divergence difference between each strategy in the feasible strategy subset and the historical strategy, dynamically generate the strategy exploration priority, which can be specifically implemented as:

[0113] Based on the KL divergence difference between the parameter distributions of each strategy in the feasible strategy subset and the historical strategy set, a strategy diversity evaluation parameter is generated. The parameter distribution is constructed by the kernel density estimation method, and the bandwidth parameter of the kernel density estimation method is adaptively adjusted according to the sample data volume. The bandwidth adjustment rule adopts the Silverman empirical rule, and the sample data volume is the sum of the strategy numbers of the feasible strategy subset and the historical strategy set. The parameter types include vibration amplitude, temperature gradient, and electromagnetic interference intensity, which are strictly consistent with the definition of the dynamic parameters in step S4. The unit of vibration amplitude is meters per second squared (m / s²), the unit of temperature gradient is degrees Celsius per minute (℃ / min), and the unit of electromagnetic interference intensity is millitesla (mT). The calculation method of the KL divergence difference is as follows: calculate the KL divergence between the parameter distribution of each strategy and the parameter distributions of all strategies in the historical strategy set one by one, and take the maximum value as the diversity evaluation parameter of this strategy. The data source of the parameter distribution is the strategy parameter space reset in step S3, and the data source of the historical strategy set is the historical iteration record within the preset time window in step S1. The storage format of the historical iteration record is a JSON file containing the timestamp, strategy parameters, and execution results.

[0114] The priority interval is divided according to the real-time health index of the device, and the real-time health index of the device is obtained by linearly regressing and predicting the vibration spectrum characteristics and temperature drift in step S4. The value range of the real-time health index of the device is from 0 to 100, the first threshold is set to 50, and the second threshold is set to 70. The threshold setting is based on the statistical relationship between the health index and the strategy failure event in the historical device maintenance record. For example, when the health index is lower than 50, the device failure probability increases to 25%. When the real-time health index of the device is lower than the first threshold, the priority interval is limited to high-diversity strategies, and only the strategies ranked in the top 30% of the KL divergence difference are allowed to enter the priority queue; when the real-time health index of the device is higher than the second threshold, the priority interval is extended to high-similarity strategies, and the strategies ranked in the top 70% of the KL divergence difference are allowed to enter the priority queue. The capacity limit of the priority queue is 50% of the number of strategies in the current strategy set. For example, when the number of strategies is 100, the queue can contain at most 50 strategies.

[0115] Dynamically sort the KL divergence difference and the historical strategy similarity. The historical strategy similarity calculates the direction consistency between policy parameter vectors through cosine similarity. The policy parameter vector consists of the normalized values of vibration amplitude, temperature gradient, and electromagnetic interference intensity. The normalization method is to divide the original parameter value by the maximum allowable value specified in the equipment technical manual. For example, if the maximum allowable value of vibration amplitude is 10 m / s², the normalized value is the original value divided by 10. The calculation method of cosine similarity is the dot product of two policy parameter vectors divided by the product of the vector norms, and the calculation result ranges from -1 to 1. The dynamic sorting rule is: when the KL divergence difference of a certain policy is greater than its cosine similarity with the historical policy, it is classified as an exploration-priority policy; otherwise, it is classified as an inheritance-priority policy. The classification results are stored as a two-dimensional label array, with the first column being the policy number and the second column being the classification identifier. The classification identifier is 0 for inheritance-priority policies and 1 for exploration-priority policies.

[0116] Generate the policy exploration priority based on the priority interval and the policy classification results. The exploration-priority policies are sorted in descending order of KL divergence difference within the interval. For example, a policy with a KL divergence difference of 0.8 has a higher priority than a policy with 0.5. The inheritance-priority policies are sorted in ascending order of historical strategy similarity. For example, a policy with a similarity of 0.2 has a higher priority than a policy with 0.6. The number of priority levels is positively correlated with the number of policies in the current policy set. When the number of policies is less than 50, 5 priority levels are set. When the number of policies is greater than 50, 1 level is extended for every additional 10 policies. The priority mapping table is generated by traversing the sorting results, and the storage format of the mapping table is a key-value pair of policy number and priority level. The update frequency of the priority mapping table is synchronized with the policy optimization period, and the policy optimization period is determined by the equipment maintenance interval in the production plan. For example, the priority update is executed at 2 am every day.

[0117] The parameter alignment process is achieved through zero-padding operations. The lengths of the parameter vectors of different policies are aligned according to the parameter importance ranking in the equipment technical manual, and zeros are padded for parameters with low importance. The optimization rule for kernel density estimation is that when the sample data volume is less than 10, it automatically switches to Gaussian mixture model to estimate the parameter distribution, and the number of Gaussian components is set to the integer part of the square root of the sample size. For example, when the sample size is 9, the number of components is set to 3. The dynamic adjustment rule for thresholds is that the first threshold and the second threshold increase by 5% annually based on the equipment usage years, which are calculated from the equipment factory date to the month. For example, when the equipment has been used for 3 years, the thresholds increase by 15%.

[0118] The rule for handling insufficient data is that if the historical strategy set is empty, the default strategy set of the same type of device is temporarily used as an alternative data source. The default strategy set is pre-stored in the device's local database, and the storage path is / opt / strategy / default.json. The rule for handling abnormal health index is that when the calculated result of the health index exceeds the range of 0 - 100, it is automatically clamped to the nearest boundary value and the data verification process is triggered. The verification process includes re-collecting sensor data and retraining the linear regression model. The rule for handling sorting conflicts is that if the KL divergence difference or similarity of multiple strategies is the same, they are sorted in ascending order of the strategy generation timestamp, which is extracted from the strategy space reset log in step S3, and the log format is Unix timestamp (millisecond precision).

[0119] The hardware dependencies include a local database storage device, such as the MySQL 8.0 Community Edition, with a memory capacity of not less than 8GB to ensure the calculation efficiency of kernel density estimation. The software dependencies include the KernelDensity class of the Python library Scikit-learn 1.2.0 used for kernel density estimation calculation, and the dot function and linalg.norm function of the NumPy 1.24.3 library used for cosine similarity calculation. The simulation step control needs to be compatible with Python 3.9 and above, and the digital twin simulation model in step S4 is called through subprocess.

[0120] The unity of dimension and unit is achieved through data standardization. The original data of vibration amplitude, temperature gradient, and electromagnetic interference intensity are converted to standard units before being input into the kernel density estimation model. The parameter range limit is forcibly constrained by the preset values in the device technical manual. For example, when the vibration amplitude exceeds 10 m / s², it is automatically truncated to 10 m / s². The calculation result of the KL divergence difference is ensured to be non-negative through natural logarithm conversion to avoid numerical overflow caused by non-overlapping probability distributions.[[ID=**7**]] [[ID=**8**]]

[0121] This step S5 solves the static trade-off defect between exploration and exploitation in traditional strategy priority allocation through the dynamic classification and sorting of KL divergence difference and historical strategy similarity. The priority interval is dynamically divided based on the device health index to ensure that the verified strategies are preferentially inherited in a high health state (ensuring stability), and new strategies are focused on exploration in a low health state (enhancing adaptability). By constructing a parameter distribution through kernel density estimation and quantifying the consistency of strategy directions using cosine similarity, it overcomes the problem of strategy rigidity caused by the existing technology relying on fixed thresholds or simple weighting. Compared with traditional methods, through the dynamic classification mechanism driven by the health index, it realizes the adaptive adjustment of the strategy optimization process and provides decision support that takes into account both stability and diversity for complex production environments.

[0122] S6. Deploy the strategy with the highest exploration priority to the actual production system and update the current set of optimized strategies. Specifically, it can be implemented as follows:

[0123] Extract the strategy with the highest priority from the strategy exploration priority ranking result generated in step S5 and verify that it meets the device real-time health index and production order process standards. The device real-time health index is obtained by linear regression prediction of the vibration spectrum characteristics and temperature drift in step S4. The value range of the health index is from 0 to 100, the first threshold is 50, and the second threshold is 70. The production order process standards are extracted from the order document, including indicators such as dimensional accuracy and surface roughness. For example, the dimensional accuracy requirement is ±0.05 mm, and the surface roughness requirement is Ra ≤ 0.8 μm. During verification, compare the strategy parameters with the standard values one by one. The parameter types include vibration amplitude, temperature gradient, and electromagnetic interference intensity, which are strictly consistent with the definition of the dynamic parameters in step S4. The unit of vibration amplitude is meters per second squared (m / s²), the unit of temperature gradient is degrees Celsius per minute (℃ / min), and the unit of electromagnetic interference intensity is millitesla (mT). If all parameters are within the allowable tolerance range, it is determined that the verification is passed.

[0124] Deploy the verified strategy to the actual production system. The deployment instruction is transmitted to the device controller through the industrial communication protocol OPC UA. The device controller model is Siemens S7-1500. After the deployment is completed, update the current set of optimized strategies according to the device real-time health index: when the health index is lower than the first threshold (50), replace the strategy with the smallest KL divergence difference in the current set; when the health index is higher than the second threshold (70), replace the strategy with the highest historical strategy similarity. The KL divergence difference is calculated by the kernel density estimation method in step S5, and the historical strategy similarity is calculated by cosine similarity. The execution frequency of the replacement operation is synchronized with the production batch. For example, trigger the update of the strategy set after each production batch is completed.

[0125] Record the execution result of the deployed strategy in the historical strategy set. The execution result includes the deviation rate between the actual production efficiency index and the simulation prediction value. The actual production efficiency index is collected in real time by sensors. For example, the processing cycle time unit is seconds, and the energy consumption value unit is kilowatt-hours; the simulation prediction value is extracted from the output of the digital twin model in step S4. The deviation rate is calculated by the root mean square error. The calculation formula is the square root of the mean square of the difference between the actual value and the predicted value, and the calculation result is normalized to a percentage form. The deviation rate is stored as an additional field in the historical strategy set for subsequent strategy optimization reference.

[0126] Dynamically adjust the length of the moving average window in step S1 according to the deviation rate. The adjustment rule is as follows: when the deviation rate exceeds the preset tolerance value (10%), shorten the window length to quickly respond to changes, for example, reduce it from 30 iterations to 20; when the deviation rate is lower than the tolerance value, extend the window length to smooth out noise interference, for example, expand it from 30 to 40. The adjustment amplitude of the window length is positively correlated with the change rate of the deviation rate, and the change rate is calculated by the difference in deviation rates between two adjacent production batches. The difference calculation method is the absolute value of the current deviation rate minus the deviation rate of the previous batch.

[0127] Generate a policy deployment execution log. The log format is the same as the storage structure of the policy space reset log in step S3, including fields such as timestamp, policy parameters, execution results, and health index. The timestamp uses the Unix timestamp format (millisecond precision), the policy parameters are stored as JSON key-value pairs, and the execution result includes the percentage of the deviation rate between the actual value and the predicted value. The log file is named by combining the device unique identification code and the timestamp. For example, if the device identification code is CNC-001 and the timestamp is 20231005120000, the log is named CNC-001_20231005120000.json. The log storage path is the / log / strategy_deployment directory of the device local database. The local database is the MySQL 8.0 Community Edition, and it is synchronously backed up to the cloud storage service AWS S3 with a backup period of 1 am every day.

[0128] The alignment of policy parameters with the production order process standards is achieved through a parameter mapping table. The mapping table records the correspondence between policy parameter names and process standard indicators. For example, "cutting speed" is mapped to "processing cycle time". During the verification process, if the process standard is updated, the mapping table is automatically reloaded and a secondary verification is triggered. The policy with the smallest KL divergence difference is calculated by traversing the current policy set, and the deployed policies are excluded during the calculation to avoid repeated replacement.

[0129] If the deployment policy verification fails (for example, the parameters exceed the tolerance range), the rollback mechanism is automatically triggered to roll back to the previous valid policy and mark the current policy as a failed policy. When the deviation rate exceeds the tolerance value three times in a row, the manual intervention process is triggered to suspend the automatic policy update and wait for the engineer to confirm. If the log storage fails (such as insufficient disk space), it is temporarily cached in memory and retried for writing. The retry interval is 5 minutes, and it is retried at most 5 times. After the retry fails, a storage exception alarm is triggered.

[0130] Hardware dependencies include industrial controllers that support the OPC UA protocol (such as Siemens S7-1500) for deploying instruction transmission, and local database storage devices (such as Raspberry Pi 4B with an SSD) for log storage. Software dependencies include the JSON log generation library (the json module in Python), cloud storage clients (AWS S3 SDK version 1.26.7), and an open-source OPC UA library (FreeOpcUa version 0.98.1) for deploying instruction transmission.

[0131] The dimensions of the actual production efficiency indicators are strictly consistent with the simulation prediction values. For example, the unit of processing cycle time is seconds, and the unit of energy consumption is kilowatt-hours. When calculating the normalized deviation rate, the denominator is the absolute value of the simulation prediction value. If the prediction value is zero, it is automatically replaced with the historical average to avoid division-by-zero errors. The adjustment step size of the moving average window length is fixed at 10 iterations, with a minimum window length of not less than 10 times and a maximum of not more than 50 times.

[0132] A closed-loop optimization system is constructed through multi-step technology linkage, deeply coupling the real-time health status of the equipment with the process of policy optimization, and breaking through the limitations of the isolated processing of policy generation, verification, and deployment in traditional methods. In steps S1 to S3, based on the policy space reset mechanism monitored by dynamic entropy values, the multi-physical field coupling effect is quantified through topological data analysis and chaotic attractor modeling, solving the industry problem that traditional simulation models are difficult to capture non-linear dynamic associations. In steps S4 to S5, a priority classification rule driven by a health index is introduced, dynamically binding the real-time state of the equipment (such as vibration, temperature, etc.) with the policy diversity evaluation parameters, replacing the traditional fixed weight or static threshold allocation method, and achieving an adaptive adjustment of exploration-exploitation balance. In step S6, the update logic of policy deployment and the historical policy set dynamically corrects the data acquisition window length through deviation rate feedback, forming a cross-step adaptive learning closed-loop. Compared with the local optimization schemes based on rule engines or empirical formulas in the prior art, through the technical collaboration of topological network construction, dynamic interval division, and cross-cycle feedback correction, a leap in global optimization ability is achieved. Deeply integrating cross-domain tools such as mathematical topology and information entropy theory with industrial optimization scenarios, for example, using persistent homology features to quantify policy stability and the asymmetric sorting of KL divergence differences and cosine similarities.

[0133] Embodiment 2: Figure 2 A structural schematic diagram of the production decision-making adaptive optimization system of the present invention is given. The production decision-making adaptive optimization system includes the following modules:

[0134] An entropy value monitoring module for real-time monitoring of the policy entropy value of the current optimization policy set, where the policy entropy value is obtained by calculating the distribution dispersion of each policy in historical iterations;

[0135] An instruction trigger module, which is used to trigger a reconstruction instruction for the digital twin simulation model when the policy entropy value is lower than a preset entropy value threshold;

[0136] A parameter reset module, which is used to reset the policy parameter space of the digital twin simulation model according to the reconstruction instruction to generate a candidate policy set containing a new policy parameter space;

[0137] A simulation verification module, which is used to perform simulation verification on the candidate policy set by using the digital twin simulation model after the policy space is reset, and screen out a feasible policy subset that meets the preset constraint conditions;

[0138] A priority generation module, which is used to dynamically generate a policy exploration priority based on the KL divergence difference between each policy in the feasible policy subset and the historical policy;

[0139] A policy deployment module, which is used to deploy the policy with the highest policy exploration priority to the actual production system and update the current optimization policy set.

[0140] The calculations involved in the embodiments are all dimensionless numerical calculations, and the preset parameters and threshold selections in the calculations are set by those skilled in the art according to the actual situation.

[0141] It should be noted that the present invention can be deployed on the device itself to achieve embedded applications, or can also run on a PC or other terminal with a user interface, so as to meet various hardware environments and usage requirements.

[0142] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on the computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that contains a set of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0143] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and modules described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0144] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or modules can be in electrical, mechanical, or other forms.

[0145] The modules described as separate components may or may not be physically separated, and the components displayed as modules may or may not be physical modules. They can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0146] In addition, in each embodiment of the present application, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module.

[0147] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or this part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0148] As described above, it is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

[0149] Finally, the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An adaptive optimization method for production decision-making, characterized in that It includes the following steps: S1. Monitor the policy entropy value of the current optimization policy set in real time. The policy entropy value is obtained by calculating the distribution dispersion of each policy in historical iterations; S2. When the policy entropy value is lower than the preset entropy value threshold, trigger the reconstruction instruction of the digital twin simulation model; S3. Reset the policy space of the digital twin simulation model according to the reconstruction instruction to generate a candidate policy set containing the new policy parameter space; S4. Use the digital twin simulation model after policy space reset to simulate and verify the candidate policy set, and screen out a feasible policy subset that meets the preset constraint conditions; S5. Dynamically generate the policy exploration priority based on the KL divergence difference between each policy in the feasible policy subset and the historical policy; S6. Deploy the policy with the highest policy exploration priority to the actual production system and update the current optimization policy set.

2. The production decision-making adaptive optimization method according to claim 1, wherein Monitor the policy entropy value of the current optimization policy set in real time. The policy entropy value is obtained by calculating the distribution dispersion of each policy in historical iterations, including: Obtain the policy parameter vector of each policy in the current optimization policy set in historical iterations; Extract the feature dimensions of all policy parameter vectors within the preset time window; Calculate the variance value of each dimension parameter of the policy parameter vector based on the numerical fluctuation range of the feature dimension; Generate the policy entropy value by weighted summation of the variance values of each dimension parameter.

3. The production decision-making adaptive optimization method according to claim 1, characterized in that When the policy entropy value is lower than the preset entropy value threshold, trigger the reconstruction instruction of the digital twin simulation model, including: Obtain the dynamic adjustment rule of the preset entropy value threshold, compare the currently monitored policy entropy value with the preset entropy value threshold, and generate a threshold deviation degree evaluation result; Generate a reconstruction instruction for the digital twin simulation model containing the reconstruction intensity level according to the judgment that the threshold deviation degree evaluation result exceeds the preset deviation threshold; Transmit the reconstruction instruction to the control interface of the digital twin simulation model, and the control interface performs the corresponding parameter initialization operation according to the reconstruction intensity level.

4. The production decision-making adaptive optimization method according to claim 1, characterized in that Reset the policy space of the digital twin simulation model according to the reconstruction instruction to generate a candidate policy set containing the new policy parameter space, including: Analyze the reconstruction intensity level in the reconstruction instruction to generate a parameter dependency relationship constraint table corresponding to the reconstruction intensity level. The parameter dependency relationship constraint table records the physical coupling rules and priority weights between process parameters; Dynamically decouple the conflict parameter dimensions and lock the independent adjustable dimensions according to the reconstruction intensity level and the parameter dependency relationship constraint table; Based on the device real-time sensor data and the parameter dependency relationship constraint table, perform multi-level boundary constraint iterative correction on the decoupled independent adjustable dimensions; Generate a candidate policy set within the corrected policy parameter space through dynamic adaptive pseudo-random sampling.

5. The production decision-making adaptive optimization method according to claim 4, characterized in that The locking condition of the independent adjustable dimension is that the parameter adjustment range does not violate the physical limit of the device and meets the process standards of the production order; The correction rule of the multi-level boundary constraint iterative correction is to preferentially relax the parameter dimension constraints with high priority weights and synchronously compress the parameter dimension constraints with low priority weights; The sampling density of the dynamic adaptive pseudo-random sampling is dynamically adjusted according to the priority weights of the parameter dimensions.

6. The production decision-making adaptive optimization method according to claim 1, characterized in that, Use the digital twin simulation model after strategy space reset to simulate and verify the candidate strategy set, and filter out a feasible strategy subset that meets the preset constraint conditions, including: Generate time series of the dynamic parameters of each strategy in the candidate strategy set. The dynamic parameters include vibration amplitude, temperature gradient, and electromagnetic interference intensity; Generate nodes of the topological network based on the local maximum and minimum points of the time series of the dynamic parameters. The absolute value of the Pearson correlation coefficient of the time series of the dynamic parameters in adjacent time windows constitutes the edges of the topological network; Extract the persistent homology features of the topological network, calculate the Betti numbers and persistence interval lengths of each dimension, and generate a set of topological invariants; Calculate the dynamic constraint satisfaction degree parameter according to the evolution trend of the set of topological invariants. The evolution trend is jointly characterized by the change rate of the persistence interval length and the stability of the Betti numbers; Filter out the strategies with dynamic constraint satisfaction degree parameters higher than the dynamic threshold to generate a feasible strategy subset. The dynamic threshold is dynamically adjusted according to the real-time health index of the device.

7. The production decision-making adaptive optimization method according to claim 1, characterized in that Based on the KL divergence difference between each strategy in the feasible strategy subset and the historical strategies, dynamically generate the strategy exploration priority, including: Generate a strategy diversity evaluation parameter based on the KL divergence difference between the parameter distributions of each strategy in the feasible strategy subset and the historical strategy set. The parameter distribution is constructed by the kernel density estimation method; Divide the priority interval according to the real-time health index of the device; Dynamically sort the KL divergence difference and the historical strategy similarity. When the KL divergence difference is greater than the historical strategy similarity, it is classified as an exploration priority strategy, otherwise it is an inheritance priority strategy; Generate the strategy exploration priority based on the priority interval and the strategy classification result. The exploration priority strategies are arranged in descending order of the KL divergence difference within the interval, and the inheritance priority strategies are arranged in ascending order of the historical strategy similarity. The number of priority levels is positively correlated with the number of strategies in the current strategy set.

8. The production decision-making adaptive optimization method according to claim 7, characterized in that When dividing the priority interval according to the real-time health index of the device, when the real-time health index of the device is lower than the first threshold, the priority interval is limited to high-diversity strategies, and when it is higher than the second threshold, the priority interval is extended to high-similarity strategies.

9. The production decision-making adaptive optimization method according to claim 1, characterized in that Deploy the strategy with the highest strategy exploration priority to the actual production system and update the current optimization strategy set, including: Extract the strategy with the highest priority from the strategy exploration priority sorting result and verify that it meets the real-time health index of the device and the process standards of the production order; Deploy the verified strategy to the actual production system and update the current optimization strategy set according to the real-time health index of the device: when the real-time health index of the device is lower than the first threshold, replace the strategy with the smallest KL divergence difference, and when it is higher than the second threshold, replace the strategy with the highest historical strategy similarity; Record the execution result in the historical strategy set, dynamically adjust the moving average window length according to the deviation rate between the actual and the simulation, and generate a strategy deployment execution log.

10. A production decision adaptive optimization system for implementing the production decision adaptive optimization method according to any one of claims 1-9, characterized in that, Including the following modules: Entropy value monitoring module, used to monitor the strategy entropy value of the current optimization strategy set in real time. The strategy entropy value is obtained by calculating the distribution dispersion of each strategy in the historical iterations; Instruction trigger module, when the strategy entropy value is lower than the preset entropy value threshold, used to trigger the reconstruction instruction of the digital twin simulation model; A parameter reset module, which is used to perform policy space reset on the policy parameter space of the digital twin simulation model according to the reconstruction instruction, and generate a candidate policy set containing the new policy parameter space; A simulation verification module, which is used to perform simulation verification on the candidate policy set by using the digital twin simulation model after policy space reset, and screen out a feasible policy subset that meets the preset constraint conditions; A priority generation module, which is used to dynamically generate policy exploration priorities based on the KL divergence differences between the policies in the feasible policy subset and the historical policies; A policy deployment module, which is used to deploy the policy with the highest policy exploration priority to the actual production system and update the current optimized policy set.

Citation Information

Patent Citations

  • Multi-objective optimization method and system for process production process based on digital twinning

    CN115423333A

  • Boiler combustion operation optimization method and system based on digital twinning

    CN118816182A

  • Industrial manufacturing process and production operation and maintenance optimization method and system based on digital twinning

    CN118884908A

  • Database adaptive data flow acquisition optimization method and system based on reinforcement learning

    CN119719783A

  • Building energy consumption dynamic optimization method and system based on BIM and reinforcement learning

    CN120145880A

Cited By

  • Laser welding parameter optimization method and device based on part characteristics

    CN121083079A

  • An AI-based API security policy optimization method and system

    CN122578343A