Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

17 results about "Double loop" patented technology

Double-loop learning entails the modification of goals or decision-making rules in the light of experience. The first loop uses the goals or decision-making rules, the second loop enables their modification, hence "double-loop". Double-loop learning recognises that the way a problem is defined and solved can be a source of the problem.

Digital twin multi-agent reinforcement learning intelligent decision-making system with secure memory playback mechanism

The invention discloses a digital twinning multi-agent reinforcement learning intelligent decision-making system and method with a secure memory playback mechanism, and the system comprises a digital twinning module which is used for constructing a virtual model and synchronizing the virtual model with a physical entity in real time; the multi-agent reinforcement learning module is used for carrying out strategy learning based on a constrained Markov decision process and balancing performance and safety through a Lagrange multiplier; the safe memory playback module is used for weighting and playing back the experience samples according to the risk and the timeliness so as to improve the learning safety; the reversible grey influence network module is used for causal modeling and reasoning and enhancing decision interpretability; the double-loop self-constraint control module ensures that a control action is always in a physical safety boundary through a barrier function and safety projection; and the convergence and stability criterion module is used for verifying strategy security convergence and system asymptotic stability. According to the method, the problems of strategy border crossing, virtual-real mismatching and the like in the high-risk manufacturing process are solved, and multi-target optimal control under the safety constraint is realized.
Owner:CHONGQING UNIV +1

Permanent magnet synchronous motor PI double-loop control method based on deep reinforcement learning

The invention discloses a permanent magnet synchronous motor PI double-loop control method based on deep reinforcement learning. The method comprises the following steps: establishing a permanent magnet synchronous motor double-loop coupling mathematical model; a traditional FOC composite controller is constructed; designing a state feedback and reward mechanism of the double-ring TD3; constructing a speed loop and current loop TD3 + PI composite controller; and carrying out system stability analysis. According to the method, a system is divided into an inner ring and an outer ring based on hybrid cooperative control fusing deep reinforcement learning and classical control, the problem that robustness is insufficient under system disturbance through PI control and the problem that tracking precision is reduced due to load sudden change and noise interference in PMSM operation are combined, nonlinear disturbance is dynamically compensated through the strategy learning ability of a TD3 algorithm, and the tracking precision is improved. Meanwhile, soft update is introduced to improve convergence stability, and a compound control strategy based on double-loop TD3 + PI is provided. According to the invention, complex non-linear processing is simplified, and the stability of the system and the accuracy of rotating speed tracking are improved at the same time.
Owner:XUZHOU NORMAL UNIVERSITY

Robot motion control method and system based on cerebellum reinforcement learning

The invention discloses a robot motion control method and system based on cerebellum reinforcement learning, and the method comprises the steps: obtaining a current environment state vector, and inputting the current environment state vector to a main strategy channel and a cerebellum compensation channel in parallel; the main strategy channel outputs a basic action based on a long-term task target, and the cerebellum compensation channel outputs a compensation action responding to real-time dynamic through an efficient query mechanism; synthesizing the basic action vector and the compensation action vector into a synthesized action vector driving robot; feeding back latest data after the robot drives the motion action vector, and determining a sensory prediction error based on the latest data; and updating the original parameters of the cerebellum compensation channel based on the sensory prediction error, and optimizing the original strategy parameters of the main strategy channel based on the latest data. According to the invention, by constructing a parallel double-channel architecture, functional decoupling of advanced decision and rapid adaptation is realized, and unification of rapid adaptation and continuous optimization is realized through a double-loop hierarchical learning system.
Owner:SINARD DIGITAL TECH (SHANGHAI) CO LTD

A robot motion control method and system based on cerebellum reinforcement learning

The application discloses a kind of robot motion control method and system based on cerebellum reinforcement learning, it includes: obtaining current environment state vector, and parallel input to main strategy channel and cerebellum compensation channel;Main strategy channel outputs the basic action based on long-term task target, and cerebellum compensation channel then outputs compensation action by efficient query mechanism to cope with real-time dynamics;The basic action vector and compensation action vector are synthesized into synthesized action vector to drive robot;After robot drive movement action vector, feedback latest data, and determine sensory prediction error based on latest data;Based on sensory prediction error, the original parameters of cerebellum compensation channel are updated, and the original strategy parameters of main strategy channel are optimized based on latest data.The application realizes the functional decoupling of high-level decision and rapid adaptation by constructing parallel double-channel architecture, and the hierarchical learning system of double loop realizes the unity of rapid adaptation and continuous optimization.
Owner:SINARD DIGITAL TECH (SHANGHAI) CO LTD

Multi-stage pid parameter regulation method for laser frequency locking system based on deep learning

The application relates to the technical field of laser system control, and discloses a multi-stage PID parameter regulation and control method for a laser frequency locking system based on deep learning, which comprises the following steps: acquiring an error signal time sequence of a PDH laser frequency locking system; outputting a frequency locking state score and fast-slow double-loop PID parameters through a deep learning model; judging whether the frequency locking state score and the fast-slow double-loop PID parameters both satisfy corresponding preset conditions; if the frequency locking state score and the fast-slow double-loop PID parameters do not both satisfy the corresponding preset conditions, iteratively optimizing the deep learning model until the frequency locking state score and the fast-slow double-loop PID parameters both satisfy the corresponding preset conditions; and outputting the fast-slow double-loop PID parameters of the optimized deep learning model. The application can realize online self-adaptive reasoning and closed-loop optimization of fast-slow double-loop multi-stage PID control parameters, and improve the intelligent level and the ability to adapt to complex working conditions of the PDH laser frequency locking system.
Owner:HANGZHOU INST FOR ADVANCED STUDY UCAS

PID parameter optimization method for reducing motion control position deviation uncertainty

ActiveCN121187221BOptimizing Control ParametersImprove the problem that optimization results may be inconsistentProgramme controlComputer controlLocal optimumAlgorithm
The application relates to a PID parameter optimization method for reducing motion control position deviation uncertainty, which comprises the following steps: firstly, determining the PID parameter optimization form and constraint conditions, and constructing an optimization objective function for the motion control position deviation uncertainty; secondly, establishing a Gaussian process regression proxy model between the PID parameters and the optimization objective function; thirdly, determining a next group of control parameters according to an expected enhancement acquisition function and measuring the corresponding optimization objective function value; finally, optimizing the PID parameters by using a double loop, wherein the inner loop uses randomly set initial control parameters to iteratively execute and obtain inner loop optimization control parameters, the outer loop is used for starting the inner loop optimization process multiple times, and finally, the final tuning parameter value is calculated by using non-convex scenario optimization; the selection of the optimization objective function of the application considers the average position deviation and the position deviation uncertainty, the optimization method adopts a double loop Bayesian optimization mode, and the method has the advantages of not being prone to local optimization and good consistency of optimization results under the action of uncertainty disturbance.
Owner:XI AN JIAOTONG UNIV

Multi-objective collaborative optimization and decision-making method for laser cladding process parameters

The invention relates to the technical field of machine learning and intelligent optimization, and discloses a multi-objective collaborative optimization and decision-making method for laser cladding process parameters, and the method comprises the steps: an initialization stage: constructing a physical correlation model and a Gaussian process agent model; in the fast loop optimization stage, the estimation performance is obtained by utilizing the physical correlation model so as to quickly update the Gaussian process agent model; in a slow loop calibration stage, concurrently obtaining foundation truth value data of the key points; and a slow loop reconstruction stage: reversely updating the physical correlation model by using the ground-based truth value, and globally reconstructing the Gaussian process proxy model. According to the method, a double-ring optimization architecture with cooperative work of fast-ring fast exploration and slow-ring concurrent calibration is constructed, so that dynamic correction of system deviation of the low-cost physical correlation model by using a high-cost foundation truth value is realized, and parameter exploration efficiency is greatly improved while optimization precision is ensured; the problem that an existing optimization method is difficult to balance among speed, cost and precision is solved.
Owner:SICHUAN LIANGSHANSHUILUOHE ELECTRICITY DEV CO LTD

Hydraulic support multi-agent autonomous learning method

This invention relates to a multi-agent autonomous learning method for hydraulic supports, belonging to the field of intelligent coal mine technology. It includes: using the scene features input by a learning-type real-loop agent and the output displacement / pressure-time action sequence as input to a learning-type dual-loop agent; the learning-type dual-loop agent corrects the displacement / pressure-time action sequence to obtain a corrected displacement / pressure-time action sequence; the corrected displacement / pressure-time action sequence is then returned to the learning-type simulation loop agent to verify for conflicts; if conflicts are found after correction, the correction continues, iterating until the corrected displacement / pressure-time action sequence is conflict-free; finally, the final actual action strategy is expanded to a preset strategy set for training a higher-quality learning-type real-loop agent. This invention achieves closed-loop data fusion analysis that combines real data and simulated data.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

Distributed collaborative learning and elastic response method, system and device based on eBPF and medium

The invention relates to the technical field of distributed collaborative learning and elastic response, in particular to an eBPF-based distributed collaborative learning and elastic response method, system, equipment and medium, which comprises the following steps of: performing global analysis in a user mode based on system behavior data acquired by a kernel mode eBPF program to construct a normal behavior baseline model; the baseline model is synchronized to a kernel mode after being lightened; in the kernel mode, deviation detection is carried out on system behaviors generated in real time based on the synchronous baseline model, and a risk score is output; determining a response level in combination with the risk score and a preset asset criticality; and in a kernel mode, executing a response action corresponding to the response level. The method has the beneficial effects that the inherent contradiction between the response speed and the detection precision in the prior art is successfully solved by constructing the fast and slow double-loop collaborative security system.
Owner:GUANGZHOU ELECTRIC POWER COMM NETWORK LTD

Auxiliary decision-making method and system for cultural relic-oriented water conservancy facility reconstruction

PendingCN122288956Aensure rigorEnsure traceabilityDesign phaseInner loop
This invention provides an auxiliary decision-making method and system for the renovation of cultural relic-related water conservancy facilities, belonging to the field of water conservancy engineering management technology. The decision-making method achieves quantitative collaborative decision-making for the dual objectives of engineering safety and cultural relic protection, transforming the qualitative requirements of cultural relic protection into structured data that can be calculated alongside engineering indicators, thus overcoming the previous decision-making dilemma of separating these two dimensions. It establishes a complete technical closed loop from assessment to design, decomposing complex decision-making problems into three logically clear stages: functional zoning, strategy optimization, and structural matching, ensuring the rigor and traceability of the decision-making logic. The introduction of a dual-loop feedback mechanism significantly improves the rationality throughout the entire lifecycle. The inner loop feedback ensures the mechanical consistency between the design scheme and the decision-making objectives, avoiding performance deviations during the design phase; the outer loop feedback uses actual operation and maintenance data to back-optimize the initial decision-making model, enabling the system to have self-learning and self-evolution capabilities.
Owner:ANHUI SURVEY & DESIGN INST OF WATER CONSERVANCY & HYDROPOWER

Multi-time scale energy optimization scheduling method

The invention relates to the technical field of automatic testing, in particular to a multi-time-scale energy optimization scheduling method. The invention relates to a multi-time scale energy optimization scheduling method, which comprises an energy system and is characterized in that the energy system comprises a data and constraint convergence layer; performing feature engineering and state estimation; generating a scene library; a timing predictor; gNN flow proxy is carried out; performing day-ahead distribution robust optimization; performing intra-day rolling consistency correction; a real-time controller; modeling health and life; consistency coordination and double-loop verification are carried out; training and distilling; and monitoring and auditing. Compared with the prior art, the multi-time scale energy optimization scheduling method provided by the invention comprises the following steps: generating a tail scene through a diffusion model and incorporating the tail scene into robust optimization; a DC-TSP is used to extract multi-scale time sequence features; the Auto-FRL is combined with the safety barrier projection to realize millisecond-level control; gNN is fast approximate to the power flow, and NeuralODE accurately describes energy storage degradation; and the global consistency of day-ahead, day-intra and real-time solutions is ensured through double-ring verification.
Owner:SHANGHAI PYTES ENERGY CO LTD

Multi-agent autonomous learning method for hydraulic support

The invention relates to a multi-agent autonomous learning method for a hydraulic support, and belongs to the technical field of coal mine intellectualization. Comprising the steps that scene features input by a learning type reality ring Agent and a displacement / pressure-time action sequence output by the learning type reality ring Agent serve as input of a learning type double-ring Agent, the learning type double-ring Agent corrects the displacement / pressure-time action sequence, and the corrected displacement / pressure-time action sequence is obtained; the corrected displacement / pressure-time action sequence is returned to the learning type simulation ring Agent again to verify whether conflicts exist or not, if conflicts exist after correction, correction is continued, loop iteration is conducted till the corrected displacement / pressure-time action sequence does not have conflicts, and a final actual action strategy is expanded to a preset strategy set. The method is used for training the learning type reality ring Agent with higher quality. According to the invention, closed-loop data fusion analysis combining real data and simulation data is realized.
Owner:TAIYUAN UNIVERSITY OF TECHNOLOGY

A beyond-visual-range air combat double-loop coupling autonomous maneuver decision-making method, device, medium and product based on situation driving

The application discloses a beyond-visual-range air combat double-loop coupling autonomous maneuver decision-making method and device based on situation driving, a medium and a product, and relates to the technical field of aerospace. The method comprises the following steps: using a trained LSTM model, respectively according to state control information of an enemy target and a local machine, performing recursive prediction to obtain multi-step track prediction information of the enemy target and the local machine; calculating the situation change gradient of both sides; inputting the track information of the local machine and the state control information of the enemy target into a trained first reinforcement learning model to generate a main action instruction; inputting the track information of the local machine, the state control information of the enemy target and the situation change gradient of both sides into a trained second reinforcement learning model to generate a preloaded action instruction; and using a null space behavior method to fuse the main action instruction and the preloaded action instruction to obtain an autonomous maneuver decision-making instruction. The application can reduce the decision-making risk under incomplete information and realize smooth tactical conversion of a combat aircraft.
Owner:RES & DEV INST OF NORTHWESTERN POLYTECHNICAL UNIV IN SHENZHEN

A real-time signaling flow intelligent arrangement system based on queue optimization

The application discloses a kind of real-time signaling stream intelligent arrangement systems based on queue optimization, and the application relates to communication network technical field, comprising: multi-dimensional feature perception and fusion module, for real-time acquisition and output the multiple dynamic attribute parameters of signaling;Adaptive scheduling decision module is connected multi-dimensional feature perception and fusion module, for receiving the dynamic attribute parameter and generating scheduling instruction.The real-time signaling stream intelligent arrangement systems based on queue optimization, by introducing the double-loop intelligent architecture of prediction and decision verification, effectively improves the foresight and reliability of signaling stream scheduling.System can make decision by comprehensively multi-dimensional dynamic factor, and after the double inspection of rule base and micro simulation, thereby still maintain the rationality of scheduling strategy and system stability when facing complex burst traffic, guarantee the quality of service of end to end.
Owner:ZHUHAI WANSI INFORMATION TECH CO LTD

Automobile door hinge production line dynamic scheduling method and system based on digital twinning

The invention discloses an automobile door hinge production line dynamic scheduling method and system based on digital twinning, and the method comprises the steps: S1, collecting the original production state data of an automobile door hinge production line, and carrying out the data cleaning preprocessing of the original production state data, and obtaining the production state data; s2, a production line digital twinborn model is obtained through combination of a preset mechanism model and a real-time data driving function, and the production state data is input into the production line digital twinborn model for simulation initialization to obtain a simulation prediction result; s3, constructing a double-loop reward agent combining a reinforcement learning agent and a double-loop reward function, inputting a simulation prediction result into the double-loop reward agent for online learning optimization, and outputting an optimal scheduling strategy parameter; and S4, the optimal scheduling strategy parameters are issued to an execution mechanism at the bottom layer of the production line, and real-time dynamic scheduling of the automobile door hinge production line is achieved.
Owner:DIJING SEMICON TECH (SUZHOU CO LTD

Automobile door hinge production line dynamic scheduling method and system based on digital twinning

The application discloses a dynamic scheduling method and system for an automobile door hinge production line based on digital twinning, and comprises the following steps: S1, collecting original production state data of the automobile door hinge production line, performing data cleaning and preprocessing on the original production state data to obtain production state data; S2, obtaining a production line digital twinning model by combining a preset mechanism model and a real-time data driving function, inputting the production state data into the production line digital twinning model for simulation initialization to obtain simulation prediction results; S3, constructing a double-loop reward intelligent agent combined with a reinforcement learning intelligent agent and a double-loop reward function, inputting the simulation prediction results into the double-loop reward intelligent agent for online learning optimization to output optimal scheduling strategy parameters; and S4, issuing the optimal scheduling strategy parameters to an execution mechanism at a bottom layer of the production line to realize real-time dynamic scheduling of the automobile door hinge production line.
Owner:DIJING SEMICON TECH (SUZHOU CO LTD

Machine learning-based denim prescoring process adaptive control method and system

This invention discloses an adaptive control method and system for denim pre-shrinking process based on machine learning, belonging to the field of fabric pre-shrinking technology. The method includes: acquiring denim production process data and real-time images, and inputting them into a pre-established lightweight physical simulation network to obtain a virtual state of the fabric surface; inputting the virtual state of the fabric surface into an adaptive control network for optimization processing to obtain pre-shrinking process parameters; based on the pre-shrinking process parameters, executing a double-loop learning process to obtain sorting control commands for controlling the pre-shrinking machine; and driving the pre-shrinking machine to perform actions according to the sorting control commands, completing the adaptive control of the denim pre-shrinking process. This invention uses a lightweight physical simulation network as a differentiable forward simulator, which can accurately deduce the virtual moisture content, shrinkage rate, and smoothness distribution of the fabric surface, providing accurate virtual quality feedback for parameter optimization. Combined with the double-loop learning mechanism, it ensures the uniformity and high standards of pre-shrinking quality from both the source and the process.
Owner:GUANGDONG HEFANG TEXTILE TECHNOLOGY CO LTD