A heavy gate dynamic torque regulation system based on reinforcement learning
By adopting a two-layer optimization framework based on reinforcement learning, neuroevolution and meta-learning, the problems of versatility and adaptability of heavy-duty gate torque control system are solved, realizing the system's intelligence and rapid adaptation, and improving deployment efficiency and stability in changing environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN JIEGUARD TECH CO LTD
- Filing Date
- 2025-07-28
- Publication Date
- 2026-05-29
Smart Images

Figure CN120871622B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence and control system technology, and more specifically, to a dynamic torque control system for heavy-duty turnstiles based on reinforcement learning. Background Technology
[0002] With the continuous advancement of artificial intelligence and reinforcement learning technologies, heavy-duty turnstiles, as crucial electromechanical equipment ensuring the safety of personnel and vehicles, are widely used in scenarios such as security checks, border inspections, and large venues. However, existing heavy-duty turnstile torque control systems generally rely on manually designed control algorithms and parameter optimization, making it difficult to meet the demands for efficient deployment and intelligent adaptation in changing environments.
[0003] Specifically, traditional control systems suffer from prominent problems such as insufficient versatility, low design efficiency, and poor adaptability. For different application scenarios, controllers often need to be redesigned and debugged specifically, leading to high system maintenance and upgrade costs. Meanwhile, system architecture and parameter optimization rely heavily on manual experience, making it difficult to achieve global optimization, and requiring substantial data and time for retraining in new environments, impacting practical application performance. Furthermore, existing methods struggle to dynamically balance exploring new strategies with utilizing historical experience, and the parameter tuning process is cumbersome, further restricting the system's intelligence and automation levels.
[0004] Therefore, a heavy-duty gate torque control system is needed that can simultaneously achieve coordinated optimization of network structure and parameters and has rapid environmental adaptability, so as to improve the system's versatility, intelligence and practical application value, and meet the requirements of efficient deployment and stable operation in complex and ever-changing scenarios. Summary of the Invention
[0005] This invention provides a dynamic torque control system for heavy-duty turnstiles based on reinforcement learning, which solves the technical problems of insufficient versatility, low design efficiency, and poor adaptability in related technologies.
[0006] This invention provides a reinforcement learning-based dynamic torque control system for heavy-duty turnstiles, comprising:
[0007] A two-layer optimization framework combining neuroevolution and meta-learning is used to simultaneously optimize neural network structure and parameters.
[0008] A neural gene encoding and decoding system for encoding controller network structures into evolvable neural genes;
[0009] A hierarchical neural architecture search space is used to provide a structured search space for evolutionary algorithms;
[0010] A neural modular evolutionary algorithm is used to achieve network structure evolution by evaluating the value of modules through functional similarity.
[0011] A hypernetwork meta-learning framework for generating task-specific controller parameters;
[0012] Evolutionary meta-learning collaborative optimization mechanism is used to enable two algorithms to promote and complement each other;
[0013] Deployment and online adaptation systems are used to achieve continuous optimization of the system during actual operation.
[0014] In a preferred embodiment, the neuroevolution and meta-learning bilayer optimization framework includes:
[0015] The structural optimization layer uses a neuroevolutionary algorithm to search for the optimal network topology.
[0016] The parameter optimization layer employs a meta-learning algorithm to achieve rapid adaptation.
[0017] The collaborative mechanism of the two-layer optimization includes a downlink information channel from structure to parameters and an uplink feedback channel from parameters to structure.
[0018] In a preferred embodiment, the neural gene encoding and decoding system includes:
[0019] The network topology coding module encodes the network's connection patterns, hierarchical structure, and neuron types into compact numerical vectors.
[0020] The activation function encoding module encodes the activation function types of neurons in each layer into discrete values;
[0021] The initial weight distribution encoding module encodes the initial distribution type and parameters of the network weights into a numerical vector;
[0022] A neural gene decoder is capable of reconstructing a complete neural network structure from neural gene representations.
[0023] In a preferred embodiment, the hierarchical neural architecture search space includes:
[0024] The macroscopic structural search space consists of a network architecture composed of feedforward networks, recurrent networks, and attention networks.
[0025] The mesoscopic module search space includes a combination of functional modules such as state encoders, policy networks, and value function networks.
[0026] The microscopic operation search space is a set of neural network operations consisting of convolution, fully connected, and recurrent units.
[0027] A hierarchical search space constraint rule system ensures that the generated network structure meets the performance requirements of the torque control system.
[0028] In a preferred embodiment, the neural modular evolutionary algorithm includes:
[0029] The functional similarity assessment module calculates the functional similarity matrix by comparing the output behavior of the network under the same input.
[0030] Functional module protection mechanism to prevent functional modules with unique functions from disappearing prematurely during the evolution process;
[0031] Novelty-based selection strategies encourage the exploration of unknown network structure spaces;
[0032] The crossover and mutation operations of functional modules include substructure exchange, connection mutation, and function replacement.
[0033] In a preferred embodiment, the hypernetwork meta-learning framework includes:
[0034] The task representation learning module encodes the features of the gate control environment into compact task embedding vectors.
[0035] The hypernetwork structure can generate all parameters of the target network based on the task embedding vector;
[0036] Hypernetwork training algorithms enable hypernetworks to quickly and accurately generate parameters that adapt to new tasks;
[0037] The parameter fine-tuning mechanism allows for rapid adaptive adjustments to the initial parameters generated by the hypernetwork.
[0038] In a preferred embodiment, the evolutionary meta-learning collaborative optimization mechanism includes:
[0039] A joint evaluation system for structural parameters comprehensively considers the potential of the structure and the adaptability of the parameters;
[0040] The two-way feedback mechanism enables the evolutionary layer and the meta-learning layer to share information and guide each other;
[0041] An adaptive scheduler that balances exploration and exploitation dynamically balances the exploratory nature of evolution and the exploitation of meta-learning based on the optimization progress.
[0042] An incremental knowledge accumulation mechanism stores valuable modules and parameter patterns discovered during evolution and learning.
[0043] In a preferred embodiment, the deployment and online adaptation system includes:
[0044] The edge cloud collaborative deployment architecture places computationally intensive evolutionary optimization in the cloud and real-time control and rapid adaptation on edge devices.
[0045] The online rapid adaptation mechanism enables the controller to quickly adjust parameters based on real-time environmental feedback;
[0046] An anomaly detection and recovery system monitors controller performance and triggers adaptive adjustments when performance degrades;
[0047] A continuous cycle of evolution and learning is used to continuously optimize the controller's structure and parameters using actual operational data.
[0048] In a preferred embodiment, a reinforcement learning-based dynamic torque control system for heavy-duty turnstiles is applied to at least one of the following scenarios:
[0049] Temporary security gates in mobile security inspection equipment;
[0050] Border inspection gates at temporary border checkpoints;
[0051] Entry gate systems for large event venues;
[0052] Emergency evacuation route turnstiles;
[0053] Industrial gate systems in smart factories.
[0054] In a preferred embodiment, a computer-readable storage medium is provided for storing computer-readable instructions that, when read by a computer, enable the operation of a reinforcement learning-based dynamic torque control system for heavy-duty gates.
[0055] The beneficial effects of this invention are as follows:
[0056] By combining a two-layer optimization mechanism of neuroevolution and meta-learning, the autonomous structural design and parameter optimization of the heavy-duty gate torque control system were achieved, improving the system's intelligence level. The system can automatically discover and adapt the optimal neural network structure according to different application scenarios and environmental conditions, eliminating the need for tedious parameter tuning and architecture design based on human experience, thus improving the control system's versatility and deployment efficiency.
[0057] In actual operation, the system demonstrates the ability to quickly adapt to new environments. Through a meta-learning mechanism, the controller can rapidly adjust parameters under limited environmental feedback, achieving efficient adaptation to new scenarios, shortening the time for system deployment and scenario switching, and meeting the needs for efficient deployment in changing environments.
[0058] Through an evolutionary meta-learning collaborative optimization mechanism, a dynamic balance is achieved between exploring new strategies and utilizing historical experience. The system can flexibly adjust the ratio of exploration to utilization based on optimization progress and environmental changes, continuously improving control performance and ensuring stability and robustness under complex operating conditions.
[0059] Furthermore, the system supports edge cloud collaborative deployment and continuous online optimization, enabling it to continuously accumulate knowledge and optimize control strategies during actual operation. Even in the face of sudden interference or environmental changes, the system can promptly detect anomalies and make adaptive adjustments, improving the stability and safety of torque control and meeting the high reliability requirements of heavy-duty turnstiles in critical scenarios. Attached Figure Description
[0060] Figure 1 This is a block diagram of a heavy-duty gate dynamic torque control system based on reinforcement learning according to the present invention.
[0061] Figure 2 This is a bar chart comparing the network structure optimization effects of the present invention;
[0062] Figure 3 This is a line graph comparing the environmental adaptability of the present invention;
[0063] Figure 4 This is a radar chart comparing the resource efficiency improvements of the present invention;
[0064] Figure 5 This is a bar chart showing the system stability test results of the present invention;
[0065] Figure 6 This is a line graph showing the application scenario adaptability test of the present invention;
[0066] Figure 7 This is an area graph showing the long-term operating performance trend of the present invention. Detailed Implementation
[0067] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, some features described in the examples may be combined in other examples.
[0068] At least one embodiment of the present invention discloses a dynamic torque control system for heavy-duty turnstiles based on reinforcement learning, such as... Figure 1 As shown, it includes:
[0069] A two-layer optimization framework combining neuroevolution and meta-learning is used to simultaneously optimize neural network structure and parameters.
[0070] Specifically, the following steps are included:
[0071] Step 1.1, Two-layer optimization objective and theoretical basis;
[0072] According to embodiments of this application, the overall goal of the two-layer optimization framework of neural evolution and meta-learning is first clarified. This framework includes a structure optimization layer and a parameter optimization layer, responsible for the evolutionary search of the neural network topology and the rapid adaptation of parameters, respectively. Its core idea is to unify structure search and weight optimization within the evolutionary-learning dual-loop optimization theory.
[0073] The optimization objective can be formalized as:
[0074]
[0075] Where g represents the candidate neural network structure, i.e., in the structure search space. A specific network topology to be evaluated; The structure search space represents the set of all neural network structures that can be explored by evolutionary algorithms; g * The optimal network structure is the one that has the best average performance under a multi-task distribution after meta-learning adaptation among all candidate structures. The initial parameters of structure g are typically the initial weights or parameter vectors of the network. Indicates in the task The above represents the parameters obtained after k steps of meta-learning adaptation; U is the meta-learning adaptation operation; k is the number of adaptation steps. For each task instance, a specific control task is sampled from the task distribution. It represents the task distribution, describing the probability distribution of all possible task instances; For task distribution All tasks The expected value calculation represents the average performance in a multi-tasking environment; This is a performance evaluation function used to measure performance on a task. The control performance of a specific network structure and its parameters (such as torque accuracy, response speed, etc.) is also considered.
[0076] The optimization objective is to find an optimal network structure g among all possible network structures g. * This allows the structure to adapt to various task environments through k-step meta-learning, resulting in an average performance evaluation value. maximize.
[0077] Step 1.2, Implementation of the structural optimization layer;
[0078] In the structural optimization layer, a neural network structure search system based on evolutionary algorithms is employed. This system includes modules for population initialization, fitness evaluation, selection, crossover, and mutation. Using the torque control performance of the heavy-duty turnstile as the fitness function, the optimal network topology is searched by simulating a biological evolutionary process.
[0079] The specific implementation is as follows:
[0080] A population of 50 to 100 network individuals is set up, and networks with different topologies are initially randomly generated. The overall fitness is calculated in a simulation environment using torque accuracy, response speed, and energy consumption as evaluation indicators.
[0081] The tournament selection method is used to select outstanding individuals, and a new generation of population is generated through crossover and mutation.
[0082] Prior knowledge can be introduced during the population initialization stage, such as incorporating known and effective control network structures into the initial population, to accelerate the evolutionary process.
[0083] Step 1.3, Implementation of the parameter optimization layer and the two-layer collaborative mechanism;
[0084] The parameter optimization layer employs a meta-learning-based rapid parameter adaptation system, including modules for task sampling, model initialization, inner loop adaptation, and outer loop update. Through meta-learning algorithms such as MAML and Reptile, it learns from tasks in different gate control environments to obtain parameter initialization methods with rapid adaptability.
[0085] During the task sampling phase, a diverse set of tasks is constructed, covering different loads, environmental disturbances, and mechanical characteristics, to ensure the broad adaptability of parameter initialization.
[0086] Furthermore, an information transmission and feedback mechanism is established between the structural layer and the parameter layer to achieve synergy in two-layer optimization. This mechanism includes a downlink information channel from structure to parameters and an uplink feedback channel from parameters to structure, promoting collaborative optimization between structure and parameters.
[0087] Optionally, an asynchronous update strategy can be adopted to enable structural optimization and parameter optimization to run independently and efficiently at different time scales, thereby improving the overall optimization efficiency.
[0088] like Figure 2 As shown, the advantages of the neuroevolutionary and meta-learning bilayer optimization framework over traditional methods in terms of network parameter quantity and control accuracy are demonstrated.
[0089] A neural gene encoding and decoding system for encoding controller network structures into evolvable neural genes;
[0090] Specifically, the following steps are included:
[0091] Step 2.1, construct the network topology coding module;
[0092] According to one embodiment of this application, an efficient neural network topology coding method is constructed to encode information such as the network connection pattern, hierarchical structure and neuron type into a compact numerical vector.
[0093] This module maps complex network topologies to fixed-length gene vectors g. topology This is so that evolutionary algorithms can process it.
[0094] In practical applications, this encoding module uses a combination of direct and indirect encoding.
[0095] For the core network layers in the torque control of heavy-duty gates, such as the state coding layer and the strategy output layer, their structure is described in detail using a direct coding method.
[0096] For the intermediate processing layer, an indirect encoding method is used to describe the generation rules, thereby reducing gene length.
[0097] For example, in a mobile security gate application case, the encoding module compressed and encoded a deep network containing 7 layers and 256 neurons into a gene vector of length 128, which improved the search efficiency of the evolutionary algorithm.
[0098] The encoding method in this application differs from the traditional direct encoding method. By introducing a hierarchical and functionally modular encoding strategy, the dimensionality of the search space is reduced, and the efficiency of the evolutionary algorithm is improved.
[0099] Step 2.2, construct the activation function encoding module;
[0100] This module designs an activation function encoding method to encode the activation function types of neurons in each layer as discrete values. It encodes various activation functions used in the network (such as ReLU, Sigmoid, Tanh, etc.) as neural gene vectors g. activation This allows combinations of different activation functions to be searched using evolutionary algorithms.
[0101] Optionally, in some implementations, the activation function encoding includes not only the function type but also the function's parameters, such as the negative slope parameter of LeakyReLU or the learnable parameter of ParametricReLU, further expanding the search space.
[0102] Step 2.3: Construct the initial weight distribution encoding module;
[0103] According to another embodiment of this application, an encoding method for a weight initialization strategy is constructed, which encodes the initial distribution type and parameters of the network weights into a numerical vector. This module encodes the weight initialization method (such as Xavier, He initialization, etc.) and its parameters into a neural gene vector g. weight Optimize the initial state of the network.
[0104] In addition, this application also provides an adaptive initialization method that automatically selects the most suitable weight initialization strategy based on the network structure and task characteristics, thereby further improving the efficiency and stability of network training.
[0105] Step 2.4: Implement the decoder for neural genes;
[0106] A neural gene decoder is constructed to reconstruct the complete neural network structure from neural gene representations. This decoder receives the neural gene vector g = [g...]. topology g activation g weight As input, the output is a complete neural network structure that can be directly used for reinforcement learning, where g represents the overall neural gene vector, containing all encoded information describing the neural network structure and parameters; g topology Network topology encoding represents structural information of a neural network, such as connection patterns, hierarchical structure, and neuron types; g activation The activation function encoding represents the type of activation function (such as ReLU, Sigmoid, Tanh, etc.) and its parameters used by neurons in each layer; g weight The weight initialization code represents the initial distribution type of the weights in each layer of the network (such as Xavier, He initialization, etc.) and its parameters.
[0107] The decoder's role is to parse out the information from each part of the neural gene vector g, and then sequentially complete the network topology reconstruction, activation function allocation, and weight initialization to generate a complete neural network structure that can be directly used for reinforcement learning tasks.
[0108] It should be noted that the decoding process includes three stages: network topology reconstruction, activation function allocation, and weight initialization, to ensure accurate mapping from neural genes to the network while maintaining decoding efficiency.
[0109] like Figure 3 As shown, this illustrates the amount of data and time required for different methods to achieve the target control performance after deployment in a new environment.
[0110] A hierarchical neural architecture search space is used to provide a structured search space for evolutionary algorithms;
[0111] Specifically, the following steps are included:
[0112] Step 3.1, define the macroscopic structure search space;
[0113] This module constructs a macroscopic structural search space for torque control systems, including various basic architecture types such as feedforward networks, cyclic networks, and attention networks. It defines the overall framework of the control network and outputs a series of macroscopic structural templates suitable for torque control.
[0114] Optionally, in some implementations, the macrostructure search space may also include hybrid architectures, such as hybrid networks combining feedforward and recurrent structures, or reinforcement learning architectures incorporating attention mechanisms, to adapt to control tasks of varying complexity.
[0115] Step 3.2, define the search space for the mesoscopic module;
[0116] According to embodiments of this application, a meso-level module search space is constructed, including various functional modules such as state encoders, policy networks, and value function networks. This module defines the functional components that constitute the control system and outputs a set of composable meso-level functional modules.
[0117] It should be noted that the meso-level modules in this application are based on functional partitioning rather than simple hierarchical partitioning. This enables the evolutionary algorithm to search at a higher level of abstraction, thereby improving search efficiency and result quality.
[0118] Step 3.3, define the search space for micro-operations;
[0119] This module constructs a search space for micro-operations, including various basic operations such as convolution, fully connected units, and recurrent units. It defines the basic computational units that constitute the functional modules and outputs a set of micro-operations that can be used to build the network.
[0120] Optionally, in some implementations, micro-operations may also include special computational units, such as gating mechanisms, residual connections, and normalization layers, to enhance the network's expressive power and training stability.
[0121] Step 3.4: Implement the constraint rules for the hierarchical search space;
[0122] According to the method provided in this application, a constraint rule system for the search space is constructed to ensure that the generated network structure meets the basic requirements of the torque control system. This rule system filters out unreasonable combinations of network structures and outputs an effective search space that conforms to both physical and computational constraints.
[0123] In addition, this application provides an adaptive constraint mechanism that dynamically adjusts constraint rules based on discoveries during the evolution process, ensuring network effectiveness without excessively restricting the emergence of innovative structures.
[0124] like Figure 4 As shown, the advantages of the method of the present invention in terms of computing resource requirements, energy consumption, and maintenance costs are demonstrated from multiple dimensions.
[0125] A neural modular evolutionary algorithm is used to achieve network structure evolution by evaluating the value of modules through functional similarity.
[0126] Specifically, the following steps are included:
[0127] Step 4.1, construct the functional similarity assessment module;
[0128] Based on the technical solution provided in this application, a functional similarity measurement method based on behavioral features is constructed to evaluate the functional equivalence of different network structures. This module calculates a functional similarity matrix by comparing the output behavior of the networks under the same input, and identifies functional modules that are similar in function but different in structure.
[0129] In some implementations, functional similarity assessment can combine multiple metrics, such as KL divergence of output distribution and similarity of behavioral trajectories, to comprehensively evaluate the degree of similarity of network functions.
[0130] Step 4.2: Implement the functional module protection mechanism;
[0131] A functional module protection mechanism is constructed to prevent functional modules with unique functions from disappearing prematurely during evolution. This mechanism ensures the breadth of evolutionary search by identifying and protecting functional diversity in the population, resulting in a more diverse network structure population.
[0132] It should be noted that the functional module protection mechanism of this application is not only based on functional diversity, but also takes into account the potential value of functional modules. Even if some functional modules are not performing well at present, functional modules with unique functions will be retained so that they can play a role in subsequent evolution.
[0133] Step 4.3: Construct a novelty-based selection strategy;
[0134] According to embodiments of this application, a novelty-based selection strategy is constructed to encourage exploration of unknown network structure spaces. This strategy combines fitness and novelty for multi-objective optimization, avoiding the evolutionary process from getting trapped in local optima and outputting innovative network structures.
[0135] In some implementations, novelty assessment can be based on the behavior space rather than the parameter space, that is, assessing the uniqueness of the behavioral trajectories generated by the network in the control task, rather than simply comparing differences in network structure.
[0136] Step 4.4: Implement the crossover and mutation operations of functional modules;
[0137] Construct crossover and mutation operations suitable for the functional modules of a neural network, including substructure exchange, connection mutation, and functional replacement. These operations can maintain the functional integrity of the network while generating structural changes, outputting a descendant network that retains its functionality but has an optimized structure.
[0138] In addition, this application also provides a smart crossover operation based on functional correlation, which can identify functionally related functional modules in the parent network and maintain the integrity of these functional modules during crossover, thereby improving the performance of the offspring network.
[0139] like Figure 5 As shown, the system failure rate and recovery time are compared under different interference conditions.
[0140] A hypernetwork meta-learning framework for generating task-specific controller parameters;
[0141] Specifically, the following steps are included:
[0142] Step 5.1, design the task representation learning module;
[0143] According to embodiments of this application, a task representation learning module is constructed to encode the features of the gate control environment into compact task embedding vectors. This module analyzes the dynamic characteristics, load patterns, and feedback features of the environment to generate task embedding vectors. As input to the hypernetwork.
[0144] Optionally, in some implementations, task representation learning can employ a contrastive learning approach, which learns more effective task embeddings by distinguishing features of different tasks, thereby improving the parameter generation accuracy of the supernetwork.
[0145] Step 5.2, construct the hypernetwork structure;
[0146] According to the technical solution provided in this application, a hypernetwork structure suitable for parameter generation is constructed, which can generate all parameters of the target network based on the task embedding vector.
[0147] This supernetwork employs a conditional generation architecture, outputting weight parameters that match the target network structure:
[0148]
[0149] in, The task embedding vector is a compact vector obtained by encoding the environmental features of the current gate control task (such as dynamic characteristics, load patterns, feedback features, etc.), and serves as the input to the hypernetwork; h φ The hypernetwork ontology is a neural network model capable of generating parameters, whose input is the task embedding vector. The output contains all parameters of the target network; φ represents the internal parameters of the supernetwork h, which are obtained through meta-learning and other methods and are used to adjust the mapping ability of the supernetwork. Indicates a specific task The generated target network parameter set includes the weight matrices and bias vectors of each layer.
[0150] In practice, the hypernetic network consists of three parts: an encoder, a processor, and a decoder.
[0151] Encoder: Embeds task vectors Mapping to a high-dimensional feature space allows for the extraction of deep-level features for the task.
[0152] Processor: Includes a multi-layer self-attention module to capture the complex relationships between task features and improve the targeting and generalization ability of parameter generation.
[0153] Decoder: Based on the structural information of the target network, it generates the corresponding weight matrix and bias vector layer by layer to ensure that the output parameters are completely matched with the target network structure.
[0154] In emergency evacuation passage gate applications, this super network can adjust control parameters in real time according to changes in personnel density and flow direction, ensuring optimal passage efficiency and safety in emergency situations.
[0155] It should be understood that the hypernetwork structure of this application is different from conventional parameter generation networks. It can dynamically adjust its output layer according to the network structure information to adapt to target networks of different sizes and structures.
[0156] Step 5.3: Implement the hypernetic network training algorithm;
[0157] An efficient training algorithm for hypernetworks is constructed, enabling them to quickly and accurately generate parameters adapted to new tasks. Based on the principle of meta-learning, this algorithm trains the hypernetwork on multiple tasks and optimizes the parameters φ, allowing the generated controller parameters to rapidly achieve high performance on new tasks.
[0158] Optionally, in some implementations, the hypernetwork training can employ a hierarchical training strategy, training the encoder and processor parts first, and then training the decoder part, to improve training efficiency and the quality of generated parameters.
[0159] Step 5.4: Construct a parameter fine-tuning mechanism;
[0160] According to another embodiment of this application, a gradient descent-based parameter fine-tuning mechanism is constructed to rapidly and adaptively adjust the initial parameters generated by the hypernetwork. This mechanism utilizes a small amount of environmental interaction data to further optimize the controller parameters and output the final task-specialized controller.
[0161] In addition, this application also provides a gradient-free parameter fine-tuning method, which is suitable for environments where it is difficult to obtain gradient information, and achieves rapid parameter adjustment through evolutionary strategies or Bayesian optimization.
[0162] like Figure 6 As shown, this demonstrates the system's adaptation process and final control accuracy in different application scenarios.
[0163] Evolutionary meta-learning collaborative optimization mechanism is used to enable two algorithms to promote and complement each other;
[0164] Specifically, the following steps are included:
[0165] Step 6.1: Construct a joint evaluation system for structural parameters;
[0166] This system designs a joint evaluation system for network structure and initial parameters, comprehensively considering the potential of the structure and the adaptability of the parameters. It evaluates the overall performance of the network by testing its performance on multiple tasks and calculating a comprehensive fitness score.
[0167] Optionally, in some implementations, the joint evaluation system may employ a multi-stage evaluation strategy, first conducting a rapid, rough evaluation to screen out potentially excellent structures, and then conducting a detailed evaluation of these structures to improve evaluation efficiency.
[0168] Step 6.2: Implement a two-way feedback mechanism;
[0169] According to embodiments of this application, a bidirectional feedback mechanism is constructed between the evolutionary layer and the meta-learning layer, enabling the two algorithms to share information and guide each other. This mechanism includes the transfer of structural information from evolution to meta-learning and the feedback of parameter performance from meta-learning to evolution, promoting the synergistic progress of the two algorithms.
[0170] It should be noted that the bidirectional feedback mechanism of this application is a closed-loop system that can adjust the search direction and optimization target of the two algorithms in real time, thereby achieving more efficient collaborative optimization.
[0171] Step 6.3: Design an adaptive scheduler that balances exploration and utilization;
[0172] An adaptive scheduler is constructed to dynamically balance the exploratory nature of evolution and the utilization of meta-learning based on the optimization progress. This scheduler monitors the optimization process, adjusting the resource allocation and influence weights of the two algorithms to achieve an optimal balance between exploring new structures and leveraging known experience.
[0173] Optionally, in some implementations, the adaptive scheduler can be based on reinforcement learning methods, taking resource allocation strategies as actions and performance improvement as rewards, to learn the optimal scheduling strategy.
[0174] Step 6.4: Implement an incremental knowledge accumulation mechanism;
[0175] Based on the technical solution provided in this application, an incremental knowledge base is constructed to store valuable modules and parameter patterns discovered during the evolution and learning process. This knowledge base is continuously updated and expanded, providing high-quality prior knowledge for subsequent network generation and parameter initialization, thus accelerating the optimization process.
[0176] In addition, this application also provides a knowledge distillation mechanism to compress and refine information in a knowledge base, thereby improving knowledge utilization efficiency and reducing storage and computational overhead.
[0177] like Figure 7 As shown, the performance trend of the system during long-term operation is demonstrated, verifying its continuous learning capability.
[0178] Deployment and online adaptation systems are used to achieve continuous optimization of the system during actual operation.
[0179] Specifically, the following steps are included:
[0180] Step 7.1, Implementation of the edge cloud collaborative deployment architecture;
[0181] This embodiment provides a layered distributed deployment architecture, including an edge control unit, a local optimization unit, and a cloud collaboration unit.
[0182] The edge control unit is deployed on the gate controller to realize real-time torque control and environmental perception;
[0183] The local optimization unit is deployed on the field server and is responsible for quickly fine-tuning control parameters and handling abnormal events locally.
[0184] The cloud-based collaborative unit is deployed on a remote server and undertakes computationally intensive tasks such as evolutionary optimization, meta-learning model training, and global knowledge management.
[0185] The three-layer unit mentioned above achieves efficient collaboration through a secure data communication protocol, ensuring that the system has both real-time response capabilities and can make full use of cloud computing power for global optimization.
[0186] Optionally, the deployment architecture can incorporate fog computing nodes to undertake data preprocessing and lightweight model optimization tasks, further improving the system's response speed and resource utilization.
[0187] Step 7.2, Construction of the online rapid adaptation mechanism;
[0188] This embodiment provides an online rapid adaptation mechanism based on meta-learning. This mechanism utilizes a pre-trained hypernetwork model to dynamically adjust controller parameters based on real-time collected environmental data.
[0189] The specific process is as follows: the edge control unit collects the current environmental status and control feedback, inputs a small amount of interactive data into the local optimization unit or the cloud meta-learning model, quickly generates adaptive parameter updates, and realizes real-time optimization of the control strategy.
[0190] This mechanism supports incremental learning, which can gradually integrate new environmental characteristics without losing historical knowledge, continuously improve controller performance, and adapt to changing real-world application scenarios.
[0191] Step 7.3, Implementation of the anomaly detection and recovery system;
[0192] This embodiment provides an integrated anomaly detection and recovery system. The system continuously monitors the controller's operating status and performance indicators, and identifies potential anomalies in real time by analyzing multi-dimensional data such as control errors and system state changes.
[0193] Upon detecting an anomaly, the system automatically triggers parameter readjustment, structural fine-tuning, or switches to a safe mode to ensure stable system operation.
[0194] Optionally, the anomaly detection module can employ predictive analytics algorithms to model future performance trends, enabling early warning and proactive intervention for anomalies, thereby further reducing the risk of system failure.
[0195] Step 7.4, Implementation of continuous evolution and learning loop;
[0196] This embodiment provides a closed-loop optimization system for continuous evolution and learning. During actual operation, the system continuously collects operational data and transmits it back to the cloud. The cloud then uses the latest data to incrementally update the evolutionary population and the meta-learning model.
[0197] The updated model and parameters are distributed to each edge control unit through a secure channel, enabling the controller to learn and evolve automatically throughout its life.
[0198] This closed-loop mechanism not only supports continuous optimization of single-point controllers, but also supports knowledge sharing and migration in multi-point deployment scenarios, forming collective intelligence and significantly improving the adaptability and optimization efficiency of the overall system.
[0199] The embodiments of the present invention have been described above. However, the embodiments are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make more equivalent embodiments under the guidance of the present embodiments, and all of them are within the protection scope of the present embodiments.
Claims
1. A dynamic torque control system for heavy-duty turnstiles based on reinforcement learning, characterized in that, include: A two-layer optimization framework combining neural evolution and meta-learning is used to simultaneously optimize the neural network structure and parameters. In the structural optimization layer, this framework uses torque accuracy, response speed, and energy consumption as evaluation metrics to calculate the overall fitness. A neural gene encoding and decoding system is used to encode the controller network structure into evolvable neural genes. The network topology encoding module in this system employs a combination of direct and indirect encoding. For core network layers in heavy-duty gate torque control, such as the state encoding layer and policy output layer, direct encoding is used to describe their structure in detail. For intermediate processing layers, indirect encoding is used to describe the generation rules. A hierarchical neural architecture search space is used to provide a structured search space for the evolutionary algorithm. A modular neural evolutionary algorithm is used to achieve network structure evolution by evaluating the value of modules based on functional similarity. A hypernetwork meta-learning framework is used to generate specialized controller parameters for specific tasks. The hypernetwork within this framework consists of an encoder, a processor, and a decoder. The encoder maps the task embedding vector to a high-dimensional feature space, extracting deep-level features of the task. The processor contains multi-layer self-attention modules to capture complex relationships between task features, improving the specificity and generalization ability of parameter generation. The decoder generates corresponding weight matrices and bias vectors layer by layer based on the structural information of the target network, ensuring that the output parameters perfectly match the target network structure. An evolutionary meta-learning collaborative optimization mechanism is used to enable two algorithms to mutually promote and complement each other. This mechanism includes: a joint structural parameter evaluation system that comprehensively considers the potential of the structure and the adaptability of the parameters; a bidirectional feedback mechanism that allows the evolutionary layer and the meta-learning layer to share information and guide each other; an adaptive scheduler that balances exploration and utilization, dynamically balancing the exploratory nature of evolution and the utilization of meta-learning based on the optimization progress; an incremental knowledge accumulation mechanism that stores valuable modules and parameter patterns discovered during evolution and learning; and a deployment and online adaptation system for continuous optimization of the system during actual operation.
2. The heavy-duty gate dynamic torque control system based on reinforcement learning according to claim 1, characterized in that, The proposed neuroevolutionary and meta-learning dual-layer optimization framework includes: a structure optimization layer, which uses a neuroevolutionary algorithm to search for the optimal network topology; a parameter optimization layer, which uses a meta-learning algorithm to achieve rapid adaptation; and a collaborative mechanism for dual-layer optimization, including a downlink information channel from structure to parameters and an uplink feedback channel from parameters to structure.
3. The heavy-duty gate dynamic torque control system based on reinforcement learning according to claim 1, characterized in that, The neural gene encoding and decoding system includes: a network topology encoding module, which encodes the network connection patterns, hierarchical structure, and neuron types into compact numerical vectors; an activation function encoding module, which encodes the activation function types of neurons in each layer into discrete values; an initial weight distribution encoding module, which encodes the initial distribution type and parameters of the network weights into numerical vectors; and a neural gene decoder, which can reconstruct the complete neural network structure from the neural gene representation.
4. The heavy-duty gate dynamic torque control system based on reinforcement learning according to claim 1, characterized in that, The hierarchical neural architecture search space includes: a macroscopic structure search space, consisting of a network architecture composed of feedforward networks, recurrent networks, and attention networks; a mesoscopic module search space, including functional module combinations of state encoders, policy networks, and value function networks; a microscopic operation search space, consisting of a set of neural network operations composed of convolutional, fully connected, and recurrent units; and a constraint rule system for the hierarchical search space to ensure that the generated network structure meets the performance requirements of the torque control system.
5. The heavy-duty gate dynamic torque control system based on reinforcement learning according to claim 1, characterized in that, The neural modular evolutionary algorithm includes: a functional similarity evaluation module, which calculates a functional similarity matrix by comparing the output behavior of the network under the same input; a functional module protection mechanism to prevent functional modules with unique functions from disappearing prematurely during the evolution process; a novelty-based selection strategy to encourage the exploration of unknown network structure space; and functional module crossover and mutation operations, including substructure exchange, connection mutation, and function replacement.
6. The heavy-duty gate dynamic torque control system based on reinforcement learning according to claim 1, characterized in that, The hypernetwork meta-learning framework includes: a task representation learning module that encodes the features of the gate control environment into compact task embedding vectors; a hypernetwork structure that can generate all parameters of the target network based on the task embedding vectors; a hypernetwork training algorithm that enables the hypernetwork to quickly and accurately generate parameters adapted to new tasks; and a parameter fine-tuning mechanism that rapidly and adaptively adjusts the initial parameters generated by the hypernetwork.
7. The heavy-duty gate dynamic torque control system based on reinforcement learning according to claim 1, characterized in that, The deployment and online adaptation system includes: an edge-cloud collaborative deployment architecture that places computationally intensive evolutionary optimization in the cloud and real-time control and rapid adaptation on edge devices; an online rapid adaptation mechanism that enables the controller to quickly adjust parameters based on real-time environmental feedback; an anomaly detection and recovery system that monitors controller performance and triggers adaptive adjustments when performance degrades; and a continuous evolution and learning loop that continuously optimizes the controller's structure and parameters using actual operating data.
8. The heavy-duty gate dynamic torque control system based on reinforcement learning according to claim 1, characterized in that, It can be applied to at least one of the following scenarios: temporary security gates in mobile security inspection equipment; border inspection gates at temporary border checkpoints; entrance gate systems for large event venues; evacuation gates for emergency evacuation routes; and industrial gate systems for smart factories.
9. A computer-readable storage medium, characterized in that, It is used to store computer-readable instructions, which, when read by a computer, enable the execution of a heavy-duty gate dynamic torque control system based on reinforcement learning as described in any one of claims 1-8.