Digital valve group impedance matching control method and device based on deep reinforcement learning

CN122544071APending Publication Date: 2026-08-11HUADIAN ZHENGZHOU MECHANICAL DESIGN INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-07
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0004]本发明的主要目的在于提供一种基于深度强化学习的数字阀组阻抗匹配控制方法及装置,旨在解决现有液压变压器控制中节流损失大、多阀协同逻辑复杂且难以适应时变负载的技术问题

Benefits of technology

[0017] This invention constructs a digital hydraulic transformer hardware system comprising parallel high-speed switching valves; models the valve group control problem as a Markov decision process, defining a reward function with energy efficiency, pressure regulation accuracy, and smoothness of action as optimization objectives; offline trains a policy network using a deep reinforcement learning algorithm incorporating maximum entropy; collects system status online in real time, generates action vectors through the policy network, and obtains the switching combination and duty cycle after discrete mapping and physical constraints, driving the valve group to adjust the equivalent flow area to achieve dynamic matching of hydraulic impedance; after executing the action, collects status feedback, calculates the reward, and updates the policy network online to adapt to time-varying loads. This method achieves high efficiency and energy saving through a throttling-free transformer, possessing advantages such as strong adaptability and low pressure pulsation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122544071A_ABST
    Figure CN122544071A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of hydraulic control technology and discloses a digital valve group impedance matching control method and device based on deep reinforcement learning. The method includes: constructing a digital hydraulic transformer hardware system containing parallel high-speed switching valves; modeling the valve group control problem as a Markov decision process, defining a reward function with energy efficiency, pressure regulation accuracy, and action smoothness as optimization objectives; offline training of a policy network using a deep reinforcement learning algorithm incorporating maximum entropy; online real-time acquisition of system status, generation of action vectors through the policy network, obtaining the switching combination and duty cycle after discrete mapping and physical constraints, driving the valve group to adjust the equivalent flow area to achieve dynamic hydraulic impedance matching; acquiring status feedback after action execution, calculating rewards, and updating the policy network online to adapt to time-varying loads. This method achieves high efficiency and energy saving through a throttling-free transformer approach, and has the advantages of strong adaptability and low pressure pulsation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of hydraulic control technology, and in particular to a digital valve group impedance matching control method and device based on deep reinforcement learning. Background Technology

[0002] Hydraulic accumulators, used as auxiliary power sources, experience a linear decrease in output pressure as the stored oil volume decreases. However, many industrial loads (such as crane lifting and injection molding machine pressure holding) require constant drive pressure. Traditional control methods involve connecting a proportional pressure reducing valve in series in the circuit. This valve throttling converts excess pressure energy into heat dissipation. This "resistive voltage division" control results in extremely low system efficiency (typically below 50%) and generates significant heat, shortening hydraulic oil lifespan. Digital hydraulic technology achieves stepless pressure variation through high-speed valve switching, avoiding throttling losses. However, due to the large number of valves, strong nonlinearity, and complex coupling, traditional PID or fuzzy control struggles to handle multi-valve collaborative decisions within milliseconds, leading to large pressure pulsations and unstable control.

[0003] The above content is only used to help understand the technical solution of the present invention and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this invention is to provide a digital valve group impedance matching control method and device based on deep reinforcement learning, which aims to solve the technical problems of large throttling losses, complex multi-valve collaborative logic, and difficulty in adapting to time-varying loads in existing hydraulic transformer control.

[0005] To achieve the above objectives, the present invention provides a digital valve group impedance matching control method based on deep reinforcement learning, the method comprising the following steps: A digital hydraulic transformer hardware system is constructed, which includes several high-speed switching valves of different diameters connected in parallel between the high-voltage accumulator and the actuator. The control problem of the high-speed switching valve is modeled as a Markov decision process. The Markov decision process defines a state space, an action space, and a reward function. The reward function is defined as a weighted sum of instantaneous energy transfer efficiency, load-side pressure tracking error, and valve group switching frequency penalty term. The policy network is trained using a deep reinforcement learning algorithm. During the training process, a maximum entropy term is introduced, and the reward value is calculated using the reward function to update the parameters of the policy network. The state of the digital hydraulic transformer hardware system is collected in real time. The state is input into the trained strategy network to generate a target action vector. The target action vector is discretized and physical constraints are applied to obtain the switching combination and duty cycle of the several high-speed switching valves. The switching combination and duty cycle are used to drive the high-speed switching valves to perform actions to adjust the equivalent flow area and achieve dynamic matching of hydraulic impedance. After the high-speed switching valve is driven to perform an action, the status feedback after the action is performed is collected, the actual reward value is calculated based on this feedback, the interaction data is stored in the experience replay pool, and the parameters of the strategy network are updated online to adapt to time-varying load characteristics.

[0006] In one embodiment, the reward function R defined in step S2 is specifically:

[0007] in, Indicates time step Instant reward value, This represents the instantaneous energy transfer efficiency, which is the ratio of load power to source power. This indicates the pressure tracking error at the load end. The switching frequency penalty term, representing the valve group's on / off state, is used to suppress high-frequency chatter. These are the weighting coefficients for each optimization objective.

[0008] In one embodiment, the method further includes: Define the state space S={P_acc,P_load,Q_req,SOC,v_valve}, where P_acc is the accumulator pressure, P_load is the load pressure, Q_req is the load demand flow, SOC is the accumulator state of charge, and v_valve is the current valve group state vector. Define the action space A={d_1,d_2,...,d_N}, where d_i∈[0,1] represents the duty cycle of the i-th switching valve; Construct a reward function based on the defined state space and action space.

[0009] In one embodiment, the step of training the policy network using a deep reinforcement learning algorithm, introducing a maximum entropy term during training, and calculating the reward value using the reward function to update the parameters of the policy network includes: Construct an actor network, a critic network, and a target critic network; A maximum entropy term is introduced into the objective function. This maximum entropy term is used by the agent to explore diverse control strategies during training to prevent it from getting trapped in local optima. The parameters of the critic network are updated by minimizing the Bellman residual loss function, wherein the target value in the Bellman residual loss function is calculated by the target critic network; The parameters of the actor network are updated by minimizing the policy loss function.

[0010] In one embodiment, the step of discretizing and mapping the target motion vector and applying physical constraints to obtain the switching combinations and duty cycles of the plurality of high-speed switching valves includes: The target action vector output by the strategy network is mapped to discrete switch combinations, and the duty cycle of each switch valve is determined based on the value of the target action vector. Dead zone compensation and minimum opening time constraints are applied to the obtained switch combinations and duty cycles to filter out excessively short pulse commands that the physical valve core cannot respond to. The switch combinations and duty cycles that do not meet the minimum opening time constraint are adjusted to obtain the switch combinations and duty cycles used to drive the plurality of high-speed switching valves to perform actions.

[0011] In one embodiment, the method further includes: Within each control cycle, the target motion vector is converted into an equivalent flow area, calculated as follows:

[0012] in, Let be the flow coefficient of the i-th high-speed switching valve. Let be the maximum flow area of ​​the i-th high-speed switching valve. The i-th element in the target action vector a* represents the duty cycle of the i-th high-speed switching valve; By dynamically adjusting the equivalent flow area, the system can be configured with variable liquid resistance, thereby maintaining a constant load pressure as the accumulator pressure decreases during the discharge process.

[0013] In one embodiment, the method further includes: Calculate the actual energy transfer efficiency at the current moment; The load pressure fluctuation rate is detected. If the load pressure fluctuation rate exceeds a preset threshold, the weight of the pressure tracking error term in the reward function is increased. The current state, action, reward, and next state are combined into a quadruple and stored in the experience replay pool. Random samples are then taken from the experience replay pool for offline policy iteration to update the parameters of the policy network. The reward is determined by the actual energy transmission efficiency and the reward function after weight adjustment.

[0014] Furthermore, to achieve the above objectives, this invention also proposes a digital valve group impedance matching control device based on deep reinforcement learning. This device is applied to the deep reinforcement learning-based digital valve group impedance matching control method described above. The device includes: A hardware building module is used to build a digital hydraulic transformer hardware system, which includes several high-speed switching valves of different diameters connected in parallel between the high-voltage accumulator and the actuator. The modeling module is used to model the control problem of the high-speed switching valve as a Markov decision process. The Markov decision process defines a state space, an action space, and a reward function. The reward function is defined as a weighted sum of instantaneous energy transfer efficiency, load-side pressure tracking error, and valve group switching frequency penalty term. The offline training module is used to train the policy network using a deep reinforcement learning algorithm. During the training process, a maximum entropy term is introduced, and the reward value is calculated using the reward function to update the parameters of the policy network. The reasoning and execution module is used to collect the state of the digital hydraulic transformer hardware system in real time, input the state into the trained strategy network to generate a target action vector, perform discrete mapping on the target action vector and apply physical constraints to obtain the switching combination and duty cycle of the several high-speed switching valves, and drive the high-speed switching valves to perform actions with the switching combination and duty cycle to adjust the equivalent flow area and realize dynamic matching of hydraulic impedance. The online update module is used to collect the status feedback after the high-speed switching valve performs the action, calculate the actual reward value based on the feedback, store the interaction data in the experience playback pool, and update the parameters of the strategy network online to adapt to time-varying load characteristics.

[0015] Furthermore, to achieve the above objectives, the present invention also proposes a digital valve group impedance matching control device based on deep reinforcement learning. The digital valve group impedance matching control device based on deep reinforcement learning includes: a memory, a processor, and a digital valve group impedance matching control program based on deep reinforcement learning stored in the memory and executable on the processor. The digital valve group impedance matching control program based on deep reinforcement learning is configured to implement the steps of the digital valve group impedance matching control method based on deep reinforcement learning as described above.

[0016] Furthermore, to achieve the above objectives, the present invention also proposes a storage medium storing a deep reinforcement learning-based digital valve group impedance matching control program, wherein when the deep reinforcement learning-based digital valve group impedance matching control program is executed by a processor, it implements the steps of the deep reinforcement learning-based digital valve group impedance matching control method described above.

[0017] This invention constructs a digital hydraulic transformer hardware system comprising parallel high-speed switching valves; models the valve group control problem as a Markov decision process, defining a reward function with energy efficiency, pressure regulation accuracy, and smoothness of action as optimization objectives; offline trains a policy network using a deep reinforcement learning algorithm incorporating maximum entropy; collects system status online in real time, generates action vectors through the policy network, and obtains the switching combination and duty cycle after discrete mapping and physical constraints, driving the valve group to adjust the equivalent flow area to achieve dynamic matching of hydraulic impedance; after executing the action, collects status feedback, calculates the reward, and updates the policy network online to adapt to time-varying loads. This method achieves high efficiency and energy saving through a throttling-free transformer, possessing advantages such as strong adaptability and low pressure pulsation. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the first embodiment of the digital valve group impedance matching control method based on deep reinforcement learning of the present invention. Figure 2 This is a topology diagram of the digital flow distribution valve group in the digital valve group impedance matching control method based on deep reinforcement learning of this invention; Figure 3 This is a schematic diagram of the deep reinforcement learning control architecture in the digital valve group impedance matching control method based on deep reinforcement learning of the present invention; Figure 4 This is a structural block diagram of the first embodiment of the digital valve group impedance matching control device based on deep reinforcement learning of the present invention.

[0019] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0021] This invention provides a digital valve group impedance matching control method based on deep reinforcement learning, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of a digital valve group impedance matching control method based on deep reinforcement learning according to the present invention.

[0022] In this embodiment, the digital valve group impedance matching control method based on deep reinforcement learning includes the following steps: Step S10: Construct the digital hydraulic transformer hardware system.

[0023] In this embodiment, the executing entity is a digital valve group impedance matching control device based on deep reinforcement learning. This device has functions such as data processing, data communication, and program execution. The device can be a computer terminal or other network device, or other devices with similar functions. This embodiment does not limit the scope of the application.

[0024] It should be noted that, as an auxiliary power source, the output pressure of a hydraulic accumulator decreases linearly with the reduction of the stored oil volume. However, many industrial loads (such as crane lifting and injection molding machine pressure holding) require constant drive pressure. Traditional control methods involve connecting a proportional pressure reducing valve in series in the circuit. This valve throttling converts excess pressure energy into heat energy for dissipation. This "resistive pressure division" control results in extremely low system efficiency (typically below 50%) and generates a large amount of heat, shortening the lifespan of the hydraulic oil. Digital hydraulic technology achieves stepless pressure variation through high-speed valve switching, avoiding throttling losses. However, due to the large number of valves, strong nonlinearity, and complex coupling, traditional PID or fuzzy control struggles to handle multi-valve collaborative decisions within milliseconds, leading to large pressure pulsations and unstable control.

[0025] To address the aforementioned technical issues, this embodiment constructs a digital hydraulic transformer hardware system comprising parallel high-speed switching valves. The valve group control problem is modeled as a Markov decision process, defining a reward function with energy efficiency, pressure regulation accuracy, and smoothness of action as optimization objectives. A deep reinforcement learning algorithm incorporating maximum entropy is used to train the policy network offline. The system status is collected online in real-time, and action vectors are generated through the policy network. After discrete mapping and physical constraints, the switching combinations and duty cycles are obtained, driving the valve group to adjust the equivalent flow area to achieve dynamic matching of hydraulic impedance. After executing the action, status feedback is collected, rewards are calculated, and the policy network is updated online to adapt to time-varying loads. This method achieves high efficiency and energy saving through a throttling-free transformer approach, possessing advantages such as strong adaptability and low pressure pulsation. Specifically, it can be implemented as follows.

[0026] In a specific implementation, the digital hydraulic transformer hardware system in this embodiment includes several high-speed switching valves of different diameters connected in parallel between the high-voltage accumulator and the actuator. For example, the digital hydraulic transformer hardware system consists of N high-speed switching valves of different diameters connected in parallel. Preferably, these diameters are in a binary ratio relationship (e.g., 1:2:4:8), thereby achieving 2 through combined switching states. N An equivalent flow area is provided to meet high-resolution regulation requirements. A high-speed switching valve connects the high-voltage accumulator and the actuator, replacing the traditional proportional pressure reducing valve. The specific topology of the digital distribution valve group in this embodiment can be found in [reference needed]. Figure 2 As shown.

[0027] Step S20: Model the control problem of the high-speed switching valve as a Markov decision process, wherein the Markov decision process defines the state space, action space and reward function.

[0028] It should be noted that the reward function is defined as a weighted sum of instantaneous energy transfer efficiency, load-side pressure tracking error, and valve group switching frequency penalty terms. The state space is defined as S={P acc ,P load Q req ,SOC,v valve}, where P acc For the accumulator pressure, P load For load pressure, Q req The load demand flow rate is given by V, and SOC is the state of charge of the energy storage device. valve This represents the current valve group state vector. The action space is defined as A = {d1, d2, ..., d...} N}, where d i ∈[0,1] represents the duty cycle of the i-th switching valve. The reward function R is as follows:

[0029] in, Indicates time step Instant reward value, This represents the instantaneous energy transfer efficiency, which is the ratio of load power to source power. This indicates the pressure tracking error at the load end. The switching frequency penalty term, representing the valve group's on / off state, is used to suppress high-frequency chatter. These are the weighting coefficients for each optimization objective.

[0030] Step S30: Train the policy network using a deep reinforcement learning algorithm, introduce a maximum entropy term during training, and calculate the reward value using the reward function to update the parameters of the policy network.

[0031] In this implementation, the Soft Actor-Critic (SAC) algorithm is used. Specifically, an actor network (Actor) π is constructed. Critic Network (Critic) Q θ And the target commentator network Q θˉ A maximum entropy term H(π(·∣st)) is introduced into the objective function to encourage the exploration of diverse control strategies and prevent the strategies from converging prematurely to local optima.

[0032] The update of the critic network is achieved by minimizing the Bellman residual loss function:

[0033] in, .

[0034] The actor network is updated by minimizing the policy loss function:

[0035] Utilizing the reparameterization trick This allows the gradient to propagate backward.

[0036] Training can employ a course-based learning strategy: initially training voltage regulation capability under a fixed accumulator pressure, introducing linear pressure decrease disturbances in the middle stage, and introducing random loads and gusts of wind disturbances in the later stage to improve the generalization performance of the policy network. Training data is generated through interaction with the environment, stored in an experience replay pool, and batch sampling is used for offline policy iteration. The deep reinforcement learning control architecture in this embodiment can refer to... Figure 3 As shown.

[0037] Step S40: Real-time acquisition of the state of the digital hydraulic transformer hardware system, inputting the state into the trained strategy network to generate a target action vector, discretizing the target action vector and applying physical constraints to obtain the switching combination and duty cycle of the several high-speed switching valves, and using the switching combination and duty cycle to drive the high-speed switching valves to perform actions to adjust the equivalent flow area and achieve dynamic matching of hydraulic impedance.

[0038] It should be noted that the target action vector a* output by the policy network is a continuous value and needs to be mapped to discrete switching combinations and duty cycles. Specifically, the continuous action vector is mapped to discrete switching states using a threshold or nearest neighbor method, and the duty cycle of each switching valve is directly determined based on the value of the action vector. For example, for the i-th valve, if a i If the value is less than 0.1, the function is turned off; otherwise, it is turned on and the duty cycle is set to a. i Next, dead-time compensation (e.g., duty cycle less than 5% is considered 0, greater than 95% is considered 1) and minimum opening time constraints (e.g., minimum pulse width 2ms) are applied. Switch combinations and duty cycles that do not meet the constraints are adjusted to obtain the final switch combinations and duty cycles used to drive the valve group.

[0039] In one embodiment, within each control cycle, the target motion vector is converted into an equivalent flow area, calculated as follows:

[0040] in, Let be the flow coefficient of the i-th high-speed switching valve. Let be the maximum flow area of ​​the i-th high-speed switching valve. The i-th element in the target action vector a* represents the duty cycle of the i-th high-speed switching valve; by dynamically adjusting the equivalent flow area, the system is configured with variable liquid resistance, thereby maintaining a constant load pressure when the accumulator pressure decreases during the discharge process.

[0041] Step S50: After the high-speed switching valve is driven to perform the action, the status feedback after the action is performed is collected, the actual reward value is calculated accordingly, the interaction data is stored in the experience playback pool, and the parameters of the strategy network are updated online to adapt to the time-varying load characteristics.

[0042] In the specific implementation, the actual energy transmission efficiency at the current moment is calculated; the load pressure fluctuation rate is detected, and if the load pressure fluctuation rate exceeds a preset threshold, the weight of the pressure tracking error term in the reward function is increased; the four-tuple consisting of the current state, action, reward, and the next state is stored in the experience replay pool, and random sampling is performed from the experience replay pool for offline policy iteration to update the parameters of the policy network, wherein the reward is determined by the actual energy transmission efficiency and the reward function after weight adjustment. For example, the actual energy transmission efficiency η at the current moment is calculated. real =Q load P load / P acc Q acc It also detects the load pressure fluctuation rate σP. If σP exceeds a preset threshold (e.g., 2%), the weight β of the pressure tracking error term in the reward function is automatically increased to enhance pressure stability. Reward value r t Calculated from the actual energy transfer efficiency and the adjusted reward function. The quadruple (s) t ,a t ,r t ,s t+1 The data is stored in the experience replay pool, and random samples are taken from it for offline policy iteration to update the policy network parameters and achieve online adaptation.

[0043] It should be noted that, through the above steps, in this embodiment, during the process of the accumulator pressure dropping from 200 bar to 160 bar, the intelligent agent can automatically increase the duty cycle of the switching valve to maintain the load end pressure at 150 bar, and the pressure fluctuation is controlled within ±2 bar. Compared with traditional throttling control, it saves more than 40% of energy and has no significant temperature rise.

[0044] This embodiment constructs a digital hydraulic transformer hardware system including parallel high-speed switching valves; the valve group control problem is modeled as a Markov decision process, and a reward function is defined with energy efficiency, pressure regulation accuracy, and action smoothness as optimization objectives; a policy network is trained offline using a deep reinforcement learning algorithm that introduces maximum entropy; the system state is collected online in real time, and action vectors are generated through the policy network. After discrete mapping and physical constraints, the switching combination and duty cycle are obtained, driving the valve group to adjust the equivalent flow area to achieve dynamic matching of hydraulic impedance; after the action is executed, state feedback is collected, rewards are calculated, and the policy network is updated online to adapt to time-varying loads. The above method achieves high efficiency and energy saving through a throttling-free transformer, and has the advantages of strong adaptability and small pressure pulsation.

[0045] Furthermore, this embodiment of the invention also proposes a storage medium storing a deep reinforcement learning-based digital valve group impedance matching control program. When the deep reinforcement learning-based digital valve group impedance matching control program is executed by a processor, it implements the steps of the deep reinforcement learning-based digital valve group impedance matching control method described above.

[0046] Reference Figure 4 , Figure 4 This is a structural block diagram of the first embodiment of the digital valve group impedance matching control device based on deep reinforcement learning of the present invention.

[0047] like Figure 4 As shown, the digital valve group impedance matching control device based on deep reinforcement learning proposed in this embodiment of the invention includes: Hardware building module 10 is used to build a digital hydraulic transformer hardware system, which includes several high-speed switching valves of different diameters connected in parallel between the high-voltage accumulator and the actuator. Modeling module 20 is used to model the control problem of the high-speed switching valve as a Markov decision process. The Markov decision process defines a state space, an action space, and a reward function. The reward function is defined as a weighted sum of instantaneous energy transfer efficiency, load-side pressure tracking error, and valve group switching frequency penalty term. The offline training module 30 is used to train the policy network using a deep reinforcement learning algorithm, introduces a maximum entropy term during the training process, and calculates the reward value using the reward function to update the parameters of the policy network. The reasoning and execution module 40 is used to collect the state of the digital hydraulic transformer hardware system in real time, input the state into the trained strategy network to generate a target action vector, perform discrete mapping on the target action vector and apply physical constraints to obtain the switching combination and duty cycle of the several high-speed switching valves, and drive the high-speed switching valves to perform actions with the switching combination and duty cycle to adjust the equivalent flow area and realize dynamic matching of hydraulic impedance. The online update module 50 is used to collect the status feedback after the high-speed switching valve performs the action, calculate the actual reward value based on the feedback, store the interaction data in the experience playback pool, and update the parameters of the strategy network online to adapt to time-varying load characteristics.

[0048] This embodiment constructs a digital hydraulic transformer hardware system including parallel high-speed switching valves; the valve group control problem is modeled as a Markov decision process, and a reward function is defined with energy efficiency, pressure regulation accuracy, and action smoothness as optimization objectives; a policy network is trained offline using a deep reinforcement learning algorithm that introduces maximum entropy; the system state is collected online in real time, and action vectors are generated through the policy network. After discrete mapping and physical constraints, the switching combination and duty cycle are obtained, driving the valve group to adjust the equivalent flow area to achieve dynamic matching of hydraulic impedance; after the action is executed, state feedback is collected, rewards are calculated, and the policy network is updated online to adapt to time-varying loads. The above method achieves high efficiency and energy saving through a throttling-free transformer, and has the advantages of strong adaptability and small pressure pulsation.

[0049] This application embodiment also provides a digital valve group impedance matching control device based on deep reinforcement learning, including a processor, a communication interface, a memory, and a communication bus. The processor, communication interface, and memory communicate with each other through the communication bus. The memory is used to store the digital valve group impedance matching control program based on deep reinforcement learning. When the processor executes the program stored in the memory, it implements the above-mentioned digital valve group impedance matching control method based on deep reinforcement learning.

[0050] The communication bus mentioned in the aforementioned deep reinforcement learning-based digital valve group impedance matching control device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc.

[0051] The communication interface is used for communication between the aforementioned deep reinforcement learning-based digital valve group impedance matching control device and other devices.

[0052] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0053] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0054] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0055] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0056] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0057] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

[0058] It should be understood that the above are merely illustrative examples and do not constitute any limitation on the technical solutions of the present invention. In specific applications, those skilled in the art can make settings as needed, and the present invention does not impose any restrictions on this.

[0059] It should be noted that the workflow described above is merely illustrative and does not limit the scope of protection of this invention. In practical applications, those skilled in the art can select some or all of the workflow to achieve the purpose of this embodiment according to actual needs, and no restrictions are imposed here.

[0060] In addition, for technical details not described in detail in this embodiment, please refer to the digital valve group impedance matching control method based on deep reinforcement learning provided in any embodiment of the present invention, which will not be repeated here.

[0061] Furthermore, it should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0062] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0063] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0064] The above are merely preferred embodiments of the present invention and do not limit the scope of the patent. Any equivalent structural or procedural transformations made based on the description and drawings of the present invention, or direct or indirect applications in other related technical fields, are similarly included within the scope of patent protection of the present invention.

[0065] It is understood that the system provided in the embodiments of the present invention corresponds to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant content can be referred to the corresponding parts of the above methods.

Claims

1. A digital valve group impedance matching control method based on deep reinforcement learning, characterized by, Includes the following steps: A digital hydraulic transformer hardware system is constructed, which includes several high-speed switching valves of different diameters connected in parallel between the high-voltage accumulator and the actuator. The control problem of the high-speed switching valve is modeled as a Markov decision process. The Markov decision process defines a state space, an action space, and a reward function. The reward function is defined as a weighted sum of instantaneous energy transfer efficiency, load-side pressure tracking error, and valve group switching frequency penalty term. The policy network is trained using a deep reinforcement learning algorithm. During the training process, a maximum entropy term is introduced, and the reward value is calculated using the reward function to update the parameters of the policy network. The state of the digital hydraulic transformer hardware system is collected in real time. The state is input into the trained strategy network to generate a target action vector. The target action vector is discretized and physical constraints are applied to obtain the switching combination and duty cycle of the several high-speed switching valves. The switching combination and duty cycle are used to drive the high-speed switching valves to perform actions to adjust the equivalent flow area and achieve dynamic matching of hydraulic impedance. After the high-speed switching valve is driven to perform an action, the status feedback after the action is performed is collected, the actual reward value is calculated based on this feedback, the interaction data is stored in the experience replay pool, and the parameters of the strategy network are updated online to adapt to time-varying load characteristics.

2. The deep reinforcement learning-based digital valve bank impedance matching control method of claim 1, wherein, The reward function R defined in step S2 is specifically: in, Indicates time step Instant reward value, This represents the instantaneous energy transfer efficiency, which is the ratio of load power to source power. This indicates the pressure tracking error at the load end. The switching frequency penalty term, representing the valve group's on / off state, is used to suppress high-frequency chatter. These are the weighting coefficients for each optimization objective.

3. The deep reinforcement learning based digital valve bank impedance matching control method of claim 1, wherein, The method further includes: Define the state space S={P_acc,P_load,Q_req,SOC,v_valve}, where P_acc is the accumulator pressure, P_load is the load pressure, Q_req is the load demand flow, SOC is the accumulator state of charge, and v_valve is the current valve group state vector. Define the action space A={d_1,d_2,...,d_N}, where d_i∈[0,1] represents the duty cycle of the i-th switching valve; Construct a reward function based on the defined state space and action space.

4. The deep reinforcement learning based digital valve bank impedance matching control method of claim 1, wherein, The process of training a policy network using a deep reinforcement learning algorithm, introducing a maximum entropy term during training, and calculating reward values ​​using the reward function to update the parameters of the policy network includes: Construct an actor network, a critic network, and a target critic network; A maximum entropy term is introduced into the objective function. This maximum entropy term is used by the agent to explore diverse control strategies during training to prevent it from getting trapped in local optima. The parameters of the critic network are updated by minimizing the Bellman residual loss function, wherein the target value in the Bellman residual loss function is calculated by the target critic network; The parameters of the actor network are updated by minimizing the policy loss function.

5. The deep reinforcement learning based digital valve bank impedance matching control method of claim 1, wherein, After discretizing and mapping the target motion vector and applying physical constraints, the switching combinations and duty cycles of the plurality of high-speed switching valves are obtained, including: The target action vector output by the strategy network is mapped to a discrete combination of switches, and the duty cycle of each switch valve is determined based on the value of the target action vector. Dead zone compensation and minimum opening time constraints are applied to the obtained switch combinations and duty cycles to filter out excessively short pulse commands that the physical valve core cannot respond to. The switch combinations and duty cycles that do not meet the minimum opening time constraint are adjusted to obtain the switch combinations and duty cycles used to drive the plurality of high-speed switching valves to perform actions.

6. The deep reinforcement learning based digital valve bank impedance matching control method of claim 1, wherein, The method further includes: Within each control cycle, the target motion vector is converted into an equivalent flow area, calculated as follows: wherein, K i is the flow coefficient of the i-th high speed on-off valve, A i is the maximum flow area of the i-th high speed on-off valve, a i is the i-th element of the target motion vector a*, representing the duty ratio of the i-th high speed on-off valve; By dynamically adjusting the equivalent flow area, the system can be configured with variable liquid resistance, thereby maintaining a constant load pressure as the accumulator pressure decreases during the discharge process.

7. The deep reinforcement learning based digital valve bank impedance matching control method of claim 1, wherein, The method further includes: Calculate the actual energy transfer efficiency at the current moment; The load pressure fluctuation rate is detected. If the load pressure fluctuation rate exceeds a preset threshold, the weight of the pressure tracking error term in the reward function is increased. The current state, action, reward, and next state are combined into a quadruple and stored in the experience replay pool. Random samples are then taken from the experience replay pool for offline policy iteration to update the parameters of the policy network. The reward is determined by the actual energy transmission efficiency and the reward function after weight adjustment.

8. A digital valve bank impedance matching control device based on deep reinforcement learning, characterized by, The deep reinforcement learning-based digital valve group impedance matching control device is applied to the deep reinforcement learning-based digital valve group impedance matching control method as described in any one of claims 1 to 7, the device comprising: A hardware building module is used to build a digital hydraulic transformer hardware system, which includes several high-speed switching valves of different diameters connected in parallel between the high-voltage accumulator and the actuator. The modeling module is used to model the control problem of the high-speed switching valve as a Markov decision process. The Markov decision process defines a state space, an action space, and a reward function. The reward function is defined as a weighted sum of instantaneous energy transfer efficiency, load-side pressure tracking error, and valve group switching frequency penalty term. The offline training module is used to train the policy network using a deep reinforcement learning algorithm. During the training process, a maximum entropy term is introduced, and the reward value is calculated using the reward function to update the parameters of the policy network. The reasoning and execution module is used to collect the state of the digital hydraulic transformer hardware system in real time, input the state into the trained strategy network to generate a target action vector, perform discrete mapping on the target action vector and apply physical constraints to obtain the switching combination and duty cycle of the several high-speed switching valves, and drive the high-speed switching valves to perform actions with the switching combination and duty cycle to adjust the equivalent flow area and realize dynamic matching of hydraulic impedance. The online update module is used to collect the status feedback after the high-speed switching valve performs the action, calculate the actual reward value based on the feedback, store the interaction data in the experience playback pool, and update the parameters of the strategy network online to adapt to time-varying load characteristics.

9. A digital valve group impedance matching control device based on deep reinforcement learning, characterized in that, The deep reinforcement learning-based digital valve group impedance matching control device includes: a memory, a processor, and a deep reinforcement learning-based digital valve group impedance matching control program stored in the memory and executable on the processor, wherein the deep reinforcement learning-based digital valve group impedance matching control program is configured to implement the steps of the deep reinforcement learning-based digital valve group impedance matching control method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium stores a digital valve group impedance matching control program based on deep reinforcement learning, which, when executed by a processor, implements the steps of the digital valve group impedance matching control method based on deep reinforcement learning as described in any one of claims 1 to 7.