Storage management method based on deep reinforcement learning under fully programmable valve array biochip

By optimizing the storage fluid path and position through deep reinforcement learning, the resource waste and path conflict problems in fluid storage management in the FPVA chip are solved, more efficient resource utilization and error recovery capabilities are achieved, and the success rate and flexibility of chip operations are improved.

CN120654896APending Publication Date: 2025-09-16FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510958512.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

In fully programmable valve array biochips, existing technologies have difficulty in effectively managing fluid storage, resulting in resource waste and path conflicts, affecting chip operation efficiency and reliability.

Method used

Using deep reinforcement learning methods, through state modeling and neural network learning, an intelligent decision path for storing fluids is generated. The Transformer network and CNN network are combined to optimize the fluid position, and an optimization objective based on the maximum rectangle reward is designed to improve resource utilization and error recovery capabilities.

Benefits of technology

Maximize the continuous available area within the chip, improve resource utilization, enhance error recovery capabilities, and optimize chip operation efficiency and success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654896A_ABST
    Figure CN120654896A_ABST
Patent Text Reader

Abstract

The invention relates to a storage management method based on deep reinforcement learning under a completely programmable valve array biochip, and belongs to the field of micro-fluidic biochip electronic automation. According to the method, in the moving path generation stage, a path generation model is adopted, and moving paths of all storage fluids are generated at a time. And screening the generated paths according to the conflict relation of the areas related to the paths, and marking part of the areas as invalid so as to avoid wrong selection of conflict areas when a new storage position is selected. In the moving distance selection stage, a position on the moving path of each storage fluid is selected as a new storage position for each storage fluid. And after redistribution of all the storage fluids is completed, the difference value of the front and back maximum continuous idle rectangular areas of the FPVA chip is calculated, and the difference value is used as an important index for evaluating the storage management optimization effect. After enough data is collected, the DRL algorithm of the moving path generation model and the moving distance selection model is updated and optimized so as to gradually improve the model performance and precision.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electronic automation of microfluidic biochips, and specifically relates to a storage management method based on deep reinforcement learning in a fully programmable valve array biochip. Background Art

[0002] Fully Programmable Valve Arrays (FPVAs) biochips, an emerging class of continuous-flow microfluidic biochips (CFMBs), combine the efficient fluid manipulation capabilities of traditional continuous-flow chips with the programmability of digital microfluidic biochips (DMFBs). In FPVA chips, valves are distributed in a regular array across a network of flow channels. Software control allows the dynamic construction of a variety of microfluidic functional modules, enabling a range of automated biochemical operations such as sample mixing, transport, storage, and reactions. The valves within the chip are typically actuated by external air pressure, forming elastic structures at specific intersections. By controlling the application and release of air pressure, the flow channels are opened and closed, allowing precise control of the flow path and behavior of liquids. These properties make FPVAs valuable for high-reliability applications such as in vitro diagnostics, drug screening, and military testing.

[0003] As chip integration continues to increase, the number of flow valves on FPVA continues to increase, and the operational process is becoming increasingly complex. To ensure the timing integrity and reliability of biochemical reaction operations, the chip needs to frequently temporarily store intermediate products during operation. For example, when the output result of an operation has not yet been scheduled for subsequent operations, the system must temporarily store the result in liquid form inside the chip for subsequent use. In addition, the limited chip resources may also make it impossible to execute some operations immediately. In this case, the corresponding input fluids also need to be stored delayed. Therefore, dynamically and effectively managing the location distribution of stored fluids to avoid resource waste and path conflicts has become one of the key issues in the operation of FPVA systems.

[0004] In existing research, most error recovery and physical design optimization methods focus on scheduling adjustments, component layout, or wiring path optimization, while research on fluid storage management is relatively lagging behind. In digital microfluidic chips (DMFBs), some studies have attempted to achieve centralized management of droplets by setting up fixed storage areas on the chip's periphery. However, in FPVA chips, due to limited chip resources, there is no strict boundary demarcation between the component and fluid usage areas, resulting in the inability to effectively adapt the above static allocation method. In addition, the disordered distribution of stored fluids on the chip will lead to the fragmentation of continuous available areas, seriously restricting the success rate of layout and path planning of subsequent operations, and reducing the overall execution efficiency of the chip.

[0005] To address these issues, existing literature has explored heuristic rule-based approaches that sequentially move stored fluids into pre-defined areas or obstacle avoidance zones to free up space. However, these approaches struggle to balance path conflicts between multiple fluids with target area optimization, and are prone to falling into local optima or causing bottlenecks. Furthermore, as the operational complexity of FPVA chips continues to increase, the limitations of traditional heuristic strategies in multi-fluid parallel operation and global resource scheduling are becoming increasingly apparent. Summary of the Invention

[0006] The purpose of the present invention is to provide a storage management method based on deep reinforcement learning in a fully programmable valve array biochip. This method realizes intelligent decision-making on the storage fluid path and final position through state modeling and neural network learning, thereby maximizing the area of ​​continuous available area within the chip, improving resource utilization and enhancing error recovery capabilities.

[0007] To achieve the above objectives, the technical solution of the present invention is: a storage management method based on deep reinforcement learning in a fully programmable valve array biochip, comprising:

[0008] Movement path generation stage: For each stored fluid to be adjusted, the current FPVA chip resource state image and the target position coordinates of the corresponding fluid are used as input states. A convolutional neural network (CNN)-based policy network outputs an action probability distribution, representing the probability of the fluid moving in the four directions of up, down, left, and right under the current state. All fluid paths are generated serially in numbered order. During the generation process, the policy network learns how to avoid other fluids and occupied areas through reinforcement learning, effectively avoiding path conflicts and logical errors, thereby ensuring that each path is a legal pressure path.

[0009] Movement distance selection stage: After the movement path is generated, the number of steps each fluid should move along its path is determined to obtain the final storage location. Using the path coordinate sequence of all fluids and the resource usage map of the current FPVA chip as input, the Transformer network and CNN network structure are combined to extract the global dependencies and spatial distribution characteristics between paths. The output result is the optimal movement distance for each fluid, that is, the number of steps it should move along the path, thereby achieving the optimal re-storage location allocation effect.

[0010] Optimization objective based on maximum rectangle reward: To train the deep reinforcement learning models in the two stages mentioned above, a reward function based on the area of ​​the maximum continuous available rectangular region is designed. Specifically, after each strategy execution, the maximum free rectangular area on the chip before and after adjustment is evaluated, and the difference between the two is fed back to the deep reinforcement learning model as a positive reward. The reward function directly quantifies the degree of improvement in resource coherence, guiding the deep reinforcement learning model to evolve towards the release of larger continuous spaces, thereby effectively improving the success rate and error recovery efficiency of subsequent physical design.

[0011] Furthermore, before the moving path generation phase, an initialization phase is also included, that is, the model parameters are initialized through uniform distribution and the environment parameters are set at the same time.

[0012] Furthermore, before the moving path generation stage and after the initialization stage, a DRL environment reset stage is also included, that is, randomly generating diverse component layouts and storage fluid distributions.

[0013] Furthermore, in the mobile path generation stage, the mobile path generation processes of multiple storage fluids are uniformly modeled as a Markov decision process; specifically, during the deep reinforcement learning model DRL training process, the mobile path generation of different storage fluids is regarded as a continuous decision-making process under the same training task, and its sampling data is integrated into a complete training trajectory; through this unified modeling method, DRL training can optimize the path generation process globally.

[0014] Furthermore, during the process of generating the movement path of the storage fluid, the areas to which the storage fluid can be moved are divided into optional areas and shielded areas. The basis for distinction is whether a part of the movement path of the corresponding storage fluid overlaps with the movement path of the storage fluid with a larger number than it. If so, these areas need to be marked as shielded areas.

[0015] Furthermore, in the moving distance selection stage, after the moving path is generated, the final position of the stored fluid on its flow path is determined by the moving distance selection module, thereby achieving the goal of redistributing the stored fluid. The overall process is formalized as a Markov decision process, in which the DRL agent decides the new position of each stored fluid in turn according to the number of the stored fluid. The agent takes the component distribution of the FPVA chip, the current distribution of the stored fluid and the generated path sequence as input, and outputs a moving distance after processing by the strategy network. This value represents the number of steps the stored fluid has moved on its moving path, and then determines its new storage position through calculation.

[0016] The present invention also provides a storage management system based on deep reinforcement learning under a fully programmable valve array biochip, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, it can implement the steps of any of the methods described above.

[0017] The present invention also provides a computer-readable storage medium on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the steps of any of the above methods can be implemented.

[0018] Compared with the existing technology, the present invention has the following beneficial effects: the method of the present invention realizes intelligent decision-making on the storage fluid path and final position through state modeling and neural network learning, thereby maximizing the area of ​​continuous available area within the chip, improving resource utilization and enhancing error recovery capability. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 Storage management effect example diagram;

[0020] Figure 2 Storage management flow chart;

[0021] Figure 3 Example diagram of movement path generation;

[0022] Figure 4 Example diagram of moving distance selection;

[0023] Figure 5 Schematic diagram of the strategy network selection of the moving distance selection module;

[0024] Figure 6 Example diagram of optimization objectives. DETAILED DESCRIPTION

[0025] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings.

[0026] The present invention provides a storage management method based on deep reinforcement learning in a fully programmable valve array biochip, comprising:

[0027] Movement path generation stage: For each stored fluid to be adjusted, the current FPVA chip resource state image and the target position coordinates of the corresponding fluid are used as input states. A convolutional neural network (CNN)-based policy network outputs an action probability distribution, representing the probability of the fluid moving in the four directions of up, down, left, and right under the current state. All fluid paths are generated serially in numbered order. During the generation process, the policy network learns how to avoid other fluids and occupied areas through reinforcement learning, effectively avoiding path conflicts and logical errors, thereby ensuring that each path is a legal pressure path.

[0028] Movement distance selection stage: After the movement path is generated, the number of steps each fluid should move along its path is determined to obtain the final storage location. Using the path coordinate sequence of all fluids and the resource usage map of the current FPVA chip as input, the Transformer network and CNN network structure are combined to extract the global dependencies and spatial distribution characteristics between paths. The output result is the optimal movement distance for each fluid, that is, the number of steps it should move along the path, thereby achieving the optimal re-storage location allocation effect.

[0029] Optimization objective based on maximum rectangle reward: To train the deep reinforcement learning models in the two stages mentioned above, a reward function based on the area of ​​the maximum continuous available rectangular region is designed. Specifically, after each strategy execution, the maximum free rectangular area on the chip before and after adjustment is evaluated, and the difference between the two is fed back to the deep reinforcement learning model as a positive reward. The reward function directly quantifies the degree of improvement in resource coherence, guiding the deep reinforcement learning model to evolve towards the release of larger continuous spaces, thereby effectively improving the success rate and error recovery efficiency of subsequent physical design.

[0030] The following is a specific implementation process of the present invention.

[0031] like Figure 2The flowchart shown in Figure 1 illustrates the execution and training process of a DRL-based storage management algorithm. During the initialization phase, model parameters are initialized using a uniform distribution, and environmental parameters such as the FPVA chip size are set. During the DRL environment reset phase, diverse component layouts and storage fluid distributions are randomly generated to improve the model's generalization and adaptability to diverse scenarios. During the movement path generation phase, a specially designed path generation model is used to generate movement paths for all storage fluids simultaneously. Subsequently, the generated paths are screened based on the conflicting regions involved in the paths, and some regions are marked as invalid to avoid incorrect selection of these conflicting regions when selecting new storage locations. During the movement distance selection phase, another model is used to select a location on the movement path for each storage fluid as its new storage location. After all storage fluids are redistributed, the difference in the area of ​​the largest continuous free rectangle before and after the FPVA chip is calculated. This difference is used as a key metric to evaluate the effectiveness of storage management optimization. After sufficient data is collected, the DRL algorithms for the movement path generation model and the movement distance selection model are updated and optimized to gradually improve model performance and accuracy. Figure 1 This is an example diagram of the storage management effect.

[0032] 1. Movement path generation phase:

[0033] In the mobile path generation module of storage management, a flow path that satisfies pressure constraints needs to be generated in order to move the stored fluid to the appropriate location. Since multiple stored fluids need to be moved to different locations, the area covered by each storage fluid's path determines its reachable location, while the maximum free area on the FPVA chip is determined by the final location of all stored fluids. Therefore, the movement paths of different storage fluids should have a certain connection, rather than being independent individuals. In order to enable the path generation module to perceive the correlation between different paths when generating all paths, this chapter makes appropriate improvements to the wiring adjustment module.

[0034] In the original routing adjustment module, only a single flow path is generated at a time, completing the path generation task without considering the global distribution. However, in storage management, when generating a path for a specific storage fluid, it is necessary to fully consider the area covered by previously generated paths and dynamically adjust the path of the current storage fluid based on this information to ensure that the final distribution of all storage fluids maximizes the maximum continuous free area of ​​the FPVA chip. To achieve this goal, this chapter unifies the path generation process for multiple storage fluids as a Markov decision process. Specifically, during DRL training, path generation for different storage fluids is treated as a continuous decision process within the same training task, and their sampled data is integrated into a complete training trajectory. This unified modeling approach allows DRL training to optimize the path generation process globally, rather than simply optimizing local paths independently. This approach improves the global coordination of path generation, avoids path conflicts to a certain extent, and achieves overall optimization of the storage fluid distribution. The main advantage of this modification is that it effectively captures the mutual influence between different storage fluid paths, generating coordinated path solutions through a global optimization method. Furthermore, not all storage fluids require or can generate movement paths. This may be because it is impossible to find a feasible movement path for it under the existing conditions, or the DRL model determines that the stored fluid does not need to move to meet the overall optimization goal, and thus chooses not to generate a movement path for it. Figure 3 As shown, no corresponding moving path is generated for the storage fluid S1, which means that its position remains unchanged. The moving paths of S2, S3 and S4 will be generated in sequence according to the order of numbers. It should be noted that when generating the moving path of a certain storage fluid, only the storage fluid positions and the positions of the actuators with numbers larger than it are considered as prohibited to reach. For example, when generating the path of S3, the position of S2 in the figure is set to be reachable, while the position of S4 is set to be prohibited to reach. The reason for this setting is that the purpose of generating the moving path is to pave the way for the subsequent selection of the moving distance, and the actual movement of the storage fluid is carried out after the moving distance is selected. It is conceivable that when S3 is subsequently moved through the path of S3, S2 may have been moved to other places, while S4 is still in its original position. Therefore, when generating the path of S3, if it passes through S2, it does not necessarily violate the constraint, but if it passes through S4, it will definitely violate the constraint.

[0035] like Figure 3In the moving path of S2 shown in the figure, the storage fluid S2 will be pushed by the air pressure from the input port In to the output port Out, so the area to which S2 can be moved is the area after S2 in the pressure path. In order to ensure the reasonable execution of the task, these areas need to be divided into two parts. One part is the position to which S2 can be moved subsequently, that is, the optional part in the figure, and the other part is the position to which S2 cannot be moved subsequently, that is, the shielded part shown in the figure. The basis for distinction is whether a part of the path overlaps with the moving path of the storage fluid with a larger number than it. If it overlaps, these areas need to be marked as shielded areas, because if they are moved to these positions, the subsequent moving path of the storage fluid will be blocked.

[0036] 2. Moving distance selection stage

[0037] After the movement path is generated, the final position of the stored fluid along its flow path is determined by the movement distance selection module, thereby completing the redistribution of the stored fluid. To achieve this goal, the problem is formalized as a Markov decision process, in which the DRL agent sequentially decides the new location of each stored fluid based on its number. In this process, the agent takes the component distribution of the FPVA chip, the current distribution of the stored fluids, and the generated path sequence as input. After processing through the policy network, the agent outputs a movement distance, which represents the number of steps the stored fluid has moved along its path. The new storage location is then determined through simple calculations.

[0038] Specifically, if Figure 4 As shown in FIG, the process of re-determining the storage fluid distribution by the DRL algorithm is demonstrated. Figure 4 (a) is a schematic diagram of the determination of the moving distance of S2. When selecting the moving distance of the current storage fluid, since the storage fluids in S3 and S4 have not yet moved, the path sequence must also include the path information of these storage fluids in order to perform more global optimization. The moving path sequence of each storage fluid is composed of the paths generated in the previous section and is represented in the form of a coordinate sequence, where (-1, -1,) is used to mark the end position of the path, which is similar to the way "END" is used to mark the end of a sentence in the field of natural language processing. In addition, the distribution of components and storage fluids in the FPVA chip is represented in matrix form, which is similar to the representation method of the layout adjustment module. For an FPVA chip of size N×N, the distribution matrix is ​​N×N. At the location containing components or storage fluids, the matrix value is set to 1, and the remaining locations are set to 0. The structure of the policy network is as follows Figure 5As shown, the movement path sequence is input into the Transformer [66-69] encoder. Through multi-layer multi-head attention, the encoder can extract high-order features from the input sequence [70,71], while the distribution matrix of components and storage fluids is extracted through CNN. Finally, the feature vectors of various types of information are fused and transformed through a multi-layer fully connected network to output the corresponding action probability distribution. This probability distribution represents the probability of each possible movement distance being selected. The maximum value of the movement distance is set to 2N in this chapter. The actual value can be flexibly set according to the usage scenario. The appropriate value can accelerate the convergence speed and training quality of the DRL model. After obtaining a specific movement distance and combining it with the movement path in the previous chapter, the new storage location of the storage fluid can be calculated. Figure 4 (b) and (c) show the moving distance selection process of S3 and S4. The path sequences that have previously selected moving distances will be eliminated in order to reduce some unnecessary model calculations.

[0039] 3. Optimization objective based on maximum rectangle reward

[0040] After generating a new storage fluid distribution through the above process, first calculate the area of ​​the largest rectangular region of the FPVA chip after redistribution and compare it with the area of ​​the initial largest rectangular region to obtain the area difference. This difference is used as a reward signal, reflecting the degree to which storage management optimization improves chip resource utilization, and will be used to guide the update and optimization process of the two models in this chapter in subsequent training. The design of the reward signal is intended to quantify the impact of storage fluid redistribution on the overall performance of the chip, so that the DRL model can focus more on optimizing global resource utilization efficiency during training. Figure 6 As shown in the figure, by rationally redistributing the storage fluid, the maximum contiguous rectangular area of ​​the FPVA chip has been increased by 8. This improvement significantly increases the chip's available resource space, enabling it to accommodate more or larger components and thus support more complex operations. This not only improves chip resource utilization but also provides greater flexibility and efficiency for task scheduling and execution of biochemical reactions, further optimizing overall biochemical reaction performance. This process fully demonstrates the key role of memory management optimization in improving chip resource allocation efficiency and supporting the execution of complex operations.

[0041] The present invention also provides a storage management system based on deep reinforcement learning under a fully programmable valve array biochip, comprising a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, it can implement the steps of any of the methods described above.

[0042] The present invention also provides a computer-readable storage medium on which computer program instructions that can be executed by a processor are stored. When the processor executes the computer program instructions, the steps of any of the above methods can be implemented.

[0043] The above are preferred embodiments of the present invention. Any changes made according to the technical solution of the present invention, as long as the resulting functions and effects do not exceed the scope of the technical solution of the present invention, shall fall within the scope of protection of the present invention.

Claims

1. A storage management method based on deep reinforcement learning in a fully programmable valve array biochip, characterized in that: include: Movement path generation stage: For each stored fluid to be adjusted, the current FPVA chip resource state image and the target position coordinates of the corresponding fluid are used as input states. The policy network is used to output the action probability distribution, which represents the probability of the fluid moving in the four directions of up, down, left, and right in the current state. All fluid paths are generated serially in numbered order. During the generation process, the policy network uses reinforcement learning to learn how to avoid other fluids and occupied areas, effectively preventing path conflicts and logical errors, thereby ensuring that each path is a legal pressure path; Movement distance selection phase: After the movement path is generated, the number of steps each fluid should take along its path to obtain the final storage location is determined. Using the path coordinate sequence of all fluids and the resource usage map of the current FPVA chip as input, the global dependencies and spatial distribution characteristics between paths are extracted. The output result is the optimal movement distance for each fluid, that is, the number of steps it should move along the path, thereby achieving the optimal re-storage location allocation effect. Optimization objective based on maximum rectangle reward: To train the deep reinforcement learning models in the two stages mentioned above, a reward function based on the area of ​​the maximum continuous available rectangular region is designed. Specifically, after each strategy execution, the maximum free rectangular area on the chip before and after adjustment is evaluated, and the difference between the two is fed back to the deep reinforcement learning model as a positive reward. The reward function directly quantifies the degree of improvement in resource coherence, guiding the deep reinforcement learning model to evolve towards the release of larger continuous spaces, thereby effectively improving the success rate and error recovery efficiency of subsequent physical design.

2. The storage management method based on deep reinforcement learning in a fully programmable valve array biochip according to claim 1 is characterized in that: In the mobile path generation stage, the policy network is a policy network based on convolutional neural network (CNN).

3. The storage management method based on deep reinforcement learning in a fully programmable valve array biochip according to claim 1 is characterized in that: In the moving distance selection stage, the global dependencies and spatial distribution characteristics between paths are extracted based on the Transformer network and CNN network structure.

4. The storage management method based on deep reinforcement learning in a fully programmable valve array biochip according to claim 1 is characterized in that: Before the moving path generation stage, an initialization stage is also included, that is, the model parameters are initialized through uniform distribution and the environmental parameters are set at the same time.

5. The storage management method based on deep reinforcement learning in a fully programmable valve array biochip according to claim 4 is characterized in that: Before the moving path generation phase and after the initialization phase, the DRL environment reset phase is also included, that is, randomly generating diverse component layouts and storage fluid distributions.

6. The storage management method based on deep reinforcement learning in a fully programmable valve array biochip according to claim 1 is characterized in that: In the movement path generation stage, the movement path generation processes of multiple storage fluids are uniformly modeled as a Markov decision process; specifically, during the deep reinforcement learning model DRL training process, the movement path generation of different storage fluids is regarded as a continuous decision process under the same training task, and its sampling data is integrated into a complete training trajectory; through this unified modeling method, DRL training can optimize the path generation process on a global scale.

7. The storage management method based on deep reinforcement learning in a fully programmable valve array biochip according to claim 6 is characterized in that: During the generation of the movement path of the storage fluid, the areas to which the storage fluid can be moved are divided into optional areas and shielded areas. The basis for distinction is whether a part of the movement path of the corresponding storage fluid overlaps with the movement path of a storage fluid with a larger number than it. If so, these areas need to be marked as shielded areas.

8. The storage management method based on deep reinforcement learning in a fully programmable valve array biochip according to claim 6 is characterized in that: In the moving distance selection stage, after the moving path is generated, the final position of the stored fluid on its flow path is determined by the moving distance selection module, thereby achieving the goal of redistributing the stored fluid. The overall process is formalized as a Markov decision process, in which the DRL agent decides the new position of each stored fluid in turn according to the number of the stored fluid. The agent takes the component distribution of the FPVA chip, the current distribution of the stored fluid, and the generated path sequence as input, and outputs a moving distance after processing by the policy network. This value represents the number of steps the stored fluid has moved on its moving path, and then determines its new storage position through calculation.

9. A storage management system based on deep reinforcement learning for a fully programmable valve array biochip, characterized in that: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the steps of the method according to any one of claims 1 to 8 can be implemented.

10. A computer-readable storage medium storing computer program instructions that can be executed by a processor, wherein when the processor executes the computer program instructions, the steps of the method according to any one of claims 1 to 8 can be implemented.