Biological information data analysis method and system based on high-throughput sequencing and storage medium

Through particle swarm optimization, genetic algorithm, Riemannian geometry, Kullback-Leibler divergence and game theory model, the uneven resource allocation and task scheduling delay in high-throughput sequencing data analysis are solved, and efficient and dynamic resource scheduling and task execution are achieved.

CN120429077APending Publication Date: 2025-08-05JINAN AIXIN ZHUOER MEDICAL LAB CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510490843.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

In the analysis of large-scale high-throughput sequencing data, uneven resource allocation, delayed task scheduling and insufficient coordination of multiple data sources lead to uneven computing resource utilization and inefficient task execution.

Method used

Particle swarm optimization and genetic algorithm are used for global optimization of resource allocation and task scheduling, combined with Riemannian geometry and Kullback-Leibler divergence to optimize data transmission paths, coordinate multi-data source resource allocation using game theory models, and combine with optimal control theory for real-time dynamic scheduling.

Benefits of technology

It realizes efficient and accurate resource allocation and task scheduling in a multi-data source environment, reduces redundant information and delays in data transmission, and ensures timely execution of high-priority tasks and maximizes resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429077A_ABST
    Figure CN120429077A_ABST
Patent Text Reader

Abstract

The invention relates to the field of biology, and discloses a biological information data analysis method and system based on high-throughput sequencing and a storage medium, and the method comprises the following steps: S1, carrying out cleaning, redundancy removal and feature selection on high-throughput sequencing data; s2, optimizing resource allocation and task scheduling of a plurality of data sources through particle swarm optimization and a genetic algorithm; s3, optimizing a data transmission path between the data sources by using Riemannian geometry and KL divergence; s4, resource allocation is optimized through the game theory, and coordination between data sources is achieved; and S5, in combination with an optimal control theory, dynamically scheduling tasks, and optimizing an execution sequence and resource use. By optimizing the data transmission path, the technical effect of effectively reducing redundant information and delay in the data transmission process is achieved, the difference between data sources and the shortest property of the data flow path are comprehensively considered, and the problem that data flow loss and delay cannot be effectively controlled in an existing method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and in particular to a bioinformatics data analysis method, system and storage medium based on high-throughput sequencing. Background Art

[0002] With the continuous development of high-throughput sequencing technology, the field of bioinformatics has become capable of processing and analyzing large amounts of genomic data. This data not only includes gene sequence information but also complex clinical data and experimental results. Existing technologies rely on a series of computational models and algorithms, typically including steps such as data cleaning, redundancy removal, and feature selection. These technical solutions have achieved some success in processing smaller data sets and can provide a certain degree of data analysis and resource scheduling optimization, especially for simple tasks, effectively allocating and scheduling computing tasks.

[0003] While existing technologies have made progress in basic data processing and small-scale task optimization, they still face significant challenges in the complex environments of large-scale data analysis and multi-resource scheduling. Traditional static scheduling methods cannot effectively cope with the dynamic changes in computing resources, resulting in uneven resource utilization and inefficient task execution. Existing resource scheduling methods also fail to adequately consider the priorities and dependencies between tasks, potentially causing high-priority tasks to be delayed due to insufficient resources, impacting overall system performance.

[0004] Furthermore, the integration of multiple data sources and task scheduling lacks an effective global optimization mechanism, which prevents maximizing computing resources and effectively guarantees the efficiency and accuracy of task execution. Traditional optimization algorithms often rely on simple local optimization strategies and fail to achieve true global coordination and dynamic adaptation at the system level. This results in resource competition and task conflicts in existing methods in multi-data source environments. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a bioinformatics data analysis method, system and storage medium based on high-throughput sequencing, which solves the problems of uneven resource allocation, task scheduling delays and insufficient coordination of multiple data sources in large-scale high-throughput sequencing data analysis in the existing technology.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A bioinformatics data analysis method based on high-throughput sequencing, comprising the following steps: S1. Preprocessing of high-throughput sequencing data, including data cleaning, redundant data removal, and feature selection; S2, using particle swarm optimization and genetic algorithms to perform global optimization, optimizing resource allocation and computing task scheduling during the integration of multiple data sources; S3, based on information geometry optimization method, uses Riemannian geometry and Kullback-Leibler divergence to optimize the data transmission path between data sources; S4. Use game theory methods to optimize resource allocation of multiple data sources and achieve coordination between data sources through game models; S5. Combined with optimal control theory, computing tasks are dynamically scheduled in real time, and the execution order of tasks is adjusted according to the usage of computing resources.

[0007] Preferably, in step S1, redundant data removal and feature selection include the following steps: Calculate the overall entropy of the data set, select features with higher information gain for data integration, and remove redundant features with lower information gain; Calculate the gain of each feature, and perform feature selection based on the size of the information gain to remove redundant features.

[0008] Preferably, in the step S2, the particle swarm optimization specifically updates the speed and position of each particle, and updates the speed and position of the particle using the following formula: v i (t+1)=wv i (t)+c1r1(p i -x i (t))+c2r2(g i -x i (t)); x i (t+1)=x i (t)+v i (t+1); Among them, v i (t) is the velocity of the particle; x i (t) is the position of the particle; p i is the best historical position of the particle; g i is the optimal position of the group; w is the inertia weight; c1, c2 are learning factors; r1, r2 are random numbers.

[0009] Preferably, in step S2, the genetic algorithm optimizes the resources based on the fitness function of each data source, and the fitness function is expressed by weighted summation: Among them, w j is the weight of each data source; x j It is the corresponding resource allocation.

[0010] Preferably, in the step S3, information geometry optimization minimizes the path length of data flow based on Riemannian geometry and minimizes the difference between data sources through Kullback-Leibler divergence to minimize redundant information in the data integration process. The path length and divergence are calculated as follows: Among them, γ is the data flow path; is the Riemannian metric; p(x i ) and q(x i ) are the probability distributions of the two data sources respectively.

[0011] Preferably, in step S4, game theory is used to model resource allocation for multiple data sources, and resource coordination between data sources is achieved through a Nash equilibrium model, ensuring that the computing tasks of each data source are optimally executed, and satisfying the following conditions: Among them, u i is the utility function of data source i; It is the optimal choice of strategy.

[0012] Preferably, the step S5 further comprises the following steps: Adopting dynamic programming and linear quadratic control algorithms, the task execution order is dynamically adjusted according to the current computing resource usage, task execution priority and system load; Select preemptive or non-preemptive scheduling strategies based on task priority, computing resource requirements, and real-time resource load, delaying the execution of low-priority tasks while high-priority tasks are executing. Use real-time feedback mechanisms to adjust task scheduling strategies based on the current system status, and use resource prediction algorithms to predict future resource requirements and adjust the order of task execution in advance. When computing resources are limited, a heuristic algorithm is used to adjust the task execution order based on the task dependencies and execution priorities, giving priority to tasks with higher resource requirements and higher priorities.

[0013] A bioinformatics data analysis system based on high-throughput sequencing, characterized by comprising: The data preprocessing module is used to clean high-throughput sequencing data, remove redundant data, select features, and calculate information entropy; the global optimization module is used to globally optimize data integration and resource allocation through particle swarm optimization and genetic algorithms; Information geometry optimization module, used to optimize data flow paths based on Riemannian geometry and Kullback-Leibler divergence; Game theory module, used to achieve resource coordination and task scheduling among multiple data sources through game models; Optimal control module, used for dynamic scheduling and optimal control of task execution based on the usage of computing resources; The system communication module is used to coordinate the data flow between modules to ensure the efficiency of the data processing process.

[0014] Preferably, the data preprocessing module and the global optimization module transmit data through the calculated feature data and information gain results, ensuring that the global optimization module uses the preprocessed data for resource optimization and scheduling; The global optimization module and the information geometry optimization module transmit the globally optimized resource allocation data to provide the optimized data flow path to the information geometry module; The information geometry optimization module and the game theory module coordinate through the optimized data flow path to ensure that data flows efficiently along the optimized path and provide coordinated data for the game model; The game theory module and the optimal control module perform task scheduling based on the coordinated resource allocation results, and dynamically allocate resources based on real-time scheduling feedback; The optimal control module performs data transmission and task scheduling feedback with each module through the system communication module, and adjusts the task sequence and resource allocation in real time.

[0015] A storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

[0016] The present invention provides a method, system, and storage medium for analyzing bioinformatics data based on high-throughput sequencing. It has the following beneficial effects: 1. The present invention adopts an optimization method based on information geometry, which optimizes the data transmission path by combining Riemannian geometry and KL divergence, achieving the technical effect of effectively reducing redundant information and delays in the data transmission process. Compared with traditional path optimization methods in the prior art, it solves the problem of the inability to effectively control data flow loss and delay in existing methods by comprehensively considering the differences between data sources and the shortest data flow path.

[0017] 2. The present invention adopts a global optimization technology that combines particle swarm optimization with genetic algorithm to optimize resource allocation and task scheduling, achieving the technical effect of efficiently and accurately allocating resources in a multi-data source environment. Compared with the single optimization method in the existing technology, it solves the problem that a single algorithm is prone to falling into local optimality in a complex environment, and effectively improves the utilization efficiency of system resources and the balance of task execution.

[0018] 3. The present invention adopts dynamic programming and linear quadratic regulation control algorithm to perform real-time dynamic scheduling of tasks, achieving the technical effect of being able to flexibly respond to different task requirements and ensure priority execution of high-priority tasks when computing resources are limited. Compared with the static scheduling method in the prior art, it can adjust the task order in real time, avoid the problem of excessive concentration or lag of resources, and ensure the timely completion of computing tasks and the maximum utilization of resources.

[0019] 4. The present invention coordinates resources and schedules tasks for multiple data sources through a game theory model, achieving the technical effect of fair and efficient resource allocation when multiple resources are shared. Compared with the existing technology that only considers the allocation of a single resource, it considers the interests of multiple parties and resource constraints through a game model, solves the problems of uneven resource allocation and unbalanced task execution in the existing technology, and improves the stability and efficiency of the overall system. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A diagram showing the steps of the method of the present invention; Figure 2 It is a system module diagram of the present invention. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0022] Example 1: Please see the attached Figure 1 The present invention provides a method for analyzing bioinformatics data based on high-throughput sequencing, comprising the following steps: S1. Preprocessing of high-throughput sequencing data, including data cleaning, redundant data removal, and feature selection; S2, using particle swarm optimization and genetic algorithms to perform global optimization, optimizing resource allocation and computing task scheduling during the integration of multiple data sources; S3, based on information geometry optimization method, uses Riemannian geometry and Kullback-Leibler divergence to optimize the data transmission path between data sources; S4. Use game theory methods to optimize resource allocation of multiple data sources and achieve coordination between data sources through game models; S5. Combined with optimal control theory, it dynamically schedules computing tasks in real time and adjusts the execution order of tasks according to the usage of computing resources; In step S1, redundant data removal and feature selection include the following steps: Calculate the overall entropy of the data set, select features with higher information gain for data integration, and remove redundant features with lower information gain; Calculate the gain of each feature, and perform feature selection based on the size of the information gain to remove redundant features.

[0023] In step S2, the particle swarm optimization specifically updates the speed and position of each particle and updates the speed and position of the particle using the following formula: v i (t+1)=wv i (t)+c1r1(p i -x i (t))+c2r2(g i -x i (t)); x i (t+1)=x i (t)+v i (t+1); Among them, v i (t) is the velocity of the particle; x i (t) is the position of the particle; p i is the best historical position of the particle; g i is the optimal position of the group; w is the inertia weight; c1, c2 are learning factors; r2, r2 are random numbers.

[0024] In step S2, the genetic algorithm optimizes resources based on the fitness function of each data source. The fitness function is expressed as a weighted sum: Among them, w j is the weight of each data source; x j It is the corresponding resource allocation.

[0025] In step S3, information geometry optimization minimizes the path length of data flow based on Riemannian geometry and minimizes the differences between data sources through Kullback-Leibler divergence to minimize redundant information during data integration. The path length and divergence are calculated as follows: Among them, γ is the data flow path; is the Riemannian metric; p(x i ) and q(x i ) are the probability distributions of the two data sources respectively.

[0026] In step S4, game theory is used to model the resource allocation of multiple data sources. The Nash equilibrium model is used to coordinate resources between data sources and ensure that the computing tasks of each data source are optimally executed, meeting the following conditions: Among them, u i is the utility function of data source i; It is the optimal choice of strategy.

[0027] The S5 step further includes the following steps: Adopting dynamic programming and linear quadratic control algorithms, the task execution order is dynamically adjusted according to the current computing resource usage, task execution priority and system load; Select preemptive or non-preemptive scheduling strategies based on task priority, computing resource requirements, and real-time resource load, delaying the execution of low-priority tasks while high-priority tasks are executing. Use real-time feedback mechanisms to adjust task scheduling strategies based on the current system status, and use resource prediction algorithms to predict future resource requirements and adjust the order of task execution in advance. When computing resources are limited, a heuristic algorithm is used to adjust the task execution order based on the task dependencies and execution priorities, giving priority to tasks with higher resource requirements and higher priorities.

[0028] Specifically, step S1 is the first step of the entire bioinformatics data analysis method and is of great significance. It provides high-quality basic data for subsequent analysis processes. In this step, the raw data needs to go through multiple processing stages, including data cleaning, redundant data removal, and feature selection. Data cleaning is to remove noise and inconsistencies in the raw data, and redundant data removal is to reduce computational complexity and retain feature data with greater information content. The goal of feature selection is to select the most representative features that can affect the analysis results from a large number of features, to ensure the efficiency and accuracy of subsequent analysis. The entire preprocessing process is aimed at cleaning and optimizing the data so that subsequent global optimization, information geometry optimization, game theory optimization, and optimal control can be performed efficiently and accurately.

[0029] In this embodiment, the main steps of data cleaning include removing missing data, outliers, and inconsistent data. When dealing with missing data, interpolation can be used to fill missing values, or samples or features with missing values can be deleted. For outlier detection, statistical methods such as Z-score and box plots can be used to identify and remove outliers.

[0030] In one possible implementation, redundant data removal is performed by calculating the information entropy and information gain of each feature.

[0031] It generally divides features with higher information gain and features with lower information gain first. The specific evaluation criteria are as follows: "Features with higher information gain" refer to features whose information gain values rank in the top of a certain proportion, where the proportion range is generally between 10% and 20%; and "features with lower information gain" refer to features whose information gain values rank below the threshold, where the threshold is generally set to 0.05 to 0.1 based on the actual situation of the data set.

[0032] Moreover, information entropy is a way to measure data uncertainty. The higher the entropy value, the more information the feature contains. Therefore, by calculating the information gain, we can determine the contribution of each feature to data classification.

[0033] Specifically, assume that the dataset D = {d1, d2, ..., d N}; where each sample d N is an M-dimensional feature vector. Information entropy H(D) is used to measure the uncertainty of a data set and is defined as follows: Among them, p(d z ) represents sample d in dataset D z probability distribution; N represents the total number of samples in the data set D; z is the index of each sample in the data set, ranging from 1 to N, that is, the number of all samples.

[0034] When calculating information gain, for each feature x j , we use information gain IG(x j ) to measure the contribution of the feature to the entropy of the data set. Information gain can be expressed as: Among them, V(x j ) represents the feature x j The value space of ; p(v) is the probability of eigenvalue v; H(D v ) is the feature x j = v when the conditional entropy of the data set.

[0035] Generally speaking, if a feature has a low information gain, it means that it contributes little to the data classification and can be removed. By calculating information gain, we can filter out features that contribute significantly to the classification, reduce redundant data, and improve computational efficiency.

[0036] Alternatively, other algorithms such as recursive feature elimination and regularized feature selection can be used to further optimize the feature set. These methods typically select the most discriminative features by evaluating their contribution to the model.

[0037] In some embodiments, to ensure that the data is suitable for subsequent analysis methods, the features can be further standardized or normalized. For example, the standardized features can be performed using the following formula: Among them, x e is the eigenvalue; μ is the mean of the characteristic; σ is the standard deviation; x ′ e is the normalized feature value. Normalization helps eliminate the scale differences between different features and prevents some features from having too large an impact on the model due to their large scale.

[0038] In this embodiment, the data after data cleaning, redundant data removal, and feature selection will serve as input for subsequent steps and enter the global optimization module for further processing and optimization. In this global optimization module, data resources are allocated and tasks are scheduled based on the cleaned and selected features, further improving computational efficiency and the accuracy of analysis results.

[0039] Step S2 uses particle swarm optimization and genetic algorithms to globally optimize the resource allocation and computing tasks of multiple data sources. The core purpose of this step is to improve the efficiency of data analysis through global optimization, ensuring efficient scheduling of computing tasks and optimal utilization of resources.

[0040] In this embodiment, particle swarm optimization is used to find the optimal solution in a wide solution space. Particle swarm optimization uses the PSO algorithm to simulate the movement of a particle group in the solution space and use the position and velocity of each particle to find the global optimal solution. Each particle represents a potential solution, and its state is determined by its position x i (t) and velocity v i (t). The particle position update process is affected by its historical best position p i and the optimal position g of the group i The update formula of particle velocity and position is as follows: i (t+1)=wv i (t)+c1r1(p i -x i (t))+c2r2(g i -x i (t)); x i (t+1)=x i (t)+v i (t+1); Among them, v i (t) is the velocity of particle i at time t; x i (t) is the position of particle i at time t; p i is the best historical position of particle i; g i is the optimal position of the group; w is the inertia weight, which controls the update amplitude of the particle velocity; c1 and c2 are learning factors used to adjust the particle's dependence on its historical best position and the group's best position; r1 and r2 are two random numbers with a value range of [0,1], which are used to increase the randomness of the search.

[0041] In general, the particle swarm optimization algorithm effectively explores the solution space and avoids being trapped in local optima. By adjusting the inertia weight w and the learning factors c1 and c2, the accuracy and efficiency of the search can be further improved. For example, a large inertia weight allows the particle swarm to search more extensively in the solution space, making it suitable for early global searches. A smaller inertia weight, on the other hand, helps particles gather near the optimal solution, enhancing the accuracy of local searches.

[0042] Alternatively, in particle swarm optimization, position and velocity constraints can be imposed based on the specific problem to ensure the feasibility of the solution. For example, in a resource allocation problem, the velocity and position of a particle can represent the allocation ratio of computing resources and be constrained to ensure that its value does not exceed a preset resource limit.

[0043] In one possible implementation, a genetic algorithm is incorporated into the global optimization process to further refine the solution. Genetic algorithms iteratively optimize the solution set by simulating the process of natural selection. The main steps of a genetic algorithm include selection, crossover, and mutation. The selection process evaluates the quality of each solution using a fitness function, which is typically defined as the objective function in resource allocation optimization problems. The fitness function is expressed as follows: Among them, f(x) is the fitness function, which indicates the quality of solution x; w j is the weight of each data source; x j is the resource allocation value associated with the jth data source; N is the total number of data sources. The larger the value of the fitness function, the better the resource allocation plan.

[0044] Specifically, genetic algorithms include selection, crossover, and mutation. The selection operation selects the best individual from the current population based on a fitness function and generates a new solution based on the selected individual. The crossover operation generates a new solution by crossing the genes (i.e., resource allocation schemes) of two solutions, thereby exploring new solution spaces. The mutation operation increases the diversity of the population by randomly changing certain genes in the solution, preventing the algorithm from falling into a local optimum.

[0045] In some embodiments, a combination of particle swarm optimization and genetic algorithms can leverage the strengths of both. Particle swarm optimization can quickly find optimal solutions in a global search, while genetic algorithms can further improve the quality of solutions by performing local optimization. This combination can effectively avoid blind search and improve the accuracy and efficiency of resource allocation.

[0046] As an option, heuristic algorithms can be incorporated into the optimization process to further restrict or modify the solution space to ensure the rationality of the solution process. For example, heuristic algorithms can filter or adjust unreasonable solutions based on prior knowledge, further improving the efficiency and feasibility of the algorithm.

[0047] In practice, the combination of particle swarm optimization and genetic algorithms effectively solves the global optimization problem in resource scheduling. Through repeated iterative optimization, not only can the optimal state of computing resource allocation be achieved, but the execution time of computing tasks can also be effectively reduced. The optimized resource allocation results serve as input for subsequent steps, and are further used and scheduled by the information geometry optimization module, the game theory optimization module, and the optimal control module.

[0048] Step S3 further refines the geometric optimization of the optimized data, using information geometry optimization methods, specifically Riemannian geometry and Kullback-Leibler (KL) divergence, to optimize the data transmission paths between data sources. This process reduces redundant information in the data flow, ensuring efficient data transmission and optimizing the flow of resources between data sources. Specifically, step S3 aims to further structure the optimized data flow obtained in the previous step, ensuring the optimal data flow path and minimizing potential losses during the computation process.

[0049] In this embodiment, the information geometry optimization method first involves optimizing data flow paths. Riemannian geometry is a method that uses curvature and metrics to measure the shortest path between data flows. During data flow path optimization, a Riemannian metric is first defined to quantify the distance between data sources. By minimizing the length of the data flow path, data transmission is optimized. Specifically, the length of the data flow path γ can be expressed as: Where L represents the length of the data flow path; γ is the trajectory of the data flow; is the Riemannian metric, which is a function that describes the distance between data sources; is the tangent vector of the path. This formula essentially calculates the shortest transmission path from one data source to another. By minimizing the path length, we ensure optimal data flow transmission, reducing redundant information and transmission delays.

[0050] In general, using Riemannian geometry to optimize data flow paths can effectively reduce data transmission losses and improve data flow efficiency. Specifically, this optimization method can help to efficiently transmit information between multiple data sources, thereby maximizing overall system resource utilization.

[0051] Alternatively, in addition to using the Riemannian geometry metric, other geometric optimization methods can be used to further reduce redundancy in data transmission. For example, other information geometry metrics can be combined to optimize the transmission paths between data sources. This can further enhance optimization effectiveness, especially when dealing with complex multi-data source problems, where combining multiple geometric optimization methods can lead to an optimal solution.

[0052] In one possible implementation, KL divergence is introduced into the information geometry optimization process to measure the difference between two probability distributions. KL divergence (Kullback-Leibler-Divergence) is a metric that measures the similarity between two probability distributions. It can effectively assess the differences between data sources and minimize these differences during data flow. The calculation formula for KL divergence is as follows: Among them, D KL (p∥q) is the KL divergence; p and q are two probability distributions; p(x i ) and q(x i ) are the probability distributions of the two data sources; x i are samples from the data source. A smaller KL divergence indicates smaller differences between data sources, leading to higher data transmission efficiency. During optimization, the goal is to minimize the KL divergence, minimizing differences between data sources and thus optimizing the data flow transmission path.

[0053] In some embodiments, in practical applications, KL divergence can effectively help adjust the weights and data distribution between data sources to minimize information loss during data transmission. Combining KL divergence with Riemannian geometry can optimize both the data flow path and the transmission process, further improving overall resource utilization.

[0054] Specifically, an optimization method combining Riemannian geometry and KL divergence can effectively address resource scheduling and data flow issues across multiple data sources. Given limited computing resources, optimizing data transmission paths and minimizing differences between data sources are crucial. The information geometry optimization method in this embodiment, by effectively integrating these two optimization techniques, enables efficient data transmission between data sources.

[0055] Step S4 uses game theory to optimize resource allocation across multiple data sources and further coordinate resource allocation and task scheduling. In this step, game theory models are used to resolve resource conflicts between data sources and ensure that tasks within each data source are executed in a reasonable order. The introduction of game theory helps optimize the coordinated use of resources, enabling different data sources to achieve optimal resource allocation and task scheduling under multiple constraints.

[0056] In this embodiment, game theory optimization models the resource allocation of each data source by establishing a game model. In the model, each data source is considered a participant, and resource allocation and task scheduling are a game between these participants. The goal of the game is to allocate resources so that each data source can obtain the maximum utility under its constraints. The utility function u iIt represents the utility of the i-th data source, which is usually calculated based on the data source's computing requirements, resource usage, and task priority. The utility function is in the following form: u i =f(resource demand, task priority, resource usage); In this model, resource allocation for all data sources is subject to certain constraints, which typically include the total available resources, task execution time, and other system constraints. The goal of the game is to maximize the utility function of each data source, that is, to achieve the best resource allocation for each data source within its resource constraints.

[0057] Typically, the Nash equilibrium model is used to solve resource allocation problems in games. In a Nash equilibrium, each data source's decision is optimal, assuming the decisions of other data sources remain unchanged. In other words, each data source chooses a strategy that maximizes its own utility, given the strategies of other data sources. During the game, the system reaches a stable state when the utility of all data sources is maximized.

[0058] Alternatively, to ensure the rationality of the game process, a penalty or incentive mechanism can be introduced into the model. For example, if a data source's resource allocation does not meet the overall system needs, a penalty can be used to adjust its utility function, thereby guiding the data source to make reasonable decisions. This mechanism can effectively prevent unreasonable resource usage and promote resource sharing and fairness in the system.

[0059] Specifically, in a game theory model, each data source not only focuses on its own resource allocation but also needs to consider the strategies of other data sources. To achieve global optimality, each data source in the game must optimize its utility while avoiding excessive competition or conflict for resources. Game theory can provide a flexible and effective optimization solution for scheduling problems in multi-task, multi-resource scenarios.

[0060] One possible implementation approach is to coordinate multiple data sources using distributed algorithms based on game theory. Each data source makes decisions independently, but through information sharing and feedback mechanisms, all data sources can make decisions towards overall optimization. For example, using the Lagrange multiplier method within a distributed algorithm to address resource constraints, task scheduling and resource allocation can be jointly optimized to further improve efficiency.

[0061] In some embodiments, in environments with intense resource competition, the application of game models can help the system fairly allocate resources across multiple data sources and avoid resource conflicts. For example, when certain data sources are overloaded, the game model can dynamically adjust task scheduling strategies to allocate resources to tasks with lower loads, thereby balancing system resources and improving overall system efficiency.

[0062] Step S5 uses optimal control theory to dynamically schedule computing tasks in real time, based on the optimization results of the previous steps. This step aims to dynamically adjust the execution order of computing tasks based on real-time computing resource usage, task priority, and load status, to achieve optimal utilization of computing resources and ensure timely task completion.

[0063] In this embodiment, task scheduling uses dynamic programming and linear quadratic regulation (LQR) control algorithms for real-time scheduling. By evaluating current resources, the control system can dynamically adjust the execution order of tasks and resource allocation at each moment. Specifically, the dynamic programming algorithm takes into account the historical usage of computing resources and the priority of tasks, and uses the Bellman equation to optimize the task scheduling problem. The basic form of this equation is: V t (x) = min u (c(x,u)+V t+1 (f(x,u))); Among them, V t (x) is the minimum cost when the system is in state x at time t; u is the control input (i.e., the task scheduling policy); c(x,u) is the instantaneous cost of executing control u in state x; and f(x,u) is the state transition function, which describes the system state transition after task scheduling. Through the iterative process of dynamic programming, the optimal task scheduling solution can be obtained.

[0064] In general, task scheduling also requires consideration of task priorities, resource constraints, and real-time changes in computing resources. To ensure that tasks with higher priorities are completed in a timely manner, the linear quadratic control algorithm (LQR) is used for further optimization. The LQR control algorithm can adjust the scheduling order of tasks based on their priorities, minimizing delays and wasted computing resources. The objective function of the LQR control problem can be expressed as: Here, x(t) is the system state (i.e., the current state of the task); u(t) is the control input (i.e., the execution order of tasks or resource allocation); Q and R are weight matrices used to adjust the relative importance of the system state and control input. By optimizing the objective function, LQR control maximizes the efficiency of computing resources while ensuring that tasks are completed on time.

[0065] Alternatively, in practical applications, task scheduling strategies can be adjusted based on real-time feedback. For example, if the system detects that certain tasks are experiencing long execution delays, the task scheduling order can be dynamically adjusted using a real-time feedback mechanism to ensure timely execution of high-priority tasks. This can be achieved by introducing a feedback control loop that adjusts the scheduling strategy in real time based on system load and task status.

[0066] Specifically, task scheduling also involves the dynamic allocation of resources. When computing resources are limited, tasks with higher computational requirements are prioritized. Resource allocation is dynamically adjusted by evaluating task execution times, resource requirements, and inter-task dependencies. For example, strategies based on model predictive control (MPC) can be used to predict system resource requirements and schedule resources in advance. MPC models can predict future resource requirements and make task scheduling decisions in advance, reducing system load imbalances or resource shortages.

[0067] In one possible implementation, dynamic scheduling also incorporates task dependencies. When certain tasks depend on the completion of previous tasks, the scheduling strategy needs to account for these dependencies, ensuring that subsequent tasks can start immediately after the predecessor task completes. Modeling and scheduling these dependencies can improve system efficiency and ensure the correct execution order between tasks.

[0068] In some embodiments, the real-time feedback mechanism for task scheduling utilizes a distributed scheduling approach. In this approach, multiple computing nodes or tasks are coordinated through a distributed algorithm. Each node makes scheduling decisions based on local information and shares this information with other nodes through a communication mechanism. Distributed scheduling can effectively address the scheduling of large-scale computing tasks, and is particularly advantageous in multi-node and large-scale data analysis.

[0069] Example 2: Please see the attached Figure 2 , a bioinformatics data analysis system based on high-throughput sequencing, comprising: The data preprocessing module is used to clean high-throughput sequencing data, remove redundant data, select features, and calculate information entropy; the global optimization module is used to globally optimize data integration and resource allocation through particle swarm optimization and genetic algorithms; Information geometry optimization module, used to optimize data flow paths based on Riemannian geometry and Kullback-Leibler divergence; Game theory module, used to achieve resource coordination and task scheduling among multiple data sources through game models; Optimal control module, used for dynamic scheduling and optimal control of task execution based on the usage of computing resources; The system communication module is used to coordinate the data flow between modules to ensure the efficiency of the data processing process.

[0070] The data preprocessing module and the global optimization module transmit data through the calculated feature data and information gain results, ensuring that the global optimization module uses the preprocessed data for resource optimization and scheduling; The global optimization module and the information geometry optimization module transmit the globally optimized resource allocation data to provide the optimized data flow path to the information geometry module; The information geometry optimization module and the game theory module coordinate through the optimized data flow path to ensure that data flows efficiently along the optimized path and provide coordinated data for the game model; The game theory module and the optimal control module perform task scheduling based on the coordinated resource allocation results, and dynamically allocate resources through real-time scheduling feedback; The optimal control module conducts data transmission and task scheduling feedback with each module through the system communication module, and adjusts the task sequence and resource allocation in real time.

[0071] Specifically, the system's operational workflow is achieved through the coordinated work of various modules. First, the data preprocessing module receives raw high-throughput sequencing data and performs cleaning, redundant data removal, feature selection, and information entropy calculation. The cleaning step removes incomplete or abnormal data. Redundant data removal selects features that contribute significantly to classification by calculating information gain, while information entropy calculation helps assess data validity and information content, ensuring input data quality. The processed data is then transferred from the data preprocessing module to the global optimization module.

[0072] The global optimization module utilizes particle swarm optimization (PSO) and genetic algorithms (GA) to optimize resource allocation and task scheduling. PSO searches for the globally optimal resource allocation solution within the solution space based on the relationship between resource requirements and computational tasks. The genetic algorithm further refines the resource allocation solution and improves the solution quality through crossover and mutation operations. The optimized resource allocation data is then transmitted to the information geometry optimization module.

[0073] The information geometry optimization module optimizes data flow paths based on Riemannian geometry and Kullback-Leibler divergence. Riemannian geometry ensures the most efficient data transfer from one source to another by minimizing the length of the data flow path. Kullback-Leibler divergence measures the differences between data sources and ensures the quality and consistency of data transmission by minimizing these differences. The optimized data flow paths are then fed into the game theory module.

[0074] The game theory module coordinates resource allocation across multiple data sources using a game model based on resource allocation and task scheduling requirements. During this game, each data source makes decisions based on its own needs and task priorities. The system ensures optimal resource allocation across all data sources through Nash equilibrium. The optimized resource allocation results are then transmitted to the optimal control module.

[0075] The optimal control module dynamically adjusts the execution order of tasks by analyzing current computing resource usage. Combined with real-time scheduling feedback, the optimal control module can adjust the execution order of tasks based on changes in system load and task priority, ensuring that high-priority tasks are completed first and dynamically allocating computing resources. Finally, the system communication module coordinates data flow between modules, ensuring timely information transmission, smooth task scheduling, and efficient system operation.

[0076] During the entire system operation, various modules work closely together, from data preprocessing to resource scheduling and task execution optimization. Each step ensures that the system's resources are maximized and tasks are completed on time.

[0077] Example 3: A storage medium stores a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

[0078] Specifically, the storage medium can be any medium capable of storing computer programs, such as a hard disk, solid-state drive, optical disk, USB flash drive, or cloud storage. The program code stored on the medium includes programs that implement functions such as preprocessing of high-throughput sequencing data, global optimization, information geometry optimization, game theory optimization, optimal control, and task scheduling.

[0079] In a specific implementation, the computer program in the storage medium includes multiple modules, each responsible for a different function. For example, the program code of the data preprocessing module is responsible for cleaning sequencing data, removing redundant data, performing feature selection, and calculating information entropy; the program code of the global optimization module optimizes resource allocation and task scheduling using particle swarm optimization (PSO) and genetic algorithms (GA); the program code of the information geometry optimization module optimizes data flow paths based on Riemannian geometry and KL divergence; the program code of the game theory module implements resource coordination and task scheduling among multiple data sources; and the program code of the optimal control module adjusts the task sequence based on real-time resource usage to ensure that tasks are completed according to priority. All of these program codes are stored in the storage medium and executed sequentially under the control of the processor.

[0080] In one possible implementation, a computer program in a storage medium completes tasks by accessing data and intermediate results stored in memory. By reading, writing, and updating data, the program controls the coordinated work between modules to ensure efficient processing from pre-processing to final dispatch.

[0081] In some embodiments, the storage medium may store the source code of the computer program, the algorithm library, the raw data to be processed, the optimized results, and the intermediate calculation results. During execution, the system reads data from the storage medium according to the needs of task scheduling and sends the data to the various processing modules based on the computing tasks and computing resources, completing an efficient data analysis and optimization process.

[0082] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A bioinformatics data analysis method based on high-throughput sequencing, characterized in that: The following steps are involved: S1. Preprocessing of high-throughput sequencing data, including data cleaning, redundant data removal, and feature selection; S2, using particle swarm optimization and genetic algorithms to perform global optimization, optimizing resource allocation and computing task scheduling during the integration of multiple data sources; S3, based on information geometry optimization method, uses Riemannian geometry and Kullback-Leibler divergence to optimize the data transmission path between data sources; S4. Use game theory methods to optimize resource allocation of multiple data sources and achieve coordination between data sources through game models; S5. Combined with optimal control theory, computing tasks are dynamically scheduled in real time, and the execution order of tasks is adjusted according to the usage of computing resources.

2. The method for analyzing bioinformatics data based on high-throughput sequencing according to claim 1, wherein: In the step S1, redundant data removal and feature selection include the following steps: Calculate the overall entropy of the data set, select features with higher information gain for data integration, and remove redundant features with lower information gain; Calculate the gain of each feature, and perform feature selection based on the size of the information gain to remove redundant features.

3. The method for analyzing bioinformatics data based on high-throughput sequencing according to claim 1, wherein: In the S2 step, the particle swarm optimization specifically updates the speed and position of each particle, and updates the speed and position of the particle using the following formula: v i (t+1)=wv i (t)+c1r1(p i -x i (t))+c2r2(g i -x i (t)); x i (t+1)=x i (t)+v i (t+1); Among them, v i (t) is the velocity of the particle; x i (t) is the position of the particle; p i is the best historical position of the particle; g i is the optimal position of the group; w is the inertia weight; c1, c2 are learning factors; r1, r2 are random numbers.

4. The method for analyzing bioinformatics data based on high-throughput sequencing according to claim 1, wherein: In step S2, the genetic algorithm optimizes resources based on the fitness function of each data source. The fitness function is expressed by weighted summation: Among them, w j is the weight of each data source; x j It is the corresponding resource allocation.

5. The method for analyzing bioinformatics data based on high-throughput sequencing according to claim 1, wherein: In the S3 step, information geometry optimization minimizes the path length of data flow based on Riemannian geometry and minimizes the differences between data sources through Kullback-Leibler divergence to minimize redundant information in the data integration process. The path length and divergence are calculated as follows: Among them, γ is the data flow path; is the Riemannian metric; p(x i ) and q(x i ) are the probability distributions of the two data sources respectively.

6. The method for analyzing bioinformatics data based on high-throughput sequencing according to claim 1, wherein: In step S4, game theory is used to model the resource allocation of multiple data sources. The Nash equilibrium model is used to coordinate resources between data sources and ensure that the computing tasks of each data source are optimally executed, meeting the following conditions: Among them, u i is the utility function of data source i; It is the optimal choice of strategy.

7. The method for analyzing bioinformatics data based on high-throughput sequencing according to claim 1, wherein: The S5 step further includes the following steps: Adopting dynamic programming and linear quadratic control algorithms, the task execution order is dynamically adjusted according to the current computing resource usage, task execution priority and system load; Select preemptive or non-preemptive scheduling strategies based on task priority, computing resource requirements, and real-time resource load, delaying the execution of low-priority tasks while high-priority tasks are executing. Use real-time feedback mechanisms to adjust task scheduling strategies based on the current system status, and use resource prediction algorithms to predict future resource requirements and adjust the order of task execution in advance. When computing resources are limited, a heuristic algorithm is used to adjust the task execution order based on the task dependencies and execution priorities, giving priority to tasks with higher resource requirements and higher priorities.

8. A bioinformatics data analysis system based on high-throughput sequencing, according to a bioinformatics data analysis method based on high-throughput sequencing according to any one of claims 1 to 7, characterized in that: include: Data preprocessing module, used for cleaning high-throughput sequencing data, removing redundant data, feature selection and information entropy calculation; Global optimization module, used to perform global optimization of data integration and resource allocation through particle swarm optimization and genetic algorithm; Information geometry optimization module, used to optimize data flow paths based on Riemannian geometry and Kullback-Leibler divergence; Game theory module, used to achieve resource coordination and task scheduling among multiple data sources through game models; Optimal control module, used for dynamic scheduling and optimal control of task execution based on the usage of computing resources; The system communication module is used to coordinate the data flow between modules to ensure the efficiency of the data processing process.

9. The bioinformatics data analysis system based on high-throughput sequencing according to claim 8, characterized in that: The data preprocessing module and the global optimization module transmit data through the calculated feature data and information gain results, ensuring that the global optimization module uses the preprocessed data for resource optimization and scheduling; The global optimization module and the information geometry optimization module transmit the globally optimized resource allocation data to provide the optimized data flow path to the information geometry module; The information geometry optimization module and the game theory module coordinate through the optimized data flow path to ensure that data flows efficiently along the optimized path and provide coordinated data for the game model; The game theory module and the optimal control module perform task scheduling based on the coordinated resource allocation results, and dynamically allocate resources based on real-time scheduling feedback; The optimal control module performs data transmission and task scheduling feedback with each module through the system communication module, and adjusts the task sequence and resource allocation in real time.

10. A storage medium, characterized in that: A computer program is stored thereon, and when executed by a processor, the program implements the steps of the method according to any one of claims 1 to 7.