Bombyx mori gene sequencing path optimization system and method based on hybrid parallel genetic algorithm

Through mixed parallel genetic algorithms, optimize the sequencing path of silkworm genes, combined with technologies such as convolutional neural networks and deep Q networks, the problems of difficulty in assembling repeat regions in the silkworm genome and accumulation of heterogeneity errors in sequencing data are solved, and efficient and accurate genome assembly is achieved.

CN120412718AInactive Publication Date: 2025-08-01YANCHENG TEACHERS UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510447294.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

There are problems in the silkworm genome where there is difficulty in assembling repeat regions with high complexity and the accumulation of errors due to the heterogeneity of sequencing data.

Method used

A silkworm gene sequencing path optimization system based on hybrid parallel genetic algorithm is adopted, including data preprocessing and feature modeling, hybrid parallel genetic algorithm optimization, dynamic path planning and resource allocation, multi-omics verification and result output, and efficient assembly is achieved using technologies such as convolutional neural network, generative adversarial network, deep Q network and Kubernetes.

Benefits of technology

Effectively avoiding the algorithm from falling into local optimality, improving the assembly success rate of silkworm gene sequencing in complex areas, reducing resource waste, improving assembly efficiency and accuracy, and reducing time and resource costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412718A_ABST
    Figure CN120412718A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of gene sequencing, in particular to a silkworm gene sequencing path optimization system and method based on a hybrid parallel genetic algorithm. According to the technical scheme, the method comprises the steps of data preprocessing and feature modeling, hybrid parallel genetic algorithm optimization, dynamic path planning and resource allocation and multi-omics verification and result output. Reference is provided for assembling path planning by recognizing the repeated area, local optimum is avoided by means of the hybrid parallel genetic algorithm, efficient assembling of the high-complexity repeated area is achieved, standardization processing is carried out on data, errors caused by data differences are reduced, and the accuracy of assembling path planning is improved. The multi-omics feedback correction module integrates transcriptome and epigenetic data, corrects an assembly result and inhibits error accumulation, and meanwhile, based on deep Q network model reinforcement learning and a dynamic resource scheduling module, assembly strategies and resource allocation are dynamically adjusted, efficient utilization of computing resources is achieved, and resource waste is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gene sequencing, and in particular to a system and method for optimizing the gene sequencing path of silkworms based on a hybrid parallel genetic algorithm. Background Art

[0002] The mulberry silkworm, also known as the domestic silkworm, is a silk-spinning insect with high economic value. By analyzing the genome of the silkworm, it is possible to deeply understand the genetic basis of important physiological processes such as the growth, development, reproduction, and cocoon silk production of the silkworm, which helps to breed silkworm varieties with higher yields, better quality, and stronger stress resistance, thereby increasing the yield and quality of cocoons and promoting the sustainable development of the sericulture industry. Due to the rich long fragment repeat sequences in the silkworm genome, such as transposons and telomere regions, there are problems in the assembly link of current silkworm gene sequencing, such as difficulty in assembling high-complexity repeat regions and error accumulation caused by the heterogeneity of sequencing data. Therefore, we propose a system and method for optimizing the gene sequencing path of silkworms based on a hybrid parallel genetic algorithm. Summary of the Invention

[0003] The purpose of the present invention is to address the problems in the background art and propose a system and method for optimizing the gene sequencing path of silkworms based on a hybrid parallel genetic algorithm.

[0004] In a first aspect, the present invention proposes a method for optimizing the gene sequencing path of silkworms based on a hybrid parallel genetic algorithm, including the following steps:

[0005] Data preprocessing and feature modeling: constructing a sequence feature map through data cleaning, noise reduction, and feature extraction;

[0006] Hybrid parallel genetic algorithm optimization: optimizing the assembly path through an intelligent algorithm that integrates genetic algorithms, simulated annealing algorithms, and particle swarm algorithms based on the preprocessed and feature-modeled data;

[0007] Dynamic path planning and resource allocation: real-time monitoring of the Q30 value, repeat density of the data, and the utilization rate of computing resources, and using the deep Q network algorithm for reinforcement learning decision-making, and allocating computing resources according to the decision results;

[0008] Multi-omics verification and result output: integrating transcriptome and epigenetic data, correcting the assembly results obtained after optimization and resource allocation, and outputting the final assembly results and visualization reports.

[0009] Optionally, the data preprocessing and feature modeling include the following steps:

[0010] Filter reads with base quality values less than 30 through the HMM model: Read the original sequencing data file from the storage device, initialize the parameters of the HMM model. For each sequencing read, calculate its probabilities in different hidden states using the HMM model. By recursively calculating the probability values at each position, concentrate and filter the reads with base quality values less than 30, and save the filtered read data to a data file. The recurrence formula for forward propagation calculation is as follows:

[0011]

[0012] where, α t (j) represents the probability of being in state j at time t and observing the sequence o1, o2, …, o t The probability of, a ij is the state transition probability, b j (o t ) is the observation probability;

[0013] Convert the processed sequence data into a k-mer frequency matrix and input it into a convolutional neural network model to output a repeat density probability map: Divide each of the filtered reads into k-mers of a fixed length, count the occurrence frequencies of each k-mer in all reads, form a matrix with all the occurrence frequency values as the input of the convolutional neural network model, and perform forward propagation calculation through the convolutional neural network model to output a repeat density probability map. The convolution operation formula for the convolutional layer of the convolutional neural network model is:

[0014]

[0015] where, x is the input feature map, w is the convolution kernel, b is the bias, and y is the output feature map;

[0016] Generate enhanced data through a generative adversarial network to expand the training set: Generate a random noise vector as the input of the generator of the generative adversarial network, generate simulated sequencing sequences with a controllable error rate, input the simulated sequencing sequences and real sequencing sequences into the discriminator for adversarial training, and merge the simulated sequencing sequences with the original sequencing data and add them to the training set.

[0017] Optionally, the hybrid parallel genetic algorithm optimization includes the following steps:

[0018] Initialize the population: Based on the preprocessed data, construct contig nodes, calculate the connection edge weights between the contig nodes. The calculation formula is as follows:

[0019]

[0020] where, OverlapLength i,jis the overlap length of contigs i and j, Length i and Length j are the lengths of contigs i and j respectively;

[0021] Randomly generate a number of chromosomes using the dynamic decision tree encoding method. Each of the chromosomes includes a plurality of the contig nodes and connection edges, and save the generated chromosomes into a population as an initial population;

[0022] CPU parallel computing fitness: Use a distributed computing framework to distribute the chromosome data in the initial population to a number of GPU nodes, calculate the N50, Consistency, Time, and Core values of each chromosome according to the multi-objective optimization function, and calculate the fitness value. The formula of the multi-objective optimization function is expressed as:

[0023]

[0024] Among them, α, β, and γ are weight coefficients. N50 represents the length of the contig corresponding when the cumulative length reaches half of the total length after sorting all contigs by length from largest to smallest. MaxLen is the maximum possible assembly length. Consistency represents the consistency between the assembly result and the original reads. Reads is the number of original reads. Time is the calculation time. Core is the number of computing cores used;

[0025] And summarize the fitness values to the master node, and analyze the fitness values through the master node to obtain the optimal chromosome;

[0026] Perform global search based on the particle swarm optimization algorithm and introduce the simulated annealing algorithm to assist mutation: Initialize each chromosome corresponding to a particle, update the speed and position of each particle according to the particle speed update formula. The particle speed update formula is:

[0027]

[0028] Among them, is the speed of particle i in the t-th generation, ω is the inertia weight, c1 and c2 are learning factors, r1 and r2 are random numbers, p best is the individual optimal position of the particle, g best is the global optimal position, is the position of particle i in the t-th generation;

[0029] Perform simulated annealing mutation operation on each particle using the simulated annealing algorithm, and decide whether to accept the mutated particle according to the simulated annealing probability formula. The simulated annealing probability formula is:

[0030] P = e-ΔE / T

[0031] Among them, ΔE is the change in fitness, and T is the temperature parameter. During the simulated annealing process, the temperature T is inversely proportional to the number of iterations. The update operation based on the particle swarm optimization algorithm and the simulated annealing mutation operation are repeated.

[0032] Optionally, the dynamic path planning and resource allocation include the following steps:

[0033] Data collection, storage, and analysis: The Q30 value, repeat density, and utilization rate of computing resources of the data are collected in real time. Among them, the resource utilization rate is obtained by using a system monitoring tool, the Q30 value and repeat density are obtained by real-time analysis of the data, the collected data information is stored in a database, and the data information is analyzed by a data analysis tool;

[0034] Reinforcement learning decision based on the deep Q-network model: The Q30 value, repeat density, and utilization rate of computing resources are used as the state input of the deep Q-network model. The deep Q-network model uses the Softmax strategy for action selection. The actions include algorithm switching and priority adjustment. The formula is expressed as follows:

[0035]

[0036] Among them, π(a|s) represents the probability of selecting action a in state s, Q(s,a) represents the Q value of taking action a in state s, and τ is the temperature parameter;

[0037] Reward calculation: A multi-dimensional reward function is constructed. The formula is expressed as follows:

[0038]

[0039] Among them, r is the reward value, ω1, ω2, ω3, ω4 are weight coefficients, ΔN50 represents the change in the N50 value before and after taking the action, ΔConsistency represents the change in the consistency between the assembly result and the original read segment before and after taking the action, and ΔResourceUtilization represents the change in the resource utilization rate before and after taking the action;

[0040] Model update: According to the reward value and action selection, the parameters of the deep Q-network model are updated by using a reinforcement learning algorithm. The Q-value update formula of the deep Q-network model is:

[0041]

[0042] Among them, s t is the state, a t is the action, r t is the reward, α is the learning rate, and γ is the discount factor;

[0043] Dynamic resource allocation: According to the decision results of the reinforcement learning algorithm, the tasks are scheduled through Kubernetes. High-priority tasks are preferentially allocated to available GPU and CPU resources, and the usage of CPU and GPU resources is monitored in real time to manage the execution status of the tasks.

[0044] Optionally, the multi-omics verification and result output include the following steps:

[0045] Verify exon continuity through the STAR alignment tool: Compare the transcriptome data with the assembly results, process and analyze the alignment result file using bioinformatics tools, and correct the assembly results according to the analysis results. The correction methods include adjusting the assembly path, adding or deleting contigs;

[0046] Merge methylation coverage constraint paths: Process the methylation site data using the bioinformatics tools, screen out the assembly paths according to the methylation site coverage, merge the screened assembly paths to form the final assembly result, and verify the final assembly result using the bioinformatics tools;

[0047] Output FASTA file and visualization report: Use the BioPython library in Python to save the final assembly result as a FASTA file, and generate a visualization report through the Matplotlib library in Python. The visualization report includes an assembly path map, N50 value, and error rate heat map, and save the visualization report as a PDF or HTML file.

[0048] In a second aspect, the present invention proposes a silkworm gene sequencing path optimization system based on a hybrid parallel genetic algorithm, including:

[0049] An input / output interface module for system-external data interaction;

[0050] A data processing module, which includes a data preprocessing sub-module, a feature extraction sub-module based on a convolutional neural network model, and a data augmentation sub-module;

[0051] An optimization module based on a hybrid parallel genetic algorithm, which includes an initialization sub-module, a fitness calculation sub-module, and a hybrid optimizer;

[0052] A dynamic resource scheduling module based on Kubernetes, which is used to allocate resources to tasks according to task priorities;

[0053] Multi-omics feedback correction module, which is used to perform feedback correction on the gene sequence assembly result based on the exon continuity alignment result of the transcriptome data and the methylation site coverage

[0054] Optionally, the data preprocessing sub-module includes cleaning and denoising the data using the HMM model, and the formula is expressed as:

[0055] P(O|λ)=∑ q P(O|q,λ)P(q|λ)

[0056] where O is the original sequencing data, q is the hidden state representing the true base sequence, λ is the model parameter, and the Baum-Welch algorithm is used to train the HMM model. By iteratively updating the model parameters, the probability P(O|λ) of the observation sequence is maximized.

[0057] Optionally, the feature extraction sub-module includes minimizing the loss function during the training process of the convolutional neural network model, and the formula of the loss function is expressed as:

[0058]

[0059] where d(x,y) is the binary mask of the labeled repeated region, and p(x,y) is the predicted probability.

[0060] Optionally, the data augmentation sub-module includes a generative adversarial network algorithm, which is used to generate simulated sequencing sequences with a controllable error rate to expand the training set. The formula of the adversarial loss function is expressed as:

[0061]

[0062] where p data (x) is the distribution of the real data, p z (z) is the distribution of the noise. The generative adversarial algorithm includes a discriminator and a generator. The goal of the discriminator is to maximize V(D,G), and the goal of the generator is to minimize V(D,G).

[0063] Optionally, the population initialization sub-module includes encoding chromosomes using a dynamic decision tree. The chromosome is represented as C={(n i ,w ij )}, where n i is the contig node, and w ij is the connection edge weight.

[0064] In summary, the present application includes at least one of the following beneficial technical effects:

[0065] The present invention uses a convolutional neural network to identify repetitive regions, providing a reference for assembly path planning. By means of a hybrid parallel genetic algorithm that combines a genetic algorithm, a simulated annealing algorithm, and a particle swarm optimization algorithm, it effectively avoids the algorithm from falling into local optima, realizes efficient assembly of repetitive regions with high complexity, and improves the assembly success rate of silkworm gene sequencing in complex regions.

[0066] Furthermore, the input-output interface module in the present invention standardizes multi-source heterogeneous sequencing data, reducing errors caused by data differences from the source. The data quality is improved through the data processing module, and the training set is expanded using a generative adversarial network to enhance the adaptability of the model to different data features. In addition, the multi-omics feedback correction module integrates transcriptome and epigenetic data to correct the assembly results, further reducing errors and effectively suppressing error accumulation caused by the heterogeneity of sequencing data.

[0067] Moreover, the present invention adopts deep Q-network model-based reinforcement learning and a dynamic resource scheduling module based on Kubernetes. It can dynamically adjust the assembly strategy and resource allocation according to the real-time data status and resource usage, realizing the efficient utilization of computing resources and reducing resource waste. At the same time, each module of the system works together, comprehensively improving the efficiency and accuracy of silkworm gene sequencing assembly from data processing, algorithm optimization to resource scheduling, reducing time and resource costs, and providing efficient and reliable technical support for silkworm gene research. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] Figure 1 A flowchart of a method for optimizing the silkworm gene sequencing path based on a hybrid parallel genetic algorithm according to the present invention is given.

[0069] Figure 2 A structural block diagram of a system for optimizing the silkworm gene sequencing path based on a hybrid parallel genetic algorithm according to the present invention is given. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] Embodiment

[0072] A system and method for optimizing the silkworm gene sequencing path based on a hybrid parallel genetic algorithm proposed by the present invention, as Figure 1 shown, includes the following steps:

[0073] I. Data preprocessing and feature modeling

[0074] 1. Data collection and input: Staff use a sequencing platform to sequence the 12th chromosome of the silkworm. The preferred sequencing platforms are Illumina HiSeq and PacBio Sequel. The collected data is stored in FASTQ and FASTA file formats and is transmitted to the input / output interface module. The input / output interface module uses a regular expression-based FASTQ / FASTA parser to standardize the data, unify the data format, and convert it into an internal data structure recognizable by the system, providing a data basis for subsequent processing;

[0075] 2. Data preprocessing and feature extraction

[0076] Data cleaning and noise reduction: The data processing module uses the HMM model to clean the standardized data. The HMM model formula is expressed as: P(O|λ) = ∑ q P(O|q,λ)P(q|λ), where O is the original sequencing data, q is the hidden state representing the true base sequence, λ is the model parameter, and the Baum-Welch algorithm is used to train the model parameter. By iteratively updating the model parameter, the probability P(O|λ) of the observation sequence is maximized;

[0077] For each sequencing read, use the HMM model to calculate its probability in different hidden states. By recursively calculating the probability values at each position, the reads with base quality values less than 30 are filtered out centrally. The filtered read data is saved to a data file. The recurrence formula for forward propagation calculation is:

[0078]

[0079] where, α t (j) represents the probability of being in state j at time t and observing the sequence o1, o2,..., o t x a ij is the state transition probability, and b j (o t ) is the observation probability.

[0080] 3. Repetitive region feature extraction: Convert the cleaned data into a sequence k-mer frequency matrix and input it into a convolutional neural network model. Use the convolution operation formula of the convolutional layer to extract data features and output a repetitive density probability map, providing key information for subsequent assembly path planning. The convolution operation formula of the convolutional layer of the convolutional neural network model is:

[0081]

[0082] where x is the input feature map, w is the convolution kernel, b is the bias, and y is the output feature map;

[0083] Data augmentation: Use a generative adversarial network for data augmentation. The generator uses an LSTM network. According to the calculation formula of the LSTM unit, it generates simulated sequencing sequences with a controllable error rate. The generator and discriminator are trained through an adversarial loss function, and the formula of the adversarial loss function is expressed as:

[0084]

[0085] Among them, p data (x) is the distribution of real data, p z (z) is the distribution of noise. The generative adversarial algorithm includes a discriminator and a generator. The goal of the discriminator is to maximize V(D, G), and the goal of the generator is to minimize V(D, G). The simulated sequences are combined with the real data to expand the training set, thereby enhancing the generalization ability of the model.

[0086] II. Optimization of the hybrid parallel genetic algorithm

[0087] 1. Population initialization: Based on the preprocessed data, construct contig nodes, calculate the connection edge weights between contig nodes according to the formula, and use the dynamic decision tree encoding method to randomly generate 100 chromosomes to construct the initial population. The formula for calculating the connection edge weights is:

[0088]

[0089] Among them, OverlapLength i,j is the overlap length of contigs i and j, Length i and Length j are the lengths of contigs i and j respectively;

[0090] 2. Fitness calculation: Distribute the chromosome data in the initial population to 10 GPU nodes, calculate the N50, Consistency, Time, and Core values of each chromosome according to the multi-objective optimization function, and substitute them into the formula of the multi-objective optimization function to calculate the fitness value. The formula of the multi-objective optimization function is expressed as:

[0091]

[0092] Among them, α, β, and γ are weight coefficients. N50 represents the contig length corresponding to when the cumulative length reaches half of the total length after sorting all contigs from largest to smallest by length. MaxLen is the maximum possible assembly length. Consistency represents the consistency between the assembly result and the original reads. Reads is the number of original reads. Time is the calculation time. Core is the number of computing cores used. The weight coefficients are dynamically adjusted through gradient descent to balance the assembly quality and resource consumption. The calculated fitness values are aggregated to the master node, and the optimal chromosome is obtained through the analysis of the fitness values by the master node;

[0093] 3. Hybrid optimizer operation: Initialize each chromosome corresponding to a particle, and update the velocity and position of each particle according to the particle velocity update formula. The particle velocity update formula is:

[0094]

[0095] Among them, is the velocity of particle i at the t-th generation, ω is the inertia weight, c1 and c2 are learning factors, r1 and r2 are random numbers, p best is the individual optimal position of the particle, g best is the global optimal position, is the position of particle i at the t-th generation;

[0096] Perform simulated annealing mutation operation on each particle using the simulated annealing algorithm, and determine whether to accept the mutated particle according to the simulated annealing probability formula. The simulated annealing probability formula is:

[0097] P = e -ΔE / T

[0098] Among them, ΔE is the change in fitness, T is the temperature parameter. During the simulated annealing process, the temperature T is inversely proportional to the number of iterations. Repeat the update operation based on the particle swarm optimization algorithm and the simulated annealing mutation operation to avoid the algorithm falling into a local optimum.

[0099] The present invention uses a convolutional neural network to identify repetitive regions, providing a reference for the assembly path planning, and by means of a hybrid parallel genetic algorithm that combines a genetic algorithm, a simulated annealing algorithm, and a particle swarm optimization algorithm, effectively avoids the algorithm falling into a local optimum, realizes the efficient assembly of highly complex repetitive regions, and improves the assembly success rate of silkworm gene sequencing in complex regions.

[0100] III. Dynamic Path Planning and Resource Allocation

[0101] 1. Real-time data monitoring: The Q30 value, repetition density, and utilization rate of computing resources of the real-time collected data are obtained. Among them, the resource utilization rate is obtained by using a system monitoring tool, the Q30 value and repetition density are obtained by performing real-time analysis on the data. The collected data information is stored in a database, and the data information is analyzed by a data analysis tool. The Q30 value, repetition density, and utilization rate of the data are used as the state input of the deep Q-network model and normalized.

[0102] 2. Reinforcement learning decision-making: Based on the current state, the deep Q-network model uses the Softmax strategy to select actions. The actions include algorithm switching and priority adjustment, and the formula is expressed as:

[0103]

[0104] where π(a|s) represents the probability of selecting action a in state s, Q(s,a) represents the Q value of taking action a in state s, and τ is the temperature parameter;

[0105] The reward value is calculated according to the multi-dimensional reward function, and the model parameters are updated. The formula is expressed as follows:

[0106]

[0107] where r is the reward value, ω1, ω2, ω3, ω4 are weight coefficients, ΔN50 represents the change in the N50 value before and after taking the action, ΔConsistency represents the change in the consistency between the assembly result and the original read segment before and after taking the action, ΔResourceUtilization represents the change in the resource utilization rate before and after taking the action. According to the reward value and action selection, the parameters of the deep Q-network model are updated using the reinforcement learning algorithm. The Q value update formula of the deep Q-network model is:

[0108]

[0109] where s t is the state, a t is the action, rx is the reward, α is the learning rate, and γ is the discount factor;

[0110] 3. Dynamic resource allocation: According to the decision result of the reinforcement learning algorithm, Kubernetes schedules tasks according to the decision result of the deep Q-network model, dynamically allocates GPU and CPU resources, and ensures that each task obtains sufficient resources and avoids resource waste by setting resource request and limit parameters.

[0111] In the present invention, a deep Q-network model-based reinforcement learning and a dynamic resource scheduling module based on Kubernetes are adopted, which can dynamically adjust the assembly strategy and resource allocation according to the real-time data status and resource usage, realizing the efficient utilization of computing resources and reducing resource waste.

[0112] IV. Multi-omics Verification and Result Output

[0113] 1. Transcriptome Correction: Prepare the transcriptome data of the 12th chromosome of the silkworm, use the STAR alignment tool to align it with the assembly result, analyze the alignment result file, check whether the splicing of gene exons is correct, and correct the assembly result according to the analysis result. The correction methods include adjusting the assembly path, adding or deleting contigs.

[0114] 2. Merging of Methylation Coverage Constraint Paths: Use bioinformatics tools to process the methylation site data, screen out the assembly paths according to the methylation site coverage rate, merge the screened assembly paths to form the final assembly result, and verify the final assembly result through bioinformatics tools.

[0115] 3. Result Output: Use the BioPython library of Python to save the final assembly result as a FASTA file, and generate a visualization report through the Matplotlib library of Python. The visualization report includes an assembly path map, N50 value, and error rate heat map, and save the visualization report as a PDF or HTML file to provide intuitive and comprehensive information for researchers.

[0116] On the other hand, the present invention provides a silkworm gene sequencing path optimization system based on a hybrid parallel genetic algorithm, as Figure 2 shown, which includes an input / output interface module, and also includes a data processing module, an optimization module based on a hybrid parallel genetic algorithm, a dynamic resource scheduling module based on Kubernetes, and a multi-omics feedback correction module. The multi-omics feedback correction module is used to feedback and correct the gene sequence assembly result based on the exon continuity alignment result of the transcriptome data and the methylation site coverage rate.

[0117] Among them, the input / output interface module is used for the system to interact with external data. The input / output interface module includes a FASTQ / FASTA parser based on regular expressions. The FASTQ file contains sequencing sequences and their quality information, while the FASTA file mainly contains sequence information. For the FASTQ file, the regular expression matches the sequence identifier line starting with "@" and the quality information separator line starting with "+", thereby extracting the sequences and quality information in the file and converting them into a unified data structure within the system for convenient subsequent processing. By standardizing multi-source heterogeneous sequencing data through the input / output interface module, errors caused by data differences are reduced at the source.

[0118] The data processing module includes a data preprocessing sub-module, a feature extraction sub-module based on a convolutional neural network model, and a data augmentation sub-module. The data preprocessing sub-module includes using an HMM model to clean and denoise the data, and the formula is expressed as:

[0119]

[0120] Among them, O is the original sequencing data, q is the hidden state, representing the true base sequence, λ is the model parameter, and the model parameter λ represents the probability distribution of the model, including the state transition probability and the observation probability. The Baum-Welch algorithm is used to train the HMM model, and by iteratively updating the model parameters, the probability P(O|λ) of the observation sequence is maximized. The data processing module realizes improving the data quality and uses a generative adversarial network to expand the training set to enhance the adaptability of the model to different data features.

[0121] The convolutional neural network model in the feature extraction sub-module automatically learns the feature patterns in the data through the combination of convolutional layers, pooling layers, and fully connected layers. In this embodiment, the convolutional neural network model is used to identify the repetitive regions in the sequencing sequences. During the training process of the convolutional neural network model, the loss function is minimized to adjust the network parameters, and the loss function formula is expressed as:

[0122]

[0123] Among them, d(x,y) is the binary mask of the labeled repetitive region, p(x,y) is the predicted probability, and the binary mask of the labeled repetitive region is the training data with labels. The training data is input into the network, the predicted probability p(x,y) is calculated, and the loss value is calculated according to the loss function, and the network parameters are updated using the backpropagation algorithm, and the iteration continues until the loss value converges.

[0124] The data augmentation sub-module includes a generative adversarial network algorithm. The generative adversarial algorithm is used to generate simulated sequencing sequences with a controllable error rate to expand the training set, and the adversarial loss function formula is expressed as:

[0125]

[0126] Among them, p data (x) is the distribution of real data, and p z (z) is the distribution of noise. The generative adversarial algorithm includes a discriminator and a generator. The goal of the discriminator is to maximize V(D, G), and the goal of the generator is to minimize V(D, G). The input of the generator is a random noise vector z, and the output is the simulated sequencing sequence G(z). During the generation process, by controlling the distribution of the noise vector and the network parameters, the control of the error rate is achieved. The discriminator D is a neural network used to determine whether the input sequence is a real sequencing sequence x or a generated simulated sequence G(z). During the training process, the generator and the discriminator compete with each other to optimize the performance.

[0127] The optimization module based on the hybrid parallel genetic algorithm includes an initialization sub-module, a fitness calculation sub-module, and a hybrid optimizer. The population initialization sub-module includes encoding chromosomes using a dynamic decision tree. The chromosome is represented as C = {(n i , w ij )}, where n i is the contig node, and w ij is the connection edge weight. The fitness calculation sub-module calculates the fitness value through a multi-objective optimization function and iterates multiple times. During the iteration process, the parameters are adjusted through the gradient descent algorithm to achieve the optimization of the fitness value. The hybrid optimizer avoids local optima through simulated annealing mutation operations and performs global search based on the particle swarm optimization algorithm, searching for the optimal solution by updating the velocity and position.

[0128] The dynamic resource scheduling module is used to allocate resources to tasks according to the task priority. The dynamic resource scheduling module adopts a priority queue and a resource preemption mechanism. Among them, the priority queue assigns priorities to each task according to the importance and computational complexity of the tasks. The computational tasks in the high-repeat regions are assigned high priorities because of their high complexity and great impact on the accuracy of the results. The resource preemption mechanism allows high-priority tasks to preempt the resources of low-priority tasks in case of resource shortage.

[0129] The multi-omics feedback correction module is used to integrate transcriptome and epigenetic data to correct the gene assembly results, thereby improving the accuracy and reliability of silkworm gene sequencing. The multi-omics feedback correction module verifies exon continuity through the STAR alignment tool and combines methylation coverage constraint paths to form the final assembly result, and finally outputs a FASTA file and a visualization report. The multi-omics feedback correction module realizes the integration of transcriptome and epigenetic data to correct the assembly results, further reducing errors and effectively suppressing the error accumulation caused by the heterogeneity of sequencing data. At the same time, each module of the system works together to comprehensively improve the efficiency and accuracy of silkworm gene sequencing and assembly from data processing, algorithm optimization to resource scheduling, reduce time and resource costs, and provide reliable technical support for silkworm gene research.

[0130] The above specific embodiments are only several optional embodiments of the present invention. Based on the technical solution of the present invention and the relevant revelations of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. An optimization method for the gene sequencing path of silkworms based on a hybrid parallel genetic algorithm, characterized in that, It includes the following steps: Data preprocessing and feature modeling: By cleaning, denoising the data and extracting features, construct a sequence feature map; Hybrid parallel genetic algorithm optimization: Based on the preprocessed and feature-modeled data, optimize the assembly path through an intelligent algorithm that integrates genetic algorithm, simulated annealing algorithm and particle swarm algorithm; Dynamic path planning and resource allocation: Monitor the Q30 value, repeat density of the data and the utilization rate of computing resources in real time, and use the deep Q network algorithm for reinforcement learning decision-making, and allocate computing resources according to the decision results; Multi-omics verification and result output: Integrate transcriptome and epigenetic data, correct the assembly results obtained after optimization and resource allocation, and output the final assembly results and visualization reports.

2. The optimized method for the gene sequencing path of silkworm based on the hybrid parallel genetic algorithm according to claim 1, characterized in that The data preprocessing and feature modeling include the following steps: Filter reads with base quality value less than 30 through the HMM model: Read the original sequencing data file from the storage device, initialize the parameters of the HMM model. For each sequencing read, calculate its probability in different hidden states using the HMM model. By recursively calculating the probability value at each position, concentrate and filter the reads with the base quality value less than 30, and save the filtered read data to a data file. The recurrence formula for the forward propagation calculation is: where α t (j) represents the probability of being in state j at time t and observing the sequence o1, o2, …, o t , a ij is the state transition probability, and b j (o t ) is the observation probability; Convert the processed sequence data into a k-mer frequency matrix and input it into a convolutional neural network model to output a repeat density probability map: Divide each of the filtered reads into k-mers of a fixed length, count the occurrence frequency of each k-mer in all the reads, form a matrix with all the occurrence frequency values as the input of the convolutional neural network model, and perform forward propagation calculation through the convolutional neural network model to output a repeat density probability map. The convolution operation formula of the convolutional layer of the convolutional neural network model is: where x is the input feature map, w is the convolution kernel, b is the bias, and y is the output feature map; Generate enhanced data through a generative adversarial network to expand the training set: Generate a random noise vector as the input of the generative adversarial network generator, generate simulated sequencing sequences with a controllable error rate, input the simulated sequencing sequences and real sequencing sequences into the discriminator for adversarial training, and merge the simulated sequencing sequences with the original sequencing data and add them to the training set.

3. The optimized method for the gene sequencing path of silkworms based on a hybrid parallel genetic algorithm according to claim 1, characterized in that, The hybrid parallel genetic algorithm optimization includes the following steps: Initialize the population: Based on the preprocessed data, construct contig nodes, and calculate the connection edge weights between the contig nodes. The calculation formula is as follows: Among them, OverlapLength i,j is the overlap length of contigs i and j, Length i and Length j are the lengths of contigs i and j respectively; Randomly generate a number of chromosomes using the dynamic decision tree encoding method. Each chromosome includes multiple contig nodes and connection edges, and save the generated chromosomes to the population as the initial population; CPU parallel calculation of fitness: Use a distributed computing framework to distribute the chromosome data in the initial population to several GPU nodes, calculate the N50, Consistency, Time and Core values of each chromosome according to the multi-objective optimization function, and calculate the fitness value. The multi-objective optimization function formula is expressed as: Among them, α, β, and γ are weight coefficients. N50 represents the contig length corresponding to when the cumulative length reaches half of the total length after sorting all contigs by length from largest to smallest. MaxLen is the maximum possible assembly length. Consistency represents the consistency between the assembly result and the original reads. Reads is the number of original reads. Time is the calculation time. Core is the number of computing cores used; And summarize the fitness values to the total node, and analyze the fitness values through the total node to obtain the optimal chromosome; Perform global search based on the particle swarm optimization algorithm and introduce the simulated annealing algorithm to assist mutation: Initialize each chromosome corresponding to a particle, and update the velocity and position of each particle according to the particle velocity update formula. The particle velocity update formula is: wherein, is the velocity of particle i in the t-th generation, ω is the inertia weight, c1 and c2 are learning factors, r1 and r2 are random numbers, p best is the individual optimal position of the particle, g best is the global optimal position, is the position of particle i in the t-th generation; Perform simulated annealing mutation operation on each particle using the simulated annealing algorithm, and determine whether to accept the mutated particle according to the simulated annealing probability formula. The simulated annealing probability formula is: P = e -ΔE / T Among them, ΔE is the change in fitness, T is the temperature parameter. During the simulated annealing process, the temperature T is inversely proportional to the number of iterations. Repeat the update operation based on the particle swarm optimization algorithm and the simulated annealing mutation operation.

4. The optimized method for the gene sequencing path of silkworms based on a hybrid parallel genetic algorithm according to claim 1, wherein The dynamic path planning and resource allocation include the following steps: Data collection, storage, and analysis: Collect the Q30 value, repeat density of the data in real time, and the utilization rate of computing resources. Among them, use the system monitoring tool to obtain the resource utilization rate, obtain the Q30 value and repeat density through real-time analysis of the data, store the collected data information in the database, and analyze the data information through the data analysis tool; Reinforcement learning decision-making based on the deep Q network model: Use the Q30 value, repeat density, and utilization rate of computing resources as the state input of the deep Q network model. The deep Q network model uses the Softmax strategy for action selection. The actions include algorithm switching and priority adjustment. The formula is as follows: Among them, π(a|s) represents the probability of selecting action a in state s, Q(s,a) represents the Q value of taking action a in state s, and τ is the temperature parameter; Reward calculation: Construct a multi-dimensional reward function. The formula is as follows: Among them, r is the reward value, ω1, ω2, ω3, and ω4 are weight coefficients, ΔN50 represents the change in the N50 value before and after taking the action, ΔConsistency represents the change in the consistency between the assembly result and the original reads before and after taking the action, and ΔResourceUtilization represents the change in the resource utilization rate before and after taking the action; Model update: Update the parameters of the deep Q network model using the reinforcement learning algorithm according to the reward value and action selection. The Q value update formula of the deep Q network model is: where s t is the state, a t is the action, r t is the reward, α is the learning rate, and γ is the discount factor; Dynamic resource allocation: According to the decision result of the reinforcement learning algorithm, schedule tasks through Kubernetes, preferentially allocate high-priority tasks to available GPU and CPU resources, and monitor the usage of CPU and GPU resources in real time to manage the execution status of tasks.

5. A method for optimizing the silkworm gene sequencing path based on a hybrid parallel genetic algorithm according to claim 1, characterized in that, The multi-omics verification and result output include the following steps: Verify exon continuity through the STAR alignment tool: Compare the transcriptome data with the assembly results, use bioinformatics tools to process and analyze the alignment result files, and correct the assembly results according to the analysis results. The correction methods include adjusting the assembly path, adding or deleting contigs; Merge methylation coverage constraint paths: Use the bioinformatics tools to process the methylation site data, screen out the assembly paths according to the methylation site coverage rate, merge the screened assembly paths to form the final assembly result, and verify the final assembly result through the bioinformatics tools; Output FASTA files and visualization reports: Use the BioPython library of Python to save the final assembly result as a FASTA file, and generate a visualization report through the Matplotlib library of Python. The visualization report includes an assembly path map, N50 value, and error rate heat map, and save the visualization report as a PDF or HTML file.

6. A silkworm gene sequencing path optimization system based on a hybrid parallel genetic algorithm, characterized in that, Including: An input / output interface module for system interaction with external data; A data processing module, which includes a data preprocessing sub-module, a feature extraction sub-module based on a convolutional neural network model, and a data augmentation sub-module; An optimization module based on a hybrid parallel genetic algorithm, which includes an initialization sub-module, a fitness calculation sub-module, and a hybrid optimizer; A dynamic resource scheduling module based on Kubernetes, which is used to allocate resources to tasks according to task priorities; A multi-omics feedback correction module, which is used to feedback and correct the gene sequence assembly results based on the exon continuity alignment results of transcriptome data and the methylation site coverage rate.

7. The optimized system for the gene sequencing path of silkworm based on the hybrid parallel genetic algorithm according to claim 6, characterized in that The data preprocessing sub-module includes using the HMM model to clean and denoise the data, and the formula is expressed as: Among them, O is the original sequencing data, q is the hidden state, representing the true base sequence, λ is the model parameter, and the Baum-Welch algorithm is used to train the HMM model. By iteratively updating the model parameters, the probability P(O|λ) of the observation sequence is maximized.

8. The optimized system for the gene sequencing path of silkworms based on the hybrid parallel genetic algorithm according to claim 6, wherein The feature extraction sub-module includes minimizing the loss function during the training process of the convolutional neural network model, and the loss function formula is expressed as: Among them, d(x,y) is the binary mask of the labeled repeated region, and p(x,y) is the predicted probability.

9. The optimized system for the gene sequencing path of silkworm based on a hybrid parallel genetic algorithm according to claim 6, characterized in that, The data augmentation sub-module includes a generative adversarial network algorithm, and the generative adversarial algorithm is used to generate simulated sequencing sequences with a controllable error rate to expand the training set. The adversarial loss function formula is expressed as: Among them, p data (x) is the distribution of real data, and p z (z) is the distribution of noise. The generative adversarial algorithm includes a discriminator and a generator. The goal of the discriminator is to maximize V(D, G), and the goal of the generator is to minimize V(D, G).

10. A system for optimizing the silkworm gene sequencing path based on a hybrid parallel genetic algorithm according to claim 6, characterized in that The population initialization sub-module includes encoding chromosomes using a dynamic decision tree, and the chromosome is represented as C = {(n i , w ij )}, where n i is a contig node, and w ij is the weight of the connecting edge.

Citation Information

Cited By

  • Vehicle virtual modification method and system based on virtual space

    CN121052144A

  • A virtual space-based vehicle virtual refitting method and system

    CN121052144B