A gene sequencing system and method for multi-gene repetitive sequences

By adopting technical means such as parallel processing, dedicated hardware acceleration, efficient process management and abnormal detection in the gene sequencing system, the problems of low efficiency of multi-gene duplicate sequence processing and poor sequencing accuracy in the existing technology are solved, and an efficient, accurate and flexible gene sequencing process is achieved.

CN119479807BActive Publication Date: 2025-06-10SHANGHAI LINGEN BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510060774.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-06-10
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

When processing multigene repeat sequences, existing gene sequencing technologies face problems such as low processing efficiency, bottlenecks in data handling, abnormal amplification products affect sequencing accuracy, and inflexible process control.

Method used

Adopt parallel processing mechanism, reduce data handling overhead, use dedicated hardware to accelerate, realize efficient process management, perform abnormal detection and correction, and improve processing efficiency and accuracy through flexible programmability adjustment processes.

Benefits of technology

It significantly improves the efficiency and accuracy of multi-gene repeat sequence processing, reduces data processing time, improves the flexibility and adaptability of the system, and ensures the quality of sequencing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479807B_ABST
    Figure CN119479807B_ABST
Patent Text Reader

Abstract

The present invention discloses a gene sequencing system and method for multi-gene repetitive sequences, belonging to the technical field of gene sequencing. A gene sequencing system for multi-gene repetitive sequences includes a high-throughput sequencing module, a nested multiplex PCR amplification module, a data processing module, a central processing module, an in-memory computing module, a storage computing unit, a process control module, an objective function module, and an anomaly detection module. The present invention solves the problem of low gene sequencing efficiency in the prior art, especially the inability to meet the gene sequencing of multi-gene repetitive sequences. Through a parallel processing mechanism, reducing data transfer overhead, dedicated hardware acceleration, efficient process management, anomaly detection and correction, and flexible programmability, the present invention significantly improves the processing efficiency and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gene sequencing, and particularly to a gene sequencing system and method for multi-gene repeat sequences. Background Art

[0002] In the field of gene sequencing, especially when it comes to the analysis of multi-gene repeat sequences, the existing technologies face many challenges. The traditional gene sequencing process usually relies on a single central processor for control, resulting in low processing efficiency and long time consumption for a single gene sequencing. In addition, with the continuous increase in the amount of sequencing data, the transportation and processing of data have become bottlenecks, and frequent data movement increases the processing delay and energy consumption. At the same time, abnormal amplification products generated during the PCR amplification process often affect the accuracy of the sequencing results, and the existing technologies have limited means for detecting and correcting these abnormal products, which are prone to introducing biases. In addition, the existing process control mechanism is not flexible enough and is difficult to be optimized and adjusted according to different sequencing requirements, resulting in a complex and difficult-to-manage processing process. Summary of the Invention

[0003] The purpose of the present invention is to provide a gene sequencing system and method for multi-gene repeat sequences. Through a parallel processing mechanism, reducing data transportation overhead, dedicated hardware acceleration, efficient process management, abnormal detection and correction, and flexible programmability, the processing efficiency and accuracy are significantly improved. The setting of dedicated hardware acceleration improves the execution efficiency of the algorithm; efficient process management reduces the time overhead of system control; abnormal detection and correction ensure data quality; flexible programmability enables the process to be adjusted according to different requirements. These advantages jointly ensure the high efficiency, accuracy, and flexibility of multi-gene repeat sequence processing, and solve the problems raised in the above background art.

[0004] To achieve the above purpose, the present invention provides the following technical solutions:

[0005] A gene sequencing system for multi-gene repeat sequences, comprising:

[0006] A high-throughput sequencing module for obtaining the original sequencing data of the target DNA fragment;

[0007] A nested multiplex PCR amplification module for specifically amplifying the repeat sequences within the target region, and the module reduces non-specific amplification by designing specific primer combinations;

[0008] A data processing module, which is based on a machine learning algorithm and is used to automatically identify and correct sequencing errors caused by repeat sequences;

[0009] A central processing module for calling and executing a central sequencing operator;

[0010] In-memory computing module, used to cache gene data and execute in-memory sequencing operators;

[0011] Storage computing unit, used to call and execute storage sequencing operators;

[0012] Process control module, used to receive process management information and execute different steps in the gene sequencing process according to management commands;

[0013] Objective function module, used to obtain independent components and noise components through ICA decomposition, and construct an objective function based on the fluorescence intensity distribution to obtain the trend term component of the PCR reaction;

[0014] Abnormality detection module, used to detect abnormal amplification cycles, and obtain and delete DNA samples with abnormal amplification according to the abnormal amplification cycles.

[0015] Preferably, the nested multiplex PCR amplification module includes a set of optimized primer pairs and a customized thermal cycling program. Each primer pair is designed to match different variants of the target repeat sequence; the thermal cycling program performs PCR reactions at multiple temperature gradients by dynamically adjusting the extension time and annealing temperature.

[0016] Preferably, the data processing module includes a deep learning module and an algorithm module. The deep learning module is based on convolutional neural networks and long short-term memory networks, and is used to learn the patterns of repeat sequences from large-scale sequencing data and distinguish real variant signals from technical errors; the algorithm module uses statistical methods to detect and correct read alignment errors caused by repeat sequences, identifies potential error sites by calculating the Hamming distance between adjacent reads, and uses Bayesian statistical methods for probability correction, thereby improving the accuracy of sequencing results.

[0017] Preferably, the central processing module includes:

[0018] General computing unit, used to execute the algorithm part not executed by the accelerator array unit, and split the instructions of the dynamic programming algorithm, and distribute the split instruction information to the instruction parsing unit;

[0019] Instruction parsing unit, used to parse the instruction information, distribute the parsed instruction information to the accelerator array unit, and support efficient interaction between the general computing unit and the accelerator array unit;

[0020] Accelerator array unit, used to perform accelerated calculations according to the parsed results; support different granularity of computing tasks and flexibly adjust the computing resource allocation;

[0021] A configuration storage unit is used to store necessary calculation data such as reference sequences, read segment sequences, and result sequences; provide parameters required by the dynamic programming algorithm; and supply the required data to the general computing unit and the accelerator array unit in a timely manner.

[0022] Preferably, the in-memory computing module includes:

[0023] A main memory unit for caching gene data in the gene sequencing process;

[0024] A logic computing unit for calling gene data in the main memory module, calling a preset in-memory sequencing operator, and executing the in-memory sequencing operator to process the called gene data.

[0025] Preferably, the in-memory computing module includes:

[0026] An operating parameter monitoring module for real-time monitoring of the operating parameters of the main memory unit; wherein, the operating parameters include main memory capacity, main memory capacity change rate, storage bandwidth, and main memory access speed;

[0027] A change rate comparison module for comparing the main memory capacity change rate with a preset main memory capacity change rate threshold;

[0028] A first data retrieval module for retrieving the current main memory capacity of the main memory unit when the main memory capacity change rate exceeds the preset main memory capacity change rate threshold;

[0029] A capacity stability evaluation coefficient acquisition module for obtaining a capacity stability evaluation coefficient by using the main memory capacity change rate and the current main memory capacity of the main memory unit;

[0030] Wherein, the capacity stability evaluation coefficient is obtained by the following formula:

[0031]

[0032] Wherein, S represents the capacity stability evaluation coefficient; n represents the number of unit times of operation of the main memory unit, and the unit time is taken as 1s; C d represents the current main memory capacity of the main memory unit; C y represents the maximum remaining main memory capacity allowed for the normal operation of the cache of the main memory unit; C i represents the main memory capacity corresponding to the main memory unit in the i-th unit time; R i represents the main memory capacity change rate in the i-th unit time; s irepresents the proportion of the redundant data amount cached in the main memory unit corresponding to the \(i\)-th unit time; \(\alpha\) represents the first adjustment coefficient, which is used to control the relationship strength between the main memory capacity of the main memory unit and the main memory capacity change rate. Moreover, the value range of the first adjustment coefficient is 2.31 - 4.68; \(\beta\) represents the second adjustment coefficient, which controls the negative impact degree of the main memory capacity change rate on the cache stability of the main memory unit. Moreover, the value range of the second adjustment coefficient is 0.18 - 0.87; \(R\) b represents the standard deviation of the main memory capacity change rate corresponding to \(n\) unit times; \(R\) fmax represents the maximum value of the change amplitude of the main memory capacity change rate between every two adjacent unit times within \(n\) unit times;

[0033] The coefficient comparison module is used to compare the capacity stability evaluation coefficient with a preset coefficient threshold;

[0034] The abnormal determination module is used to determine whether there is an operation abnormality in the main memory unit by using the storage bandwidth and the main memory access speed when the capacity stability evaluation coefficient is lower than the preset coefficient threshold.

[0035] Preferably, the abnormal determination module includes:

[0036] The second data retrieval module is used to retrieve the storage bandwidth and the main memory access speed when the capacity stability evaluation coefficient is lower than the preset coefficient threshold;

[0037] The operation evaluation coefficient acquisition module is used to acquire the operation evaluation coefficient of the main memory unit by using the storage bandwidth and the main memory access speed;

[0038] Among them, the operation evaluation coefficient of the main memory unit is obtained through the following formula:

[0039]

[0040] Among them, \(G\) represents the operation evaluation coefficient of the main memory unit; \(n\) represents the number of unit times of the main memory unit operation, and the unit time takes the value of 1s; \(B\) i represents the storage bandwidth corresponding to the main memory unit in the \(i\)-th unit time; \(R\) i represents the main memory capacity change rate in the \(i\)-th unit time; \(V\) i represents the main memory access speed corresponding to the main memory unit in the \(i\)-th unit time; \(P\) Bi represents the storage bandwidth change rate corresponding to the main memory unit in the \(i\)-th unit time; \(P\) Vi represents the main memory access speed change rate corresponding to the main memory unit in the \(i\)-th unit time;

[0041] An integrated operation evaluation coefficient acquisition module, configured to integrate the operation evaluation coefficient of the main memory unit and the capacity stability evaluation coefficient to obtain an integrated operation evaluation coefficient of the main memory unit;

[0042] Among them, the integrated operation evaluation coefficient is obtained through the following formula:

[0043]

[0044] Among them, H represents the integrated operation evaluation coefficient; w 01 and w 02 represent the weight values corresponding to the capacity stability evaluation coefficient and the operation evaluation coefficient of the main memory unit; S represents the capacity stability evaluation coefficient; G represents the operation evaluation coefficient of the main memory unit;

[0045] An integrated coefficient comparison module, configured to compare the integrated operation evaluation coefficient with a preset integrated coefficient threshold;

[0046] An anomaly determination and alarm module, configured to determine that there is an anomaly in the operation of the main memory unit and perform an anomaly alarm when the integrated operation evaluation coefficient is lower than the preset integrated coefficient threshold.

[0047] Preferably, the objective function module is further configured to:

[0048] Obtain the sine wave, its frequency and phase in the residual curve through short-time Fourier transform;

[0049] Obtain the component estimation factor of each type of phase, obtain the optimal phase and the corresponding several error sine waves, and form an error component;

[0050] Obtain the overall Gaussian kurtosis of the error component and construct an objective function according to the Gaussian kurtosis.

[0051] Preferably, the anomaly detection module includes:

[0052] An isolation forest tree construction unit, configured to construct an isolation forest tree, identify an abnormal amplification cycle, and recursively divide all fluorescence intensity amplitudes into left and right subsets until no further division is possible;

[0053] An outlier calculation unit, configured to calculate the number of paths of each node and use it as an outlier. The node with fewer paths is more likely to represent an abnormal amplification cycle;

[0054] An anomaly threshold setting unit, configured to set a threshold for anomaly scoring to identify an abnormal amplification cycle, preset an anomaly threshold, and perform linear normalization on the outlier provided by the outlier calculation unit to obtain an anomaly score. If the anomaly score of a certain amplification cycle is lower than the threshold, it is determined as an abnormal amplification cycle;

[0055] A differential image analysis unit, which is used to analyze abnormal amplification cycles by the frame difference method. The unit uses a real-time fluorescence PCR instrument to capture cumulative images in each amplification cycle, and performs a difference between the cumulative image corresponding to the abnormal amplification cycle and the cumulative image of its previous cycle by the frame difference method to identify DNA samples with abnormal amplification;

[0056] An abnormal sample deletion unit, which is used to delete DNA samples with abnormal amplification. It deletes these samples according to the information of abnormal amplification samples provided by the differential image analysis unit.

[0057] A gene sequencing method for multi-gene repeat sequences, which is implemented based on a gene sequencing system for multi-gene repeat sequences, and includes the following steps:

[0058] Step 1: Use a high-throughput sequencing module to obtain the original sequencing data of the target DNA fragment;

[0059] Step 2: Use a nested multiplex PCR amplification module to amplify the target repeat sequence;

[0060] Step 3: Apply a data processing module based on machine learning to analyze the sequencing data and identify and correct sequencing errors;

[0061] Step 4: Parallel process gene data through a central processing unit, an in-memory computing unit, and a storage computing unit;

[0062] Step 5: Apply an objective function module to detect and correct abnormal amplification;

[0063] Step 6: Apply an anomaly detection module to obtain and delete DNA samples with abnormal amplification;

[0064] Step 7: Compare the corrected sequencing results with a known reference genome to determine the specific positions and variations of the repeat sequences.

[0065] Compared with the prior art, the beneficial effects of the present invention are:

[0066] 1. Through the parallel processing mechanism of the central processing unit, the in-memory computing unit, and the storage computing unit, the present invention improves the processing efficiency of the gene sequencing process. This parallel processing ability helps to speed up the sequencing speed of multi-gene repeat sequences when dealing with a large amount of data.

[0067] 2. Through the design of the in-memory computing module and the storage computing unit, the present invention reduces the distance and time overhead of data transfer and improves the performance of data processing. Especially for data that needs to be accessed frequently, the in-memory computing module can achieve fast data processing and avoid the delay caused by multiple data reads.

[0068] 3. Through the use of the accelerator array module and the customized operation function module, the present invention provides hardware acceleration for specific algorithms in gene sequencing, improving the algorithm execution efficiency. Especially when dealing with multi-gene repetitive sequences, the calculation time can be significantly reduced.

[0069] 4. In the present invention, the abnormal detection module can effectively identify and correct abnormalities during the PCR amplification process, ensuring the quality of sequencing data. In addition, the abnormal detection module can also obtain and delete abnormally amplified DNA samples to prevent these samples from interfering with subsequent analysis.

[0070] 5. Through the programmable gene sequencing process control of the present invention, the gene sequencing process can be adjusted according to different types of multi-gene repetitive sequences, increasing the flexibility and adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 It is a schematic diagram of the data flow structure of the present invention;

[0072] Figure 2 It is a schematic diagram of the control flow structure of the present invention;

[0073] Figure 3 It is a schematic diagram of the information flow structure of the present invention;

[0074] Figure 4 It is a flowchart of the operation of the sequencing system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0075] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0076] To solve the problem of low gene sequencing efficiency in the prior art, especially the inability to meet the gene sequencing of multi-gene repetitive sequences, please refer to Figures 1-4 , the following technical solutions are provided in this embodiment:

[0077] A gene sequencing system for multi-gene repetitive sequences, comprising:

[0078] A high-throughput sequencing module for obtaining the original sequencing data of the target DNA fragment;

[0079] A nested multiplex PCR amplification module for specifically amplifying repetitive sequences within the target region, and the module reduces non-specific amplification by designing specific primer combinations;

[0080] A data processing module, which is based on machine learning algorithms and is used to automatically identify and correct sequencing errors caused by repetitive sequences;

[0081] A central processing module, which is used to call and execute central sequencing operators;

[0082] An in-memory computing module, which is used to cache gene data and execute in-memory sequencing operators;

[0083] A storage computing unit, which is used to call and execute storage sequencing operators;

[0084] A process control module, which is used to receive process management information and execute different steps in the gene sequencing process according to management commands;

[0085] An objective function module, which is used to obtain independent components and noise components through ICA decomposition, and construct an objective function based on the fluorescence intensity distribution to obtain the trend term component of the PCR reaction;

[0086] An anomaly detection module, which is used to detect abnormal amplification cycles, and obtain and delete DNA samples with abnormal amplification according to the abnormal amplification cycles.

[0087] The nested multiplex PCR amplification module includes a set of optimized primer pairs and a customized thermal cycling program. Each of the primer pairs is designed to match different variants of the target repetitive sequence to improve amplification efficiency and specificity; the thermal cycling program performs PCR reactions at multiple temperature gradients by dynamically adjusting the extension time and annealing temperature to improve the amplification efficiency of repetitive sequences in complex templates.

[0088] The operation process of the primer pairs is as follows:

[0089] Using the bioinformatics tool Primer3, design a pair of primers according to the sequence information of the target region. The primer length is usually between 18 and 30 bases, and the Tm value is approximately between 55 and 65 °C; by adjusting parameters such as primer length, GC content, and Tm value, ensure the specificity and stability of the primers; use the BLAST tool to check the binding possibility of the primers to non-target regions and reduce non-specific amplification; set positive and negative controls, observe whether there are non-specific amplification bands, and select primer pairs with no non-specific amplification bands and high amplification efficiency; through gradient dilution of the primer solution, conduct PCR amplification experiments to find the optimal concentration, and the concentration range is 0.2 μM to 1.3 μM.

[0090] The operation process of the thermal cycling program includes an initial denaturation process, a cycling denaturation process, an annealing and extension process, and a cooling process. During the operation of the thermal cycling program, known amplified products are set as positive controls, and reactions without template DNA are set as negative controls. The PCR products are added to a 1% agarose gel and analyzed by electrophoresis to observe whether there are bands of the expected size. Qubit is used for DNA concentration determination to ensure that the concentration of the amplified products meets the requirements and better evaluate the amplification efficiency.

[0091] A set of optimized primer pairs and a customized thermal cycling program work together to significantly improve the amplification efficiency and specificity of multi-gene repeat sequences through specific primer design and optimized PCR conditions, reduce non-specific amplification products, and ensure the accuracy of subsequent sequencing analysis.

[0092] The data processing module includes a deep learning module and an algorithm module. The deep learning module is based on a convolutional neural network (CNN) and a long short-term memory network (LSTM) and is used to learn the patterns of repeat sequences from large-scale sequencing data and distinguish real variant signals from technical errors. The algorithm module uses statistical methods to detect and correct read alignment errors caused by repeat sequences, identifies potential error sites by calculating the Hamming distance between adjacent reads, and uses Bayesian statistical methods for probability correction to improve the accuracy of sequencing results.

[0093] The central processing module includes:

[0094] A general computing unit, which is used to execute the algorithm part that cannot be executed by the accelerator array unit, split the instructions of the dynamic programming algorithm, and distribute the split instruction information to the instruction parsing unit;

[0095] An instruction parsing unit, which is used to parse the instruction information, distribute the parsed instruction information to the accelerator array unit, and support the efficient interaction between the general computing unit and the accelerator array unit to ensure the correct execution of the instructions;

[0096] An accelerator array unit, which is used to perform accelerated calculations according to the parsed results to improve the execution efficiency of the dynamic programming algorithm; support different granularity computing tasks and flexibly adjust the computing resource allocation;

[0097] A configuration storage unit, which is used to store necessary computing data such as reference sequences, read sequences, and result sequences; provide the parameters required by the dynamic programming algorithm to ensure the correctness of the calculation; and supply the required data to the general computing unit and the accelerator array unit in a timely manner.

[0098] Through the collaborative work among the general computing unit, instruction parsing unit, accelerator array unit, and configuration storage unit, the central processing module realizes the efficient execution of algorithms and the effective management of data, ensuring that the key steps in the gene sequencing process can be completed quickly and accurately.

[0099] Specifically, the in-memory computing module includes:

[0100] The main memory unit is used to cache gene data in the gene sequencing process. It adopts volatile storage devices such as SDRAM and is used to temporarily store intermediate data used in the gene sequencing process.

[0101] The logic computing unit is used to call gene data in the main memory module, call preset in-memory sequencing operators, and execute the in-memory sequencing operators to process the called gene data. The operator types include FM index operators or Hash index operators, which are used to process intermediate data.

[0102] Specifically, the in-memory computing module includes:

[0103] The operating parameter monitoring module is used to monitor the operating parameters of the main memory unit in real time. Among them, the operating parameters include main memory capacity, main memory capacity change rate, storage bandwidth, and main memory access speed.

[0104] The change rate comparison module is used to compare the main memory capacity change rate with a preset main memory capacity change rate threshold.

[0105] The first data retrieval module is used to retrieve the current main memory capacity of the main memory unit when the main memory capacity change rate exceeds the preset main memory capacity change rate threshold.

[0106] The capacity stability evaluation coefficient acquisition module is used to obtain a capacity stability evaluation coefficient by using the main memory capacity change rate and the current main memory capacity of the main memory unit.

[0107] Among them, the capacity stability evaluation coefficient is obtained through the following formula:

[0108]

[0109] Among them, S represents the capacity stability evaluation coefficient; n represents the number of unit time of the main memory unit operation, and the unit time takes the value of 1s; C d represents the current main memory capacity of the main memory unit; C y represents the maximum remaining main memory capacity allowed for the normal operation of the main memory unit cache; C i represents the main memory capacity corresponding to the main memory unit in the i-th unit time; R i represents the main memory capacity change rate in the i-th unit time; s iIt represents the proportion of the redundant data amount cached in the main memory unit corresponding to the \(i\)-th unit time; \(\alpha\) represents the first adjustment coefficient, which is used to control the relationship strength between the main memory capacity of the main memory unit and the main memory capacity change rate, and the value range of the first adjustment coefficient is 2.31 - 4.68; \(\beta\) represents the second adjustment coefficient, which controls the negative impact degree of the main memory capacity change rate on the cache stability of the main memory unit, and the value range of the second adjustment coefficient is 0.18 - 0.87; \(R\) b It represents the standard deviation of the main memory capacity change rate corresponding to \(n\) unit times; \(R\) fmax It represents the maximum value of the change amplitude of the main memory capacity change rate between every two adjacent unit times in \(n\) unit times;

[0110] A coefficient comparison module, which is used to compare the capacity stability evaluation coefficient with a preset coefficient threshold;

[0111] An abnormal determination module, which is used to determine whether there is an operation abnormality in the main memory unit by using the storage bandwidth and the main memory access speed when the capacity stability evaluation coefficient is lower than the preset coefficient threshold.

[0112] The technical effects of the above technical solution are as follows: By using a volatile storage device (such as SDRAM) in the main memory unit to cache gene data in the gene sequencing process, intermediate data can be temporarily stored, thereby improving the data processing efficiency of the gene sequencing process. This is particularly important for large-scale and highly complex gene data processing tasks. The logic calculation unit can call the gene data in the main memory module and execute preset in-memory sequencing operators (such as FM index operator or Hash index operator) to process this data. This in-memory computing method reduces the data transmission overhead and improves the processing speed. The operation parameter monitoring module can monitor the operation parameters of the main memory unit in real time, including the main memory capacity, the main memory capacity change rate, the storage bandwidth, and the main memory access speed. This provides the possibility to detect and handle potential problems in a timely manner. By comparing the main memory capacity change rate with the preset main memory capacity change rate threshold and calculating the capacity stability evaluation coefficient using the above formula, this technical solution can comprehensively evaluate the capacity stability of the main memory unit. This helps to prevent performance degradation or data loss problems caused by insufficient or unstable capacity. When the capacity stability evaluation coefficient is lower than the preset coefficient threshold, the abnormal determination module can further determine whether there is an operation abnormality in the main memory unit by using the storage bandwidth and the main memory access speed. This intelligent determination method improves the reliability and stability of the system. Although not directly mentioned in the technical solution, based on the results of real-time monitoring and evaluation, the system can dynamically adjust resource allocation (such as increasing the main memory capacity, optimizing the storage structure, etc.) to meet the changing data processing requirements.

[0113] In summary, through real-time monitoring, evaluation, and optimization of the operating status of the main memory unit, this technical solution improves the data processing efficiency of the gene sequencing process and the reliability and stability of the system. This is of great significance for promoting the development and application of gene sequencing technology.

[0114] Specifically, the abnormal determination module includes:

[0115] A second data retrieval module, configured to retrieve the storage bandwidth and the main memory access speed when the capacity stability evaluation coefficient is lower than a preset coefficient threshold;

[0116] An operating evaluation coefficient acquisition module, configured to obtain the operating evaluation coefficient of the main memory unit by using the storage bandwidth and the main memory access speed;

[0117] Among them, the operating evaluation coefficient of the main memory unit is obtained through the following formula:

[0118]

[0119] Among them, G represents the operating evaluation coefficient of the main memory unit; n represents the number of unit times of the main memory unit operation, and the unit time is taken as 1 s; B i represents the storage bandwidth corresponding to the main memory unit in the i-th unit time; R i represents the main memory capacity change rate in the i-th unit time; V i represents the main memory access speed corresponding to the main memory unit in the i-th unit time; P Bi represents the storage bandwidth change rate corresponding to the main memory unit in the i-th unit time; P Vi represents the main memory access speed change rate corresponding to the main memory unit in the i-th unit time;

[0120] A comprehensive operating evaluation coefficient acquisition module, configured to integrate the operating evaluation coefficient of the main memory unit and the capacity stability evaluation coefficient to obtain the comprehensive operating evaluation coefficient of the main memory unit;

[0121] Among them, the comprehensive operating evaluation coefficient is obtained through the following formula:

[0122]

[0123] Among them, H represents the comprehensive operating evaluation coefficient; w 01 and w 02 represent the weight values corresponding to the capacity stability evaluation coefficient and the operating evaluation coefficient of the main memory unit; S represents the capacity stability evaluation coefficient; G represents the operating evaluation coefficient of the main memory unit;

[0124] A comprehensive coefficient comparison module, configured to compare the comprehensive operating evaluation coefficient with a preset comprehensive coefficient threshold;

[0125] Anomaly determination and alarm module, which is used to determine that there is an anomaly in the operation of the main memory unit and perform anomaly alarm when the comprehensive operation evaluation coefficient is lower than the preset comprehensive coefficient threshold.

[0126] The technical effects of the above technical solution are as follows: Through the second data retrieval module, the storage bandwidth and the main memory access speed are obtained, and then the operation evaluation coefficient of the main memory unit is calculated using these parameters. This not only considers the capacity stability of the main memory unit but also comprehensively evaluates its data transmission and processing capabilities, thus providing more comprehensive performance monitoring and evaluation. Through the calculation of the comprehensive operation evaluation coefficient (H), the capacity stability evaluation coefficient (S) and the operation evaluation coefficient (G) of the main memory unit are integrated. This integration method enables the system to more accurately reflect the overall operation status of the main memory unit, providing a more reliable basis for subsequent anomaly determination. In the calculation of the comprehensive operation evaluation coefficient, w 01 and w 02 As the weight values corresponding to the capacity stability evaluation coefficient and the operation evaluation coefficient of the main memory unit, they can be flexibly set according to actual needs. This weight setting method enables the system to perform targeted weight allocation for each evaluation index according to different application scenarios and performance requirements, thereby improving the accuracy and practicality of the evaluation. When the comprehensive operation evaluation coefficient is lower than the preset comprehensive coefficient threshold, the anomaly determination and alarm module can automatically determine that there is an anomaly in the operation of the main memory unit and perform anomaly alarm. This intelligent determination method not only improves the response speed of the system but also can timely remind the management personnel to conduct fault troubleshooting and handling, thus avoiding problems such as performance degradation or data loss caused by faults. By real-time monitoring and evaluating the operation status of the main memory unit and performing anomaly determination and alarm according to the evaluation results, this technical solution can timely detect and handle potential performance problems. This helps to improve the reliability and stability of the system and ensure that key tasks such as gene sequencing processes can proceed smoothly.

[0127] In summary, through the technical effects such as comprehensive performance monitoring and evaluation, introduction of the comprehensive operation evaluation coefficient, flexible weight setting, intelligent anomaly determination and alarm, and improvement of the reliability and stability of the system, this technical solution provides strong support for key tasks such as gene sequencing processes.

[0128] The operation process of the in-memory computing module is as follows:

[0129] S01. Data storage

[0130] The main memory unit receives gene data transmitted from the central processing unit or other units; these gene data are temporarily stored in the SDRAM for subsequent processing.

[0131] S02, data call

[0132] The logic computing unit initiates a data request to the main memory module according to the needs of the current processing task; the main memory unit transfers the required gene data to the logic computing unit.

[0133] S03. Data Processing

[0134] The logic computing unit calls the preset in-memory sequencing operator according to the current task requirements; the logic computing unit executes the corresponding operator to process the gene data obtained from the main memory unit.

[0135] S04. Data storage and return

[0136] The processed gene data is stored back to the main memory unit; if necessary, the data in the main memory unit can be returned to the central processing module or other computing modules through the DIMM interface.

[0137] The in-memory computing module stores genetic data in volatile memory devices and uses a dedicated hardware structure that supports hash computing to execute high-concurrency sequencing operators, reducing the overhead of data movement and improving data processing performance in the genetic sequencing process. Through the close cooperation between the main memory unit and the logical computing unit, efficient processing and management of genetic data is achieved.

[0138] The central processing module, in-memory computing module and storage computing unit form a parallel processing mechanism, which realizes the efficient execution of different steps in the gene sequencing process. The central processing module is responsible for global control, task scheduling and execution of advanced algorithms. The in-memory computing module is used for fast processing of intermediate data and supports high-concurrency index calculations, while the storage computing unit focuses on the storage of long-term data and data compression and decompression. The three realize the efficient execution of the gene sequencing process through parallel operation and division of labor, reduce the time overhead of data movement, and improve the overall processing performance.

[0139] The process control module is mainly used to manage the gene sequencing process in the in-memory computing unit and the storage computing unit. Specifically, the main functions of the process detection module include:

[0140] Receive process management information from external sources such as process management software;

[0141] Arbitrate and parse the received information, obtain the message queue number, and write the management command into the corresponding management message queue;

[0142] Call the management command in the management message queue according to the data information, and execute the delay controlled by the data information according to the command;

[0143] Delete the management command after it is called, so that the command executed in the next gene sequencing process is at the top of the management message queue, facilitating quick invocation.

[0144] In this way, the process detection module realizes the efficient control of the gene sequencing process, reduces the time overhead of the system in the gene sequencing process control, and improves the performance of the gene sequencing processing process.

[0145] The objective function module optimizes the extraction of independent components through a series of complex calculation steps and finally obtains the trend term component of the PCR reaction. The specific process of the operation of the objective function module is as follows:

[0146] S11. Calculate the overall Gaussian kurtosis of the error component:

[0147] This module first obtains the frequency of each error sine wave, divides the error sine waves into several frequency categories according to the frequency, and for each frequency category, calculates the ratio of the number of error sine waves in this category to the total number of error sine waves, that is, the distribution probability of this frequency category. Based on these distribution probabilities, the Gaussian kurtosis is calculated, and this Gaussian kurtosis reflects the overall characteristics of the error component.

[0148] S12. ICA decomposition of the residual curve:

[0149] Use the FASTICA algorithm to decompose the residual curve. Each decomposition will generate two independent components. The module classifies all data points according to their fluorescence intensity to form several amplitude categories, calculates the ratio of the number of data points in this category to the total number of data points in this independent component, that is, the fluorescence intensity distribution probability corresponding to this amplitude category, and calculates the Gaussian kurtosis of this independent component. Select the independent component with the smallest Gaussian kurtosis as the independent component, and the one with the largest Gaussian kurtosis as the noise component.

[0150] S13. Construct the objective function:

[0151] Construct the objective function based on the fluorescence intensity distribution probability corresponding to the amplitude category in the independent component, the overall Gaussian kurtosis of the error component, and the residual curve. The objective function takes into account the number of amplitude categories in the independent component, the distribution probability corresponding to the amplitude category, the mean and standard deviation of the distribution probability, the overall Gaussian kurtosis of the error component, and the mean square error between the independent component and the residual curve.

[0152] S14. Determine the optimal decomposition result:

[0153] The module will calculate the output value of the objective function corresponding to the result of each decomposition, and finally select the decomposition result corresponding to the minimum output value of the objective function as the optimal decomposition result, and use the independent component in this result as the trend term component of the PCR reaction.

[0154] Through the above steps, the objective function module can effectively identify and remove abnormal amplification cycles in the PCR reaction, ensuring that the finally obtained trend term component is accurate, thereby improving the accuracy of gene sequencing results.

[0155] The anomaly detection module includes:

[0156] An isolation forest tree construction unit, which is used to construct an isolation forest tree to identify abnormal amplification cycles. By means of recursion, all fluorescence intensity amplitudes are divided into left and right subsets until no further division is possible;

[0157] An outlier calculation unit, which is used to calculate the number of paths of each node and use it as an outlier. The node with fewer paths is more likely to represent an abnormal amplification cycle;

[0158] An anomaly threshold setting unit, which is used to set the threshold of the anomaly score to identify abnormal amplification cycles. A default anomaly threshold is set, and the outliers provided by the outlier calculation unit are linearly normalized to obtain the anomaly score. If the anomaly score of a certain amplification cycle is lower than the threshold, it is determined as an abnormal amplification cycle;

[0159] A differential image analysis unit, which is used to analyze abnormal amplification cycles by the frame difference method. Using the cumulative images taken by a real-time fluorescence PCR instrument in each amplification cycle, the cumulative image corresponding to the abnormal amplification cycle is differentiated from the cumulative image of its previous cycle by the frame difference method to identify the DNA samples with abnormal amplification;

[0160] An abnormal sample deletion unit, which is used to delete the DNA samples with abnormal amplification. According to the information of abnormal amplification samples provided by the differential image analysis unit, these samples are deleted to ensure the accuracy of subsequent whole-genome resequencing.

[0161] In order to better show the operation process of a gene sequencing system for multi-gene repetitive sequences, this embodiment proposes a gene sequencing method for multi-gene repetitive sequences, including the following steps:

[0162] Step 1: Before sequencing, first ensure the quality of the DNA samples, including purity, concentration, integrity, etc.; fragment the DNA and add appropriate adapter sequences to construct a library suitable for sequencing; use the Illumina MiSeq high-throughput sequencing platform for sequencing to obtain raw sequencing data; conduct a preliminary quality assessment on the sequencing data to ensure data quality;

[0163] Step 2: Design nested primers according to the target region to improve the specificity of amplification. Use nested multiplex PCR technology to amplify the target repetitive sequence, monitor the amplification process by a real-time fluorescence PCR instrument, collect the fluorescence data generated during the amplification process, and compare it with the theoretical amplification curve to identify potential abnormal amplifications;

[0164] Step 3: Preprocess the original sequencing data, including quality trimming and adapter removal, and use the DeepVariant machine learning model to train a model to identify sequencing errors; use the trained model to correct the sequencing errors in the sequencing data.

[0165] Step 4: The central processing unit is responsible for scheduling, allocating tasks to the in-memory computing unit and the storage computing unit. The in-memory computing unit processes intermediate data, and the storage computing unit is responsible for long-term storage and auxiliary computing. Integrate the results processed by each unit to form a complete processing flow.

[0166] Step 5: Use the objective function module to detect abnormal amplification through methods such as short-time Fourier transform, perform ICA decomposition on the abnormal data, optimize the independent components and correct the anomalies, and obtain the trend term component of the PCR reaction to reflect the abnormal part in the actual process.

[0167] Step 6: Use the anomaly detection module to identify abnormal amplification cycles through algorithms such as isolation forest trees, and delete DNA samples with abnormal amplification to reduce the impact on subsequent analysis and ensure the data quality after removing abnormal samples.

[0168] Step 7: Prepare the corrected sequencing data, perform format conversion before alignment, use the BWA alignment tool to align the sequencing data with the reference genome, annotate variations such as SNPs and Indels in the test samples to provide information such as gene function and disease relevance, and finally output the alignment results, including information such as variant positions and types, for subsequent research use.

[0169] Working principle: Obtain the original sequencing data of the target DNA fragment through high-throughput sequencing technology; use nested multiplex PCR technology to amplify the target repetitive sequences to increase the copy number of the target region; perform quality control on the sequencing data through a machine learning-based data processing module, and identify and correct sequencing errors to ensure data accuracy; utilize the central processing unit, in-memory computing unit, and storage computing unit to process gene data in parallel, and improve the speed and efficiency of data processing through efficient data management and task scheduling; use the objective function module to detect and correct possible anomalies during the PCR amplification process to ensure the authenticity and reliability of the amplification results; identify and delete DNA samples with abnormal amplification through the anomaly detection module to further improve the accuracy of the sequencing results; finally, align the corrected sequencing results with the known reference genome to determine the specific positions and variations of the repetitive sequences, providing basic data for subsequent genetic analysis. This series of steps ensures the accuracy and efficiency from data acquisition to the final analysis results.

[0170] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.

[0171] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention.

Claims

1. A gene sequencing system for multi-gene repetitive sequences, characterized in that: include: High-throughput sequencing module, used to obtain raw sequencing data of target DNA fragments; A nested multiplex PCR amplification module, used for specifically amplifying repetitive sequences in a target region, wherein the module reduces nonspecific amplification by designing a specific primer combination, and the nested multiplex PCR amplification module includes a set of optimized primer pairs and a customized thermal cycle program, wherein each primer pair is designed to match a different variant of a target repetitive sequence; the thermal cycle program performs a PCR reaction under multiple temperature gradients by dynamically adjusting the extension time and the annealing temperature; A data processing module, which is based on a machine learning algorithm and is used to automatically identify and correct sequencing errors caused by repeated sequences, including a deep learning module and an algorithm module; The deep learning module is based on convolutional neural networks and long short-term memory networks, and is used to learn the patterns of repeated sequences from large-scale sequencing data and distinguish between real mutation signals and technical errors; The algorithm module uses statistical methods to detect and correct read segment alignment errors caused by repeated sequences, identifies potential error sites by calculating the Hamming distance between adjacent read segments, and uses Bayesian statistical methods for probability correction; The central processing module is used to call and execute the central sequencing operator; An in-memory computing module, used for caching gene data and executing in-memory sequencing operators, including an operation parameter monitoring module, a change rate comparison module, a first data retrieval module, a capacity stability evaluation coefficient acquisition module, a coefficient comparison module, an abnormality determination module, a main memory unit, and a logic computing unit; The operating parameter monitoring module is used to monitor the operating parameters of the main memory unit in real time; wherein the operating parameters include main memory capacity, main memory capacity change rate, storage bandwidth and main memory access speed; The change rate comparison module is used to compare the main memory capacity change rate with a preset main memory capacity change rate threshold; The first data retrieval module is used to retrieve the current main memory capacity of the main memory unit when the main memory capacity change rate exceeds a preset main memory capacity change rate threshold; The capacity stability evaluation coefficient acquisition module is used to acquire the capacity stability evaluation coefficient by using the main memory capacity change rate and the current main memory capacity of the main memory unit; The coefficient comparison module is used to compare the capacity stability evaluation coefficient with a preset coefficient threshold; The abnormality determination module is used to determine whether the main memory unit has an operation abnormality by using the storage bandwidth and the main memory access speed when the capacity stability evaluation coefficient is lower than the preset coefficient threshold, and the abnormality determination module includes a second data retrieval module, an operation evaluation coefficient acquisition module, a comprehensive operation evaluation coefficient acquisition module, a comprehensive coefficient comparison module and an abnormality determination and alarm module; The second data retrieval module is used to retrieve the storage bandwidth and the main memory access speed when the capacity stability evaluation coefficient is lower than a preset coefficient threshold, the operation evaluation coefficient acquisition module is used to obtain the operation evaluation coefficient of the main memory unit by using the storage bandwidth and the main memory access speed, the comprehensive operation evaluation coefficient acquisition module is used to integrate the operation evaluation coefficient of the main memory unit with the capacity stability evaluation coefficient to obtain the comprehensive operation evaluation coefficient of the main memory unit, the comprehensive coefficient comparison module is used to compare the comprehensive operation evaluation coefficient with a preset comprehensive coefficient threshold, and the abnormality determination and alarm module is used to determine that the operation of the main memory unit is abnormal and issue an abnormality alarm when the comprehensive operation evaluation coefficient is lower than a preset comprehensive coefficient threshold; The main memory unit is used to cache gene data in the gene sequencing process; The logic computing unit is used to call the gene data in the main memory module, call the preset in-memory sequencing operator, and execute the in-memory sequencing operator to process the called gene data; A storage computing unit, used to call and execute the storage sequencing operator; A process control module is used to receive process management information and execute different steps in the gene sequencing process according to management commands; The objective function module is used to obtain independent components and noise components through ICA decomposition, and to construct an objective function based on the fluorescence intensity distribution to obtain the trend term component of the PCR reaction; obtain the sine wave and its frequency and phase in the residual curve through short-time Fourier transform; obtain the component estimation factor of each type of phase, obtain the optimal phase and the corresponding error sine waves to form the error component; obtain the overall Gaussian kurtosis of the error component, and construct the objective function based on the Gaussian kurtosis; The abnormal detection module is used to detect abnormal amplification cycles, and obtain and delete abnormally amplified DNA samples according to the abnormal amplification cycles.

2. A gene sequencing system for multiple gene repeat sequences according to claim 1, characterized in that: The central processing module comprises: A general computing unit is used to execute the algorithm part executed by the non-accelerator array unit, split the instructions of the dynamic programming algorithm, and distribute the split instruction information to the instruction parsing unit; An instruction parsing unit, used to parse instruction information, distribute the parsed instruction information to the accelerator array unit, and support the interaction between the general computing unit and the accelerator array unit; The accelerator array unit is used to perform accelerated calculations based on the results of the analysis; it supports computing tasks of different granularities and adjusts the allocation of computing resources; A storage unit is configured to store the computational data of the reference sequence, the read sequence, and the result sequence; provide the parameters required by the dynamic programming algorithm; and supply the required data to the general computing unit and the accelerator array unit.

3. A gene sequencing system for multi-gene repeat sequences according to claim 2, characterized in that: The capacity stability evaluation coefficient is obtained by the following formula: Where S represents the capacity stability evaluation coefficient; n represents the number of unit times of the main memory unit operation, and the unit time is 1s; C d Indicates the current main memory capacity of the main memory unit; C y Indicates the maximum remaining main memory capacity allowed for the normal operation of the main memory unit cache; C i represents the main memory capacity corresponding to the main memory unit in the i-th unit time; R i represents the main memory change rate per unit time; s i represents the proportion of redundant data cached in the main memory unit corresponding to the i-th unit time; α represents the first adjustment coefficient, which is used to control the strength of the relationship between the main memory capacity and the main memory change rate of the main memory unit, and the value range of the first adjustment coefficient is 2.31-4.68; β represents the second adjustment coefficient, which controls the negative impact of the main memory change rate on the cache stability of the main memory unit, and the value range of the second adjustment coefficient is 0.18-0.87; R b Represents the standard deviation of the main memory change rate corresponding to n unit time; R fmax It represents the maximum value of the change rate of main memory between every two adjacent unit times in n unit times.

4. A gene sequencing system for multiple gene repeat sequences according to claim 3, characterized in that: The operation evaluation coefficient of the main memory unit is obtained by the following formula: Wherein, G represents the operation evaluation coefficient of the main memory unit; n represents the number of unit time of the main memory unit operation, and the unit time is 1s; B i R represents the storage bandwidth corresponding to the main memory unit in the i-th unit time; i V represents the main memory change rate per unit time i; i represents the main memory access speed corresponding to the main memory unit in the i-th unit time; P Bi represents the storage bandwidth change rate corresponding to the main memory unit in the i-th unit time; P Vi represents the change rate of the main memory access speed corresponding to the main memory unit in the i-th unit time; The comprehensive operation evaluation coefficient of the main memory unit is obtained by the following formula: H=w 01 ·S·log 10 (1+w 02 ·G) Wherein, H represents the comprehensive operation evaluation coefficient; w 01 and w 02 Represents the weight value corresponding to the capacity stability evaluation coefficient and the operation evaluation coefficient of the main memory unit; S represents the capacity stability evaluation coefficient; G represents the operation evaluation coefficient of the main memory unit.

5. A gene sequencing system for multiple gene repeat sequences according to claim 4, characterized in that: The anomaly detection module comprises: An isolated forest tree construction unit is used to construct an isolated forest tree, identify abnormal amplification cycles, and recursively divide all fluorescence intensity amplitudes into left and right subsets until further division is impossible; An outlier calculation unit, used to calculate the number of paths of each node and use it as an outlier; The abnormal threshold setting unit is used to set the threshold of the abnormal score, identify the abnormal amplification cycle, preset an abnormal threshold, and perform linear normalization processing on the abnormal value provided by the abnormal value calculation unit to obtain the abnormal score. If the abnormal score of a certain amplification cycle is lower than the threshold, it is identified as an abnormal amplification cycle; A differential image analysis unit is used to analyze abnormal amplification cycles by using a frame difference method, using a real-time fluorescence PCR instrument to capture cumulative images in each amplification cycle, and using the frame difference method to differentiate the cumulative image corresponding to the abnormal amplification cycle from the cumulative image of the previous cycle to identify abnormally amplified DNA samples; The abnormal sample deletion unit is used to delete the abnormally amplified DNA sample and delete the sample according to the abnormally amplified sample information provided by the differential image analysis unit.

6. A gene sequencing method for multiple gene repetitive sequences, implemented based on a gene sequencing system for multiple gene repetitive sequences according to any one of claims 1 to 5, characterized in that: The following steps are involved: Step 1: Use a high-throughput sequencing module to obtain raw sequencing data of the target DNA fragment; Step 2: amplify the target repetitive sequence using a nested multiplex PCR amplification module; Step 3: Analyze sequencing data using a machine learning-based data processing module to identify and correct sequencing errors; Step 4: Processing the gene data in parallel through the central processing unit, the in-memory computing unit and the storage computing unit; Step 5: Apply the objective function module to detect and correct abnormal amplification; Step 6: Use the anomaly detection module to obtain and delete abnormally amplified DNA samples; Step 7: Compare the corrected sequencing results with the known reference genome to determine the specific location and variation of the repeated sequences.

Citation Information

Patent Citations

  • Gene sequencing system and sequencing method

    CN113241120A

  • Whole genome re-sequencing analysis method and system

    CN117912560A