Gene sequencing and data analysis method, device and system, and storage medium
By placing the sample tag before the gene fragment in the gene sequencing object, identifying the sample source and determining the target sequencing requirements using fluorescence image detection, the problem in the prior art that needs to wait for all sequencing to be completed before analysis is solved, and efficient gene sequencing and data analysis are achieved.
Patent Information
- Application Number
- PCT/CN2024/119767
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-19
- Filing Date
- 2024-09-19
- Publication Date
- 2025-07-24
AI Technical Summary
The existing gene sequencing technology needs to wait until all sequencing objects are completed before the sample source can be identified, resulting in too long sequencing and data analysis.
In the sequencing object, the sample tag and gene fragment are detected successively through fluorescence images, the sample source is identified based on the sample tag and the target sequencing requirements are determined, and the sequencing results that meet the target sequencing requirements are split in advance and analyzed.
It realizes the early splitting and analysis of sequencing results during gene sequencing, shortening user waiting time and improving sequencing and data analysis efficiency.
Smart Images

Figure CN2024119767_24072025_PF_FP_ABST
Abstract
Description
Gene sequencing and data analysis methods, equipment, systems and storage media
[0001] This disclosure claims priority to a Chinese patent application filed with the Patent Office of China on January 19, 2024, with application number 202410079454.6 and application name “Nucleic Acid Detection and Data Analysis Methods, Equipment, Systems and Storage Medium,” the entire contents of which are incorporated herein by reference. Technical Field
[0002] The present disclosure relates to the field of gene sequencing technology, and in particular to a gene sequencing and data analysis method, device, system and storage medium. Background Art
[0003] With the rapid development of second-generation gene sequencing technology, it has been applied to various scenarios to solve corresponding problems, such as solving biological problems. Second-generation gene sequencing technology is based on a balance between sequencing time and sequencing costs. It usually sequences multiple sequencing objects derived from multiple sequencing samples simultaneously. In order to distinguish the sample sources of the sequencing objects, the sequencing objects contain not only the gene fragments to be sequenced, but also sample tags. The sample tags are used to indicate the sample from which the corresponding sequencing object originated.
[0004] During gene sequencing, gene fragments are usually sequenced first, followed by sample tag sequencing. It is necessary to wait until all sequencing objects are sequenced before all sequencing objects from the same sequencing sample can be identified based on the detected sample tags, and further data analysis can be performed on the gene fragment sequences of all sequencing objects from the same sequencing sample.
[0005] Summary of the Invention
[0006] The embodiments of the present application at least provide a gene sequencing and data analysis method, device, system and storage medium.
[0007] In a first aspect, a gene sequencing and data analysis device is provided, comprising: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, the machine-readable instructions being used to perform gene sequencing and data analysis on sequencing objects derived from multiple gene sequencing samples, the sequencing objects comprising sample tags and gene fragments, the sample tags being used to indicate the sample source of the sequencing objects, and the sequencing position of the sample tags in the sequencing objects being located before the gene fragments, and the sequencing objects having target sequencing requirements that match their sample sources; when the gene sequencing and data analysis device is in operation, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the following method is performed:
[0008] sequentially detecting sample tags and gene fragments in the sequencing object according to the fluorescent image of the sequencing object;
[0009] Identifying the sample source of the sequencing object according to the detected sample tags, and determining the target sequencing requirements of the sequencing object according to the identified sample source;
[0010] During the gene sequencing process, if the current sequencing results of the sequencing objects of any gene sequencing sample among the multiple gene sequencing samples meet the target sequencing requirements, the current sequencing results of the sequencing objects of the gene sequencing samples that have met the target sequencing requirements are split out, and data analysis is performed based on the current sequencing results of the sequencing objects of the gene sequencing samples that have met the target sequencing requirements.
[0011] In an optional embodiment, the target sequencing requirement includes at least one of sequencing length and number of sequencing ends.
[0012] In an optional embodiment, in the method executed by the processor, the detecting of the sample tags and gene fragments in the sequencing object in sequence according to the fluorescent image of the sequencing object includes:
[0013] Detecting an i-th base in the sequencing object according to the i-th fluorescent image of the sequencing object;
[0014] Determining a current sequencing quality of the sequencing object according to the i-th base in the sequencing object;
[0015] When the current sequencing quality of the sequencing object meets the preset quality requirement, continue sequencing the (i+1)th fluorescence image of the sequencing object;
[0016] If the current sequencing quality of the sequencing object does not meet the preset quality requirement, if the number of consecutive occurrences of the sequencing quality not meeting the preset quality requirement reaches a preset threshold, sequencing of the i+1th fluorescent image of the sequencing object is stopped; if the number of consecutive occurrences of the sequencing quality not meeting the preset quality requirement is less than the preset threshold, sequencing of the i+1th fluorescent image of the sequencing object continues; where i is a positive integer starting from 1.
[0017] In an optional embodiment, in the method executed by the processor, the detecting of the sample tags and gene fragments in the sequencing object in sequence according to the fluorescent image of the sequencing object includes:
[0018] extracting a fluorescent image of a sample label and a fluorescent image of a gene fragment from the fluorescent image of the sequencing object;
[0019] performing noise reduction processing on the fluorescence image of the sample label to generate a noise-reduced fluorescence image of the sample label;
[0020] The de-noised fluorescent image of the sample label and the fluorescent image of the gene fragment are sequenced respectively to obtain the sample label and the gene fragment in the sequencing object.
[0021] In an optional embodiment, in the method executed by the processor, performing noise reduction processing on the fluorescence image of the sample label to generate a noise-reduced fluorescence image of the sample label includes:
[0022] Determining a value of a weight calculation parameter of a filter kernel based on the length of the sample tag detected in the current sequencing result;
[0023] Determining weight values of a plurality of to-be-determined weights of the filter core based on a value of a weight calculation parameter of the filter core, wherein the to-be-determined weight at a center position in the filter core is determined based on a sum of other to-be-determined weights except the to-be-determined weight at the center position;
[0024] Based on the filter kernel, a convolution operation is performed on the fluorescence image of the sample label to generate a noise-reduced fluorescence image of the sample label.
[0025] In an optional embodiment, the sequencing object further has target analysis requirements that match its sample source;
[0026] In the method executed by the processor, performing data analysis on the sequencing object of the gene sequencing sample based on the data of the current sequencing result includes:
[0027] According to the data of the current sequencing results and the target analysis requirements, data analysis is performed on the sequencing object of the gene sequencing sample.
[0028] In an optional embodiment, the method executed by the processor includes, before sequentially sequencing the sample tags and gene fragments in the sequencing object based on the fluorescent image of the sequencing object, the following steps:
[0029] Based on at least one of the sample data volume of the multiple gene sequencing samples and the target sequencing requirements of the sequencing object of each gene sequencing sample, the size of the first computing resource required for the gene sequencing process is determined, and based on the size of the first computing resource, the first computing resource is allocated from the total computing resources for use in the gene sequencing process.
[0030] In an optional embodiment, the method executed by the processor further includes, before performing data analysis on the sequencing object of the gene sequencing sample based on the data of the current sequencing result:
[0031] Based on the data volume of the current sequencing results, the target sequencing requirements of the sequencing objects of each gene sequencing sample, and at least one of the target analysis requirements, the size of the second computing resources required for the data analysis process is determined, and based on the size of the second computing resources, the second computing resources are allocated to the data analysis process from the remaining computing resources, where the remaining computing resources are the computing resources after deducting the first computing resources from the total computing resources.
[0032] In a second aspect, a gene sequencing and data analysis method is provided, characterized in that it is applied to a sequencing object derived from multiple gene sequencing samples, the sequencing object comprising a sample tag and a gene fragment, the sample tag being used to indicate the sample source of the sequencing object, and the sequencing position of the sample tag in the sequencing object being located before the gene fragment, and the sequencing object having a target sequencing requirement that matches its sample source; the method comprising:
[0033] sequentially detecting sample tags and gene fragments in the sequencing object according to the fluorescent image of the sequencing object;
[0034] Identifying the sample source of the sequencing object according to the detected sample tags, and determining the target sequencing requirements of the sequencing object according to the identified sample source;
[0035] During the gene sequencing process, if the current sequencing result of the sequencing object of any gene sequencing sample among the multiple gene sequencing samples meets the target sequencing requirements, the current sequencing result of the sequencing object of the gene sequencing sample that has met the target sequencing requirements is split out, and data analysis is performed based on the current sequencing result of the sequencing object of the gene sequencing sample that has met the target sequencing requirements.
[0036] In an optional embodiment, the target sequencing requirement includes at least one of sequencing length and number of sequencing ends.
[0037] In an optional embodiment, the detecting the sample tags and gene fragments in the sequencing object in sequence according to the fluorescent image of the sequencing object includes:
[0038] Detecting an i-th base in the sequencing object according to the i-th fluorescent image of the sequencing object;
[0039] Determining a current sequencing quality of the sequencing object according to the i-th base in the sequencing object;
[0040] When the current sequencing quality of the sequencing object meets the preset quality requirement, continue sequencing the (i+1)th fluorescence image of the sequencing object;
[0041] If the current sequencing quality of the sequencing object does not meet the preset quality requirement, if the number of consecutive occurrences of the sequencing quality not meeting the preset quality requirement reaches a preset threshold, sequencing of the i+1th fluorescent image of the sequencing object is stopped; if the number of consecutive occurrences of the sequencing quality not meeting the preset quality requirement is less than the preset threshold, sequencing of the i+1th fluorescent image of the sequencing object continues; where i is a positive integer starting from 1.
[0042] In an optional embodiment, the sequentially sequencing the sample tags and gene fragments in the sequencing object according to the fluorescent image of the sequencing object includes:
[0043] extracting a fluorescent image of a sample label and a fluorescent image of a gene fragment from the fluorescent image of the sequencing object;
[0044] performing noise reduction processing on the fluorescence image of the sample label to generate a noise-reduced fluorescence image of the sample label;
[0045] The de-noised fluorescent image of the sample label and the fluorescent image of the gene fragment are sequenced respectively to obtain the sample label and the gene fragment in the sequencing object.
[0046] In an optional embodiment, performing noise reduction processing on the fluorescence image of the sample label to generate a noise-reduced fluorescence image of the sample label includes:
[0047] Determining a value of a weight calculation parameter of a filter kernel based on the length of the sample tag detected in the current sequencing result;
[0048] Determining weight values of a plurality of to-be-determined weights of the filter core based on a value of a weight calculation parameter of the filter core, wherein the to-be-determined weight at a center position in the filter core is determined based on a sum of other to-be-determined weights except the to-be-determined weight at the center position;
[0049] Based on the filter kernel, a convolution operation is performed on the fluorescence image of the sample label to generate a noise-reduced fluorescence image of the sample label.
[0050] In an optional embodiment, the sequencing object further has target analysis requirements that match its sample source;
[0051] The performing data analysis on the sequencing object of the gene sequencing sample according to the data of the current sequencing result includes:
[0052] According to the data of the current sequencing results and the target analysis requirements, data analysis is performed on the sequencing object of the gene sequencing sample.
[0053] In an optional embodiment, before successively detecting the sample tags and gene fragments in the sequencing object based on the fluorescent image of the sequencing object, the method includes:
[0054] Based on at least one of the sample data volume of the multiple gene sequencing samples and the target sequencing requirements of the sequencing object of each gene sequencing sample, the size of the first computing resource required for the gene sequencing process is determined, and based on the size of the first computing resource, the first computing resource is allocated from the total computing resources for use in the gene sequencing process.
[0055] In an optional embodiment, before performing data analysis on the sequencing object of the gene sequencing sample based on the data of the current sequencing result, the method further includes:
[0056] Based on the data volume of the current sequencing results, the target sequencing requirements of the sequencing objects of each gene sequencing sample, and at least one of the target analysis requirements, the size of the second computing resources required for the data analysis process is determined, and based on the size of the second computing resources, the second computing resources are allocated to the data analysis process from the remaining computing resources, where the remaining computing resources are the computing resources after deducting the first computing resources from the total computing resources.
[0057] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the gene sequencing and data analysis method described in any of the above methods are executed.
[0058] In a fourth aspect, a gene sequencing and data analysis system is provided, comprising: a gene sequencing device and a server, the gene sequencing device being communicatively connected to the server, the gene sequencing device being configured to sequence sequencing objects derived from a plurality of gene sequencing samples, the sequencing objects comprising sample tags and gene fragments, the sample tags being configured to indicate the sample source of the sequencing objects, the sequencing position of the sample tags in the sequencing objects being located before the gene fragments, the sequencing objects having target sequencing requirements and target analysis requirements that match their sample sources; wherein:
[0059] The gene sequencing device is configured to sequentially detect sample tags and gene fragments in the sequencing object based on the fluorescent image of the sequencing object; identify the sample source of the sequencing object based on the detected sample tags, and determine the target sequencing requirements of the sequencing object based on the identified sample source; during the gene sequencing process, if the current sequencing result of the sequencing object of any gene sequencing sample among the multiple gene sequencing samples has met the target sequencing requirements, split the current sequencing result of the sequencing object of the gene sequencing sample that has met the target sequencing requirements, and send the current sequencing result and sample source of the sequencing object of the gene sequencing sample that has met the target sequencing requirements to the server;
[0060] The server is configured to determine the target analysis requirements of the sequencing object of the genetic sequencing sample that has met the target sequencing requirements based on the sample source, and perform data analysis on the current sequencing results of the sequencing object of the genetic sequencing sample that has met the target sequencing requirements based on the target analysis requirements.
[0061] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the technical solutions of the present application.
[0062] The sequencing gene sequencing and data analysis method provided in the embodiments of the present application is applied to sequencing objects derived from a variety of sequencing gene sequencing samples. Each sequencing object includes a sample tag and a nucleic acid sequence. The sample tag is used to indicate the sample source of the sequencing object, and the sequencing position of the sample tag in the sequencing object is located before the nucleic acid sequence. Each sequencing object has a target sequencing requirement that matches its sample source to meet different sequencing needs.
[0063] Since the sequencing position of the sample label in the sequencing object in the present application is located before the nucleic acid sequence, after the sample label and the nucleic acid sequence in the sequencing object are sequenced in sequence according to the fluorescent images of each sequencing object, the sample source of the sequencing object can be identified according to the sequenced sample label, and the target sequencing requirements of each sequencing object can be determined according to the identified sample source. In this way, in the process of sequencing gene sequencing, based on the current sequencing results of each sequencing object, if the current sequencing result of the sequencing object of any sequencing gene sequencing sample among the multiple sequencing gene sequencing samples meets the target sequencing requirements, data analysis can be performed on the sequencing object of the sequencing gene sequencing sample based on the data of the current sequencing result. There is no need to wait until the nucleic acid sequences of all sequencing objects participating in the sequencing gene sequencing are all sequenced before performing data analysis, thereby improving the efficiency of sequencing gene sequencing and data analysis, and realizing sequencing and analysis of multiple sequencing gene sequencing samples.
[0064] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application. It should be understood that the following drawings only illustrate certain embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without inventive effort.
[0066] FIG1 is a schematic diagram showing the positional relationship between sample tags and nucleic acid sequences in a sequencing object provided by some embodiments of the present application;
[0067] FIG2 shows a schematic flow chart of the gene sequencing and data analysis methods provided in some embodiments of the present application;
[0068] FIG3 is a schematic diagram showing the sequencing lengths set in the gene sequencing and data analysis methods provided in some embodiments of the present application;
[0069] FIG4 is a schematic diagram showing the internal software of a sequencer device in the gene sequencing and data analysis methods provided in some embodiments of the present application;
[0070] FIG5 shows a schematic structural diagram of a gene sequencing and data analysis device provided in some embodiments of the present application;
[0071] FIG6 shows a schematic structural diagram of a gene sequencing and data analysis system provided in some embodiments of the present application. DETAILED DESCRIPTION
[0072] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application is not intended to limit the scope of the application for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work fall within the scope of protection of the present application.
[0073] As described in the background technology, the second-generation gene sequencing technology usually sequences gene fragments first and then sequences sample tags. It is necessary to wait until all sequencing objects are sequenced before identifying all sequencing objects derived from the same sequencing sample based on the detected sample tags, and further perform data analysis on the gene fragment sequences of all sequencing objects derived from the same sequencing sample.
[0074] However, this results in users having to wait longer for the equipment to complete gene sequencing and data analysis. For example, sequencing a conventional biological sample for 150 base pairs (pairs) can take up to 48 hours. After sequencing is complete, specialized data analysis is required, further increasing the time required to produce the data analysis report.
[0075] However, the inventors of this application discovered during their research that, based on actual user application needs, the gene sequencing requirements for sequencing objects derived from different types of biological sequencing samples may differ. For example, for sequencing objects derived from one or several types of biological sequencing samples, in order to meet specific sequencing length requirements for gene fragments, all sequencing objects may need to undergo 150bp of paired-end sequencing, while for sequencing objects derived from other types of biological sequencing samples, only 75bp of single-end sequencing may be required. For another example, for sequencing objects derived from one or several types of biological sequencing samples, only 75bp of single-end sequencing may be required, while for sequencing objects derived from other types of biological sequencing samples, only 30-50bp of single-end sequencing may be required.
[0076] Based on the different sequencing requirements, for sequencing objects that can complete gene sequencing in advance compared to other types of biological sequencing samples, if they can be separated from the sequencing objects of all biological sequencing samples in advance and enter the data analysis process in advance, data analysis can be achieved while gene sequencing is being performed, which can greatly shorten the user's waiting time.
[0077] Based on the above research, the present application provides a gene sequencing and data analysis method, which is applied to sequencing objects derived from multiple gene sequencing samples. Each sequencing object contains a sample label and a gene fragment. The sample label is used to indicate the sample source of the sequencing object, and the sequencing position of the sample label in the sequencing object is located before the gene fragment. Each sequencing object has a target sequencing requirement that matches its sample source to meet different sequencing needs.
[0078] Since the sequencing position of the sample label in the sequencing object in the present application is located before the gene fragment, after the sample label and gene fragment in the sequencing object are detected in sequence according to the fluorescent image of each sequencing object, the sample source of the sequencing object can be first identified according to the detected sample label, and the target sequencing requirement of each sequencing object can be determined according to the identified sample source. In this way, during the gene sequencing process, based on the current sequencing results of each sequencing object, if the current sequencing results of all sequencing objects of any gene sequencing sample among multiple gene sequencing samples meet the target sequencing requirement, the current sequencing results of the sequencing objects of the gene sequencing sample that has met the target sequencing requirement can be split out in advance, and data analysis can be performed in advance based on the current sequencing results of the sequencing objects of the gene sequencing sample that has met the target sequencing requirement. There is no need to wait until the sequencing objects of all sequencing samples participating in the gene sequencing are all sequenced before performing unified data analysis, thereby improving the efficiency of gene sequencing and data analysis, and realizing data analysis while sequencing multiple gene sequencing samples.
[0079] It should be noted that the defects in the existing technology are the results obtained by the inventor after practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by this application for the above problems below should be the contributions made by the inventor to this application during the application process.
[0080] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0081] To facilitate understanding of this embodiment, we first introduce in detail a gene sequencing and data analysis method disclosed in an embodiment of this application. This method can be applied to gene sequencing and data analysis equipment, for example, to achieve simultaneous sequencing and data analysis of multiple gene sequencing samples.
[0082] The method of the present application is applied to fluorescent images of sequencing objects derived from multiple gene sequencing samples. The sequencing objects include sample tags and gene fragments. The sample tags are used to indicate the sample source of the sequencing object, and the sequencing position of the sample tags in the sequencing object is located before the gene fragment. Each sequencing object has a target sequencing requirement that matches its sample source, and the target sequencing requirement here is not limited to at least one of the sequencing length and the number of sequencing ends.
[0083] Among them, in some implementations of the present application, the gene sequencing sample can be, for example, the blood, cell tissue, etc. of an organism, and a DNA (Deoxyribo Nucleic Acid, deoxyribonucleotide) single chain or RNA single chain (Ribo Nucleic Acid, ribonucleotide) can be extracted and separated from the gene sequencing sample. In other implementations of the present application, the sequencing object can be, for example, a DNA single chain or RNA single chain cut into multiple fragments, and then the DNA fragments or RNA fragments are chemically modified and a linker is added, that is, the sequencing object can be a nucleic acid library obtained by a library preparation process, or the sequencing object can also be a whole DNA single chain or a whole RNA single chain, wherein RNA can be obtained by DNA transcription. Through the above-mentioned processing process, each sequencing object includes a sample tag and a gene fragment, and the sample tag is used to indicate the sample source of the sequencing object. In some implementations of the present application, the sample tag can be, for example, a short oligonucleotide sequence containing multiple bases.
[0084] In the implementation of the present application, the sequencing position of the sample tag in the sequencing object is located before the gene fragment, that is, when sequencing the sequencing object, the fluorescent image of the sample tag can be first collected and the sample tag can be detected from the fluorescent image of the sample tag, and then the fluorescent image of the gene fragment can be collected and the nucleic acid sequence can be detected from the fluorescent image of the gene fragment.
[0085] Exemplarily, see the schematic diagram of the positional relationship between the sample label and the gene fragment in the sequencing object shown in Figure 1 (the figure does not illustrate the sequencing primers of the sample label and the gene fragment). Figure 1-a illustrates the situation where the sequencing object includes only one sample label. When the sequencing object shown in Figure 1-a is sequenced, in the implementation of the present application, the sample label (index) is detected first, and then the gene fragment is detected. Moreover, for the gene fragment, by setting, only one end of the gene fragment can be detected, or both ends of the gene fragment can be detected. Figure 1-b illustrates the situation where the sequencing object includes two sample labels. When the sequencing object shown in Figure 1-a is sequenced, in the implementation of the present application, sample label 1 (index1) is detected first, then sample label 2 (index2) is detected, and finally the gene fragment is detected. For the gene fragment, by setting, only one end of the gene fragment can be detected, or both ends of the gene fragment can be detected.
[0086] In some implementations of the present application, the target sequencing requirement includes at least one of the sequencing length and the number of sequencing ends. The sequencing length is used to indicate the number of genes (or base pairs) to be sequenced, and the unit is bp. The sequencing length of each sequencing object can be one or more. For example, the sequencing length of a sequencing object can be: 35bp, 50bp, 75bp, 100bp, etc. When the target sequencing requirement includes that the sequencing length is 35bp, when the current sequencing results of all sequencing objects of a certain gene sequencing sample are 35bp in length, that is, the target sequencing requirement of all sequencing objects of the gene sequencing sample is met, the current sequencing results of all sequencing objects of the gene sequencing sample can be data split, and data analysis of the gene sequencing sample can be performed based on the current sequencing results of all sequencing objects of the gene sequencing sample.
[0087] The number of sequencing ends includes single-end (corresponding to single-end sequencing) and double-end (corresponding to double-end sequencing). The gene sequencing and data analysis method is described in detail below. Referring to FIG2 , a flow chart of the gene sequencing and data analysis method provided in an embodiment of the present application is shown. The method includes S210-S230, wherein:
[0088] S210 , sequentially detecting sample tags and gene fragments in the sequencing object according to the fluorescent image of the sequencing object.
[0089] S220 , identifying the sample source of the sequencing object according to the detected sample label, and determining the target sequencing requirement of the sequencing object according to the identified sample source.
[0090] S230, during the gene sequencing process, if the current sequencing result of the sequencing object of any gene sequencing sample among the multiple gene sequencing samples meets the target sequencing requirements, the current sequencing results of the sequencing object of the gene sequencing sample that has met the target sequencing requirements are split out, and data analysis is performed based on the current sequencing results of the sequencing object of the gene sequencing sample that has met the target sequencing requirements.
[0091] In S210, after obtaining fluorescence images of sequencing objects from multiple gene sequencing samples, the fluorescence images of the sequencing objects are subjected to image analysis, base recognition, and other processing to detect the sample tags and gene fragments in the sequencing objects. For example, in each sequencing cycle of the sequencing process, a fluorescently labeled base is paired and synthesized with the sample tag on the sequencing object and a base on the gene fragment, so that the synthesized base of the sample tag and the gene fragment also carries a fluorescent label. The optical illumination system in the gene sequencing device illuminates the fluorescent label carried by the base to stimulate a fluorescent signal. The optical detection system in the gene sequencing device collects the fluorescent signal. After each sequencing cycle, a fluorescence image of the sample tag or a base on the gene fragment of all sequencing objects can be obtained. After multiple sequencing cycles, fluorescence images of the sample tags and multiple bases on the gene fragment to be sequenced can be obtained for all sequencing objects, so that the base on the sample tag and the base on the gene fragment in any sequencing object can be detected in sequence based on the multiple fluorescence images of the sequencing object.
[0092] It should be noted that the lengths of sample tags in different sequencing targets can vary, and / or the lengths of gene fragments in different sequencing targets can vary. For example, the length of the sample tag in sequencing target one can be 8 bp, while the length of the sample tag in sequencing target two can be 6 bp. For another example, the sequencing lengths and number of sequencing ends of gene fragments in sequencing targets can include: single-end 50-75 bp, single-end 35-50 bp, double-end 2*75 = 150 bp, etc.
[0093] In specific implementations, a human-computer interface can be provided on the gene sequencing and data analysis equipment, through which users can set corresponding target sequencing requirements for all sequencing objects of each gene sequencing sample. For example, the sequencing length and / or number of sequencing ends can be set for all sequencing objects of a gene sequencing sample.
[0094] Since the method proposed in this application is applied to sequencing objects derived from multiple genetic sequencing samples, and the target sequencing requirements of sequencing objects from different sample sources are not exactly the same, it is possible that sequencing object one has met the sequencing requirements, while sequencing object two still does not meet the sequencing requirements and still needs to be sequenced. However, sequencing object one still exists on the genetic sequencing and data analysis equipment (for example, a genetic sequencer). In this case, it is easy to cause the subsequent sequencing object one to follow the continuous sequencing of sequencing object two, which will affect the sequencing quality, increase the detection amount of redundant fluorescence signals, and waste the computing resources of the equipment.
[0095] To alleviate these issues, quality metrics can be set, such as base quality scores (Qscore30, Q30), Q20, and N content. These metrics can then be used to calculate the current sequencing quality based on the current sequencing results. This allows sequencing to be stopped if the current sequencing quality does not meet the preset quality requirements. Stopping sequencing here means ceasing any base-calling-related processing of the acquired fluorescence images.
[0096] In some implementations of the present application, sequentially detecting sample labels and gene fragments in the sequencing object based on the fluorescent image of the sequencing object may include: detecting the i-th base in the sequencing object based on the i-th fluorescent image of the sequencing object; determining the current sequencing quality of the sequencing object based on the i-th base in the sequencing object; if the current sequencing quality of the sequencing object meets a preset quality requirement, continuing sequencing the i+1-th fluorescent image of the sequencing object; if the current sequencing quality of the sequencing object does not meet the preset quality requirement, if the number of consecutive occurrences of the sequencing quality not meeting the preset quality requirement reaches a preset threshold, stopping sequencing the i+1-th fluorescent image of the sequencing object; if the number of consecutive occurrences of the sequencing quality not meeting the preset quality requirement is less than the preset threshold, continuing sequencing the i+1-th fluorescent image of the sequencing object; wherein i is a positive integer starting from 1.
[0097] During implementation, the first base in the sequencing object is detected based on the first fluorescence image of the sequencing object. The first fluorescence image may be a fluorescence image acquired at the first acquisition time. After detecting the first base, the current sequencing quality of the sequencing object can be calculated based on the first base. For example, the current sequencing quality can be characterized using quality indicators such as Q30 and Q20.
[0098] Determine whether the current sequencing quality meets the preset quality requirements, for example, by determining whether the calculated Q30 index value is greater than a preset quality threshold. If so, the current sequencing quality is determined to meet the preset quality requirements. If the current sequencing quality meets the preset quality requirements, sequencing the next fluorescent image of the sequencing object continues. If the current sequencing quality does not meet the preset quality requirements, if the number of consecutive occurrences of the sequencing quality failing to meet the preset quality requirements reaches a preset threshold, sequencing the next fluorescent image of the sequencing object is discontinued.
[0099] It should be noted that the above-mentioned preset number threshold can be set according to actual needs. For example, if the preset number threshold is 1, then if the sequencing quality of a sequencing object fails to meet the preset quality requirement once, sequencing of the next fluorescent image of the sequencing object can be stopped. For another example, if the preset number threshold is 3, then if the sequencing quality of a sequencing object fails to meet the preset quality requirement three times in a row, sequencing of the next fluorescent image of the sequencing object can be stopped. Of course, the above is only an example, and this application does not impose any specific restrictions on the value of the preset number threshold. This approach can also reduce redundant calculations during the sequencing process and reduce the consumption of computing resources.
[0100] In some implementations of the present application, when the current sequencing quality of a sequencing object does not meet the preset quality requirements, the sequencing object can be marked first, and sequencing of multiple subsequent fluorescence images of the marked sequencing object can be continued, and the sequencing quality of the subsequent multiple fluorescence images can be determined. If the sequencing quality of multiple consecutive (for example, 3, 5, etc.) fluorescence images of the sequencing object does not meet the preset quality requirements, sequencing of all subsequent fluorescence images of the sequencing object is stopped.
[0101] Since gene sequencing and data analysis equipment have poor stability at the beginning of sequencing, sequencing errors are usually larger than sequencing errors after the equipment is stable. Considering that the sample labels are detected first after being pre-placed, the sequencing errors of gene sequencing and data analysis equipment for sample labels are usually larger. In order to improve the sequencing accuracy of sample labels, so that the fluorescence images of the subsequent gene fragments can be corrected through the fluorescence images of the sample labels, so that the position of the gene fragment in the fluorescence image can be located more accurately, and the gene fragment can be detected more accurately, the fluorescence image of the sample label can be denoised in advance.
[0102] In a specific implementation, the sample labels and gene fragments in the sequencing object are detected in sequence according to the fluorescence image of the sequencing object, including: extracting the fluorescence image of the sample label from the fluorescence image of the sequencing object; performing noise reduction processing on the fluorescence image of the sample label to generate a noise-reduced fluorescence image of the sample label; sequencing the noise-reduced fluorescence image of the sample label and the fluorescence image of the gene fragment of the nucleic acid to obtain the sample labels and gene fragments in the sequencing object.
[0103] The fluorescence image of the sequencing target includes the fluorescence image of the sample tag and the fluorescence image of the gene fragment. The fluorescence image of the sample tag can be extracted from the fluorescence image of the sequencing target. For example, the fluorescence image of the sample tag can be determined based on characteristics of the sample tag, such as length and base combination. Noise reduction processing is then performed on the fluorescence image of the sample tag to generate a noise-reduced fluorescence image of the sample tag.
[0104] Then, the base analysis program can sequence the fluorescence image of the sample label after noise reduction and the fluorescence image of the gene fragment respectively to obtain the sample label and gene fragment in the sequencing object; for example, the base analysis program can establish a template image based on the fluorescence image of the sample label after noise reduction, use the template image to perform image correction on the fluorescence image of the sample label and the fluorescence image of the gene fragment, obtain the corrected fluorescence image of the sample label and the fluorescence image of the gene fragment, and perform base analysis on the fluorescence image of the sample label after image correction and the fluorescence image of the gene fragment to obtain the sample label and the gene fragment in the sequencing object. Since the fluorescence image of the sample label may have the problem of uneven light spot, the established template image is inaccurate. This application can alleviate the above problem by performing noise reduction on the fluorescence image of the sample label, improve the accuracy of the template image and the accuracy of subsequent image correction, and thus ensure the recognition accuracy of the gene fragment. Moreover, this application performs noise reduction on the fluorescence image of the sample label. Compared with the related art of performing noise reduction on the entire fluorescence image, while ensuring the sequencing accuracy of the gene fragment, it can reduce computing resources and improve sequencing efficiency.
[0105] The specific process of noise reduction processing is exemplified below. The noise reduction processing of the fluorescence image of the sample label to generate the noise-reduced fluorescence image of the sample label includes: determining the value of the weight calculation parameter of the filter kernel based on the length of the sample label sequenced in the current sequencing result; determining the weight values of multiple weights to be determined of the target filter kernel based on the value of the weight calculation parameter of the filter kernel, wherein the weight to be determined at the center position of the filter kernel is determined based on the sum of other weights to be determined except the weight to be determined at the center position; and performing noise reduction processing on the fluorescence image of the sample label based on the filter kernel to generate the noise-reduced fluorescence image of the sample label.
[0106] Considering the presence of noise in sequencing scenarios, research has found that the main types are Gaussian noise and salt and pepper noise. Furthermore, the longer the sample label, the richer the base types it contains, and the greater the noise. To effectively mitigate the aforementioned noise, the value of the filter kernel weight calculation parameter can be determined based on the length of the sample label detected in the current sequencing result, where the value of the filter kernel weight calculation parameter can be a value between 0 and 1. Considering that when the value of the filter kernel weight calculation parameter is small, the weight value at the center of the filter kernel is also small, and thus the noise data can be processed into the background, making the filter kernel determined based on the filter kernel weight calculation parameter more suitable for the mixed gene sequencing scenario of multi-gene sequencing samples of gene sequencing and data analysis equipment, thereby improving the image noise reduction effect in this scenario. Therefore, the length of the sample label can be set to be negatively correlated with the value of the filter kernel weight calculation parameter. For example, a mapping relationship can be preset that negatively correlates the length of the sample label with the value of the filter kernel weight calculation parameter, so that according to this mapping relationship, the value of the filter kernel weight calculation parameter is determined by the length of the sample label.
[0107] Based on the values of the weight calculation parameters of the filter kernel, weight values of multiple to-be-determined weights of the filter kernel are determined, wherein the filter kernel includes multiple to-be-determined weights related to the weight calculation parameters of the filter kernel, and the to-be-determined weight at the center of the filter kernel is determined based on the sum of the other to-be-determined weights excluding the to-be-determined weight at the center. Furthermore, after the weight values of the multiple to-be-determined weights of the filter kernel are determined, a convolution operation can be performed on the fluorescence image of the sample label using the filter kernel to implement noise reduction processing on the fluorescence image of the sample label, thereby obtaining a noise-reduced fluorescence image of the sample label.
[0108] Generally, the filter kernel can be constructed by using the expression of the fractional differential of the one-dimensional function f(t). The expression of the fractional differential of the one-dimensional function f(t) is shown in the following formula (1):
[0109] For example, the size of the filter kernel can be selected as 5×5, that is, the filter kernel is a 5×5 matrix, so the template of the 5×5 filter kernel is as follows:
[0110] Among them, the filter kernel includes multiple weights to be determined, that is, Aa0, a1, and a2 are weights to be determined. Then according to formula (1), we can set a0=1, a1=-v, Substituting a0, a1, and a2 into the template of the 5×5 filter kernel, we can obtain the specific content of the 5×5 filter kernel, as shown below:
[0111] In order to make the 5×5 filter kernel more suitable for gene sequencing scenarios, the value of the weight to be determined at the center position of the filter kernel can be determined, that is, the value of A can be determined, and finally the filter kernel for convolution processing can be obtained. The sum of all weights contained in the filter kernel is -12v+4v 2 +A, in order to more accurately sharpen the fluorescent image of the sample label, the sum of the above weights can be made 1, that is, -12v+4v 2 +A=1, and further reasoning can be obtained: A=12v-4v 2 +1, the final filter kernel is as follows:
[0112] Among them, v is the weight calculation parameter of the filter kernel. It can be seen that the filter kernel includes multiple weights to be determined, namely -v, 12v-4v 2 +1, and the weight values of the multiple weights to be determined are determined by the value of the weight calculation parameter of the filter kernel. Therefore, after determining the value of the weight calculation parameter v of the filter kernel, the weight values of the multiple weights to be determined in the filter kernel can be determined according to the value of the weight calculation parameter of the filter kernel. For example, if the value of v is 0.2, the weight values of the multiple weights to be determined in the filter kernel are:
[0113] In some implementations of the present application, the weight to be determined at the center position of the filter kernel is determined based on the sum of other weights to be determined except for the weight to be determined at the center position. Since the weight values of other weights to be determined are related to the parameter values of the weight calculation parameters of the filter kernel, the weight value of the weight to be determined at the center position is also related to the value of the weight calculation parameters of the filter kernel, that is, the weight value of the weight to be determined at the center position will change with the change of the value of the weight calculation parameters of the filter kernel, and the value of the weight calculation parameters of the filter kernel is negatively correlated with the length of the sample label, so as to ensure the image sharpening effect in different situations and improve the image denoising effect.
[0114] After detecting and obtaining the sample tags and gene fragments of the sequencing target, the process may further include S220: identifying the sample source of the sequencing target based on the detected sample tags, and determining the target sequencing requirements for the sequencing target based on the identified sample source. Different sequencing targets may correspond to different target sequencing requirements, and the target sequencing requirements for the sequencing targets can be set when logging in.
[0115] Target sequencing requirements include sequencing length and / or number of sequencing ends.
[0116] If the target sequencing requirement includes sequencing length, one or more sequencing lengths can be set so that data analysis can be performed when the current sequencing result reaches the sequencing length indicated by the target sequencing requirement.
[0117] Referring to FIG. 3 , multiple sequencing lengths are shown, for example, including single-ended 35 cycles, single-ended 50 cycles, single-ended 75 cycles, single-ended 100 cycles, single-ended 150 cycles, double-ended 75 cycles, double-ended 100 cycles, and double-ended 150 cycles. In this application, the data of the current sequencing results of different sequencing lengths can be subjected to different data analyses, or the same data analysis can be performed.
[0118] If the target sequencing requirement includes the number of sequencing ends, single-end sequencing or paired-end sequencing can be determined based on the number of sequencing ends. It is understood that if single-end sequencing is used, second-end sequencing and analysis will not be performed after the first-end sequencing is completed; if paired-end sequencing is used, the second-end sequencing process will continue after the first-end sequencing is completed.
[0119] In actual implementation, if the user only sets the sequencing length of a sequencing target but does not set the number of sequencing ends, a default number (e.g., single-end) can be used as the number of sequencing ends for the sequencing target. Furthermore, if the user only sets the number of sequencing ends for a sequencing target but does not set the sequencing length, the default setting can be that the sequencing data of each sequencing target can only be split and further analyzed after all sequencing targets have been fully sequenced.
[0120] In S230, during the gene sequencing process, for each sequencing object of a gene sequencing sample, a determination is made based on the current sequencing result of the sequencing object to determine whether the current sequencing result of the sequencing object meets the target sequencing requirement for the sequencing object. If so, the current sequencing results of all sequencing objects of the gene sequencing sample that meet the target sequencing requirement are separated from the current sequencing results of all sequencing objects of the multiple gene sequencing samples. Based on the separated current sequencing results, data analysis is performed on the gene sequencing samples from which all sequencing objects of the gene sequencing sample are derived. The analysis content can be set as needed, such as species richness analysis, miRNA analysis, etc.
[0121] In some implementations of the present application, the current sequencing results of the sequencing objects that meet the target sequencing requirements can be split from the current sequencing results of the sequencing objects of the multiple gene sequencing samples according to the sample tags of the sequencing objects; the sequencing objects are subjected to data analysis based on the data of the current sequencing results of the gene sequencing samples. For example, if the target sequencing requirements of the sequencing object include single-end 50 cycles and single-end 100 cycles, when the current sequencing result reaches the single-end 50 cycles, the current sequencing result detected by the single-end 50 cycles can be split according to the sample tags of the sequencing object, and data analysis is performed based on the data of the current sequencing results obtained by the split to obtain an analysis result. And when the current sequencing result reaches the single-end 100 cycles, the current sequencing result detected by the single-end 100 cycles can be further split according to the sample tags of the sequencing object, and nucleic acid data analysis is performed based on the data of the current sequencing results obtained by the split to obtain another analysis result. The data of the current sequencing result can be recorded in, for example, a fastq sequence file.
[0122] If the target sequencing requirement includes sequencing length, the number of sequencing cycles corresponding to the current sequencing result of each sequencing object can be recorded during the sequencing process. For each sequencing object, it is determined whether the number of sequencing cycles corresponding to the current sequencing result of the sequencing object meets the sequencing length included in the target sequencing requirement. If so, data splitting and analysis are performed. In conjunction with Figure 3, it can be seen that if the number of sequencing cycles corresponding to the current sequencing result of the sequencing object is 35 cycles, the target sequencing requirement is met, and the data sequenced for 35 cycles are split and analyzed; for another example, if the number of sequencing cycles corresponding to the current sequencing result of the sequencing object is 50 cycles, the target sequencing requirement is again met, and the data sequenced for 50 cycles are split and analyzed.
[0123] In order to more accurately and comprehensively split the data of the current sequencing result, a floating cycle number can be set. For example, two floating cycle numbers can be set. In conjunction with Figure 3, it can be seen that when the sequencing cycle number corresponding to the current sequencing result of the sequencing object is 37 sequencing cycles, it is determined that the target sequencing requirement has been met; when the sequencing cycle number corresponding to the current sequencing result of the sequencing object is 52 sequencing cycles, the target sequencing requirement has been met again. In conjunction with Figure 3, it can be seen that because the sample label has been pre-sequenced, at cycle 35, the data of the current sequencing result includes the sample label and the gene fragments sequenced from 1 to 35 sequencing cycles; at cycle 50, the data of the current sequencing result includes the sample label and the gene fragments sequenced from 1 to 50 sequencing cycles.
[0124] In some implementations of the present application, after splitting out the data of the current sequencing results of the sequencing object that meets the target sequencing requirements, the data of the current sequencing results of the sequencing object can also be stored in the target storage address, and the analysis program can be controlled to start, and the analysis program can be controlled to obtain the data of the current sequencing results from the target storage address, and the data of the current sequencing results can be analyzed to obtain the analysis results.
[0125] Combined with the analysis of Figure 3, at 35 cycles, the data of the current sequencing result obtained by splitting (i.e., the nucleic acid sequence obtained by sequencing from cycles 1 to 35) can be stored in the target storage address 1, and at 50 cycles, the data of the current sequencing result obtained by splitting (i.e., the gene fragment obtained by sequencing from cycles 36 to 50) can be stored in the target storage address 2. When analyzing the data of the current sequencing result of 50 cycles, the sample label can be obtained from the preset storage address, and the gene fragment can be obtained from the target storage address 1 and the target storage address 2. The obtained data is used as the data of the current sequencing result to perform data analysis on the object to be sequenced of the gene sequencing sample.
[0126] The sequencing object also has a target analysis requirement that matches its sample source; performing data analysis on the sequencing object of the gene sequencing sample based on the data of the current sequencing result includes: performing data analysis on the sequencing object of the gene sequencing sample based on the data of the current sequencing result and the target analysis requirement.
[0127] The target analysis requirements can include analysis content corresponding to each sequencing target. The analysis content for different sequencing targets can differ, or the analysis content for different sequencing targets can be partially identical. If the target sequencing requirement for a sequencing target includes a sequencing length, and there are multiple sequencing lengths, the analysis content for that sequencing target can include analysis content matching each sequencing length. For example, if the target sequencing requirements for sequencing target one include single-ended 35 cycles and single-ended 50 cycles, the analysis content for sequencing target one can include analysis content for sequencing target one at single-ended 35 cycles and single-ended 50 cycles, etc.
[0128] Then, the target analysis requirements of the sequencing object can be used to analyze the data of the current sequencing result to obtain the analysis results of the sequencing object.
[0129] A variety of analysis programs can be set up on the gene sequencer device, and before sequencing, an analysis program that matches each sequencing length of the sequencing object can be set according to the target analysis requirements of the sequencing object. Then, when the data of the current sequencing result of the sequencing object reaches the sequencing length included in the target sequencing requirements, the analysis program that matches the sequencing length can be controlled to start, and the analysis program is controlled to obtain the data of the current sequencing result from the target storage address, and the data of the current sequencing result is analyzed to generate an analysis report. The mapping relationship between the analysis program and the sequencing length can be set according to actual business needs. For example, non-invasive prenatal gene sequencing (NIPT) analysis can be performed at 35 cycles on a single end, and miRNA analysis can be performed at 50 cycles on a single end.
[0130] In an optional embodiment, before sequentially sequencing the sample labels and nucleic acid sequences in the sequencing object based on the fluorescent image of the sequencing object, the method includes: determining the size of a first computing resource required for the gene sequencing process based on at least one of the sample data volume of the multiple gene sequencing samples and the target sequencing requirements of the sequencing object of each gene sequencing sample, and allocating the first computing resource from the total computing resources to the gene sequencing process based on the size of the first computing resource.
[0131] In order to achieve a reasonable allocation of resources and improve the utilization of computing resources, a load balancer is provided in this application. Before sequencing, the load balancer determines the size of the first computing resources required for the sequencing process based on the sample data volume of multiple gene sequencing samples and / or the target sequencing requirements of the sequencing object of each gene sequencing sample. For example, the first resource size may include but is not limited to: memory usage, external memory usage, number of threads, number of cores of the central processing unit (CPU), etc. And according to the size of the first computing resource, the first computing resource is allocated from the total computing resources to the gene sequencing process to ensure that the sequencing process can proceed normally. The total computing resources are all the computing resources possessed by the gene sequencer device. The resource size of the total computing resources can be determined according to the model of the gene sequencer device, etc.
[0132] Furthermore, the first computing resource can be allocated in real time according to instruction information in the gene sequencing process.
[0133] In an optional embodiment, before performing data analysis on the gene sequencing sample based on the data of the current sequencing result, it also includes: determining the size of the second computing resource required for the data analysis process based on the data volume of the current sequencing result, the target sequencing requirements of the sequencing object of each gene sequencing sample, and at least one of the target analysis requirements, and allocating the second computing resource from the remaining computing resources to the data analysis process based on the size of the second computing resource, where the remaining computing resource is the computing resource after deducting the first computing resource from the total computing resource.
[0134] Since both the sequencing process and the analysis process in the present application can be performed on a gene sequencer device, the computing resources of the gene sequencer device are relatively tight. In order to improve resource utilization, before performing data analysis, the size of the second computing resource required for the analysis process can be determined based on the data volume of the current sequencing result, the target sequencing requirements of the sequencing object of each gene sequencing sample, and at least one of the target analysis requirements. The size of the second computing resource can be the size of all computing resources required by the analysis program for the entire analysis process. If the analysis program needs to perform multiple analysis steps when analyzing the data of the current sequencing result, the size of the computing resources required for each analysis step can be determined, and the maximum computing resource size required in the multiple analysis steps can be determined as the size of the second computing resource.
[0135] In a specific implementation, if the target sequencing requirement of a sequencing object includes a sequencing length, and there are multiple sequencing lengths, when launching a second analysis program that matches the second sequencing length, if the first analysis program for the first sequencing length is still in progress, the second analysis program can be launched after the first analysis program completes. In this case, more second computing resources can be allocated to the first analysis program. Alternatively, the second analysis program can be launched simultaneously, that is, the first and second analysis programs are both running. In this case, fewer second computing resources can be allocated to the first analysis program so that the second analysis program can start and run normally.
[0136] According to the size of the second computing resources, the second computing resources used in the analysis process are allocated from the remaining computing resources, so as to achieve reasonable allocation of analysis resources and ensure the stability of real-time sequencing and real-time analysis without interfering with the sequencing process.
[0137] In specific implementation, when splitting the current sequencing data, computing resources can be allocated to the splitting process in real time to ensure the normal progress of the splitting process.
[0138] Referring to the internal software schematic diagram of the gene sequencer device shown in FIG4 , first, gene sequencing is performed on the machine, that is, multiple gene sequencing samples are sequenced on the machine, and target sequencing requirements are set for each sequencing object in the multiple gene sequencing samples. The target sequencing requirements may include, for example: single-end 35 cycles (SE35), single-end 50 cycles (SE50), single-end 75 cycles (SE735), double-end 100 cycles (PE100), double-end 150 cycles (PE150). When sequencing the sequencing objects of multiple gene sequencing samples, the control software can perform real-time resource monitoring and dynamic resource allocation, that is, allocate the first computing resources required for the sequencing process, the computing resources required for the splitting process, and the second computing resources required for the analysis process, etc., thereby reasonably allocating the total computing resources on the gene sequencer device. The image acquisition and processing software can collect the fluorescence image of the sequencing object during the sequencing process, and perform noise reduction processing on the fluorescence image of the sample label to obtain the fluorescence image of the sample label after noise reduction. The analysis software sequences the fluorescent images of gene fragments and the de-noised fluorescent images of sample labels, generating a fastq file containing the sample labels and gene fragments of the sequencing target. The data splitting and conversion software determines whether the current sequencing result of the sequencing target meets the target sequencing requirements. If so, it splits the data of the current sequencing result of the sequencing target and stores the resulting data in a target storage address. The analysis program then retrieves the current sequencing result data from the target storage address, performs data analysis, and generates an analysis report. For example, at SE35, analysis program one can analyze the sample labels and gene fragments from single-end cycles 1 to 35, outputting analysis report one. At SE50, analysis program two can analyze the sample labels and gene fragments from single-end cycles 1 to 50, outputting analysis report two. Similarly, at SE75, analysis program three can output analysis report three, at PE100, analysis program four can output analysis report four, and at PE150, analysis program five can output analysis report five.
[0139] Based on the same inventive concept, an embodiment of the present application also provides a gene sequencing and data analysis system, which is applied to sequencing objects derived from multiple gene sequencing samples. The sequencing objects include sample tags and gene fragments. The sample tags are used to indicate the sample source of the sequencing objects, and the sequencing position of the sample tags in the sequencing objects is located before the gene fragments. The sequencing objects have target sequencing requirements and target analysis requirements that match their sample sources. The system includes: a gene sequencing device and a server, and the gene sequencing device and the server are communicatively connected.
[0140] The gene sequencing device is configured to sequentially detect sample tags and gene fragments in the sequencing object based on the fluorescent image of the sequencing object; identify the sample source of the sequencing object based on the sequenced sample tags, and determine the target sequencing requirements of the sequencing object based on the identified sample source; during the gene sequencing process, if the current sequencing result of the sequencing object of any gene sequencing sample among the multiple gene sequencing samples has met the target sequencing requirements, the current sequencing result of the sequencing object of the gene sequencing sample that has met the target sequencing requirements will be split out, and the current sequencing result and sample source of the sequencing object of the gene sequencing sample that has met the target sequencing requirements will be sent to the server.
[0141] The server is configured to determine the target analysis requirements of the sequencing object of the genetic sequencing sample that has met the target sequencing requirements based on the sample source, and perform data analysis on the current sequencing results of the sequencing object of the genetic sequencing sample that has met the target sequencing requirements based on the target analysis requirements.
[0142] The process by which a gene sequencing device performs gene sequencing on a sequencing object to obtain data on the current sequencing result can be described with reference to the above method description and will not be described in detail here. After splitting the data on the current sequencing result of the sequencing object, the gene sequencing device can send the data on the current sequencing result of the sequencing object and the sample label to the server so that the server can analyze the data on the current sequencing result and the sample label. That is, the server can determine the sample source of the sequencing object based on the sample label and determine the target analysis requirements of the sequencing object based on the sample source of the sequencing object. Then, based on the target analysis requirements of the sequencing object, data analysis is performed on the data on the current sequencing result of the sequencing object of the genetic sequencing sample to obtain an analysis result.
[0143] Since the above-mentioned system is applied to sequencing objects derived from a variety of gene sequencing samples, each sequencing object contains a sample tag and a nucleic acid sequence. The sample tag is used to indicate the sample source of the sequencing object, and the sequencing position of the sample tag in the sequencing object is located before the nucleic acid sequence. Each sequencing object has a target sequencing requirement that matches its sample source, so that the gene sequencing equipment can perform gene sequencing according to the target sequencing requirements of the sequencing object to meet different sequencing needs.
[0144] Since the sequencing position of the sample tag in the sequencing object in the present application is located before the nucleic acid sequence, the gene sequencing device can first identify the sample source of the sequencing object based on the detected sample tag after detecting the sample tag and the nucleic acid sequence in the sequencing object based on the fluorescent image of each sequencing object, and determine the target sequencing requirements for each sequencing object based on the identified sample source. Therefore, during the gene sequencing process, based on the current sequencing results of each sequencing object, if the current sequencing result of the sequencing object of any gene sequencing sample among the multiple sequenced gene sequencing samples meets the target sequencing requirements, the sequencing object of the gene sequencing sample can be split according to the data of the current sequencing result, and the current sequencing result data and sample tags of the sequencing objects of the split gene sequencing samples can be sent to the server so that the server can perform data analysis on the current sequencing result data and sample tags of the sequencing object. There is no need to wait until the nucleic acid sequences of all sequencing objects participating in the gene sequencing are fully sequenced before performing data analysis, thereby improving the efficiency of gene sequencing and data analysis.
[0145] Based on the same technical concept, the embodiment of the present application also provides a gene sequencing and data analysis device. Referring to Figure 5, it is a structural diagram of the gene sequencing and data analysis device provided by the embodiment of the present application, including a processor 501, a memory 502, and a bus 503. Among them, the memory 502 is used to store execution instructions, including a memory 5021 and an external memory 5022; the memory 5021 here is also called an internal memory, which is used to temporarily store the calculation data in the processor 501, as well as the data exchanged with the external memory 5022 such as a hard disk. The processor 501 exchanges data with the external memory 5022 through the memory 5021. When the gene sequencing and data analysis device 500 is running, the processor 501 communicates with the memory 502 through the bus 503, so that the processor 501 executes the following instructions:
[0146] sequentially detecting a sample tag and a nucleic acid sequence in the sequencing object according to the fluorescent image of the sequencing object;
[0147] Identifying the sample source of the sequencing object according to the detected sample tags, and determining the target sequencing requirements of the sequencing object according to the identified sample source;
[0148] During the gene sequencing process, if the current sequencing results of the sequencing objects of any gene sequencing sample among the multiple gene sequencing samples meet the target sequencing requirements, the current sequencing results of the sequencing objects of the gene sequencing samples that have met the target sequencing requirements are split out, and data analysis is performed based on the current sequencing results of the sequencing objects of the gene sequencing samples that have met the target sequencing requirements.
[0149] In addition, embodiments of the present application further provide a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program executes the steps of the gene sequencing and data analysis methods described in the above method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.
[0150] An embodiment of the present application also provides a computer program product, including a computer program / instruction, which, when executed by a processor of the computer program / instruction, implements the gene sequencing and data analysis methods provided in the embodiments of the present application.
[0151] The methods in the embodiments of the present application can be implemented in whole or in part through software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in this application are performed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, a core network device, an OAM, or other programmable device.
[0152] The computer program or instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer program or instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired or wireless method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; an optical medium, such as a digital video disk; or a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.
[0153] Based on the same technical concept, in one implementation of the present application, the gene sequencing process can be implemented by a gene sequencing device, while the nucleic acid data analysis process can be implemented by a backend server, with the gene sequencing device and the backend server being communicatively connected. As shown in FIG6 , the gene sequencing and data analysis system 60 includes a gene sequencing device 61 and a server 62 communicatively connected to the server, wherein:
[0154] Gene sequencing equipment 61 is configured to sequentially detect sample tags and nucleic acid sequences in a sequencing object based on a fluorescent image of the sequencing object; identify the sample source of the sequencing object based on the detected sample tags, and determine a target sequencing requirement for the sequencing object based on the identified sample source; during the gene sequencing process, if the current sequencing result of the sequencing object of any gene sequencing sample among the multiple gene sequencing samples has met the target sequencing requirement, separate the current sequencing result of the sequencing object of the gene sequencing sample that has met the target sequencing requirement, and send the current sequencing result and sample source of the sequencing object of the gene sequencing sample that has met the target sequencing requirement to the server 62;
[0155] Server 62 is configured to determine the target analysis requirements of the sequencing object of the genetic sequencing sample that has met the target sequencing requirements based on the sample source, and perform data analysis on the data of the current sequencing results of the sequencing object of the genetic sequencing sample that has met the target sequencing requirements based on the target analysis requirements.
[0156] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present application, which are used to illustrate the technical solutions of the present application, rather than to limit them. The scope of protection of the present application is not limited thereto. Although the present application has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed in the present application, or perform equivalent replacements for some of the technical features thereof. These modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application shall be subject to the scope of protection of the claims.
Claims
1. A gene sequencing and data analysis device, characterized in that, Comprising: A processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. The machine-readable instructions are used to perform gene sequencing and data analysis on sequencing objects derived from multiple gene sequencing samples. The sequencing objects include sample tags and gene fragments. The sample tags are used to indicate the sample sources of the sequencing objects, and the sequencing positions of the sample tags in the sequencing objects are before the gene fragments. The sequencing objects have target sequencing requirements matching their sample sources. When the gene sequencing and data analysis device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the following method is performed: Detect the sample tags and gene fragments in the sequencing objects successively according to the fluorescence images of the sequencing objects; Identify the sample sources of the sequencing objects according to the detected sample tags, and determine the target sequencing requirements of the sequencing objects according to the identified sample sources; During the gene sequencing process, if the current sequencing result of the sequencing object of any gene sequencing sample among the multiple gene sequencing samples reaches the target sequencing requirement, split out the current sequencing result of the sequencing object of the gene sequencing sample that has reached the target sequencing requirement, and perform data analysis according to the current sequencing result of the sequencing object of the gene sequencing sample that has reached the target sequencing requirement.
2. The device according to claim 1, characterized in that, The target sequencing requirements include at least one of sequencing length and the number of sequencing ends.
3. The device according to claim 1 or 2, characterized in that, In the method executed by the processor, the detecting the sample tags and gene fragments in the sequencing objects successively according to the fluorescence images of the sequencing objects includes: Detect the i-th base in the sequencing object according to the i-th fluorescence image of the sequencing object; Determine the current sequencing quality of the sequencing object according to the i-th base in the sequencing object; When the current sequencing quality of the sequencing object meets the preset quality requirement, continue sequencing the (i + 1)-th fluorescence image of the sequencing object; When the current sequencing quality of the sequencing object does not meet the preset quality requirement, if the number of consecutive occurrences of the sequencing quality not meeting the preset quality requirement reaches the preset threshold, stop sequencing the (i + 1)-th fluorescence image of the sequencing object. If the number of consecutive occurrences of the sequencing quality not meeting the preset quality requirement is less than the preset threshold, continue sequencing the (i + 1)-th fluorescence image of the sequencing object. Here, i is a positive integer starting from 1.
4. The device according to claim 1 or 2, characterized in that In the method executed by the processor, the detecting the sample tags and gene fragments in the sequencing objects successively according to the fluorescence images of the sequencing objects includes: Extract the fluorescence image of the sample tag and the fluorescence image of the gene fragment from the fluorescence image of the sequencing object; Perform noise reduction processing on the fluorescence image of the sample tag to generate a noise-reduced fluorescence image of the sample tag; Perform sequencing on the noise-reduced fluorescence image of the sample tag and the fluorescence image of the gene fragment respectively to obtain the sample tag and the gene fragment in the sequencing object.
5. The device according to claim 4, characterized in that, In the method executed by the processor, the denoising process of the fluorescence image of the sample label to generate a denoised fluorescence image of the sample label includes: Determining the weight calculation parameter Value of the filter kernel based on the length of the sample label detected in the current sequencing result; Determining the weight values of multiple undetermined weights of the filter kernel based on the value of the weight calculation parameter of the filter kernel, wherein the undetermined weight located at the central position in the filter kernel is determined based on the sum of the other undetermined weights except the undetermined weight at the central position; Performing a convolution operation on the fluorescence image of the sample label based on the filter kernel to generate a denoised fluorescence image of the sample label.
6. The device according to claim 1 or 2, characterized in that, The sequencing object also has a target analysis requirement matching its sample source; In the method executed by the processor, the data analysis of the sequencing object of the gene sequencing sample according to the data of the current sequencing result includes: Performing data analysis on the sequencing object of the gene sequencing sample according to the data of the current sequencing result and the target analysis requirement.
7. The device according to claim 1 or 2, characterized in that, In the method executed by the processor, before sequencing the sample label and gene fragment in the sequencing object according to the fluorescence image of the sequencing object successively, it includes: Determining the size of the first computing resource required for the gene sequencing process according to at least one of the sample data volumes of the multiple gene sequencing samples and the target sequencing requirements of the sequencing objects of each gene sequencing sample, and allocating the first computing resource from the total computing resources to the gene sequencing process for use according to the size of the first computing resource.
8. The device according to claim 7, characterized in that, In the method executed by the processor, before performing data analysis on the sequencing object of the gene sequencing sample according to the data of the current sequencing result, it further includes: Determining the size of the second computing resource required for the data analysis process according to at least one of the data volume of the current sequencing result, the target sequencing requirements of the sequencing objects of each gene sequencing sample, and the target analysis requirement, and allocating the second computing resource from the remaining computing resources to the data analysis process for use, where the remaining computing resources are the computing resources obtained by deducting the first computing resource from the total computing resources.
9. A gene sequencing and data analysis method, characterized in that, Applied to sequencing objects derived from multiple gene sequencing samples, the sequencing object includes a sample label and a gene fragment, the sample label is used to indicate the sample source of the sequencing object, and the sequencing position of the sample label in the sequencing object is before the gene fragment, and the sequencing object has a target sequencing requirement matching its sample source; the method includes: Successively detecting the sample label and the gene fragment in the sequencing object according to the fluorescence image of the sequencing object; Identifying the sample source of the sequencing object according to the detected sample label, and determining the target sequencing requirement of the sequencing object according to the identified sample source; During the sequencing process of sequencing genes, if the current sequencing result of the sequencing object of any gene sequencing sample among the multiple gene sequencing samples meets the target sequencing requirements, the current sequencing result of the sequencing object of the gene sequencing sample that has met the target sequencing requirements is split out, and data analysis is performed based on the current sequencing result of the sequencing object of the gene sequencing sample that has met the target sequencing requirements.
10. The method according to claim 1, characterized in that, The target sequencing requirements include at least one of the sequencing length and the number of sequencing ends.
11. The method according to claim 9 or 10, characterized in that, The detecting the sample tag and the gene fragment in the sequencing object successively according to the fluorescence image of the sequencing object includes: Detecting the i-th base in the sequencing object according to the i-th fluorescence image of the sequencing object; Determining the current sequencing quality of the sequencing object according to the i-th base in the sequencing object; When the current sequencing quality of the sequencing object meets the preset quality requirements, continue sequencing the (i + 1)-th fluorescence image of the sequencing object; When the current sequencing quality of the sequencing object does not meet the preset quality requirements, if the number of consecutive times that the sequencing quality does not meet the preset quality requirements reaches the preset threshold, stop sequencing the (i + 1)-th fluorescence image of the sequencing object, and if the number of consecutive times that the sequencing quality does not meet the preset quality requirements is less than the preset threshold, continue sequencing the (i + 1)-th fluorescence image of the sequencing object; where i is a positive integer starting from 1.
12. The method according to claim 9 or 10, characterized in that, The sequencing the sample tag and the gene fragment in the sequencing object successively according to the fluorescence image of the sequencing object includes: Extracting the fluorescence image of the sample tag and the fluorescence image of the gene fragment from the fluorescence image of the sequencing object; Performing noise reduction processing on the fluorescence image of the sample tag to generate a noise-reduced fluorescence image of the sample tag; Sequencing the noise-reduced fluorescence image of the sample tag and the fluorescence image of the gene fragment respectively to obtain the sample tag and the gene fragment in the sequencing object.
13. The method according to claim 12, wherein The performing noise reduction processing on the fluorescence image of the sample tag to generate a noise-reduced fluorescence image of the sample tag includes: Determining the value of the weight calculation parameter of the filter kernel based on the length of the sample tag detected in the current sequencing result; Determining the weight values of multiple undetermined weights of the filter kernel based on the value of the weight calculation parameter of the filter kernel, where the undetermined weight located at the central position in the filter kernel is determined based on the sum of the other undetermined weights except the undetermined weight at the central position; Performing a convolution operation on the fluorescence image of the sample tag based on the filter kernel to generate a noise-reduced fluorescence image of the sample tag.
14. The method according to claim 9 or 10, characterized in that, The sequencing object also has target analysis requirements matching its sample source; The performing data analysis on the sequencing object of the gene sequencing sample according to the data of the current sequencing result includes: Performing data analysis on the sequencing object of the gene sequencing sample according to the data of the current sequencing result and the target analysis requirements.
15. The method according to claim 9 or 10, characterized in that, Before detecting the sample tag and the gene fragment in the sequencing object successively according to the fluorescence image of the sequencing object, the method includes: Determine the size of the first computing resource required for the gene sequencing process according to at least one of the sample data volume of the multiple gene sequencing samples and the target sequencing requirements of the sequencing object of each gene sequencing sample, and allocate the first computing resource from the total computing resources to the gene sequencing process for use according to the size of the first computing resource.
16. The method according to claim 15, characterized in that, Before performing data analysis on the sequencing object of the gene sequencing sample based on the data of the current sequencing result, it further includes: Determine the size of the second computing resource required for the data analysis process according to at least one of the data volume of the current sequencing result, the target sequencing requirements of the sequencing object of each gene sequencing sample, and the target analysis requirements, and allocate the second computing resource from the remaining computing resources to the data analysis process for use according to the size of the second computing resource. The remaining computing resources are the computing resources obtained by deducting the first computing resource from the total computing resources.
17. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, it executes the steps of the gene sequencing and data analysis method according to any one of claims 9 to 17.
18. A gene sequencing and data analysis system, characterized in that, The system includes: a gene sequencing device and a server, which are communicatively connected between the gene sequencing device and the server. The gene sequencing device is used to sequence the sequencing objects derived from multiple gene sequencing samples. The sequencing object includes a sample label and a gene fragment. The sample label is used to indicate the sample source of the sequencing object, and the sequencing position of the sample label in the sequencing object is before the gene fragment. The sequencing object has target sequencing requirements and target analysis requirements that match its sample source; wherein: The gene sequencing device is configured to sequentially detect the sample label and the gene fragment in the sequencing object according to the fluorescence image of the sequencing object; identify the sample source of the sequencing object according to the detected sample label, and determine the target sequencing requirements of the sequencing object according to the identified sample source; during the gene sequencing process, if the current sequencing result of the sequencing object of any gene sequencing sample in the multiple gene sequencing samples has reached the target sequencing requirements, split out the current sequencing result of the sequencing object of the gene sequencing sample that has reached the target sequencing requirements, and send the current sequencing result and the sample source of the sequencing object of the gene sequencing sample that has reached the target sequencing requirements to the server; The server is configured to determine the target analysis requirements of the sequencing object of the gene sequencing sample that has reached the target sequencing requirements according to the sample source, and perform data analysis on the current sequencing result of the sequencing object of the gene sequencing sample that has reached the target sequencing requirements according to the target analysis requirements.
Citation Information
Patent Citations
Sequencing data result analysis method and device and sequencing library construction and sequencing method
CN107368706A
Gene sequencing method, device, equipment and medium
CN115612722A
Nucleic acid detection and data analysis method, equipment, system and storage medium
CN117893512A
Methods and systems for accurate genotyping of repeat polymorphisms
US20230162815A1