Gene sequencing sample data matching method based on micro-fluidic chip
By employing multi-strategy comparison and candidate set screening, combined with microfluidic chip-based sample purification, standardization, and data preprocessing, a standardized reference database was established. This solved the problems of uneven purification, imprecise control of reaction conditions, susceptibility to signal capture interference, and messy data formats in gene sequencing sample data matching, achieving high-precision and efficient gene sequencing sample data matching.
Patent Information
- Application Number
- CN202511278382.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-12-19
AI Technical Summary
Existing gene sequencing sample data matching suffers from problems such as uneven sample purification leading to data bias, coarse control of microfluidic chip reaction conditions, susceptibility of signal capture to background noise interference, messy data formats, difficulty in integrating reference databases, and a single alignment strategy, all of which affect sequencing accuracy and efficiency.
We employ a multi-strategy alignment and candidate set screening approach, combining sequence, variant, and functional multi-dimensional alignment. Through quantitative indicators and biological validation, we dynamically optimize the alignment strategy. We also utilize microfluidic chips for sample purification, standardization, signal capture, and data preprocessing, and establish a standardized reference database to improve matching reliability.
It significantly improves the accuracy and reliability of gene sequencing sample data matching, meets the diverse needs of clinical and scientific research, reduces interference factors, improves response efficiency and repeatability, and ensures a high degree of matching between features and analysis targets.
Smart Images

Figure CN121171371A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of gene sequencing sample matching, in particular to a gene sequencing sample data matching method based on a microfluidic chip. BACKGROUND
[0002] There are many problems in the existing gene sequencing sample data matching: sample purification often causes data deviation due to impurity residues and uneven concentration; microfluidic chip reaction conditions are controlled in a rough way, and temperature and fluid stability are insufficient, affecting sequencing repeatability; signal capture is easily disturbed by background noise, and the signal-to-noise ratio is low and the data format is chaotic; data preprocessing lacks a systematic method, and noise filtering and feature extraction are weak in pertinence; reference database multi-source data integration is difficult, the structure is not standardized, the matching strategy is single, the matching accuracy and reliability are insufficient, and there is a lack of effective evaluation and optimization mechanism, which seriously restricts the precision and efficiency of gene sequencing sample data matching. SUMMARY
[0003] The purpose of the present application is to provide a gene sequencing sample data matching method based on a microfluidic chip, which adopts multi-strategy alignment and candidate set screening, combines sequence, variation, and function multi-dimensional alignment to improve accuracy, dynamically optimizes alignment strategies through quantitative indicators and biological verification evaluation, enhances the adaptability to complex samples, significantly improves matching reliability and practical value, meets the diversified needs of clinical and scientific research, adapts to sample types such as blood and cells through differential lysis strategies, removes impurities such as proteins and salts through step-by-step purification, and unifies nucleic acid concentration and fragment state through standardized processing, thereby reducing interference from the source, and solving the problems in the prior art.
[0004] To achieve the above-mentioned purpose, the present application provides the following technical scheme: The gene sequencing sample data matching method based on a microfluidic chip comprises: Purifying and standardizing the gene sequencing original sample; loading the purified and standardized original sample into the microfluidic chip for preparation and processing of the microfluidic chip, while controlling the biochemical reaction conditions in the microfluidic chip; capturing the gene sequencing signal of the detection area of the microfluidic chip after biochemical reaction of the microfluidic chip, and converting it into processable raw data; performing data preprocessing on the converted raw data; extracting gene features from the preprocessed raw data; establishing a standardized reference gene sample database, and performing alignment and matching of the extracted gene features with the reference gene sample database; and finally evaluating and optimizing the alignment and matching results.
[0005] Preferably, the purification and standardization of the gene sequencing original sample comprises: First, the cracking method is selected according to the original sample type. When the original sample type is a blood sample, the red blood cells are removed by centrifugation, and then the white blood cells are treated with a lysis buffer. When the original sample type is a cell or tissue sample, a lysis buffer containing a detergent is used to destroy the cell membrane and cell wall, release intracellular nucleic acids, and add proteinase K to degrade the proteins in the sample. The original sample after lysis is subjected to nucleic acid separation and purification. The nucleic acid separation and purification process is as follows: the original sample is subjected to protein removal, nucleic acid precipitation, salt removal, nucleic acid dissolution, and residual removal in sequence. The original sample after nucleic acid separation and purification is subjected to standardization treatment. The standardization treatment is as follows: the absorbance ratio of the nucleic acid solution at 260 nm and 280 nm is detected using a UV spectrophotometer, the nucleic acid fragment size is analyzed by agarose gel electrophoresis, the concentration of the purified nucleic acid is determined by fluorescence quantification or UV spectrophotometry, and finally the original sample is diluted or concentrated with nuclease-free water according to the loading requirements of the microfluidic chip. Finally, the purification and standardization of the original sample for gene sequencing are completed.
[0006] Preferably, the purified and standardized original sample is loaded into a microfluidic chip for microfluidic chip preparation and processing, while the biochemical reaction conditions in the microfluidic chip are controlled, including: First, the microfluidic chip is subjected to structural inspection and cleaning. First, the microchannels, reaction chambers, sample inlets, and detection zones of the microfluidic chip are inspected for physical damage. Then, nuclease-free water or a special chip cleaning solution is injected into the channel through the sample inlet, and after standing for 5-10 minutes, the channel is blown dry with nitrogen or clean air to remove residual impurities, dust, or protective agents left over from the factory; After the structural inspection and cleaning of the microfluidic chip are qualified, the loading method of the original sample is confirmed. The loading method includes manual loading and automatic loading. Manual loading is used for small chips or low-throughput scenarios, while automatic loading is used for high-throughput or precise control scenarios. After the loading method is confirmed, the original sample is loaded onto the microfluidic chip; After the original sample is loaded, the sample is processed in the microfluidic chip according to the requirements of gene sequencing. The processing flow is as follows: the original sample is mixed with reagents, and the mixed original sample is sealed. At the same time, the sealed microfluidic chip is fixed on the chip seat. The biochemical reaction in the microfluidic chip fixed on the chip seat is controlled. The biochemical reaction control includes temperature control, fluid environment control, reaction time control, and environmental atmosphere control. Among them, the temperature control includes constant temperature reaction control and variable temperature reaction control; the fluid environment control includes fluid and pressure regulation and bubble and evaporation prevention and control. Finally, the biochemical reaction control of the microfluidic chip is completed.
[0007] Preferably, the microfluidic chip captures the gene sequencing signal of the detection area of the microfluidic chip after the biochemical reaction of the microfluidic chip, and converts it into processable raw data, including: The microfluidic chip with completed biochemical reaction control is taken out, and the surface of the chip is wiped with dust-free paper. At the same time, the position of the detection area is confirmed by using a microscope or a chip positioning mark; After the position of the detection area is confirmed, the detection equipment is confirmed according to the signal type. The signal type includes a fluorescence signal, an electrochemical signal and a scattering signal. The detection equipment for the fluorescence signal is a fluorescence microscope or a confocal laser scanning microscope, which is equipped with excitation light filters and emission light filters matched with fluorescent markers. The electrochemical signal is connected to an electrochemical workstation, and the working electrode and the reference electrode of the chip detection area are connected to the electrode interface of the workstation. The detection equipment for the scattering signal is a microspectrometer; The detection area of the microfluidic chip is scanned by using the detection equipment. When scanning the signal, the blank area on the microfluidic chip without sample is scanned, and the background signal value is recorded. At the same time, the non-specific signal on the microfluidic chip is removed by using signal filtering software, and the effective signal of the reaction characteristics is retained; The retained effective signal is dynamically captured, and the change curve of the signal with time is recorded during dynamic signal capture; The captured dynamic signal is converted into a digital signal by using an analog-to-digital converter built in the equipment, and the digital signal is associated according to time or space dimensions. After association, structured data is formed; The structured data is standardized according to the sequencing universal format. The standardization processing includes sample number, chip batch, detection time, equipment model and signal channel; The structured data after standardization processing is stored, and processable raw data is obtained after storage.
[0008] Preferably, the converted raw data is preprocessed, including: The stored processable raw data is read, and after reading, key information is parsed. The key information includes sample ID, chip detection area coordinates, signal intensity value, detection time, channel identification and equipment parameter log; The parsed key information is checked for integrity. The integrity check is to check whether the meta information of the raw data is complete. If there is a missing, the information is supplemented in the signal conversion link; After the integrity check is completed, the signal quality is quantitatively analyzed, the quantitative analysis is to calculate the statistical parameters of the effective signal, the statistical parameters include the average value, the standard deviation, the maximum value and the minimum value of the signal strength, and after the statistical parameters are calculated, the signal-to-noise ratio analysis is performed, the signal-to-noise ratio analysis is to divide the average strength of the effective signal by the average strength of the background signal, if SNR≥3, it indicates that the signal quality is qualified; if SNR<3, it needs to be marked as low-quality data; The data after quantitative analysis is sequentially subjected to data cleaning, data standardization, data noise reduction, signal enhancement, data fragmentation and data alignment; Among them, data cleaning is to identify outliers by box plot or threshold method, manually or automatically remove abnormal outliers, at the same time, the regions without loading samples in the microfluidic chip detection area and the regions with signal value of 0 are deleted, and the background signal value is corrected; data standardization is to unify the scale by using normalization processing for the data differences of different samples or different chip batches, convert the structured data into a standardized format, and for multi-channel signals, split by channel and realign the dimension; data noise reduction is to process the data with random noise by using sliding window smoothing method, and for spatial dimension data, use median filter method to remove isolated noise points; signal enhancement is to improve the signal recognition degree of weak signal region by contrast enhancement algorithm; data fragmentation is to split the data into several fragments according to the preset length if the original data corresponds to long fragment gene sequence, and the start and end position information of the fragments is retained during splitting; data alignment is for spatial dimension data, if the chamber position is slightly offset due to chip processing error, the signal points are mapped to the standard grid through coordinate correction, and for time series data, if there is a difference in the reaction start time of different samples, the time axis is unified through alignment algorithm; Finally, the data preprocessing process is completed.
[0009] Preferably, the gene features are extracted from the preprocessed raw data, including: According to the purpose of gene sequencing, the core features are confirmed, wherein the target of gene sequencing includes whole genome sequencing, targeted gene sequencing and RNA expression sequencing, and the core features include detection range, core detection content, data characteristics, technical advantages and typical applications; The feature extraction dimension in the preprocessed raw data is confirmed, including spatial dimension, time dimension and signal strength dimension; After the feature extraction dimension is confirmed, the signal in the preprocessed raw data is converted correspondingly; The conversion of the fluorescence signal is to convert the pre-processed signal intensity into a base type according to the signal channel identifier, and to concatenate the bases in the arrangement order according to the time or space sequence to form the original base sequence; the conversion of the electrochemical signal and the scattering signal is to match the characteristic map of the known nucleotide according to the signal fluctuation characteristics, and to convert into a base sequence; After the corresponding conversion is completed, the gene fragment boundary and length are confirmed, wherein the start and end positions of the gene fragment signal in the pre-processed original data are confirmed, and at the same time, the length of each fragment is confirmed by the number of bases or the spatial range covered by the signal, to obtain the basic gene sequence; The core sequence characteristics and fragmentation characteristics in the basic gene sequence are extracted, and the original base sequence and the fragment gene sequence are obtained after extraction; Finally, the original base sequence and the fragment gene sequence are used as the gene feature data extracted from the original data.
[0010] Preferably, a standardized reference gene sample database is established, and the extracted gene features are matched with the reference gene sample database, including: The establishment of the standardized reference gene sample database is: integrating multi-source reference data, the multi-source reference data including standard gene sequences, known functional gene annotations, variation databases in public databases, and self-built specific sample data, and then quality screening the multi-source reference data, the quality screening being low-quality sequence elimination, incomplete data annotation, and high-confidence data retention of the collected multi-source reference data which are verified by experiments or multiple iterations; After the multi-source reference data integration is completed, the database structure of the standardized reference gene sample database is constructed, the database structure construction being to construct the database in layers according to the data types and application scenarios, including a basic layer, a feature layer, and an application layer, wherein the basic layer stores original gene sequences and core annotations; the feature layer stores the extracted gene features; the application layer constructs a sub-database for a specific scenario; at the same time, a multi-dimensional index is established, including a sequence index, a feature index, and a classification index; According to the extracted gene features, a comparison strategy is selected, including sequence overall comparison, local feature comparison, and multi-feature joint comparison; After the comparison strategy is selected, the multi-dimensional index in the reference gene sample database is used to preliminarily screen the extracted gene features, the preliminary screening being to match the sample sequence fragments with the reference sequences through the k-mer index, and to screen out the reference sequences sharing ≥50% k-mer fragments as the candidate set, and if the sample contains a known functional variation, the reference sequence containing the variation in the database is directly searched and included in the candidate set; The reference sequence in the candidate set is compared with the extracted gene feature in each dimension, including sequence layer comparison, variation layer comparison and function layer comparison, wherein the sequence layer comparison is to calculate the consistency percentage of the sample sequence and the reference sequence, and record the difference site; the variation layer comparison is to compare the position and type of the sample variation and the reference variation, and count the proportion of the matching variation number to the total variation of the sample; the function layer comparison is to compare whether the function domain type and position of the sample and the reference sequence overlap, and calculate the matching degree of the function feature; According to the comparison matching result sorting result of the comparison in each dimension, the highest sorting result is selected as the final comparison matching result.
[0011] Preferably, the reference sequence in the candidate set is compared with the extracted gene feature in each dimension, including: The sample sequence and the reference sequence are split, and the consistency of each sample subsequence and the corresponding reference subsequence after splitting is compared to obtain the sequence consistency percentage of each subsequence; Based on the consistency comparison result of each sample subsequence and the corresponding reference subsequence after splitting, the difference site with sequence difference is determined; Based on the preset matching rule, the variation position and variation type of each difference site in the same subsequence of the sample sequence are determined to obtain the first variation result of each subsequence; The ratio of the number of variation features in each subsequence of the sample sequence to the number of corresponding variation sites is taken as the second variation result of the corresponding subsequence; Each subsequence is annotated with a function domain, and the function domain coordinates of each subsequence of the reference sequence and each subsequence of the sample sequence are mapped to the same genome coordinate system, and the function domain overlap is detected to obtain the first function result; The ratio of the number of function domain features in each subsequence of the sample sequence to the number of function domains in the current subsequence of the sample sequence is taken as the second function result of the corresponding subsequence; The consistency percentage of all subsequences in the sample sequence is taken as the sequence alignment result of the sample sequence in the sequence layer; The first variation result and the corresponding second variation result of all subsequences in the sample sequence are integrated to obtain the variation alignment result of the sample sequence in the variation layer, and the function alignment result of the sample sequence in the function layer is obtained in the same way; Based on the sequence alignment result, the variation alignment result and the function alignment result, the corresponding alignment weight is combined for hierarchical integration to obtain the first alignment result of the sample sequence; The sequence layer comparison, variation layer comparison and function layer comparison are integrated in parallel, and the second alignment result of the sample sequence is based on the parallel integration result; The first comparison result and the second comparison result are sorted respectively, and the sorting reliabilities of the first comparison result and the second comparison result are weighted to obtain a comparison matching result of the sample sequence.
[0012] Preferably, the comparison matching result is finally evaluated and optimized, including: The comparison matching result is first evaluated in multiple dimensions, including quantitative index evaluation, biological rationality verification and abnormal matching case analysis. The quantitative index evaluation is to calculate core matching quality indexes, including matching accuracy, feature coverage, variation consistency and repeat stability, and low-quality matching data is marked at the same time; the biological rationality verification is to check the biological logic of the matching result, and if high-frequency variations of irrelevant genes are matched, non-matching rationality is marked, and the clinical sample is checked to see whether the comparison matching result is consistent with the clinical phenotype of the sample; the experimental sample is checked to see whether the comparison matching result is consistent with the processing background of the sample, and whether the experimental design target is verified; the abnormal matching case analysis is to analyze cases with high quantitative scores but biological irrationality or low quantitative scores but potential effectiveness in the comparison matching result. According to the multi-dimensional evaluation result, the comparison strategy is optimized, including increasing the weight of the core feature or tightening the sequence similarity threshold for low-accuracy matching results, expanding the comparison range or increasing the dimension of feature extraction for low-coverage matching results, and calling reference samples of the same population in the database in priority for different population samples during comparison for population-specific bias results. Finally, the optimization of the comparison matching result is completed, and the final optimization result is converted into visual data and transmitted to a display terminal for display of the optimization and evaluation results.
[0013] For low-accuracy matching results, the weight of the core feature is increased or the sequence similarity threshold is tightened, including: For low-accuracy matching results, the weight of the corresponding feature is increased based on the accuracy of the core feature and the accuracy threshold to obtain an optimized weight; at the same time, the similarity threshold of the corresponding feature is tightened to obtain an optimized similarity threshold. The comparison strategy is optimized according to the optimized weight and the optimized similarity threshold.
[0014] Compared with the prior art, the present application has the following advantages: 1.The gene sequencing sample data matching method based on a microfluidic chip provided by the present application, which is adapted to sample types such as blood and cells through a differential lysis strategy, removes impurities such as proteins and salts through step-by-step purification, and unifies nucleic acid concentration and fragment state through standardized processing to reduce interference from the source; the microfluidic chip is strictly cleaned and structurally inspected, and is loaded manually and automatically to adapt to different scenarios, accurately controls reaction conditions such as temperature and fluid, improves reaction efficiency and repeatability, reduces experimental bias, and provides high-quality samples and stable reaction environments for subsequent analysis.
[0015] 2.The gene sequencing sample data matching method based on a microfluidic chip provided by the present application, which adopts a differential conversion strategy for multiple types of signals such as fluorescence and electrochemistry to preserve original information of gene sequences; preprocessing simplifies data through operations such as noise reduction and alignment, extracts core features in combination with sequencing purposes, preserves global sequence information and refines local fragment features, ensures high matching of features and analysis targets, effectively solves the problems of multiple noises and disordered formats of original data, and provides accurate feature support for comparison and matching.
[0016] 3.The gene sequencing sample data matching method based on a microfluidic chip provided by the present application, which integrates multiple source data to construct a hierarchical database and improves retrieval efficiency in combination with multi-dimensional indexing; adopts multi-strategy comparison and candidate set screening, improves accuracy in combination with multi-dimensional comparison of sequences, variations, and functions; dynamically optimizes comparison strategies through quantitative indicators and biological verification evaluation, enhances adaptability to complex samples, significantly improves matching reliability and practical value, and meets diversified needs of clinical and scientific research. BRIEF DESCRIPTION OF DRAWINGS
[0017] Figure 1 The figure is a schematic diagram of the gene sequencing sample data matching steps of the present application. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0019] To solve the problems in the prior art, such as insufficient standardization of gene sequencing sample purification, inaccurate control of microfluidic chip processing and reaction conditions, easy interference of signal capture, and disordered data conversion format, which affect sequencing accuracy and consistency, please refer to Figure 1 The present embodiment provides the following technical solutions: The gene sequencing sample data matching method based on a microfluidic chip comprises: The gene sequencing raw sample is purified and standardized first; the purified and standardized raw sample is loaded into a microfluidic chip for preparation and processing of the microfluidic chip, while the biochemical reaction conditions in the microfluidic chip are controlled; the gene sequencing signal of the detection area of the microfluidic chip after biochemical reaction is captured and converted into processable raw data; the converted raw data is preprocessed; gene features are extracted from the preprocessed raw data; a standardized reference gene sample database is established, and the extracted gene features are matched with the reference gene sample database; and finally, the matching result is evaluated and optimized.
[0020] Specifically, through purification and standardization preprocessing, the interference substances such as proteins and impurity nucleic acids in the raw sample can be effectively removed, and the sample concentration and fragment length parameters are unified to a standard range, which greatly reduces the influence of sample individual differences on subsequent analysis, lays a foundation for data reliability, reduces the matching deviation caused by uneven sample quality, and significantly reduces reagent consumption and saves costs due to the microchannel structure of the chip, which reduces the reaction system volume to microliters or even nanoliters; at the same time, the chip can accurately control the biochemical reaction conditions such as temperature and flow rate to ensure that the reaction is carried out in the optimal environment, improve the reaction efficiency and repeatability, reduce the variation coefficient of experimental results, and make the data more consistent; the high-sensitivity sensor integrated in the detection area of the microfluidic chip can specifically capture the gene sequencing signal, and with the help of a special conversion algorithm, it can effectively filter out background noise and convert the signal into high signal-to-noise ratio raw data, providing a high-quality data source for subsequent analysis and reducing the interference of invalid data on the matching result; the preprocessing step simplifies the data amount through noise reduction and correction operations, the feature extraction focuses on the key sites and mutation information in the gene sequence, reduces the interference of redundant data on the matching process, improves the matching efficiency, and enhances the relevance of the features and target genes, providing a guarantee for accurate matching; the establishment of the standardized reference database unifies the matching standards of different samples and improves the comparability of cross-experiment and cross-platform data; combined with the evaluation and optimization mechanism, the feature weight and matching parameters can be dynamically adjusted according to the matching result, the matching model is continuously optimized, and the matching adaptability to rare mutations and complex gene structure samples is stronger, finally realizing high-precision and high-robustness gene sequencing sample data matching.
[0021] The gene sequencing raw sample is purified and standardized, including: First, the lysis method is selected according to the type of the raw sample. When the raw sample is a blood sample, the red blood cells are removed by centrifugation, and then the white blood cells are treated with a lysis buffer; when the raw sample is a cell or tissue sample, a lysis buffer containing a detergent is used to destroy the cell membrane and cell wall, release the intracellular nucleic acids, and add proteinase K to degrade the proteins in the sample; The original sample after lysis is subjected to nucleic acid separation and purification, and the nucleic acid separation and purification process is: the original sample is sequentially subjected to protein removal, nucleic acid precipitation, salt washing, nucleic acid dissolution and residual removal; The original sample after nucleic acid separation and purification is subjected to standardization treatment, and the standardization treatment is: the absorbance ratio of the nucleic acid solution at 260 nm and 280 nm is detected by using an ultraviolet spectrophotometer, then the nucleic acid fragment size is analyzed by agarose gel electrophoresis, the concentration of the purified nucleic acid is determined by fluorescence quantitative method or ultraviolet spectrophotometry, and finally the original sample is diluted or concentrated with nuclease-free water according to the loading requirements of the microfluidic chip; Finally, the purification and standardization treatment of the original sample for gene sequencing are completed.
[0022] Specifically, according to the difference of sample types (blood, cells / tissues), the lysis strategy is selected: the red blood cells are removed by centrifugation to avoid the contamination and interference of hemoglobin on nucleic acid, and then the white blood cells are lysed to release nucleic acid; the cell or tissue sample is treated by a detergent to destroy the membrane structure and proteinase K to degrade protein, so as to realize efficient release of nucleic acid. This "tailor-made" lysis design can maximize the preservation of target nucleic acid and reduce the release of non-specific impurities (such as red blood cell components and extracellular proteins), laying a pure foundation for subsequent purification. The nucleic acid separation and purification are realized by the step-by-step operation of "protein removal-nucleic acid precipitation-salt washing-nucleic acid dissolution-residual removal", which realizes multi-dimensional impurity removal: proteinase K pre-degrades protein, and residual protein is further removed in the subsequent steps; salt ions and small molecule impurities such as detergents in the lysis buffer are effectively removed by salt washing; residual removal specifically eliminates inhibitors (such as metal ions and organic residues) that may affect subsequent reactions. The nucleic acid obtained finally has high purity, avoiding the interference of impurities on biochemical reactions (such as PCR amplification and sequencing enzyme activity) of the microfluidic chip. The 260 nm / 280 nm absorbance ratio can be quickly judged by ultraviolet spectrophotometry, which can quickly judge the nucleic acid purity (pure DNA ratio is about 1.8, and pure RNA is about 2.0), ensuring that the nucleic acid quality meets the standard; the fragment size can be directly evaluated by agarose electrophoresis analysis, avoiding the influence of degraded nucleic acid on sequencing results; the concentration is accurately determined by fluorescence quantitative method or ultraviolet spectrophotometry, and then adjusted to the standard concentration according to the loading requirements of the microfluidic chip, realizing the unification of nucleic acid concentration and fragment state among different samples. This standardization treatment eliminates individual differences (such as nucleic acid concentration fluctuation and fragment integrity difference) of the original sample, ensures the consistency of subsequent microfluidic chip reaction conditions, reduces experimental bias caused by sample state, improves data repeatability and reliability, and provides high-purity, high-integrity and high-consistency original sample for gene sequencing from the source, reducing interference factors, and laying a core foundation for the accuracy of subsequent microfluidic chip reaction and data matching.
[0023] The purified and standardized raw sample is loaded into the microfluidic chip for microfluidic chip preparation and processing, while the biochemical reaction conditions in the microfluidic chip are controlled, including: First, the microfluidic chip is structurally inspected and cleaned. First, the microchannels, reaction chambers, sample inlets, and detection zones of the microfluidic chip are inspected for physical damage. Then, nuclease-free water or a special chip cleaning solution is injected into the channel through the sample inlet, and the channel is left to stand for 5-10 minutes before being blown dry with nitrogen or clean air to remove residual impurities, dust, or protective agents from the factory; After the structural inspection and cleaning of the microfluidic chip are both qualified, the loading method of the raw sample is confirmed. The loading method includes manual loading and automatic loading. Manual loading is used for small chips or low-throughput scenarios, while automatic loading is used for high-throughput or precision control scenarios; After the loading method is confirmed, the raw sample is loaded onto the microfluidic chip; After the raw sample is loaded, the sample is processed in the microfluidic chip according to the requirements of gene sequencing. The processing flow is as follows: the raw sample is mixed with reagents, and the mixed raw sample is sealed. At the same time, the sealed microfluidic chip is fixed on the chip seat; The biochemical reaction in the microfluidic chip fixed on the chip seat is controlled. The biochemical reaction control includes temperature control, fluid environment control, reaction time control, and environmental atmosphere control; Among them, the temperature control includes constant temperature reaction control and variable temperature reaction control; the fluid environment control includes fluid and pressure regulation and bubble and evaporation prevention and control; Finally, the biochemical reaction control of the microfluidic chip is completed.
[0024] Specifically, the structure inspection can identify physical defects such as microchannel blockage and chamber damage in advance, avoiding uneven fluid distribution or reaction failure caused by chip damage; the cleaning process removes impurities and protectants with nuclease-free water or special reagents, and combines with nitrogen drying to ensure channel cleanliness, eliminating pollution risks (such as exogenous nucleic acids and chemical residues) from the source, providing a pure microenvironment for subsequent reactions, reducing background interference, and selecting manual loading for small and low-throughput scenarios to simplify the operation process; for high-throughput and precision requirements, automatic loading is adopted to achieve uniformity and repeatability of sample distribution through mechanical control, reducing human error. This "scenario-based" selection ensures the convenience of small batch experiments and meets the standardization requirements of large-scale detection, broadening the application scope of the technology, and ensuring consistent reaction component concentrations through uniform mixing of samples and reagents; sealing prevents evaporation of reaction liquid and invasion of external pollutants, avoiding changes in system volume or cross-contamination; chip fixation eliminates positional deviation caused by vibration during the reaction process, ensuring the stability of signal capture in the detection area and providing multiple safeguards for reaction uniformity; temperature control covers constant and variable temperature scenarios, accurately matching the reaction requirements of PCR amplification and other gradient temperature control processes to ensure enzyme activity and reaction efficiency; fluid control maintains a stable flow field by adjusting flow rate and pressure, and cooperates with bubble and evaporation prevention and control to avoid fluid disturbance to the reaction interface; time and environmental atmosphere control provide suitable conditions for specific reactions (such as anaerobic amplification and enzyme digestion) to achieve precise regulation of the entire reaction process, ultimately improving reaction repeatability and data reliability, laying a solid foundation for high-quality signal capture and data analysis.
[0025] After the biochemical reaction of the microfluidic chip, the gene sequencing signal of the detection area of the microfluidic chip is captured and converted into processable raw data, including: The microfluidic chip after completing the biochemical reaction control is taken out, and the surface of the chip is wiped with a dust-free paper, at the same time, the position of the detection area is confirmed by using a microscope or a chip positioning marker; After the position of the detection area is confirmed, the detection equipment is confirmed according to the signal type, and the signal type includes a fluorescence signal, an electrochemical signal and a scattering signal, wherein the detection equipment of the fluorescence signal is a fluorescence microscope or a confocal laser scanning microscope, and is equipped with an excitation light filter and an emission light filter matched with a fluorescence marker; the electrochemical signal is connected with an electrochemical workstation, and the working electrode and the reference electrode of the chip detection area are connected with the electrode interface of the workstation; the detection equipment of the scattering signal is a microspectrometer; The detection area of the microfluidic chip is scanned by using the detection equipment, and when the signal is scanned, the blank area without sample on the microfluidic chip is scanned, and the background signal value is recorded, at the same time, the non-specific signal on the microfluidic chip is removed by using a signal filtering software, and the effective signal of the reaction characteristics is reserved; The reserved effective signal is subjected to dynamic signal capture, and a signal change curve over time is recorded during dynamic signal capture; The captured dynamic signal is converted into a digital signal by an analog-to-digital converter built in the device, and the digital signal is associated according to time or space dimensions, and structured data is formed after association; The structured data is subjected to standardization processing according to a sequencing universal format, and the standardization processing includes sample number, chip batch, detection time, device model and signal channel; The structured data after standardization processing is stored, and processed raw data is obtained after storage.
[0026] Specifically, chip surface wiping and detection area positioning ensure that the optical path is clean, and signal distortion caused by surface stains or positioning deviation is avoided; the background signal of the blank area is recorded to provide a reference for subsequent signal filtering, and the software is used to remove non-specific signals, which significantly improves the signal-to-noise ratio of effective signals, reduces false positive data interference, and matches special equipment according to different signal characteristics such as fluorescence, electrochemistry and scattering, such as a fluorescence microscope equipped with a special filter group to enhance the specificity of fluorescence signals, an electrochemical workstation to accurately capture electrode reaction signals, and a miniature spectrometer to adapt to scattered light detection, to ensure that various signals can be efficiently recognized, to expand the application scenarios of the technology, to dynamically record the change curve of the signal over time, to completely retain the reaction kinetics characteristics, and to provide a basis for analyzing the reaction process; the digital signal is associated according to time and space dimensions to form structured data, to avoid information fragmentation, to facilitate subsequent feature extraction and comparison, to improve data utilization efficiency, to integrate metadata such as sample number and device model according to a sequencing universal format, to realize standardized management of different experimental data, to enhance cross-platform data compatibility, to completely store information to trace experimental process parameters, to facilitate the investigation of matching deviation reasons, to provide data support for result verification and optimization, and to ultimately improve the reliability and analysis value of raw data.
[0027] In order to solve the problems in the prior art that the raw data of gene sequencing has many noises, the format is not unified, the signal quality difference is large, the feature extraction lacks pertinence, and it is difficult to effectively convert into reliable gene features, which affects the accuracy of subsequent data analysis, please refer to Figure 1 The embodiment provides the following technical solutions: The converted raw data is subjected to data preprocessing, including: The stored processed raw data is read, and key information is parsed after reading, and the key information includes sample ID, chip detection area coordinates, signal intensity value, detection time, channel identifier and device parameter log; The parsed key information is subjected to integrity check, and the integrity check is to check whether the metadata of the raw data is complete, and if there is a missing, the information is supplemented in the signal conversion link; After the integrity check, the signal quality is quantitatively analyzed, the quantitative analysis is to calculate the statistical parameters of the effective signal, the statistical parameters include the average value, the standard deviation, the maximum value and the minimum value of the signal strength, after the statistical parameter calculation, the signal signal-to-noise ratio analysis is carried out, the signal-to-noise ratio analysis is to divide the average strength of the effective signal by the average strength of the background signal, if SNR≥3, it indicates that the signal quality is qualified; if SNR<3, it needs to be marked as low-quality data; The data after quantitative analysis is sequentially subjected to data cleaning, data standardization, data noise reduction, signal enhancement, data fragmentation and data alignment; Among them, the data cleaning is to identify outliers by box plot or threshold method, manually or automatically remove abnormal outliers, at the same time, the areas without loading samples in the microfluidic chip detection area and the areas with signal value of 0 are deleted, and the background signal value is corrected; the data standardization is to unify the scale by using normalization processing for the data difference of different samples or different chip batches, convert the structured data into standardized format, and for multi-channel signal, split by channel and realign the dimension; the data noise reduction is to process the data with random noise by using sliding window smoothing method, for spatial dimension data, use median filter method to remove isolated noise points; the signal enhancement is to improve the signal recognition degree of weak signal area by contrast enhancement algorithm; the data fragmentation is to split the data into several fragments according to the preset length if the original data corresponds to long fragment gene sequence, and the start and end position information of the fragment is retained during splitting; the data alignment is for spatial dimension data, if the chamber position is slightly offset due to chip processing error, the signal point is mapped to the standard grid through coordinate correction, for time sequence data, if there is difference in the reaction starting time of different samples, the time axis is unified through alignment algorithm; Finally, the data preprocessing process is completed.
[0028] Specifically, according to different targets such as whole genome, targeted gene, RNA expression sequencing, etc., the core feature dimensions (such as detection range, technical advantage, etc.) are determined, the feature extraction is focused on the key information required for the sequencing purpose, the irrelevant data interference is avoided, the extracted gene features are highly matched with the analysis target, and the foundation is laid for subsequent accurate matching. For different signal types such as fluorescence, electrochemistry, scattering, etc., differential conversion strategies are adopted: fluorescence signals are combined with time and space sequence to concatenate bases, electrochemical and scattering signals are converted through feature map matching, accurate conversion of various signals to base sequences is realized, the original information of gene sequences is completely retained, sequence loss or misjudgment caused by signal type difference is avoided, the starting / ending position and length of the gene fragment are determined, the physical boundary of each gene fragment is accurately defined, the integrity of the basic gene sequence and the consistency of the fragment division are ensured, the feature confusion caused by ambiguous fragment boundary is avoided, and a clear structure foundation is provided for subsequent fragmented feature extraction. Extracting the original base sequence (complete sequence information) and the fragment gene sequence (split feature) not only retains the global sequence features of the gene, but also refines the local fragment features, meets the needs of different matching scenarios (such as overall sequence comparison or local fragment analysis), enhances the flexibility and applicability of feature data, and finally provides high-quality, multi-level feature support for efficient comparison between gene features and reference database.
[0029] Extracting gene features from preprocessed raw data includes: According to the purpose of gene sequencing, the core features are confirmed, wherein the target of gene sequencing includes whole genome sequencing, targeted gene sequencing and RNA expression sequencing, and the core features include detection range, core detection content, data characteristics, technical advantage and typical application; Confirming the feature extraction dimensions in the preprocessed raw data, including spatial dimension, time dimension and signal intensity dimension; After confirming the feature extraction dimensions, the signals in the preprocessed raw data are converted correspondingly; Among them, the conversion of fluorescence signal is to convert the signal intensity after preprocessing into base type according to the signal channel identifier, and to concatenate the bases in the order of arrangement to form the original base sequence by combining time or space sequence; the conversion of electrochemical signal and scattering signal is to match the feature map of known nucleotides according to the signal fluctuation characteristics, and to convert into base sequence; After corresponding conversion, the boundaries and lengths of gene fragments are confirmed, wherein the starting and ending positions of gene fragment signals in the preprocessed raw data are confirmed, and at the same time, the length of each fragment is confirmed by the number of bases or the spatial range covered by the signal, to obtain the basic gene sequence; Extracting the core sequence features and fragmented features in the basic gene sequence, and obtaining the original base sequence and the fragment gene sequence after extraction; The original base sequence and the fragment gene sequence are finally extracted as gene feature data in the original data.
[0030] Specifically, according to different targets such as whole genome, targeted gene, RNA expression sequencing, etc., the core feature dimension (such as detection range, technical advantage, etc.) is determined, the feature extraction is focused on the key information required for the sequencing purpose, the irrelevant data interference is avoided, the extracted gene features are highly matched with the analysis target, the foundation is laid for subsequent accurate matching, and different signal types such as fluorescence, electrochemistry and scattering are adopted for differential conversion strategy: the fluorescence signal is combined with the space-time sequence of bases, the electrochemical and scattering signals are converted by feature spectrum matching, the accurate conversion of various signals to base sequences is realized, the original information of the gene sequence is completely retained, the sequence loss or misjudgment caused by the difference in signal types is avoided, the starting / ending position and length of the gene fragment are determined, the physical boundary of each gene fragment is accurately defined, the integrity of the basic gene sequence and the consistency of the fragment division are ensured, the feature confusion caused by the fuzzy fragment boundary is avoided, and a clear structure foundation is provided for subsequent fragmentation feature extraction. Extracting the original base sequence (complete sequence information) and the fragment gene sequence (split feature) retains the global sequence features of the gene and refines the local fragment features, meets the needs of different matching scenarios (such as overall sequence comparison or local fragment analysis), enhances the flexibility and applicability of feature data, and finally provides high-quality, multi-level feature support for efficient comparison of gene features and reference database.
[0031] In order to solve the problems in the prior art that multi-source gene data integration is difficult, the quality is uneven, the reference database structure is not standardized, the comparison strategy is single, the matching accuracy and reliability are insufficient, and there is a lack of systematic evaluation and optimization mechanism, please refer to Figure 1 The embodiment provides the following technical solutions: A standardized reference gene sample database is established, and the extracted gene features are compared and matched with the reference gene sample database, including: The establishment of the standardized reference gene sample database is: integrating multi-source reference data, the multi-source reference data including standard gene sequences in public databases, known functional gene annotations, variation databases, and self-built specific sample data, and then performing quality screening on the multi-source reference data, the quality screening being low-quality sequence elimination, incomplete data annotation, and retaining high-credibility data verified by experiments or multiple iterations on the collected multi-source reference data; After the multi-source reference data integration is completed, the database structure of the standardized reference gene sample database is constructed, the database structure is constructed in layers according to data types and application scenarios, including a basic layer, a feature layer and an application layer, wherein the basic layer stores original gene sequences and core annotations; the feature layer stores extracted gene features; the application layer constructs a sub-database for a specific scene; at the same time, multi-dimensional indexes are established, including sequence indexes, feature indexes and classification indexes; According to the extracted gene features, a comparison strategy is selected, including sequence overall comparison, local feature comparison and multi-feature joint comparison; After the comparison strategy is selected, the multi-dimensional indexes in the reference gene sample database are used to preliminarily screen the extracted gene features, the preliminary screening is to match sample sequence fragments with reference sequences through k-mer index, and reference sequences sharing ≥50% k-mer fragments are screened out as a candidate set, if the sample contains a known functional variation, the reference sequence containing the variation in the database is directly searched and included in the candidate set; The reference sequences in the candidate set are compared with the extracted gene features in each dimension, including sequence level comparison, variation level comparison and function level comparison, wherein the sequence level comparison is to calculate the consistency percentage of the sample sequence and the reference sequence, and record the difference sites; the variation level comparison is to compare whether the position and type of the sample variation and the reference variation are consistent, and to calculate the proportion of the number of matched variations in the total sample variations; the function level comparison is to compare whether the functional domain type and position of the sample and the reference sequence overlap, and to calculate the matching degree of the functional features; According to the comparison result of each dimension, the comparison matching result is sorted, and the highest one is selected as the final comparison matching result.
[0032] Specifically, by integrating public databases, variant data and self-built data, high reliability information is retained through quality screening to ensure reliable data foundation; multi-dimensional index is matched with hierarchical database structure (basic layer, feature layer and application layer) to realize ordered data storage and meet the needs of rapid retrieval in different scenarios, improve the compatibility and calling efficiency of cross-source data, select overall alignment, local alignment or multi-feature joint alignment according to the type of gene feature, and match different needs such as full sequence analysis and specific variant detection, avoid the limitations of single strategy, enhance the adaptability of the alignment system to complex gene features, quickly narrow down the candidate range through k-mer index matching and known variant retrieval, reduce invalid alignment calculation, while ensuring that the candidate set covers high similarity reference sequences, improving efficiency while avoiding missed detection, laying a foundation for accurate matching, aligning dimension by dimension from sequence consistency, variant type to functional domain feature, covering the structure and functional information of genes comprehensively, rather than relying on a single sequence indicator, which can effectively identify genes with similar sequences but different functions, or homologous sequences with specific variants, significantly improving the accuracy and biological significance of the matching results, selecting the optimal matching result through comprehensive scoring and sorting, reducing subjective judgment errors, while providing clear priority basis for subsequent evaluation and optimization, ensuring that the final matching result not only meets the sequence characteristics, but also meets the functional expectations, providing strong support for accurate identification of gene sequencing samples.
[0033] The reference sequences in the candidate set and the extracted gene features are aligned dimension by dimension, including: The sample sequence and the reference sequence are split, and the sequence consistency percentage of each sample subsequence and the corresponding reference subsequence is obtained based on the consistency comparison of each sample subsequence and the corresponding reference subsequence after splitting; Based on the consistency comparison results of each sample subsequence and the corresponding reference subsequence after splitting, the difference sites with sequence differences are determined; Based on the preset matching rule, the variation position and variation type of each difference site in the same subsequence of the sample sequence are determined, and the first variation result of each subsequence is obtained; The ratio of the number of variation features in each subsequence of the sample sequence to the number of corresponding variation sites is taken as the second variation result of the corresponding subsequence; Each subsequence is annotated with functional domains, and the functional domain coordinates of each subsequence of the reference sequence and each subsequence of the sample sequence are mapped to the same genome coordinate system, the functional domain overlap is detected, and the first functional result is obtained; The ratio of the number of functional domain features in each subsequence of the sample sequence to the number of functional domains in the current subsequence of the sample sequence is taken as the second functional result of the corresponding subsequence; The consistency percentage of all subsequences in the sample sequence is taken as the sequence alignment result of the sample sequence at the sequence level; The first variation result and the corresponding second variation result of all sub-sequences in the sample sequence are integrated to obtain a variation alignment result of the sample sequence at the variation level, and similarly, a functional alignment result of the sample sequence at the functional level is obtained; The sequence alignment result, the variation alignment result and the functional alignment are integrated based on the corresponding alignment weight to obtain a first alignment result of the sample sequence; The sequence level alignment, the variation level alignment and the functional level alignment are integrated in parallel, and a second alignment result of the sample sequence is obtained based on the parallel integration result; The first alignment result and the second alignment result are sorted respectively, and the sorting reliability of the first alignment result and the second alignment result is weighted to obtain an alignment matching result of the sample sequence.
[0034] In this embodiment, sequence splitting is to split a long gene sequence (sample sequence or reference sequence) into multiple shorter sub-sequences according to a fixed length or a specific rule. For example, the length of the sample sequence is 10,000 bp, and the length of the reference sequence is 10,000 bp. Then, according to the splitting of every 1,000 bp, the sample sub-sequences S1 (1-1000 bp), S2 (1001-2000 bp) … S10 (9001-10000 bp) are obtained, and the reference sub-sequences R1, R2, …, R10 are obtained.
[0035] In this embodiment, the percentage of consistency refers to the proportion of the number of bases that are completely matched in each sub-sequence of the sample sequence to the total number of bases in the sub-sequence. For example, after aligning the sample sub-sequence S1 (1000 bp) with the reference sub-sequence R1, it is found that 950 bases are completely matched, and the percentage of consistency is 95%.
[0036] In this embodiment, the difference site refers to the site corresponding to the position where the bases of the sample sub-sequence and the reference sub-sequence are inconsistent. The difference site includes variation or sequencing error. For example, at the 150 bp position of S1, the sample is “A” and the reference is “G”, and the 150 bp is a difference site.
[0037] In this embodiment, to determine the difference site, the position of the site, the sample base, the reference base and the alignment quality value need to be determined.
[0038] In this embodiment, the first variation result is the variation position (chromosome coordinate) and type (such as SNP, Indel) determined based on the difference site, which reflects the variation characteristics of the corresponding sub-sequence. For example, there is a difference site 150 bp (A→G) as SNP in S1, and the corresponding first variation result is (150, SNP).
[0039] In this embodiment, the second variation result refers to the ratio of the number of variation features in the subsequence of the sample sequence to the total number of difference sites, which is used to measure the density of variation detection. For example, in S1, 2 variations (SNP+Indel) are detected, and the total number of difference sites is 5. Therefore, the second variation result is 0.4, and 40% of the difference sites are valid variations.
[0040] In this embodiment, the functional domain annotation is annotated by a preset database to annotate the functional domain and its position in the gene sequence. For example, the preset database includes Pfam, InterPro, etc., and the functional domain includes kinase domain, DNA binding domain, etc. For example, a functional domain is annotated in a subsequence of the sample sequence, and the same functional domain at the same position is annotated in the corresponding subsequence of the reference sequence.
[0041] In this embodiment, the functional domain coordinate mapping is to map the functional domain positions of the sample and reference subsequences to the same genome coordinate system to eliminate splitting bias. The mapping to the same genome coordinate system can determine whether the coordinate positions overlap. For example, the genome coordinate system can be human GRCh38, etc.
[0042] In this embodiment, the first functional result refers to the overlap of the functional domain of the sample subsequence and the reference subsequence, including the overlap length, type consistency, etc.
[0043] In this embodiment, the second functional result refers to the ratio of the number of functional domains matched in the sample subsequence to the total number of functional domains in the sample, reflecting the retention degree of the corresponding function of the gene.
[0044] In this embodiment, the sequence alignment result refers to a set of consistency percentages of all subsequences in the sample sequence. For example, the consistency percentages of 10 subsequences of the sample sequence are [95%, 92%, 98%, …, 95%], and the sequence alignment result is the average value of the consistency percentages, which can be 94%, for example.
[0045] In this embodiment, the variation alignment result is a comprehensive variation alignment result of the first variation result and the second variation result of all subsequences in the sample sequence. For example, subsequence 1: variation 1 (SNP), variation 2 (Indel), second variation result = 0.4; subsequence 2: no variation, second variation result = 0; variation alignment result: total variation = 2, average second variation result = 0.2.
[0046] In this embodiment, the functional alignment result is a comprehensive functional alignment result of the first functional result and the second functional result of all subsequences in the sample sequence. For example, subsequence 1: functional domain matching ratio = 1.0; subsequence 2: functional domain matching ratio = 0.8 (part of the functional domain is missing); functional alignment result: average functional matching ratio = 0.9.
[0047] In this embodiment, the hierarchical integration is based on sequence, variation, functional alignment results and preset weight to calculate the comprehensive score for the ranking of sample sequences; for example, the preset weight is 0.5, 0.3 and 0.2, and the sequence alignment score is 94.2, the variation alignment score is 0.2, and the functional alignment score is 0.9, then the first alignment result = 0.5x94.2 + 0.3x0.2 + 0.2x0.9 = 47.28.
[0048] In this embodiment, the parallel integration is to perform parallel processing on the sub-sequences after sequence splitting of sequence alignment, variation alignment and functional alignment.
[0049] In this embodiment, the second alignment result refers to the integrated alignment result after parallel integration; for example, after parallel processing, the sequence score is 95, the variation score is 0.25, and the functional score is 0.95. The second alignment result = 0.5x95 + 0.3x0.25 + 0.2x0.95 = 47.865.
[0050] In this embodiment, the weighted ranking is to assign weights to the ranking reliability of the first and second alignment results, and the sum of the weights corresponding to the first alignment result and the second alignment result is 1; for example, the ranking reliability can be determined by variance or confidence score; for example, the ranking reliability of the first alignment result is 0.9, the ranking reliability of the second alignment result is 0.7, the first alignment result is 47.28, and the second alignment result is 47.865, and the weighted score is 47.51.
[0051] In this embodiment, the alignment matching result is determined according to the weighted score of the sample sequence and the reference sequence, and the higher the weighted score, the more matched the sequence; for example, the weighted score of the sample sequence and the reference sequence A is the highest, and A is the best matching result.
[0052] Specifically, the entire gene sequence is split into multiple short gene sub-sequences, and the consistency percentage of each sub-sequence with each sub-sequence of the reference sequence is calculated to quickly locate the difference site; then, combined with the variation detection rule, the effective variation site is extracted from the difference site, and the variation density of the effective variation is calculated to form the variation spectrum feature; then, through the functional domain annotation and coordinate mapping, the functional structure overlap and functional retention proportion of each sub-sequence in the sample sequence and the reference sequence are analyzed, and the hierarchical integration strategy is adopted to fuse the alignment results of the sequence, variation and function three levels according to the preset weight to generate the score of the first alignment result; at the same time, the gene data of each level is integrated in parallel and weighted to obtain the score of the second alignment result. Finally, the corresponding weight is dynamically adjusted combined with the ranking reliability of the first alignment result and the second alignment result, and the optimal matching sequence after weighted ranking is output to improve the matching efficiency of large-scale genomic data.
[0053] The beneficial effects of the above technical solutions are: the sample sequence of gene sequencing is matched by the multi-dimensional hierarchical comparison method, wherein the sequence splitting and parallel processing greatly shorten the matching calculation time of the gene sequence; meanwhile, the difference site detection is combined with the variation characteristic quantization, which can more efficiently distinguish effective variations and sequencing errors, and improve the reliability of variation identification; meanwhile, the functional domain annotation and coordinate mapping are used to evaluate the functional similarity from the structure level, which makes up for the limitations of sequence alignment, and finally the sequence alignment, variation alignment and functional alignment results are integrated in a hierarchical and parallel manner for multi-level feature comparison and fusion to generate a more comprehensive matching result of the sample sequence, thereby providing an efficient and accurate genome matching scheme for gene sequencing.
[0054] Finally, the alignment matching result is evaluated and optimized, including: First, the alignment matching result is evaluated in multiple dimensions, including quantitative index evaluation, biological rationality verification, and abnormal matching case analysis; The quantitative index evaluation is to calculate the core matching quality indicators, including matching accuracy, feature coverage, variation consistency and repeat stability, and to mark low-quality matching data; the biological rationality verification is to check the biological logic of the matching result, and if high-frequency variations of irrelevant genes are matched, the non-matching rationality is marked, and the clinical sample is checked to see if the alignment matching result is consistent with the sample's clinical phenotype; the experimental sample is checked to see if the matching result meets the experimental design goal; the abnormal matching case analysis is to analyze the cases with high quantitative score but biological irrationality or low quantitative score but potential effectiveness in the alignment matching result; According to the multi-dimensional evaluation result, the alignment strategy is optimized, including: for low-accuracy matching results, increasing the weight of core features or tightening the sequence similarity threshold; for low-coverage matching results, expanding the alignment range or increasing the dimension of feature extraction; for population-specific bias results, preferentially calling reference samples of the same population in the database when aligning different population samples; Finally, the optimized alignment matching result is obtained, and the final optimization result is converted into visual data and transmitted to the display terminal for display of the optimization and evaluation results.
[0055] For low-accuracy matching results, the weight of core features is increased or the sequence similarity threshold is tightened, including: For low-accuracy matching results, based on the accuracy of core features and the accuracy threshold, the weight of the corresponding features is increased to obtain the optimized weight T; at the same time, the similarity threshold of the corresponding features is tightened to obtain the optimized similarity threshold S; The comparison strategy is optimized according to the optimized weight and the optimized similarity threshold.
[0056] The calculation formula of the optimized weight T is obtained as follows: ; Wherein, T is the optimized weight of the core feature, is the real-time weight of the core feature, Q is the matching accuracy of the core feature in the sample sequence, is the accuracy threshold of the core feature, is the weight adjustment coefficient, in the data matching of gene sequencing, the weight adjustment coefficient is used to control the influence degree of the key feature on the data matching result, for example, the key features include SNP hotspot area, random sequence, etc., The value range of T is 0-1.
[0057] The calculation formula of the optimized similarity threshold S is obtained as follows: ; Wherein, S is the optimized similarity threshold of the core feature, is the real-time similarity threshold of the core feature, is the threshold tightening step, wherein the threshold tightening step is determined based on the data matching accuracy of gene sequencing, and the value range of the threshold tightening step is (0, 0.5).
[0058] The beneficial effects of the above technical solution are that by increasing the core feature weight, the false matching caused by sequencing noise or local similarity can be reduced; at the same time, the similarity threshold is tightened to avoid false judgment of irrelevant sequences as matching due to random similarity, so that the priority of high similarity correct matching can be improved, and the interference of fuzzy matching can be eliminated, and the detection and analysis performance of variation detection or function analysis can be improved.
[0059] Specifically, the quantifiable index evaluation realizes the quantifiable judgment of the result quality by matching the core parameters such as the accuracy rate and the variation consistency, and the low-quality data marking facilitates the accurate filtering of invalid information; the biological rationality verification avoids the deviation of biological significance caused by the simple dependence on data indicators through cross verification from the angles of gene function logic, clinical phenotype association, experimental design target, etc.; the abnormal case analysis takes into account the cases of "high score unreasonable" and "low score potentially effective", reduces the misjudgment of special samples, and comprehensively ensures the accuracy of the evaluation; the feature weight or threshold is adjusted for the low accuracy rate result, directly improving the matching contribution of the core features; the comparison range and feature dimension are expanded for the low coverage problem, making up for the information loss; the reference samples of the same population are preferentially called for the population-specific deviation, enhancing the adaptability of cross-population data. This "problem-oriented" optimization mechanism can dynamically correct the comparison model, continuously improve the matching ability of complex samples, and the visual conversion and terminal display make the evaluation optimization results more intuitive, facilitating researchers to quickly interpret the matching quality and abnormal points; the optimized results not only retain the support of quantitative indicators, but also integrate biological logic verification, providing a high-credibility data basis for subsequent clinical diagnosis, experimental analysis and other applications, realizing the precise connection from data matching to actual application, and finally improving the reliability and practical value of the genetic sequencing sample data matching.
[0060] It should be noted that the relational terms herein such as first and second and the like are used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0061] Although the embodiments of the present application have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made thereto without departing from the principles and spirit of the present application.
Claims
1. A method for matching sample data of gene sequencing based on a microfluidic chip, characterized in that, The method comprises the following steps: Purification and standardization are performed on a gene sequencing raw sample; the purified and standardized raw sample is loaded into a microfluidic chip for preparation and processing of the microfluidic chip, while biochemical reaction conditions in the microfluidic chip are controlled; gene sequencing signals in a detection area of the microfluidic chip after biochemical reaction of the microfluidic chip are captured and converted into processable raw data; the converted raw data are preprocessed; gene features are extracted from the preprocessed raw data; A standardized reference gene sample database is established, and the extracted gene features are compared and matched with the reference gene sample database; and finally, the comparison and matching result is evaluated and optimized; The comparison and matching of the extracted gene features with the reference gene sample database comprises the following steps: The reference sequences in the candidate set are compared with the extracted gene features in each dimension, including sequence level comparison, variation level comparison and function level comparison; the sequence level comparison is to calculate the consistency percentage of the sample sequence and the reference sequence, and record the difference sites; the variation level comparison is to compare whether the position and type of the sample variation and the reference variation are consistent, and to calculate the proportion of the matched variation number to the total variation of the sample; the function level comparison is to compare whether the functional domain type and position of the sample and the reference sequence overlap, and to calculate the matching degree of the functional features; According to the comparison result in each dimension, the comparison and matching result is sorted, and the highest one is selected as the final comparison and matching result. 2.The microfluidic chip-based genetic sequencing sample data matching method according to claim 1, wherein, The comparison and matching of the extracted gene features with the reference gene sample database further comprises the following steps: The establishment of the standardized reference gene sample database comprises the following steps: integrating multi-source reference data, the multi-source reference data comprising standard gene sequences, known functional gene annotations, variation databases and self-built specific sample data in public databases, and performing quality screening on the multi-source reference data, the quality screening comprising removing low-quality sequences, incomplete data annotation and retaining high-confidence data verified by experiments or multiple iterations from the collected multi-source reference data; After the multi-source reference data are integrated, a database structure of the standardized reference gene sample database is constructed, the database structure being constructed in layers according to data types and application scenarios, comprising a basic layer, a feature layer and an application layer, wherein the basic layer stores original gene sequences and core annotations; the feature layer stores the extracted gene features; the application layer constructs a sub-database for a specific scene; meanwhile, multi-dimensional indexes are established, including sequence indexes, feature indexes and classification indexes; According to the extracted gene features, a comparison strategy is selected, comprising sequence overall comparison, local feature comparison and multi-feature joint comparison. After the comparison strategy selection is completed, the extracted gene features are preliminarily screened by using the multidimensional index in the reference gene sample database. The preliminary screening is to match the sample sequence fragments with the reference sequences by the k-mer index, and the reference sequences sharing >=50% k-mer fragments are screened out as the candidate set. If the sample contains a known functional variation, the reference sequence containing the variation in the database is directly searched and included in the candidate set. 3.The microfluidic chip-based gene sequencing sample data matching method of claim 2, wherein, The reference sequences in the candidate set and the extracted gene features are compared dimension by dimension, including: The sample sequence and the reference sequence are split, and the consistency of each sample subsequence and the corresponding reference subsequence after splitting is compared to obtain the sequence consistency percentage of each subsequence; Based on the consistency comparison results of each sample subsequence and the corresponding reference subsequence after splitting, the difference sites with sequence differences are determined; Based on the preset matching rule, the variation position and variation type of each difference site in the same subsequence of the sample sequence are determined to obtain the first variation result of each subsequence; The ratio of the number of variation features in each subsequence of the sample sequence to the number of corresponding variation sites is taken as the second variation result of the corresponding subsequence; Each subsequence is annotated with a functional domain, and the functional domain coordinates of each subsequence of the reference sequence and each subsequence of the sample sequence are mapped to the same genomic coordinate system to detect the functional domain overlap to obtain the first functional result; The ratio of the number of functional domain features in each subsequence of the sample sequence to the number of functional domains in the current subsequence of the sample sequence is taken as the second functional result of the corresponding subsequence; The consistency percentage of all sub-sequences in the sample sequence is taken as the sequence alignment result of the sample sequence at the sequence level; The first variation result and the corresponding second variation result of all sub-sequences in the sample sequence are combined to obtain the variation alignment result of the sample sequence at the variation level. Similarly, the functional alignment result of the sample sequence at the functional level is obtained; Based on the sequence alignment result, the variation alignment result, and the functional alignment combined with the corresponding alignment weight, hierarchical integration is performed to obtain the first alignment result of the sample sequence; The sequence level alignment, variation level alignment, and functional level alignment are integrated in parallel, and the second alignment result of the sample sequence is obtained based on the parallel integration result; The first alignment result and the second alignment result are sorted respectively, and the sorting reliability of the first alignment result and the second alignment result is weighted to obtain the alignment matching result of the sample sequence. 4.The microfluidic chip-based gene sequencing sample data matching method of claim 2, wherein, The raw gene sequencing sample is purified and standardized, including: First, select the lysis method according to the type of the raw sample. When the raw sample is a blood sample, centrifuge to remove red blood cells, and then treat white blood cells with lysis buffer. When the raw sample is a cell or tissue sample, use a lysis buffer containing a detergent to destroy the cell membrane and cell wall, release the intracellular nucleic acids, and add proteinase K to degrade the proteins in the sample; The lysed raw sample is subjected to nucleic acid separation and purification. The nucleic acid separation and purification process is as follows: the raw sample is subjected to protein removal, nucleic acid precipitation, salt removal, nucleic acid dissolution, and residual removal in sequence; The original sample after nucleic acid separation and purification is standardized, and the standardization treatment is: using a UV spectrophotometer to detect the absorbance ratio of the nucleic acid solution at 260nm and 280nm, then analyzing the nucleic acid fragment size by agarose gel electrophoresis, using fluorescence quantitative method or UV spectrophotometer to determine the concentration of the purified nucleic acid, and finally diluting or concentrating the original sample with nuclease-free water according to the loading requirements of the microfluidic chip; Finally, the purification and standardization of the original sample for gene sequencing are completed. 5.The microfluidic chip-based genetic sequencing sample data matching method of claim 4, wherein, The purified and standardized original sample is loaded into the microfluidic chip for the preparation and processing of the microfluidic chip, and the biochemical reaction conditions in the microfluidic chip are controlled, including: First, the microfluidic chip is structurally inspected and cleaned, wherein the microchannels, reaction chambers, sample inlets and detection zones of the microfluidic chip are first inspected for physical damage, then nuclease-free water or special chip cleaning solution is injected into the channel through the sample inlet, and after standing for 5-10 minutes, nitrogen or clean air is used to blow dry, removing the residual impurities, dust or protective agent at the factory; After the structural inspection and cleaning of the microfluidic chip are qualified, the loading method of the original sample is confirmed, including manual loading and automatic loading, wherein manual loading is used for small chips or low-throughput scenarios; automatic loading is used for high-throughput or precise control scenarios; After the loading method is confirmed, the original sample is loaded and distributed on the microfluidic chip; After the original sample is loaded, the sample is processed in the microfluidic chip according to the requirements of gene sequencing, and the processing flow is: the original sample is mixed with reagents, and the mixed original sample is sealed, and at the same time, the sealed microfluidic chip is fixed on the chip seat; The biochemical reaction in the microfluidic chip fixed on the chip seat is controlled, including temperature control, fluid environment control, reaction time control and environmental atmosphere control; Among them, the temperature control includes constant temperature reaction control and variable temperature reaction control; the fluid environment control includes fluid and pressure regulation and bubble and evaporation prevention and control; Finally, the biochemical reaction control of the microfluidic chip is completed. 6.The microfluidic chip-based genetic sequencing sample data matching method of claim 5, wherein, The gene sequencing signal of the detection zone of the microfluidic chip is captured after the biochemical reaction of the microfluidic chip, and is converted into processable raw data, including: The microfluidic chip after biochemical reaction control is taken out, and the chip surface is wiped with dust-free paper, and at the same time, the position of the detection zone is confirmed by using a microscope or chip positioning mark; After the position of the detection zone is confirmed, the detection equipment is confirmed according to the signal type, including fluorescence signal, electrochemical signal and scattering signal, wherein the detection equipment of the fluorescence signal is a fluorescence microscope or a confocal laser scanning microscope, and is equipped with excitation light filters and emission light filters matched with fluorescence markers; the electrochemical signal is connected to an electrochemical workstation, and the working electrode, reference electrode and working electrode of the chip detection zone are connected to the electrode interface of the workstation; the detection equipment of the scattering signal is a miniature spectrometer; The detection area of the microfluidic chip is scanned by a detection device, and the blank area on the microfluidic chip without sample is scanned during the signal scanning, and the background signal value is recorded. At the same time, the non-specific signal on the microfluidic chip is removed by using signal filtering software, and the effective signal of the reaction characteristics is reserved; The reserved effective signal is dynamically captured, and the change curve of the signal with time is recorded during the dynamic signal capture; The captured dynamic signal is converted into a digital signal by an analog-to-digital converter built in the device, and the digital signal is associated according to the time or space dimension, and the structured data is formed after the association; The structured data is standardized according to the sequencing universal format, and the standardization processing includes sample number, chip batch, detection time, device model and signal channel; The structured data after standardization processing is stored, and the processed raw data is obtained after storage. 7.The microfluidic chip-based genetic sequencing sample data matching method of claim 6, wherein, The converted raw data is preprocessed, including: The stored processable raw data is read, and the key information is parsed after reading, including sample ID, chip detection area coordinates, signal intensity value, detection time, channel identification and device parameter log; The parsed key information is checked for integrity, and the integrity check is to check whether the meta information of the raw data is complete. If there is a missing, the information supplement is performed in the signal conversion link; After the integrity check is completed, the signal quality is quantitatively analyzed, and the quantitative analysis is to calculate the statistical parameters of the effective signal, including the average value, standard deviation, maximum value and minimum value of the signal intensity. After the statistical parameter calculation, the signal-to-noise ratio analysis is performed, and the signal-to-noise ratio analysis is to divide the average intensity of the effective signal by the average intensity of the background signal. If SNR≥3, the signal quality is qualified. If SNR<3, it needs to be marked as low-quality data; The data after quantitative analysis is sequentially subjected to data cleaning, data standardization, data noise reduction, signal enhancement, data fragmentation and data alignment; The data cleaning is to identify outliers by a box plot or a threshold method, manually or automatically eliminate abnormal outliers, and at the same time, delete the area where the sample is not loaded in the microfluidic chip detection area and the area where the signal value is continuously 0, and correct the background signal value; the data standardization is to unify the scale by normalizing processing for the data difference of different samples or different chip batches, convert the structured data into a standardized format, and split the multi-channel signal and then realign the dimension; the data denoising is to process the data with random noise by using a sliding window smoothing method, and remove isolated noise points by using a median filtering method for spatial dimension data; the signal enhancement is to enhance the signal recognition degree by a contrast enhancement algorithm for the weak signal area; the data fragmentation is to split the data into several fragments according to a preset length if the original data corresponds to a long fragment gene sequence, and the start and end position information of the fragments is retained during the splitting; the data alignment is to map the signal points to a standard grid by coordinate correction for spatial dimension data if the chamber position is slightly offset due to chip processing error, and to unify the time axis by an alignment algorithm for time sequence data if there is a difference in the reaction start time of different samples. Finally, the data preprocessing process is completed. 8.The microfluidic chip-based gene sequencing sample data matching method of claim 7, wherein, Gene features are extracted from the preprocessed raw data, including: According to the purpose of gene sequencing, the core features are confirmed, wherein the target of gene sequencing includes whole genome sequencing, targeted gene sequencing and RNA expression sequencing, and the core features include detection range, core detection content, data characteristics, technical advantages and typical applications; The feature extraction dimension in the preprocessed raw data is confirmed, including spatial dimension, time dimension and signal intensity dimension; After the feature extraction dimension confirmation, the signals in the preprocessed raw data are converted correspondingly; The conversion of the fluorescence signal is to convert the signal intensity after preprocessing into base types according to the signal channel identifier, and to concatenate the bases in the order of arrangement to form the original base sequence in combination with the time or spatial order; the conversion of the electrochemical signal and the scattering signal is to match the feature map of the known nucleotide according to the signal fluctuation characteristics, and to convert it into a base sequence; After the corresponding conversion, the gene fragment boundary and length are confirmed, wherein the start and end positions of the gene fragment signal in the preprocessed raw data are confirmed, and at the same time, the length of each fragment is confirmed by the number of bases or the spatial range covered by the signal, to obtain the basic gene sequence; The core sequence features and fragmentation features in the basic gene sequence are extracted, and the original base sequence and fragment gene sequence are obtained after extraction. Finally, the original base sequence and fragment gene sequence are taken as the gene feature data extracted from the raw data. 9.The microfluidic chip-based gene sequencing sample data matching method of claim 8, wherein, Finally, the comparison and matching results are evaluated and optimized, including: First, the comparison and matching results are evaluated in multiple dimensions, including quantitative index evaluation, biological rationality verification and abnormal matching case analysis; Among them, the quantitative index evaluation is to calculate the core matching quality index, including the calculation of matching accuracy, feature coverage, variation consistency and repeat stability, and the low-quality matching data is marked; the biological rationality verification is to check the biological logic of the matching result, if the high-frequency variation of irrelevant gene is matched, the non-matching rationality is marked, and the clinical sample is compared with the matching result and the clinical phenotype of the sample to check whether they are consistent; the experimental sample is checked with the matching result and the processing background of the sample to verify whether it meets the experimental design target; the abnormal matching case analysis is to analyze the cases with high quantitative score but biological irrationality or low quantitative score but potential effectiveness in the matching result; According to the multi-dimensional evaluation results, the comparison strategy is optimized, the comparison strategy optimization is to increase the weight of the core feature or tighten the sequence similarity threshold for the low accuracy matching result; for the low coverage matching result, the comparison range is expanded or the dimension of feature extraction is increased; for the population-specific bias result, the reference sample of the same population in the database is called first in the comparison for different population samples; Finally, the comparison and matching results are optimized, and the final optimization results are converted into visual data, and transmitted to the display terminal for display of the optimization and evaluation results. 10.The microfluidic chip-based genetic sequencing sample data matching method of claim 9, wherein, For low accuracy matching results, increase the weight of core features or tighten sequence similarity thresholds, including: For low accuracy matching results, based on the accuracy of core features and the accuracy threshold, the weight of the corresponding feature is increased to obtain the optimized weight; at the same time, the similarity threshold of the corresponding feature is tightened to obtain the optimized similarity threshold; According to the optimization weight and the optimization similarity threshold, the comparison strategy is optimized.
Citation Information
Cited By
Intelligent identification method for transgenic crops based on high-throughput sequencing
CN121544285A
Intelligent identification method for transgenic crops based on high-throughput sequencing
CN121544285B