Genomic DNA automated extraction and purification system based on algorithm model
An automated genomic DNA extraction system that dynamically adjusts operating parameters through an algorithm model solves the problems of unstable DNA extraction and poor adaptability to high throughput in traditional methods, achieving a highly efficient and precise DNA extraction process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGDONG VOCATIONAL COLLEGE OF SCI & TRADE
- Filing Date
- 2025-05-20
- Publication Date
- 2026-05-05
AI Technical Summary
Traditional DNA extraction methods cannot dynamically adjust operating parameters, resulting in unstable DNA yield and purity in complex samples. They are difficult to adapt to high-throughput experimental scenarios and lack real-time monitoring capabilities, making them prone to experimental failure due to human intervention and operational deviations.
An automated genomic DNA extraction and purification system based on an algorithm model is adopted. The system collects data through a sample processing module, constructs a sub-model through a parameter calculation module, and executes the optimal operating parameters through an extraction module. By combining real-time feedback control and model optimization iteration, key operating parameters such as lysis, centrifugation, and washing are dynamically adjusted.
It significantly improves the automation and accuracy of DNA extraction, reduces operational errors, allows for rapid adaptation to new samples, reduces experimental pre-optimization time, and ensures the stability and efficiency of extraction results.
Smart Images

Figure CN120526844B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of biotechnology and automation, specifically to an automated genomic DNA extraction and purification system based on an algorithm model. Background Technology
[0002] Genomic DNA extraction and purification are fundamental steps in molecular biology experiments and are widely used in fields such as gene sequencing, disease diagnosis, species identification, and bioengineering. With the rapid development of high-throughput sequencing technology, laboratories have an increasingly urgent need for automation, standardization, and efficiency in DNA extraction. However, traditional DNA extraction methods have significant limitations.
[0003] Regarding parameter settings, existing automated equipment mostly uses fixed operating parameters, such as lysis time and centrifugation speed, which cannot dynamically adjust the process according to sample type (e.g., high-protein blood). This leads to unstable DNA yield and purity in complex samples. From an operational experience perspective, when facing novel or complex samples (e.g., difficult-to-lyse microorganisms), one can only rely on numerous manual pre-experiments to optimize parameters, which is not only time-consuming and labor-intensive but also difficult to standardize. In terms of real-time monitoring, traditional systems lack the ability to monitor key indicators (e.g., lysis buffer turbidity and washing buffer conductivity) in real time, making it impossible to correct operating strategies promptly during experiments. This can easily lead to experimental failures due to impurity residues or operational deviations. Furthermore, due to frequent manual intervention and cumbersome operating procedures, traditional methods are difficult to adapt to high-throughput experimental scenarios, resulting in low overall efficiency. Summary of the Invention
[0004] To address the problems existing in the prior art, this invention provides an automated genomic DNA extraction and purification system based on an algorithm model. The system includes: a sample processing module that collects initial biomass data and historical experimental data to construct real sample features and historical sample features; a parameter calculation module that constructs sub-models corresponding to each operational stage of DNA extraction and outputs the optimal operational parameters for each stage; each sub-model includes a training unit and a prediction unit; the training unit uses historical sample features as a training set for training; the prediction unit inputs real sample features into the trained sub-model to obtain the optimal operational parameters; and the extraction module receives the optimal operational parameters, sets the optimal operational parameters for the automated equipment, and then performs DNA extraction. This invention can significantly improve the automation level of DNA extraction, increasing efficiency while meeting the precision requirements of modern biotechnology.
[0005] The present invention adopts the following technical solution: an automated genomic DNA extraction and purification system based on an algorithm model, the system comprising: a sample processing module, a parameter calculation module, and an extraction module;
[0006] The sample processing module is used to collect initial biomass data and construct real sample features; acquire experimental data of historical DNA extraction and construct historical sample features;
[0007] The parameter calculation module is used to construct sub-models corresponding to each operation stage of DNA extraction based on the gradient boosting tree algorithm, and output the optimal operation parameters corresponding to each operation stage of DNA extraction through the sub-models corresponding to each operation stage; the sub-model includes a training unit and a prediction unit.
[0008] The training unit is used to train the sub-models corresponding to each operation stage of DNA extraction using historical sample features as the training set, so as to obtain the trained sub-models.
[0009] The prediction unit is used to input real sample features into the trained sub-model to obtain the optimal operation parameters corresponding to each operation stage of DNA extraction.
[0010] The extraction module is used to receive the optimal operating parameters corresponding to each stage of DNA extraction, and to perform DNA extraction after setting the optimal operating parameters for the automated equipment.
[0011] Furthermore, the system also includes: a model optimization iteration module;
[0012] The model optimization and iteration module is used to detect the extracted DNA and generate detection results, construct an experimental report based on the optimal operating parameters of the current sub-model and the detection results, and use the experimental report as the training set of the training unit to optimize and iterate the sub-models corresponding to each operation stage of DNA extraction until the set conditions are met.
[0013] Furthermore, the initial biomass data includes: optical density at a wavelength of 600 nm, absorbance at a wavelength of 405 nm, and sample type.
[0014] Furthermore, the various operational stages of the DNA extraction include: lysis stage, centrifugation stage, washing stage, and elution stage.
[0015] Furthermore, the historical sample features are used as the training set to train the sub-models corresponding to each operational stage of DNA extraction, specifically:
[0016] The average value of the parameters corresponding to each operational stage of DNA extraction in the historical sample features is used as the initial prediction value.
[0017] The sub-models corresponding to each operation stage are used to build decision trees through iteration, and the initial predicted values and the true values are fitted with residuals until the set conditions are met, so as to obtain the trained sub-models corresponding to each operation stage.
[0018] Furthermore, the optimal operating parameters corresponding to each operating stage include:
[0019] Pyrolysis stage: pyrolysis solution dosage, pyrolysis temperature, and pyrolysis time;
[0020] Centrifugation stage: centrifugation speed and centrifugation time;
[0021] Washing stage: amount of rinsing solution used, number of washes;
[0022] Elution phase: eluent temperature and elution volume.
[0023] The beneficial effects of this invention are as follows: This invention dynamically adjusts key operational parameters such as lysis, centrifugation, and washing through the gradient boosting tree algorithm, solving the "one-size-fits-all" problem in traditional methods and greatly improving the purity of DNA extraction from samples; from sample identification and parameter calculation to equipment control and result feedback, no manual intervention is required, significantly reducing operational errors and making it suitable for high-throughput scenarios; the model can quickly adapt to new samples through training, reducing experimental pre-optimization time, and the model can be further automatically corrected by the optimization iteration mechanism to ensure the stability of extraction results. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a schematic diagram of an automated genomic DNA extraction and purification system based on an algorithm model, according to an embodiment of the present invention.
[0026] Figure 2 This is a schematic diagram illustrating an automated extraction and purification process for genomic DNA based on an algorithm model, according to an embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] A schematic diagram of an automated genomic DNA extraction and purification system based on an algorithm model according to an embodiment of the present invention is shown below. Figure 1As shown, it includes: a sample processing module, a parameter calculation module, an extraction module, and a model optimization and iteration module;
[0029] The sample processing module is used to collect initial biomass data and construct real sample characteristics; and to acquire experimental data of historical DNA extraction and construct historical sample characteristics.
[0030] In this embodiment of the invention, the sample processing module can identify the sample type by scanning the barcode or the label set on the sample shell, and obtain initial biomass data for samples such as bacteria and blood, such as OD600 value / A405 value, and construct real sample characteristics by combining the reagent kit type, etc.
[0031] Initial biomass data refers to the relevant data on the initial content or quantity of biological matter in a sample;
[0032] OD600 value: OD is an abbreviation for Optical Density. OD600 refers to the optical density of a sample measured at a wavelength of 600nm. In microbiology, OD600 is used to measure the cell density of bacterial cultures because bacterial cells absorb light of specific wavelengths, and optical density and cell number are linearly related within a certain range.
[0033] A405 value: A stands for Absorbance. A405 refers to the absorbance at a wavelength of 405nm. The absorbance of the reaction product is detected by measuring the wavelength of 405nm, thereby reflecting the content of the target substance in the sample.
[0034] Kit type: A kit is a combination of reagents and materials used for a specific detection or experiment, such as a kit for detecting specific bacteria or a kit for detecting a certain marker in blood. The type of kit can affect the principle, sensitivity, and specificity of the detection.
[0035] In this embodiment of the invention, a barcode is affixed to the sample, and the barcode stores the sample type (e.g., E. coli, bacteria, blood). The sample type is determined by scanning the sample barcode using a camera or other device with scanning and input functions. The initial biomass data of the sample is obtained in different ways depending on the sample type. For bacterial samples, an ultra-micro nucleic acid protein analyzer is used to detect the initial turbidity (OD600) or cell concentration; for blood samples, a blood analyzer is used to detect the hemoglobin concentration (A405 value).
[0036] The sample processing module is also used to acquire historical sample data, which includes two aspects: first, the "optimal operating parameters" for different samples, representing the relevant parameters that achieve the best results when extracting samples from historical data, such as temperature, time, pressure, and solvent volume; second, the "corresponding yield and purity." "Yield" refers to the ratio of the actual yield to the theoretical yield of the target product during extraction, reflecting extraction efficiency; "purity" refers to the proportion of pure substances in the target product, reflecting the quality of the extracted product. In this embodiment of the invention, the historical sample data needs to systematically collect standardized data of different sample types, operating parameters, and corresponding results. The sample processing module acquires all operating steps and test results recorded in the laboratory information management system to avoid manual omissions. This data is stored in three categories: input features, operating parameters, and result labels. The data contained in each category is shown below:
[0037] (1) Input features:
[0038] Sample attributes: Sample type (0=bacteria, 1=blood), species (e.g., "chicken", "E. coli"), tissue location (e.g., "whole blood");
[0039] Initial parameters: biomass (OD600 / A405 value), reagent kit type (e.g., "DP302" or "DP104"). "DP302" and "DP104" are reagent kit numbers used as examples in this embodiment of the invention. In actual operation, the actual numbers of existing reagent kits can be referred to.
[0040] (2) Operating parameters:
[0041] Pyrolysis stage: pyrolysis solution dosage, pyrolysis temperature, and pyrolysis time;
[0042] Centrifugation stage: centrifugation speed and centrifugation time;
[0043] Washing stage: amount of rinsing solution used, number of washes;
[0044] Elution phase: eluent temperature and eluent volume.
[0045] (3) Result Labels:
[0046] Key labels: DNA yield (ng), purity (A260 / A280);
[0047] Auxiliary labels: electrophoresis integrity score (0-10 points), operation time (min).
[0048] The sample processing module further standardizes the data in the collected historical sample data, normalizes continuous variables (such as pyrolysis temperature) by scaling to the 0-1 range, encodes categorical variables (such as sample type) using one-hot encoding, and uses the Z-score method to remove outliers from the historical samples. This allows the construction of historical sample features based on the processed historical sample data. The Z-score method is a commonly used statistical method for identifying outliers in a dataset. Its specific implementation can be found in existing technologies, and will not be elaborated upon in this embodiment of the invention.
[0049] After collecting and processing the historical sample data, continuous incremental collection is required. This means that after each automated extraction (including manual operation), the actual operation parameters and results are recorded and added to the training set as new data. When processing new sample types, 20-30 initial data points are obtained through manual orthogonal experiments before being incorporated into the model training to improve generalization ability. Data is stored using CSV / Parquet files or a MySQL database, with fields corresponding to the data structure, so that it can be imported into the machine learning framework in batches later.
[0050] The parameter calculation module is used to construct sub-models corresponding to each operation stage of DNA extraction based on the gradient boosting tree algorithm, and output the optimal operation parameters corresponding to each operation stage of DNA extraction through the sub-models corresponding to each operation stage.
[0051] Gradient Boosting Decision Tree (GBDT) is an ensemble learning algorithm that iteratively trains multiple decision trees and sums their results to obtain the final prediction. GBDT excels in handling regression and classification problems and can capture complex nonlinear relationships in data. This invention uses historical sample data as a basis to train independent GBDT models for each of the lysis, centrifugation, washing, and elution stages. For the lysis stage, a GBDT model is trained using sample features from the training set and the optimal operating parameters for the lysis stage to output the lysis buffer volume, lysis temperature, and lysis time for the actual sample characteristics. Independent GBDT models are constructed for each of the four different stages (lysis, centrifugation, washing, and elution). Each model focuses only on the relationship between the data features and results of the corresponding stage. By learning from historical sample data for that stage, a prediction model suitable for that stage is established, allowing for more accurate modeling and prediction for each stage.
[0052] The sub-model includes a training unit and a prediction unit;
[0053] The training unit is used to train the sub-models corresponding to each operational stage of DNA extraction using historical sample features as the training set, so as to obtain the trained sub-models.
[0054] The training unit first uses the mean values of parameters corresponding to each operational stage of DNA extraction from historical sample features as initial prediction values. Then, it iteratively constructs a decision tree to fit the residual between the initial prediction value and the true value of the current model (using mean squared error MS as the loss function). Each new tree focuses on correcting the prediction bias of the previous model. To prevent overfitting, this embodiment of the invention introduces a regularization strategy, including limiting the depth of the decision tree (usually 3-5 layers), setting a minimum number of leaf node samples, adding a learning rate (to control the contribution of each tree), and subsampling (randomly selecting some samples / features for training), to ensure that the model has strong generalization ability on historical data. Finally, the hyperparameters (number of trees, learning rate, maximum depth, etc.) are systematically tuned through 5-fold cross-validation to minimize the prediction error, ultimately forming a set of efficient prediction models for different stage parameters. This model can independently optimize the operation of each stage and achieve full-process parameter coordination through the unification of input features, forming a closed loop of "initial prediction - residual iterative fitting - loss function optimization".
[0055] The prediction unit is used to input real sample features into the trained sub-model to obtain the optimal operation parameters corresponding to each operation stage of DNA extraction.
[0056] The real sample features are input into the trained GBDT model, and the model at each stage outputs the optimal operating parameters for that stage. The optimal operating parameters include:
[0057] Pyrolysis stage: pyrolysis solution dosage, pyrolysis temperature, and pyrolysis time;
[0058] Centrifugation stage: centrifugation speed and centrifugation time;
[0059] Washing stage: rinsing solution dosage and number of washes; (after determining the optimal number of washes, the number of washes can be dynamically increased based on impurity detection results).
[0060] Elution phase: elution buffer temperature (e.g., 60℃ to increase DNA solubility), elution volume.
[0061] The extraction module receives the optimal operating parameters for each stage of DNA extraction, and performs DNA extraction after setting the optimal operating parameters for the automated equipment.
[0062] In this embodiment of the invention, the extraction module inputs the optimal operating parameters output by the Gradient Boosting Tree (GBDT) algorithm model into automated equipment such as centrifuges, temperature control modules, and pipetting robots to achieve DNA extraction and purification. Simultaneously, a real-time feedback control system is constructed based on turbidity and conductivity sensors. During the lysis stage, the lysis time and lysis buffer volume are dynamically adjusted using the OD600 value as an indicator. During the washing stage, the washing strategy is optimized based on the column membrane flow rate, and an abnormal termination mechanism is set. The adjusted parameters and abnormal data are then incorporated back into the training data to optimize the GBDT model, forming a data-driven automated optimization closed loop that ensures the DNA extraction process is efficient, accurate, and controllable.
[0063] In this embodiment of the invention, the constructed real-time feedback control system can achieve dynamic control of the lysis and washing stages through sensor data, ensuring that key steps meet the standards. After lysis, the clarity of the solution is monitored by a turbidity sensor, and the clarity of the bacterial lysate (OD600 value) and the hemoglobin concentration of the blood lysate (A405 value) are monitored by a turbidity sensor / 405nm spectrophotometry dedicated sensor feedback control algorithm to ensure that the cells are fully broken and the DNA is completely released. After washing, the washing degree is monitored by a conductivity sensor, and the solution conductivity is monitored by a conductivity sensor feedback control algorithm to ensure that the washing is thorough and does not damage the DNA (a conductivity sensor detection value below 10 μS / cm indicates impurity residue). Examples of specific feedback control strategies for some operation stages are given in this embodiment of the invention:
[0064] Turbidity sensor / 405nm spectrophotometry dedicated sensor feedback control algorithm: During the lysis phase, for bacterial samples, the clarity of the lysis buffer (OD600 value) is monitored to ensure complete cell lysis and DNA release. An infrared turbidity sensor (detection wavelength 600nm) is used to reflect the concentration of unlysed cells / fragments in the solution through light absorbance, with an accuracy of ±0.02 OD units. Real-time sampling is performed every second. The trigger condition for real-time data is that the sensor data is adjusted immediately based on threshold judgment to ensure that the current sample lysis meets the standard. The adjustment strategy for bacterial samples is shown in Table 1:
[0065] Table 1. Schematic diagram of feedback control strategy during bacterial lysis phase
[0066]
[0067] When the lysis result is slightly below standard, maintain the current lysis temperature (e.g., 56℃) and extend the lysis time until the OD600 of the test result is less than 40% of the initial value. When the lysis result is severely below standard, use a high-precision pipette to add lysis buffer (error ±0.5μl) and vortex mix for 10 seconds to ensure uniform distribution of enzyme solution, and extend the lysis time until the OD600 of the test result is less than 40% of the initial value. When abnormal termination occurs, mark the sample as "difficult to lyse type" (record sample ID and initial characteristics) and automatically import the abnormal data into the historical database for GBDT model optimization. If the OD600 after lysis is less than 40% of the initial value, the washing stage will use the "standard washing program" by default (2 washes, 600μl rinsing buffer). If the lysis stage triggers the addition of lysis buffer, the washing stage will automatically add 1 wash (to prevent residual lysis buffer from affecting subsequent experiments).
[0068] For blood samples, the lysis status is indicated by monitoring the A405 value (hemoglobin concentration) in the lysis buffer. A dedicated 405nm spectrophotometric sensor is used, and the absorbance value reflects the hemoglobin concentration in the blood sample. Samples are collected in real time every second. The trigger condition for real-time data is that the sensor data is adjusted immediately based on a threshold judgment to ensure that the current sample lysis meets the standard. The adjustment strategy for blood samples is shown in Table 2.
[0069] Table 2. Schematic diagram of feedback control strategy during bacterial sample lysis phase
[0070]
[0071] When the result is slightly substandard, maintain the current lysis temperature and extend the lysis time until the A405 value of the test result is less than 20% of the initial value. When the result is severely substandard, use a high-precision pipette to add proteinase K (error ±0.5 μl) and vortex mix for 10 seconds to ensure uniform distribution of the enzyme solution. Extend the lysis time until the A405 value of the test result is less than 20% of the initial value. When the abnormal termination occurs, mark the sample as "difficult to lyse" (record sample ID and initial characteristics) and automatically import the abnormal data into the historical database for GBDT model optimization. If the A405 value after lysis is less than 20% of the initial value, the washing stage will use the "standard washing program" (2 washes, 600 μl of rinsing buffer) by default. If the lysis stage triggers the addition of lysis buffer, the washing stage will automatically add 1 wash (to prevent residual lysis buffer from affecting subsequent experiments).
[0072] It should be noted that if the lysis strategy is adjusted during the extraction process, the adjusted lysis time, lysis buffer volume and sample characteristics should be correlated and used as new training data for the GBDT model to optimize the initial parameter prediction of "difficult-to-lyse samples".
[0073] Conductivity sensor feedback control algorithm: The conductivity sensor assesses the ion concentration in the solution by measuring the conductivity of the solution. During the washing stage, impurities affect the conductivity of the washing liquid. The adjustment strategies for different ranges are shown in Table 3.
[0074] Table 3. Schematic diagram of feedback control strategy during the washing stage
[0075]
[0076] When a low-speed warning is triggered, automatically replenish 600 μl of wash buffer (containing 70% ethanol), increase the centrifugation speed to 13000 rpm (10% higher than the standard procedure), and extend the washing time until the conductivity sensor reading is below 10 μS / cm. When severe blockage occurs, inject 300 μl of wash buffer in reverse, allow to stand for 30 seconds to soften impurities, and perform one high-intensity wash (centrifugation at 14000 rpm, 700 μl of wash buffer each time), but the total number of washes should be ≤5. When the process terminates safely, record the blockage type (e.g., protein residue) and use the abnormal data to optimize the washing parameters for the corresponding sample type.
[0077] In this embodiment of the invention, experiments have verified that when the number of washing cycles is greater than 5, the DNA yield decreases by more than 10%. In order to prevent DNA loss, the total number of washing cycles should not exceed 5. In order to ensure that there is no obvious ethanol residue, an empty column centrifugation step (12,000 rpm, 2 minutes) is added after each washing. The conductivity sensor is used to confirm that the conductivity sensor detection value is less than 10 μS / cm.
[0078] Similarly, if the washing strategy is adjusted during the extraction process, the abnormal flow rate type of the clogged sample (such as protein type) can be associated with the sample type (blood), thereby updating the GBDT model's prediction of the number of washes.
[0079] In this embodiment of the invention, the extraction module adopts an automated execution method, which means that the optimal operating parameters output from the algorithm are input into the automated equipment, and the automated equipment operates automatically to obtain the extracted and purified DNA. The automated equipment includes a centrifuge, a temperature control module (controlling the reaction temperature), a pipette robot with a waste liquid aspiration module (precisely adding each solution and discarding waste liquid), a vortex mixer, an automatic flipper (for complete sample lysis), and other automated extraction equipment. The pipette robot has a built-in negative pressure pump tip, which extends to the bottom of the collection tube through the negative pressure tip to accurately aspirate waste liquid and retain the adsorption column.
[0080] The model optimization and iteration module is used to detect the extracted DNA and generate detection results. Based on the optimal operating parameters of the current sub-model and the detection results, an experimental report is constructed. The experimental report is used as the training set of the training unit to optimize and iterate the sub-models corresponding to each operation stage of DNA extraction until the set conditions are met.
[0081] The purified sample, extracted by the model optimization iteration module, is then analyzed for concentration and purity using an ultra-micro nucleic acid and protein analyzer via the DNA detection module. Electrophoresis integrity data is acquired using an agarose gel imaging system. Subsequently, the model input features of the real sample, the model output optimal operating parameters, and the detection results are integrated to output an experimental report. Finally, using the detection data as feedback, the gradient boosting tree algorithm model parameters are continuously optimized to improve the accuracy of predicting the optimal operating parameters, forming a complete closed loop from detection and reporting to model optimization. A schematic diagram of the overall implementation process is shown below. Figure 2 As shown, this drives the continuous iteration and upgrading of the DNA extraction automation process.
[0082] In a specific experimental embodiment of the present invention:
[0083] The sample processing module extracts 12 μl of whole chicken blood sample and affixes a barcode containing sample information (label code: BLOOD-20240601-001, storing information such as "chicken blood", "blood sample", and "DP348 reagent kit"). The system automatically identifies the sample as blood by scanning the barcode with the module's camera and retrieves the corresponding reagent kit parameter template (Tiangen DP348, suitable for blood genomic DNA extraction). Buffer GS is added to the sample to bring the total volume to 200 μl. After mixing, the hemoglobin concentration is detected using an ultra-micro nucleic acid protein analyzer (A405 value = 0.9), assessing the level of impurities after erythrocyte lysis in the blood (A405 > 0.8 indicates a high protein background). This generates the actual sample characteristics: sample type (blood, uniquely encoded [0,1]), species (chicken), initial biomass (A405 = 0.9), and reagent kit type (DP348), which are then input into the parameter calculation module.
[0084] The parameter calculation module can receive multi-dimensional feature input: The model receives real sample features: [sample type = blood, hemoglobin concentration = 0.9, reagent kit = DP348, target DNA = genomic DNA], calls pre-trained blood sample lysis sub-model, washing sub-model, and other independent GBDT models, and outputs optimized parameters for each stage based on the input multi-dimensional features:
[0085] Lysis phase: Proteinase K concentration = 30 μl (higher than the usual 20 μl, because A405 = 0.9 indicates high protein, requiring enhanced protein digestion), lysis temperature = 56℃, time = 45 min (15 min longer than the usual 30 min to ensure full rupture of erythrocyte nuclear membranes).
[0086] Centrifugation stage: The first centrifugation speed is 14,000 rpm (normally 12,000 rpm) for 2 minutes to promote cell debris precipitation.
[0087] Washing stage: The amount of rinsing solution is increased to 700 μl / wash (normally 600 μl), and the number of washes is set to 3 (normally 2 times, but high protein samples need to be washed more thoroughly).
[0088] Elution stage: Preheat the elution buffer to 60°C, use 50 μl of buffer, and extend the standing time to 5 min to improve DNA elution efficiency.
[0089] The extraction module inputs the optimal operating parameters for each operating stage obtained above into the automated equipment for the following operations:
[0090] Lysis procedure: 30 μl proteinase K and 200 μl buffer GB were precisely added by the pipette robot. After vortexing for 10 seconds, the temperature control module raised the temperature to 56°C and maintained it for 45 min. During this period, the automatic inverter was used to invert and mix the mixture once every 5 min (to ensure uniform digestion).
[0091] Centrifugation and washing: Centrifuge at 14,000 rpm for 2 min, collect the supernatant and transfer it to the adsorption column; add 700 μl of PW washing solution sequentially using a pipette, perform 3 washes, each centrifuged at 13,000 rpm (10% higher than normal) for 60 seconds.
[0092] Elution and collection: After the elution buffer TB is preheated to 60°C, it is added to the adsorption column, allowed to stand at room temperature for 2 min, and centrifuged at 12000 rpm for 2 min. The DNA solution is automatically collected into a 1.5 ml centrifuge tube.
[0093] Lysis Adjustment: After lysis, the A405 value detected by the dedicated sensor of 405nm spectrophotometry is 0.15 (<20% of the initial value of 0.9, which meets the standard), and the standard washing process is initiated; if the detected value is abnormal (e.g., 0.4), the system automatically triggers the strategy of "supplementing 5μl of proteinase K+ to extend the lysis time until the detected A405 value is less than 20% of the initial value" (not triggered in this embodiment of the invention).
[0094] Washing adjustment: The conductivity sensor detects that the conductivity of the solution after washing is 8μS / cm (the threshold is <10μS / cm), indicating that the impurity residue is within the standard and no additional washing is triggered (if the conductivity is >100μS / cm, one high-intensity wash will be added automatically).
[0095] The model optimization and iteration module acquires the detection results: After extraction, the ultra-micro nucleic acid analyzer shows a concentration of 65 ng / μl (compared to approximately 50 ng / μl using traditional methods), A260 / A280 = 1.86 (compared to approximately 1.75 using traditional methods, representing a 6% increase in purity). Agarose gel electrophoresis shows a single bright main band (molecular weight approximately 23 kb) without tailing. Gel imaging analysis shows a band intensity of 1800 (out of 2000, with an integrity score of 9). Sample characteristics (sample type: chicken blood, initial biomass: A405 = 0.9, kit type: ...) are also considered. The experimental report was generated by integrating the Tiangen DP348 kit, operating parameters (proteinase K = 30 μl, lysis temperature 56℃, lysis time 45 min, centrifugation speed 14000 rpm, centrifugation time 2 min, rinsing buffer volume 700 μl / wash, 3 washes, elution buffer volume 50 μl and preheated to 60℃) and detection results (DNA concentration 65 ng / μl, purity 1.86, electrophoresis integrity score 9 points). This data was then added to historical sample data as a training set to optimize the plant sample washing parameter prediction model.
[0096] In another specific experimental embodiment of the present invention:
[0097] Genomic DNA extraction from E. coli (using the Tiangen DP104 kit)
[0098] The sample processing module collects 1 ml of E. coli culture medium, identifies it as a bacterial sample by barcode scanning, and detects its OD600 value as 1.2 (high initial biomass) using an ultra-micro nucleic acid and protein analyzer, before inputting it into the model.
[0099] The parameter calculation module, taking into account the high biomass characteristics of bacterial samples, outputs the optimal operating parameters for each stage of the operation:
[0100] Lysis stage: Use 15 μl of lysis buffer (usually 10 μl), lysis temperature 56℃, time 13 min, to ensure full lysis of the cell wall.
[0101] Centrifugation stage: The first centrifugation speed is 15,000 rpm (normally 12,000 rpm) for 2 minutes to separate bacterial cell fragments.
[0102] Washing stage: 600 μl of rinsing solution per wash, 2 washes (normally 2 washes, but not increased due to fewer bacteria and impurities), but the centrifugation speed was increased to 13,000 rpm to improve washing efficiency.
[0103] Elution phase: Elution buffer volume 30 μl (normally 50 μl, for concentrated DNA), temperature 60℃, standing time 2 min.
[0104] The extraction module inputs the optimal operating parameters into the automated equipment for the following operations:
[0105] Lysis procedure: Add 15 μl of lysis buffer and 200 μl of GA buffer using a pipette, maintain the temperature at 56°C using a temperature control module, and mix for 10 min using an automatic inverter.
[0106] Real-time monitoring: After pyrolysis, the OD600 value is 0.1 (meets the standard), and the standard washing process begins; during the washing stage, the conductivity sensor detects a value below 10 μS / cm (normal), and the process proceeds according to preset parameters.
[0107] The model optimization iteration module obtained the following detection results: DNA concentration 100 ng / μl (approximately 80 ng / μl using traditional methods), A260 / A280 = 1.88 (purity meets the standard), clear electrophoresis bands, and an integrity score of 9.5. The optimized parameters (lysis time 13 min, centrifugation speed 15,000 rpm) and results of the high biomass bacterial sample were entered into the database to enhance the generalization ability of the bacterial sample model.
[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An automated genomic DNA extraction and purification system based on an algorithm model, characterized in that, include: The module includes a sample processing module, a parameter calculation module, an extraction module, and a model optimization and iteration module. The sample processing module is used to construct real sample characteristics by measuring the sample optical density value OD600 at a wavelength of 600 nm or the absorbance at a wavelength of 405 nm, and combining the sample type and reagent kit type identified by the barcode; and to acquire experimental data of historical DNA extraction and construct historical sample characteristics. The parameter calculation module is used to construct independent sub-models for the four operation stages of lysis, centrifugation, washing, and elution according to the gradient boosting tree algorithm. The sub-models for each operation stage output the optimal operation parameters for each operation stage of DNA extraction. The sub-models include training units and prediction units. The training unit is used to train the sub-models corresponding to each operation stage of DNA extraction using historical sample features as the training set, and obtain the trained sub-models. The mean values of the parameters corresponding to each operation stage of DNA extraction in the historical sample features are used as the initial predicted values. The initial predicted values are fitted to the true values by iteratively constructing a decision tree. A regularization strategy is introduced to limit the depth of the decision tree and set the minimum number of leaf node samples. The prediction unit is used to input real sample features into the trained sub-model to obtain the optimal operation parameters corresponding to each operation stage of DNA extraction. The extraction module is used to receive the optimal operating parameters corresponding to each stage of DNA extraction, and to extract DNA after setting the optimal operating parameters for the automated equipment. During the lysis stage, the lysis time and the amount of lysis buffer are dynamically adjusted based on the OD600 value. During the washing stage, the washing strategy is optimized based on the column membrane flow rate, and an abnormal termination mechanism is set. The adjusted parameters and abnormal data will be incorporated into the training data again to optimize the GBDT model and form a data-driven automated optimization closed loop. The model optimization and iteration module is used to detect the extracted DNA and generate detection results, construct an experimental report based on the optimal operating parameters of the current sub-model and the detection results, and use the experimental report as the training set of the training unit to optimize and iterate the sub-models corresponding to each operation stage of DNA extraction until the set conditions are met.
2. The automated genomic DNA extraction and purification system based on an algorithm model according to claim 1, characterized in that: The optimal operating parameters corresponding to each operating stage include: Pyrolysis stage: pyrolysis solution dosage, pyrolysis temperature, and pyrolysis time; Centrifugation stage: centrifugation speed and centrifugation time; Washing stage: amount of rinsing solution used, number of washes; Elution phase: eluent temperature and elution volume.
Citation Information
Patent Citations
Method and system for realizing DNA trace material evidence identification based on deep learning
CN118038997A
Technological process optimization system and method for intelligently extracting plant calcium
CN119225318A
Gene detection system for algae biological genes
CN119380815A
Automatic model adjusting and optimizing system based on AI training intelligent workbench
CN119398112A