Method and system for nucleic acid integrity determination
By employing a deep learning model-based method for target localization and quantitative assessment in gel mapping, the qualitative judgment problem of existing nucleic acid integrity determination is solved, enabling accurate and automated quantitative detection of nucleic acid integrity, which is suitable for longitudinal analysis and comparison of a large number of samples.
Patent Information
- Application Number
- CN202210316284.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2042-03-28
AI Technical Summary
Existing methods for determining nucleic acid integrity mainly rely on manual judgment, which has problems such as difficulty in qualitative judgment, high cost, low throughput, unsuitability for large-scale promotion, and difficulty in longitudinal analysis and comparison.
A method for localizing and quantitatively evaluating targets in film images based on a deep learning model is adopted. By training the film image samples and scoring their integrity, a film image recognition model is established, and deep learning networks such as ResNet34 are used for automated detection.
It enables accurate and automated quantitative detection of nucleic acid integrity, reduces detection costs, increases detection throughput, and is suitable for longitudinal analysis and comparison of a large number of samples.
Smart Images

Figure CN116883308B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of biotechnology, and more particularly provides a method and system for determining nucleic acid integrity. BACKGROUND
[0002] Integrity is an important indicator for determining the quality of nucleic acid (DNA, RNA) samples, and provides important standard support for verifying sample storage quality and the success of subsequent genetic sequencing. Currently, the industry's commonly used method for determining nucleic acid is the agarose gel electrophoresis method, that is, by "charge effect" and "molecular sieve effect", DNA of different fragment sizes is separated, and then skilled experimenters determine sample integrity according to the gel image taken by an imager.
[0003] At present, the sample integrity can only be qualitatively determined by viewing the gel image with the naked eye, that is, whether the sample is degraded can be determined, but the degradation degree cannot be quantitatively determined. This brings difficulties to the quality comparison between samples, such as: it is difficult to compare the integrity between samples of the same batch, and it is impossible to analyze the changes in the integrity of historical samples. Moreover, the brightness, exposure, contrast and other interference factors during the imaging of the imager also bring difficulties to the naked eye judgment, so the professional degree and proficiency of the experimenters are required to be high. At the same time, the existing integrity determination is qualitative determination, and the standard is relatively vague, but the sample gel image is complex, and the existing standard cannot guide all operations, and the determination results of the same gel image by experimenters with different proficiency may be different, so the determination results have to be rechecked at present, which wastes manpower.
[0004] In view of this, the Agilent 2100 system uses microfluidic technology to separate nucleic acid samples on the supporting chip of the system by capillary electrophoresis, and then uses fluorescence detection to quantify the integrity of nucleic acid. However, the system has the following disadvantages: 1) the learning model of the system is obtained by manual scoring of 1300 RNA samples by experts, the sample quantity is small, and the manual scoring itself is inaccurate; 2) the learning model of the system is obtained according to the size of the peaks of 18S and 28S ribosomes in the graph, which is not applicable to devices that only produce electrophoresis gel images and do not produce ribosome peak graphs; 3) the detection of nucleic acid integrity by the system is different from the traditional gel electrophoresis method, but the supporting chip is expensive, the cost of detecting a single sample is high, the number of samples detected at a time is small, and the throughput is low, which is not suitable for wide use; 4) most of the historical samples are detected by the gel electrophoresis method, and the difference in the used technology cannot be used for longitudinal analysis and comparison of the integrity of the historical samples.
[0005] Therefore, the existing technology needs a method and system for quantitatively detecting nucleic acid integrity with high throughput, low cost, accurate results and reduced manpower. SUMMARY
[0006] In order to perform nucleic acid integrity determination, the present application aims to provide a gel image target positioning and gel image quantitative evaluation method based on a deep learning model.
[0007] Therefore, in a first aspect, the present application provides a method for establishing a gel image recognition model, the method comprising:
[0008] 1) obtaining the degradation degree of each training set gel image sample and the main band position of the band region of each training set gel image sample based on the image of the training set gel image sample;
[0009] 2) sorting all training set gel image samples according to the degradation degree and the main band position, and scoring the integrity of the training set gel image samples according to the sorting;
[0010] 3) inputting the image and integrity score of the training set gel image sample as training data to train the model to obtain the trained model.
[0011] In one embodiment, in 1), the main band position of the band region of the training set gel image sample is determined by using a marker.
[0012] In one embodiment, in 1), the method comprises locating the sample from the electropherogram based on the position of the notch and the marker on the electropherogram to obtain the image of the training set gel image sample.
[0013] In one embodiment, in 1), the image of the training set gel image sample is a grayscale image.
[0014] In one embodiment, in 1), each training sample is divided into two intervals, and the degradation degree is represented by the relative brightness of the band region of the rear interval.
[0015] In one embodiment, the relative brightness of the band region of the rear interval is: (the total brightness of the band region of the rear interval - the total brightness of the band region of the front interval) / the total brightness of the band region of the front and rear intervals.
[0016] In one embodiment, the positions of the front and rear intervals are determined by using a marker.
[0017] In one embodiment, the front and rear intervals are bounded by 6557bp-9416bp, preferably 6557bp.
[0018] In one embodiment, the front and rear intervals are 6557bp-250bp and 23130bp-6557bp.
[0019] In one embodiment, in 2), the integrity score is a degradation level of 1-100.
[0020] In one embodiment, in 3), the model is a network for image classification, such as ResNet34, VGG16, VGG19, ResNet18, ResNet50, ResNet101 or DenseNet model.
[0021] In a second aspect, the present application provides a system for establishing a gel image recognition model, comprising:
[0022] An information obtaining module configured to obtain the degradation degree of each training set gel image sample and the main band position of the band region of each training set gel image sample based on the images of the training set gel image samples;
[0023] A scoring module configured to sort all training set gel image samples according to the degradation degree and the main band position, and score the integrity of the training set gel image samples according to the sorting;
[0024] A training module configured to input the images and integrity scores of the training set gel image samples as training data to train the model, and obtain the trained model.
[0025] In a third aspect, the present application provides a gel image recognition model established by the method of the first aspect or the system of the second aspect.
[0026] In a fourth aspect, the present application provides a system for gel image nucleic acid integrity detection by a gel image recognition model, comprising the gel image recognition model of the third aspect.
[0027] In a fifth aspect, the present application provides a method for gel image nucleic acid integrity detection by the gel image recognition model of the third aspect, comprising:
[0028] 4) inputting the image of the to-be-tested gel image sample into the trained model to obtain the integrity score of the to-be-tested gel image sample.
[0029] The method and system of the present application can more accurately and more automatically detect nucleic acid integrity based on gel images. BRIEF DESCRIPTION OF DRAWINGS
[0030] The present application is described below with reference to the following drawings:
[0031] Figure 1 is a gel image with marker (e.g. including markers M1 and M2) labeling sites and lane number annotations.
[0032] Figure 2 shows a schematic diagram of the annotation information of the gel image.
[0033] Figure 3 shows a detection model for preliminary positioning according to one embodiment.
[0034] Figure 4 A schematic diagram of initial positioning detection results is shown according to an embodiment.
[0035] Figure 5 A general flow of fine positioning implementation is shown according to an embodiment.
[0036] Figure 6 A ResNet34 network structure diagram is shown.
[0037] Figure 7 A quantitative evaluation effect is shown according to an example. DETAILED DESCRIPTION
[0038] In the present application, the marker is used to determine the main band position of the training set gel sample band region, and the position of the front and rear two intervals of the training set gel sample band region. In the gel sample, one or two markers can be included to determine the molecular weight of the position in the gel.
[0039] In the present application, for determining the gel sample band region on the gel map, the sample band can be manually determined, or an algorithm can be used for recognition, such as with the help of a computer program, and recognition by software. In a specific embodiment, mobilenetV3_YOLOv5 is used for gel sample initial positioning to identify the position of the notch and the marker in the image, which lays the foundation for the next precise positioning. The gel sample initial positioning algorithm can include two processes of training model and algorithm detection, and the trained model is used for detection. Before training the model, the sample needs to be labeled (label), that is, the category and position information of the target is written into the label file according to the YOLO label format, and then the label file and the image are input into the model mobilenetV3_YOLOv5 for iterative training. When the loss converges, the training is stopped, and the model file is output. The trained model can be used for algorithm detection. When performing algorithm detection, first input the image containing the sample, and perform Z-score normalization preprocessing on the image; then input into the trained mobilenetV3_YOLOv5 for target detection.
[0040] In the present application, any suitable method can be used to accurately position the sample strip area of the gel image. In one embodiment, the gray scale information of the sample can be statistically calculated using traditional image algorithms, and the sample can be accurately positioned in combination with the prior position of the marker. First, sample segmentation is required, and the image ROI operation is performed according to the position information of the target detected by the mobilenetV3_YOLOv5 in the previous step to segment each independent sample. Then, accurate positioning can be performed, and the gray scale and gradient information of the segmented sample is calculated to find the bright area scale of the marker (e.g., including markers M1 and M2) to accurately position the sample strip. In one example, the calculation process is as follows: convert each segmented M1 or M2 into a gray scale image, calculate the gray scale in the y-axis direction, and the gradient change information of the gray scale, then calculate each wave peak according to the gray scale curve, and combine the prior position information of the M1 and M2 bright area scale to calculate the position of the marker 23130, the marker 6557, and the marker 250 in the image. According to the marker M1 and M2 scale positioning, the strip area of all samples in the whole image can be accurately positioned, and after the region of interest of all samples is cropped, it is input to the next step for quantitative analysis.
[0041] In the present application, quantitative analysis can be performed as follows: after accurate positioning and cropping of the sample image, the sample image is input to the model ResNet34 for sample quantitative analysis. According to the quantitative requirements, 100 kinds of identification results can be output, with scores of 1.0-10.0 (accurate to one decimal place), and the higher the score, the lower the degree of degradation. The deep learning model needs to be pre-trained according to the existing samples, and the trained model is used for automatic evaluation of sample scores.
[0042] The model training stage includes: before model training, all samples need to be labeled (Label), and the degradation degree of each sample is labeled with 1.0-10.0 (accurate to one decimal place), and the higher the score, the lower the degree of degradation, and the better the sample quality. For the label making of the initial set of samples, at least 1000 (preferably 10000) samples of each type are guaranteed to ensure the credibility of the model trained using this initial sample for scoring unknown test samples, and to reduce the workload of reviewing the labels of unknown samples. Due to the lack of reference lines for the complexity of the gel image, manual visual labeling cannot exclude interference factors such as brightness, exposure, and contrast, and it is also impossible to accurately and meticulously divide the sample differences. In order to greatly improve the accuracy and avoid errors caused by manual labeling during model training, a digital method can be used to label the samples in combination with the brightness sum ratio of the strip area and the distribution of the strip area in the marker interval. For example, the specific steps are as follows:
[0043] 1) Calculate the brightness sum ratio of the strip region of each independent sample (denoted as score), which is calculated as follows: (brightness sum of the 23130bp-6557bp interval of the marker - brightness sum of the 6557bp-250bp interval of the marker) / brightness sum of the 23130bp-250bp interval. Through large sample statistical analysis, the success rate of DNA library construction of fragments with a size between 23130bp and 6557bp is > 99%, so this interval is used as the DNA fragment interval with better quality, and the 6557bp-250bp interval is used as the DNA fragment interval with poor quality. Then, the ratio of each sample is normalized to avoid interference caused by image factors such as exposure and contrast. The calculation process of normalization is as follows: divide the ratio of each sample in the same image by the average ratio of M1 and M2 in the same image, and the obtained value is used as the normalized score (denoted as score norm) of the sample.
[0044] 2) Calculate the distribution of the sample strip region in the marker interval. For each sample, convert it to a grayscale image. For a picture, take the lower left corner of the image as the origin to establish a plane rectangular coordinate system, and each point in the image can be determined by the plane rectangular coordinate system. To avoid infection caused by image noise, the x-axis 1 / 3, 1 / 2 and 2 / 3 positions are selected respectively, and the gray scale and gradient change information of the y-axis direction at the corresponding positions are calculated to obtain the gray scale curve. Then, according to the gray scale curve, each wave peak is calculated, and the x position information of the maximum wave peak of the gray scale value is extracted; then, the positions of the markers 23130 and 250 corresponding to M1 and M2 in the image are used as the starting and ending positions of the strip, and the relative position of the maximum wave peak position of the sample to the markers 23130 and 250 is calculated as the position score (score pos).
[0045] 3) Sort all samples, the larger the score norm, the better the sample quality, and when the score norm is the same, the higher the score pos score, the better the sample quality. According to the above standard, all samples are equally divided into 100 intervals, each interval corresponds to a label (label) from 1.0 to 10.0, representing 100 degradation levels, and then the label information of all samples is written into a label file. Finally, the label file and the image are input into the model ResNet34 for iterative training, and when the loss converges, the training is stopped, and the model file is output.
[0046] After the model is trained, the trained model can be used for quantitative evaluation. In an embodiment, the evaluation process is as follows: load the trained ResNet34 model, normalize the image and calculate the score, wherein the preprocessing process includes calculating the score and normalization, and the calculation method is referred to the training stage; perform forward inference on the model, output the score of each category to which the to-be-quantitatively-evaluated sample belongs, and then perform argmax operation to obtain the final label category to which it belongs.
[0047] For the to-be-evaluated sample, the brightness sum ratio of the strip region of each sample and the position score of the strip are calculated, and the comprehensive ranking is sorted, and finally divided into 100 intervals, corresponding to 100 degradation levels. Then, the deep learning technology is used to learn the rating experience, so as to solve the rating problem of the to-be-evaluated sample.
[0048] In the present application, the method for establishing a gel image recognition model can be realized by a system for establishing a gel image recognition model.
[0049] Embodiment:
[0050] Obtain nucleic acid gel electrophoresis gel image: select lambda-Hind as Marker 1, select D2000 as Marker 2, and select DNA samples produced by Shenzhen National Gene Library Sample Library as training samples; add the training samples and Marker 1 and Marker 2 to the prepared and formed agarose gel holes, respectively, run the gel using a DYY-6C electrophoresis instrument, and take a gel image using a BIO-RAD gel imaging system. The amount of training samples should not be less than 10,000. Before the image of the gel sample is digitally represented, the image is usually converted to a grayscale image.
[0051] Embodiment one, constructing an initial position detection model.
[0052] Use mobilenetV3_YOLOv5 to perform sample initial positioning, identify the positions of the black slot, M1 and M2 in the image, and lay the foundation for the next precise positioning. The algorithm is divided into two processes of model training and inference detection, and the trained model is used for detection.
[0053] 1. Constructing a training sample data set.
[0054] Before training, the gel image with marker (for example, including markers M1 and M2) marking position and lane number annotation needs to be annotated, specifically, the class and position information of the target are written into a label file, 80% of the samples in the training set are used to train the deep neural network, and 20% of the samples are used for verification to calculate the accuracy.
[0055] The gel image with marker (for example, including markers M1 and M2) marking position and lane number annotation and its corresponding annotation information are as followsFigure 1 and Figure 2 As shown in Table 1, where class 0 is black slot; 1 is marker M1; 2 is marker M2. The initial positioning detection model sample class number table is shown in Table 1.
[0056] Table 1
[0057] Sample Class Total Number of Samples (N) Number of Samples in Training Set (N*80%) Number of Samples in Test Set (N*20%) 0 (Black Slot) 100000 80000 20000 1(M1) 10000 8000 2000 2(M12) 10000 8000 2000
[0058] 2. Constructing the initial positioning detection model.
[0059] The initial positioning detection model Figure 3 As shown in Table 1, where class 0 is black slot; 1 is marker M1; 2 is marker M2. The initial positioning detection model sample class number table is shown in Table 1.
[0060] 3. Training the initial positioning detection model.
[0061] Loss function setting: total loss function Loss = GIOU Loss +Loss conf +Loss class , wherein GIOU Loss is the Bounding Box loss function, Loss conf is the confidence loss function, and Loss class is the classification loss function.
[0062] Bounding Box loss function: GIOU Loss = 1 - GIOU = 1 - (IOU - |Q| / C), wherein C represents the smallest bounding rectangle between the detection frame and the previous frame, and Q represents the difference between the joint of the smallest bounding rectangle and the two bounding boxes.
[0063] Confidence loss function:
[0064]
[0065] Classification loss function:
[0066] wherein, represents that there is a target in the i-th grid of the j detection frames, represents that there is no target in the i-th grid of the j detection frames; λ noobj represents the loss weight of the positioning error; and is the training value, and is the predicted value.
[0067] Model training hyperparameter settings: set the maximum iteration epochs of the training initial positioning detection model to 300,000, and the sample batch size mini_batch to 8. The optimization algorithm of gradient descent uses the Adam algorithm, the initial learning rate (LearningRate) is set to 0.01, the learning rate is changed according to the cosine annealing function, the cosine annealing hyperparameter is set to 0.2, the learning rate momentum is set to 0.937, the weight decay coefficient is 0.0005, the preheating learning epoch is 3.0, the preheating learning rate momentum is 0.8, the preheating learning rate is 0.1, the coefficient of giou loss is 0.05, the coefficient of classification loss is 0.5, the weight of positive samples in classification BCELoss is 1, and the coefficient of object loss is 1.
[0068] Model training iteration stopping index requirements: if the model training times reach the set maximum round requirement, the model stops training, and if the model loss does not decrease for 1000 consecutive epochs, the model stops training.
[0069] The label file and the image are input into the model for iterative training, and when the training indicators meet the requirements, the training is stopped and the model file is output.
[0070] 4. Initial positioning detection model inference prediction.
[0071] The trained initial positioning detection model is used for detection, and the model inference prediction process includes: inputting an image containing samples, preprocessing the image, and then inputting it into the trained mobilenetV3_YOLOv5 to obtain black slots, M1 and M2, and their position information through inference. The preprocessing method is: calculate the overall gray average μ and standard deviation σ of the image to be processed, and the gray value X of each point on the normalized image is X=(x-μ) / σ, x is the original gray value of each point on the image to be processed, and σ represents the gray standard deviation of the image to be processed. The prediction output of the detection is as shown in Figure 4 , and the evaluation indicators of the model meet the use requirements, wherein precision: 0.99, MAP@5: 0.995.
[0072] Example two, accurate positioning example.
[0073] The overall process of fine positioning implementation is as shown in Figure 5 A-D, which can include the following steps:
[0074] 1) Sample segmentation according to the position information of the target detected by Mobilenetv3-yolo-v5, as shown in Figure 5 A is the segmented M1;
[0075] 2) In the segmented sample, the gray scale and gradient information are counted, as shown in Figure 5 B, the corresponding peak points are calculated, and the key positions are selected by geometric position;
[0076] 3) Find out the bright area scale of M1 and M2 to realize accurate positioning of the region of interest, as shown in Figure 5 C is the position of the positioning scale 23130.
[0077] 4) According to the scale positioning, the region of interest of all samples in the whole image can be accurately positioned, and the effect is as follows Figure 5 D shows that after cutting the region of interest of all samples, it is input to the next step for quantitative analysis.
[0078] Example three: quantitative analysis example.
[0079] After accurate positioning, the cut image is input to the deep learning model for sample quantitative analysis. The model can output 10 kinds of results according to the quantitative requirements, and the score is 1-10 points. The higher the score, the lower the degradation degree. The deep learning model needs to be pre-trained according to the existing samples. The trained model is used to automatically evaluate the sample score.
[0080] 1. Construct a classification sample data set.
[0081] The sample data set of the lane image is annotated and processed as follows: lane region of interest brightness value statistics; lane region of interest brightness value ratio; normalization; sample sorting and partitioning.
[0082] The specific description is as follows:
[0083] 1) Distribute the total brightness of the Marker 23130bp-6557bp interval, the total brightness of the Marker 6557bp-250 interval, and the total brightness of the 23130bp-250bp interval;
[0084] 2) Calculate the brightness sum ratio (denoted as score) of the region of interest of each lane image sample, and the calculation method is: (Marker 23130bp-6557bp interval brightness sum-Marker 6557bp-250 interval brightness sum) / 23130bp-250bp interval brightness sum;
[0085] 3) Normalization of score for each sample; the calculation process of normalization is: the score of each sample in the same image is divided by the average score of M1 and M2 in the same image, and the obtained value is taken as the normalized score (denoted as score norm) of the sample;
[0086] 4) Then sort all sample score norms, divide all score norms into 10 intervals equally, each interval corresponds to label 1-10, and then write all sample label information to the label file.
[0087] The samples used in each category are shown in Table 2.
[0088] Table 2
[0089] Sample Class Total Number of Samples (N) Number of Samples in Training Set (N*80%) Number of Samples in Test Set (N*20%) 1 10000 8000 2000 2 10000 8000 2000 3 10000 8000 2000 4 10000 8000 2000 5 10000 8000 2000 6 10000 8000 2000 7 10000 8000 2000 8 10000 8000 2000 9 10000 8000 2000 10 10000 8000 2000
[0090] 2. Construct a classification model.
[0091] The classification model ResNet34 network structure is shown in Figure 6 . Wherein, each green rectangular box represents a Stage of ResNet, from top to bottom, Stage1 (Conv2_x), Stage2 (Conv3_x), Stage3 (Conv4_x), Stage4 (Conv5_x). Each yellow rectangular box in the green rectangular box represents one or more standard residual units, and the number on the left side of the yellow rectangular box represents the number of cascaded residual units, such as 5x representing 5 cascaded residual units. In the classification model:
[0092] Channel number change: the input channel is 3, and the channel numbers of the 4 Stages are 64, 128, 256, and 512 in turn, that is, the channel number is doubled after passing through a Stage;
[0093] Layer calculation: the number of residual units contained in each Stage is 3, 4, 6, and 3 in turn, each residual unit contains 2 convolution layers, and the total number of layers = (3+4+6+3)*2+1+1 = 34, plus the first 7x7 convolution layer and the 3x3 max pooling layer;
[0094] Downsampling: the purple part in the yellow rectangular box represents the downsampling operation, that is, the feature map size is halved, and the right arrow mark is the feature map size after downsampling (input 224x224 is taken as an example); the orange rectangular box in the first green rectangular box represents max pooling, and the first downsampling occurs here;
[0095] Convolution layer parameter interpretation: take Conv 3x3, c512, s2, p1 as an example, 3x3 represents the size of the convolution kernel, c512 represents the number of convolution kernels / output channels is 512, s2 represents the convolution step is 2, p1 represents the padding of the convolution is 1;
[0096] Pooling layer parameter interpretation: Max_pool 3x3, c64, s2, p1, 3x3 represents the size of the pooling region (similar to the size of the convolution kernel), c64 represents the input and output channels are 64, s2 represents the pooling step is 2, p1 represents the padding is 1.
[0097] The downsampling operation in ResNet occurs in the first residual unit or the maximum pooling layer of each stage, and the implementation is that the operation is performed by taking the step size as 2 in convolution or pooling.
[0098] 3. Classification model training.
[0099] The loss function loss is a multi-class cross-entropy loss function, and the number of categories is set to 10 here;
[0100] Model training hyperparameter setting: the maximum number of iterations of the model epochs = 120, the sample batch size mini_batch = 256, the initial learning rate is 0.1, the learning rate is reduced to 0.1 every 30 epochs, the learning rate momentum is set to 0.9, the weight decay coefficient is set to 0.0001, and the RMSprop training optimization method is used;
[0101] Model training iteration stopping index requirement: when the number of model training iterations reaches the set maximum number of rounds, or the loss on the validation set does not decrease for 30 rounds.
[0102] The label file and the image are input into the model for iterative training, and when the training index meets the requirement, the training is stopped, and the model file is output.
[0103] 4. Classification model inference prediction.
[0104] The trained model is used for quantitative evaluation, and the evaluation process includes preprocessing the sample region of interest, then inputting the deep learning model, and obtaining the evaluation result by inference. The preprocessing process includes calculating the score and normalization processing, and the evaluation result is the sample integrity score. The evaluation effect is as follows Figure 7 The numbers on each sample lane are the integrity scores of the samples.
[0105] The application quantifies the degradation degree of samples, outputs the integrity score, facilitates the comparison and analysis between samples, and provides more data support for sample science. The application directly uses the gel map output by the imager, forms a deep learning model, and can be applied to all gel electrophoresis methods and maintain the accuracy of the learning model. The method of the application sorts enough sample integrity detection results and divides them into 100 intervals of 1.0-10.0 scores, learns and corrects according to a large amount of data, changes the previous fuzzy judgment standard and the inaccuracy of manual judgment, and outputs accurate results, which can further guide the downstream operation and reduce the failure rate of library construction. The application can be embodied as an independent analysis system, which is connected to existing gel electrophoresis detection equipment of various brands and models, can further quantitatively analyze the gel map of historical samples, and quantitatively evaluates the output gel map, maintaining the high-throughput and low-cost characteristics of the existing equipment and methods.
Claims
1. A method for establishing a gel image recognition model, the method comprising: 1) obtaining degradation degree of each training set gel image sample and main band position of each training set gel image sample band region based on image of the training set gel image sample, wherein the main band position is determined by converting the image of the training set gel image sample into a gray scale image, and determining the band position at the maximum peak of the gray scale value of the gray scale image as the main band position; 2) sorting all training set gel image samples according to degradation degree and main band position, and scoring the training set gel image samples according to the sorting based on integrity, wherein the lower the degradation degree, the better the sample quality, and when the degradation degree is the same, the greater the main band position, the better the sample quality; 3) inputting the image of the training set gel image sample and the integrity score as training data to train the model to obtain the trained model.
2. The method of claim 1, in 1), each training sample is divided into two intervals, and the degradation degree is represented as the relative brightness of the band region of the rear interval.
3. The method of claim 2, the relative brightness of the band region of the rear interval is: (total brightness of the band region of the rear interval - total brightness of the band region of the front interval) / total brightness of the band region of the front and rear intervals.
4. The method of claim 2, the positions of the front and rear intervals are determined by markers.
5. The method of claim 2, the front and rear intervals are bounded by 6557bp-9416bp and 6557bp.
6. The method of claim 2, the front and rear intervals are 6557bp-250bp and 23130bp-6557bp.
7. The method of any one of claims 1-6, in 1), the main band position of the band region of the training set gel image sample is determined by markers.
8. The method of any one of claims 1-6, in 1), the method comprises locating the sample on the electropherogram based on the positions of the notches and markers on the electropherogram to obtain the image of the training set gel image sample.
9. The method of any one of claims 1-6, in 1), the image of the training set gel image sample is a gray scale image.
10. The method of any one of claims 1-6, in 2), the integrity score is a degradation level of 1-100.
11. The method of any one of claims 1-6, in 3), the model is a network for image classification.
12. The method of claim 11, the model is selected from any one of the following models: ResNet34, VGG16, VGG19, ResNet18, ResNet50, ResNet101, and DenseNet model.
13. A system for establishing a gel image recognition model, the system comprising: an information acquisition module configured to obtain degradation degree of each training set gel image sample and main band position of each training set gel image sample band region based on image of the training set gel image sample, wherein the main band position is determined by converting the image of the training set gel sample into a gray scale image, and determining the band position at the maximum wave crest of the gray scale value of the gray scale image as the main band position; a scoring module configured to sort all training set gel samples according to degradation degree and main band position, and score the training set gel samples according to integrity according to the sorting, wherein the sorting of the samples comprises: the lower the degradation degree, the better the sample quality, and when the degradation degree is the same, the greater the main band position, the better the sample quality; a training module configured to input the image and integrity score of the training set gel sample as training data to train the model, and obtain the trained model.
14. A system for detecting nucleic acid integrity of a gel image using a gel image recognition model, the system comprising a gel image recognition model established according to the method of any one of claims 1-12 or the system of claim 13.
15. A method for detecting nucleic acid integrity of a gel image using a gel image recognition model established according to the method of any one of claims 1-12 or the system of claim 13, the method comprising: inputting the image of a to-be-tested gel sample into the trained model to obtain the integrity score of the to-be-tested gel sample.
Citation Information
Patent Citations
Sequencing of nucleic acids by emergence
CN111566211A
Method for detecting integrity of genome DNA based on multiplex PCR
CN113130007A