Crop variety analysis method, device and equipment based on simple repetitive sequence
By extracting features from crop fluorescence signal images and inputting them into a pre-trained random forest classifier, the accuracy and reliability issues of crop variety analysis in existing technologies are solved, achieving efficient and accurate variety analysis.
Patent Information
- Application Number
- CN202511780081.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-02-06
AI Technical Summary
In the existing technology, crop variety analysis methods based on simple repeating sequences lack efficient and accurate image recognition and data processing methods, which affects the accuracy and reliability of variety analysis.
By extracting features from the target fluorescence signal image and inputting it into a pre-trained random forest classifier, accurate, efficient, and reliable analysis of crop varieties can be achieved.
It enables efficient, accurate, and reliable analysis of crop varieties, improving the accuracy and reliability of variety analysis.
Smart Images

Figure CN121472386A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of crop analysis technology, and in particular to a method, apparatus and equipment for crop variety analysis based on simple repeating sequences. Background Technology
[0002] Simple Sequence Repeats (SSRs) are molecular markers widely found in eukaryotic genomes. They have advantages such as high polymorphism, codominant inheritance, and good reproducibility, and have important application value in variety identification, genetic diversity analysis, and gene mapping.
[0003] Currently, SSR-based variety analysis methods mainly include SSR site screening, primer design, polymerase chain reaction (PCR) amplification, capillary electrophoresis, and data analysis.
[0004] The key aspect lies in the automatic collection and comparison analysis of SSR genotypes, especially the analysis of capillary electrophoresis images. The lack of efficient and accurate image recognition and data processing methods affects the accuracy and reliability of variety analysis. Summary of the Invention
[0005] This invention provides a method, apparatus, and device for crop variety analysis based on simple repeating sequences, which addresses the shortcomings of existing technologies that lack efficient and accurate image recognition and data processing methods, thus affecting the accuracy and reliability of variety analysis. The technical solution of this invention extracts features from the target fluorescence signal image and inputs the target image features into a pre-trained random forest classifier, thereby enabling accurate, efficient, and reliable analysis of crop varieties.
[0006] This invention provides a method for analyzing crop varieties based on simple repeating sequences, comprising the following steps.
[0007] The target primer set was determined based on the simple repetitive sequence sites pre-screened for the target crop; The target primer set was amplified by polymerase chain reaction to obtain the amplification product corresponding to the target primer set; The amplification products were separated by electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products; Extract the target image features from the target fluorescence signal image; The target image features are input into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier; the random forest classifier is trained based on the training image features and the corresponding variety identifiers.
[0008] According to a method for analyzing crop varieties based on simple repetitive sequences provided by the present invention, the step of performing polymerase chain reaction amplification on the target primer set to obtain the amplification product corresponding to the target primer set includes: Determine the polymerase chain reaction curves corresponding to the target primer set under each preset parameter condition; The preset parameter conditions corresponding to the polymerase chain reaction curve with the highest similarity to the standard polymerase chain reaction curve among all the polymerase chain reaction curves are determined as the optimal conditions; The target primer set was amplified by polymerase chain reaction under optimal conditions to obtain the amplification product corresponding to the target primer set.
[0009] According to the present invention, a method for analyzing crop varieties based on simple repeating sequences is provided, wherein the preset parameters include: magnesium ion concentration, primer concentration, deoxyribonucleoside triphosphate concentration, aquatic thermophilic bacteria enzyme activity, and annealing temperature.
[0010] According to a method for analyzing crop varieties based on simple repeating sequences provided by the present invention, the step of electrophoretically separating the amplified products to obtain a target fluorescence signal image corresponding to the amplified products includes: The amplification products were separated by electrophoresis using primer fluorescent labeling dyes and a capillary electrophoresis apparatus to obtain the target fluorescence signal image corresponding to the amplification products.
[0011] According to the present invention, a method for analyzing crop varieties based on simple repeating sequences includes the electrophoretic separation of the amplification products using primer fluorescent labeling dyes and capillary electrophoresis to obtain target fluorescence signal images corresponding to the amplification products, comprising: The amplification products are separated by electrophoresis using the primer fluorescent labeling dye and the capillary electrophoresis apparatus to obtain the initial fluorescence signal image corresponding to the amplification products. The initial fluorescence signal image is calibrated by segment size to obtain the intermediate fluorescence signal image corresponding to the initial fluorescence signal image; The intermediate fluorescence signal image is converted to grayscale to obtain the grayscale fluorescence signal image of the intermediate fluorescence signal image; The grayscale fluorescence signal image is subjected to median filtering to obtain the target fluorescence signal image corresponding to the grayscale fluorescence signal image.
[0012] According to the present invention, a method for analyzing crop varieties based on simple repeating sequences is provided. The target image features include: peak area of bands, mobility, and information entropy. The step of inputting the target image features into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier includes: The peak area of the strip, the mobility, and the information entropy are respectively input into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier.
[0013] The present invention also provides a crop variety analysis device based on simple repeating sequences, comprising the following modules: The determination module is used to determine the target primer set based on the simple repetitive sequence sites pre-screened for the target crop; An amplification module is used to perform polymerase chain reaction amplification on the target primer set to obtain the amplification product corresponding to the target primer set; The image module is used to perform electrophoretic separation of the amplification products to obtain the target fluorescence signal image corresponding to the amplification products; The feature extraction module is used to extract target image features from the target fluorescence signal image; The analysis module is used to input the target image features into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier; the random forest classifier is trained based on the training image features and the corresponding variety identifiers.
[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the crop variety analysis method based on simple repeating sequences as described above.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the crop variety analysis method based on simple repeating sequences as described above.
[0016] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the crop variety analysis method based on simple repeating sequences as described above.
[0017] This invention provides a method, apparatus, and equipment for crop variety analysis based on simple repeating sequences. The method involves determining a target primer set based on pre-screened simple repeating sequence sites in the target crop; performing polymerase chain reaction (PCR) amplification on the target primer set to obtain the corresponding amplification products; separating the amplification products by electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products; extracting target image features from the target fluorescence signal image; and inputting the target image features into a pre-trained random forest classifier to obtain the crop variety analysis results output by the random forest classifier. The random forest classifier is trained based on training image features and corresponding variety identifiers. This invention, by extracting features from the target fluorescence signal image and inputting these features into a pre-trained random forest classifier, can accurately, efficiently, and reliably analyze crop varieties. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating the crop variety analysis method based on simple repeating sequences provided by the present invention.
[0020] Figure 2 This is a schematic diagram of the structure of the crop variety analysis device based on simple repeating sequences provided by the present invention.
[0021] Figure 3 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0023] To address the aforementioned problems in the prior art, this invention provides a method for analyzing crop varieties based on simple repeating sequences. Figure 1 This is a flowchart illustrating the crop variety analysis method based on simple repeating sequences provided by the present invention, as shown below. Figure 1 As shown, the method includes the following steps 110 to 150.
[0024] Step 110: Determine the target primer set based on the simple repetitive sequence sites pre-screened for the target crop.
[0025] Specifically, SSR loci of target crops can be screened using bioinformatics analysis tools. Screening criteria can include repeat unit length, repeat count, and distribution uniformity. Then, a target primer set can be determined based on Primer 3 or BatchPrimer 3 and the SSR loci, where each primer includes a target simple repetitive sequence locus. For example, the target primer set can cover 10 chromosomes and includes multiple primers. The spacing between primers can be set to less than or equal to 5 cm to ensure uniform locus distribution. Primer specificity is assessed using Primer-BLAST alignment to avoid non-specific amplification. Simultaneously, the influence of secondary structures (such as dimers, hairpin structures, etc.) is calculated using OligoAnalyzer. Primer specificity assessment refers to evaluating whether there are interactions between primers; if no interactions exist, they can be used.
[0026] Step 120: Perform polymerase chain reaction amplification on the target primer set to obtain the amplification product corresponding to the target primer set.
[0027] Specifically, multiplex PCR (4-plex) technology can be used to amplify the target primer set. The length of the amplified fragment can be set to 100-400 bp (bp represents the pairing unit of double-stranded nucleic acid) to ensure primer specificity and amplification efficiency. After amplification, the amplified product corresponding to the target primer set can be obtained.
[0028] In one embodiment, the step of performing polymerase chain reaction amplification on the target primer set to obtain the amplification product corresponding to the target primer set includes: Determine the polymerase chain reaction curves corresponding to the target primer set under each preset parameter condition; The preset parameter conditions corresponding to the polymerase chain reaction curve with the highest similarity to the standard polymerase chain reaction curve among all the polymerase chain reaction curves are determined as the optimal conditions; The target primer set was amplified by polymerase chain reaction under optimal conditions to obtain the amplification product corresponding to the target primer set.
[0029] Specifically, multiple preset parameter conditions can be set in advance, thereby determining the PCR curves corresponding to the target primer set under each preset parameter condition. Furthermore, all PCR curves can be compared with a standard PCR curve to determine the similarity between each PCR curve and the standard PCR curve. The preset parameter conditions corresponding to the PCR curve with the highest similarity are determined as the optimal conditions. Then, based on the optimal conditions, PCR amplification of the target primer set is performed to obtain the amplification product.
[0030] In the above embodiments, the target primer set is amplified by PCR. PCR amplification can amplify trace amounts to visible amounts, making subsequent analysis of crop varieties more accurate.
[0031] In one embodiment, the preset parameters include: magnesium ion concentration, primer concentration, deoxyribonucleoside triphosphate concentration, aquatic thermophilic bacteria enzyme activity, and annealing temperature.
[0032] Specifically, the preset parameters can be within the following ranges: magnesium ion concentration of 2.0-3.0 mM (millomoles per liter), primer concentration of 0.1-0.5 μM (micromoles per liter), deoxyribonucleoside triphosphate (dNTP) concentration of 200 μM, Thermus aquaticus (Taq) activity of 1 international unit per microliter (U / 20 μL), and annealing temperature of 55 to 65 degrees Celsius.
[0033] In the above embodiments, multiple preset parameter conditions are set, each of which includes multiple parameters, so that the optimal conditions for PCR amplification can be found among the above conditions, which can improve the reliability of subsequent variety analysis.
[0034] Step 130: Perform electrophoretic separation on the amplification products to obtain the target fluorescence signal image corresponding to the amplification products.
[0035] Optionally, before electrophoretic separation of the amplification products, unbound dNTPs and primers can be removed by ExoSAP-IT enzyme treatment, and the amplification products can be purified using the QIAquick PCR Purification Kit to improve detection sensitivity. The purified amplification products can then be separated by electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products.
[0036] In one embodiment, the step of electrophoretically separating the amplification products to obtain a target fluorescence signal image corresponding to the amplification products includes: The amplification products were separated by electrophoresis using primer fluorescent labeling dyes and a capillary electrophoresis apparatus to obtain the target fluorescence signal image corresponding to the amplification products.
[0037] Specifically, the amplification products can be separated by electrophoresis using primer fluorescent labeling dyes and capillary electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products.
[0038] For example, the capillary can be filled with POP-7 polymer, and the electrophoresis parameters can be set to a voltage of 15 kV and a runtime of 1800 seconds to optimize fragment migration resolution. The resolution of the target fluorescence signal image can be 300 dpi, and it can be stored in TIFF format to ensure high fidelity of the original data. The primer fluorescent labeling dyes can be FAM, VIC, NED, or PET dyes, and the capillary electrophoresis apparatus can be an ABI 3730 series.
[0039] In the above embodiments, the amplification products are separated by electrophoresis based on primer fluorescent labeling dyes and capillary electrophoresis. The resulting target fluorescence signal image has high single-molecule sensitivity and ultra-high spatial resolution. The fluorescence signal image can capture "weak, dynamic, multi-color, three-dimensional, and living" biological information into high signal-to-noise ratio, quantitative, and traceable image data in one go.
[0040] In one embodiment, the electrophoretic separation of the amplification products based on primer fluorescent labeling dyes and capillary electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products includes: The amplification products are separated by electrophoresis using the primer fluorescent labeling dye and the capillary electrophoresis apparatus to obtain the initial fluorescence signal image corresponding to the amplification products. The initial fluorescence signal image is calibrated by segment size to obtain the intermediate fluorescence signal image corresponding to the initial fluorescence signal image; The intermediate fluorescence signal image is converted to grayscale to obtain the grayscale fluorescence signal image of the intermediate fluorescence signal image; The grayscale fluorescence signal image is subjected to median filtering to obtain the target fluorescence signal image corresponding to the grayscale fluorescence signal image.
[0041] Specifically, the amplification products can be separated by electrophoresis based on primer fluorescent labeling dyes and capillary electrophoresis to obtain the initial fluorescence signal image corresponding to the amplification product. Then, the fragment size of the initial fluorescence signal image can be calibrated to obtain the intermediate fluorescence signal image corresponding to the initial fluorescence signal image. For example, the LIZ-500 molecular weight internal standard can be used for fragment size calibration to correct the error caused by electrophoretic fluctuations.
[0042] Optionally, before performing grayscale processing on the intermediate fluorescence signal image, GeneMapper 5.0 software can be used to identify the peak value of the intermediate fluorescence signal image, retain the effective signal with a peak value ≥ 150 RFU, and then perform grayscale processing.
[0043] Furthermore, grayscale processing can convert the RGB three-channel image into an 8-bit grayscale image to simplify calculations, thereby obtaining a grayscale fluorescence signal image. Then, an adaptive median filtering algorithm can be used to perform median filtering on the grayscale fluorescence signal image to eliminate salt-and-pepper noise and preserve stripe edge details. The filtering window size for median filtering can be, for example, 5x5.
[0044] Optionally, the positions of the bands of different samples in the target fluorescence signal image can be aligned using the Dynamic Time Warping (DTW) algorithm to correct the mobility deviation caused by fluctuations in electrophoresis conditions.
[0045] In the above embodiments, multiple image processing steps are performed on the initial fluorescence signal image to facilitate the extraction of subsequent image features.
[0046] Step 140: Extract the target image features from the target fluorescence signal image.
[0047] Specifically, target image features can be extracted from the target fluorescence signal image.
[0048] Step 150: Input the target image features into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier; the random forest classifier is trained based on the training image features and the corresponding variety identifier.
[0049] Specifically, a random forest classifier can be pre-built. For example, the number of decision trees can be set to 500, with a maximum depth of 10 layers per tree. Input features include peak area of bands, mobility, and information entropy. The random forest classifier can then be trained based on training image features and corresponding variety identifiers. The training image features include peak area of bands, mobility, and information entropy. Optionally, 5-fold cross-validation can be used to optimize the parameters of the random forest classifier (such as feature subset ratio and node splitting criterion), with the classification threshold set to a confidence level ≥ 0.95. The model performance can be evaluated using a confusion matrix and ROC curve (AUC = 0.993) to achieve automated classification and identification of genetic differences among varieties.
[0050] Furthermore, the target image features can be input into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier.
[0051] In one embodiment, the target image features include: peak area of the bands, mobility, and information entropy; the step of inputting the target image features into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier includes: The peak area of the strip, the mobility, and the information entropy are respectively input into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier.
[0052] Specifically, the target image features include: peak area of the bands, mobility, and information entropy. The peak area of the bands is calculated by integrating the area of the fluorescence intensity curve, reflecting the abundance of the amplified product. Mobility measures fragment size. Information entropy uses Shannon entropy to measure the complexity of the fluorescence signal to distinguish SSR patterns of different varieties. By inputting the peak area of the bands, mobility, and information entropy into a pre-trained random forest classifier, the variety analysis results of the target crop output by the random forest classifier can be obtained.
[0053] In the above embodiments, a random forest classifier was used to achieve computerized identification of differences in crop varieties, providing an efficient and reliable analytical method for SSR marker-based variety and seed identification and molecular-assisted breeding.
[0054] This invention provides a method for crop variety analysis based on simple repeating sequences. The method involves determining a target primer set based on pre-screened simple repeating sequence sites in the target crop; performing polymerase chain reaction (PCR) amplification on the target primer set to obtain the corresponding amplification products; separating the amplification products by electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products; extracting target image features from the target fluorescence signal image; and inputting the target image features into a pre-trained random forest classifier to obtain the crop variety analysis results output by the random forest classifier. The random forest classifier is trained based on the training image features and corresponding variety identifiers. This invention, by extracting features from the target fluorescence signal image and inputting these features into a pre-trained random forest classifier, can accurately, efficiently, and reliably analyze crop varieties.
[0055] The following describes the crop variety analysis device based on simple repeating sequences provided by the present invention. The crop variety analysis device based on simple repeating sequences described below and the crop variety analysis method based on simple repeating sequences described above can be referred to and correspond to each other.
[0056] Figure 2 This is a schematic diagram of the crop variety analysis device based on simple repeating sequences provided by the present invention, as shown below. Figure 2 As shown, the crop variety analysis device 200 based on simple repeating sequences includes the following modules: Module 210 is used to determine the target primer set based on the simple repetitive sequence sites pre-screened for the target crop; The amplification module 220 is used to perform polymerase chain reaction amplification on the target primer set to obtain the amplification product corresponding to the target primer set; Image module 230 is used to perform electrophoretic separation on the amplification products to obtain the target fluorescence signal image corresponding to the amplification products; Feature extraction module 240 is used to extract target image features from the target fluorescence signal image; The analysis module 250 is used to input the target image features into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier; the random forest classifier is trained based on the training image features and the corresponding variety identifiers.
[0057] In one embodiment, the amplification module 220 is specifically used for: Determine the polymerase chain reaction curves corresponding to the target primer set under each preset parameter condition; The preset parameter conditions corresponding to the polymerase chain reaction curve with the highest similarity to the standard polymerase chain reaction curve among all the polymerase chain reaction curves are determined as the optimal conditions; The target primer set was amplified by polymerase chain reaction under optimal conditions to obtain the amplification product corresponding to the target primer set.
[0058] In one embodiment, the preset parameters include: magnesium ion concentration, primer concentration, deoxyribonucleoside triphosphate concentration, aquatic thermophilic bacteria enzyme activity, and annealing temperature.
[0059] In one embodiment, the image module 230 is specifically used for: The amplification products were separated by electrophoresis using primer fluorescent labeling dyes and a capillary electrophoresis apparatus to obtain the target fluorescence signal image corresponding to the amplification products.
[0060] In one embodiment, the image module 230 is further configured to: The amplification products are separated by electrophoresis using the primer fluorescent labeling dye and the capillary electrophoresis apparatus to obtain the initial fluorescence signal image corresponding to the amplification products. The initial fluorescence signal image is calibrated by segment size to obtain the intermediate fluorescence signal image corresponding to the initial fluorescence signal image; The intermediate fluorescence signal image is converted to grayscale to obtain the grayscale fluorescence signal image of the intermediate fluorescence signal image; The grayscale fluorescence signal image is subjected to median filtering to obtain the target fluorescence signal image corresponding to the grayscale fluorescence signal image.
[0061] In one embodiment, the target image features include: peak area of the bands, mobility, and information entropy; the analysis module 250 is specifically used for: The peak area of the strip, the mobility, and the information entropy are respectively input into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier.
[0062] This invention provides a crop variety analysis device based on simple repeating sequences. The device determines a target primer set based on pre-screened simple repeating sequence sites of the target crop; it performs polymerase chain reaction (PCR) amplification on the target primer set to obtain the corresponding amplification products; it separates the amplification products by electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products; it extracts the target image features from the target fluorescence signal image; and it inputs the target image features into a pre-trained random forest classifier to obtain the crop variety analysis results output by the random forest classifier. The random forest classifier is trained based on the training image features and corresponding variety identifiers. This invention, by extracting features from the target fluorescence signal image and inputting these features into a pre-trained random forest classifier, can accurately, efficiently, and reliably analyze crop varieties.
[0063] Figure 3 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 3 As shown, the electronic device may include: a processor 310, a communications interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communications interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 can call logical instructions in the memory 330 to execute a crop variety analysis method based on simple repeating sequences, the method including: The target primer set was determined based on the simple repetitive sequence sites pre-screened for the target crop; The target primer set was amplified by polymerase chain reaction to obtain the amplification product corresponding to the target primer set; The amplification products were separated by electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products; Extract the target image features from the target fluorescence signal image; The target image features are input into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier; the random forest classifier is trained based on the training image features and the corresponding variety identifiers.
[0064] Furthermore, the logical instructions in the aforementioned memory 330 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0065] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the crop variety analysis method based on simple repeating sequences provided by the above methods, the method comprising: The target primer set was determined based on the simple repetitive sequence sites pre-screened for the target crop; The target primer set was amplified by polymerase chain reaction to obtain the amplification product corresponding to the target primer set; The amplification products were separated by electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products; Extract the target image features from the target fluorescence signal image; The target image features are input into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier; the random forest classifier is trained based on the training image features and the corresponding variety identifiers.
[0066] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the crop variety analysis method based on simple repeating sequences provided by the above methods, the method comprising: The target primer set was determined based on the simple repetitive sequence sites pre-screened for the target crop; The target primer set was amplified by polymerase chain reaction to obtain the amplification product corresponding to the target primer set; The amplification products were separated by electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products; Extract the target image features from the target fluorescence signal image; The target image features are input into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier; the random forest classifier is trained based on the training image features and the corresponding variety identifiers.
[0067] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0068] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for analyzing crop varieties based on simple repeating sequences, characterized in that, include: The target primer set was determined based on the simple repetitive sequence sites pre-screened for the target crop; The target primer set was amplified by polymerase chain reaction to obtain the amplification product corresponding to the target primer set; The amplification products were separated by electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products; Extract the target image features from the target fluorescence signal image; The target image features are input into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier; the random forest classifier is trained based on the training image features and the corresponding variety identifiers.
2. The method for analyzing crop varieties based on simple repeating sequences according to claim 1, characterized in that, The polymerase chain reaction amplification of the target primer set to obtain the amplification product corresponding to the target primer set includes: Determine the polymerase chain reaction curves corresponding to the target primer set under each preset parameter condition; The preset parameter conditions corresponding to the polymerase chain reaction curve with the highest similarity to the standard polymerase chain reaction curve among all the polymerase chain reaction curves are determined as the optimal conditions; The target primer set was amplified by polymerase chain reaction under optimal conditions to obtain the amplification product corresponding to the target primer set.
3. The method for analyzing crop varieties based on simple repeating sequences according to claim 2, characterized in that, The preset parameters include: magnesium ion concentration, primer concentration, deoxyribonucleoside triphosphate concentration, aquatic thermophilic bacteria enzyme activity, and annealing temperature.
4. The method for analyzing crop varieties based on simple repeating sequences according to claim 1, characterized in that, The step of separating the amplification products by electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products includes: The amplification products were separated by electrophoresis using primer fluorescent labeling dyes and a capillary electrophoresis apparatus to obtain the target fluorescence signal image corresponding to the amplification products.
5. The method for analyzing crop varieties based on simple repeating sequences according to claim 3, characterized in that, The electrophoretic separation of the amplification products based on primer fluorescent labeling dyes and capillary electrophoresis to obtain the target fluorescence signal image corresponding to the amplification products includes: The amplification products are separated by electrophoresis using the primer fluorescent labeling dye and the capillary electrophoresis apparatus to obtain the initial fluorescence signal image corresponding to the amplification products. The initial fluorescence signal image is calibrated by segment size to obtain the intermediate fluorescence signal image corresponding to the initial fluorescence signal image; The intermediate fluorescence signal image is converted to grayscale to obtain the grayscale fluorescence signal image of the intermediate fluorescence signal image; The grayscale fluorescence signal image is subjected to median filtering to obtain the target fluorescence signal image corresponding to the grayscale fluorescence signal image.
6. The method for analyzing crop varieties based on simple repeating sequences according to any one of claims 1 to 5, characterized in that, The target image features include: peak area of the bands, mobility, and information entropy; the step of inputting the target image features into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier includes: The peak area of the strip, the mobility, and the information entropy are respectively input into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier.
7. A crop variety analysis device based on simple repeating sequences, characterized in that, include: The determination module is used to determine the target primer set based on the simple repetitive sequence sites pre-screened for the target crop; An amplification module is used to perform polymerase chain reaction amplification on the target primer set to obtain the amplification product corresponding to the target primer set; The image module is used to perform electrophoretic separation of the amplification products to obtain the target fluorescence signal image corresponding to the amplification products; The feature extraction module is used to extract target image features from the target fluorescence signal image; The analysis module is used to input the target image features into a pre-trained random forest classifier to obtain the variety analysis results of the target crop output by the random forest classifier; the random forest classifier is trained based on the training image features and the corresponding variety identifiers.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the crop variety analysis method based on simple repeating sequences as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the crop variety analysis method based on simple repeating sequences as described in any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the crop variety analysis method based on simple repeating sequences as described in any one of claims 1 to 6.