A method for rapidly predicting the virus-carrying rate of a rice sawtooth leaf stunt disease population

CN122503552APending Publication Date: 2026-08-04ZHEJIANG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG UNIV
Filing Date
2026-07-08
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

其缺点在于:症状出现具有严重滞后性(通常需30天以上),不能反映早期隐性感染情况;秧苗期感染植株通常尚未表现出明显可识别症状,难以用于移栽前风险判断;受水稻品种、生育期、环境条件和其他病害影响较大;主要反映已出现症状的植株比例,不能准确反映分子水平上的RRSV实际带毒率,不利于病害早期预警

Benefits of technology

[0030] 1. Traditional methods require testing at least 30 individual plants to predict the virus infection rate of the entire field. This invention uses a pooled sample of leaves from 30 rice plants to reduce the number of samples to 1/30 of the original, significantly increasing the detection throughput. It is especially suitable for large-scale field sample screening and rapid assessment before seedling transplanting, significantly reducing the workload and improving detection efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122503552A_ABST
    Figure CN122503552A_ABST
Patent Text Reader

Abstract

The present application relates to rice virus rate prediction technical field, especially a kind of rice sawtooth leaf stunt disease population virus rate fast prediction method, its technical scheme includes: by collecting the mixed sample consisting of N rice leaves, after RNA extraction, reverse transcription and RT-qPCR detection, obtain the 2 ‑ΔCt value of rice sawtooth leaf stunt disease virus RRSV P10 fragment and internal reference gene;The value is input into the random forest prediction model trained in advance, and the population prediction virus rate can be output;Wherein, the random forest model is trained with the characteristics of the mixed sample 2 ‑ΔCt value of known actual virus rate population, and the actual virus rate determined by single-plant RT-PCR is used as label to obtain;Realize the fast prediction of rice population virus rate by mixed sample detection combined with machine learning, suitable for different growth periods and geographical regions, can assist in judging transplanting risk, with the characteristics of high efficiency and low cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rice virus infection rate prediction technology, and in particular to a rapid prediction method for the virus infection rate of rice spur leaf dwarf disease populations. Background Technology

[0002] Rice ragged dwarf disease is an important viral disease of rice caused by Rice Ragged dwarf Virus (RRSV), which is mainly transmitted by brown planthoppers. Infected rice plants exhibit symptoms such as stunted growth, dark green leaves, serrated leaf margins, wrinkled leaves, and inhibited growth, severely affecting heading and grain filling. Especially when rice is infected with RRSV during the seedling stage, early symptoms are often lacking and difficult to detect visually in the field, leading to missed opportunities for optimal replanting or transplanting. However, infection at this stage has a sustained negative impact on post-transplant growth and development, potentially causing severe stunting, failure to head, or inability to properly fill grains after heading, resulting in significant yield reduction or even crop failure. Therefore, accurately determining the RRSV carrier level in seedling populations or test fields before transplanting or in the early stages of disease occurrence is crucial for assessing subsequent disease risk, guiding transplanting decisions, and implementing early control measures.

[0003] Currently, the main methods for detecting RRSV in rice include field symptom observation, single-plant RT-PCR detection, single-plant RT-qPCR quantitative detection, dot-ELISA immunological detection, and conventional pooled RT-PCR or RT-qPCR detection.

[0004] The field symptom survey method mainly involves observing whether rice plants exhibit typical symptoms such as stunted growth, serrated leaf margins, and wrinkled leaves, counting the number of diseased plants, and calculating the field incidence rate. Its disadvantages include: symptom onset is severely delayed (usually requiring more than 30 days), failing to reflect early latent infection; infected seedlings typically do not yet show obvious identifiable symptoms, making it difficult to use for pre-transplant risk assessment; it is significantly affected by rice variety, growth stage, environmental conditions, and other diseases; and it primarily reflects the proportion of plants with symptoms, failing to accurately reflect the actual RRSV carrier rate at the molecular level, thus hindering early disease warning.

[0005] The single-plant RT-PCR detection method involves extracting RNA from individual rice samples collected in the field, reverse transcribing it to obtain cDNA, and then performing PCR amplification using RRSV-specific primers. The amplification results are used to determine whether each rice plant carries the virus, and the actual virus-carrying rate of the population is calculated. Its advantages are that the detection results are intuitive and can provide the actual virus-carrying rate; however, it has significant disadvantages: it requires testing a large number of plants individually, resulting in a large workload; it involves multiple RNA extraction, reverse transcription, and PCR testing steps, leading to high reagent and labor costs; the detection cycle is long, making it unsuitable for large-scale rapid field screening; and when the number of samples reaches hundreds or even thousands of plants, the detection efficiency is insufficient to meet the needs of actual production management.

[0006] Single-plant RT-qPCR quantitative detection method detects the accumulation of specific RRSV fragments in individual rice plants using RT-qPCR, providing relatively sensitive molecular quantitative results. Its disadvantages are similar to those of single-plant RT-PCR detection: it still requires testing a large number of plants individually, resulting in high detection costs, heavy workload, and low throughput; large-scale detection involves high experimental and data processing costs; it is difficult to efficiently serve rapid risk assessment of entire fields or batches of seedlings; and it is not conducive to quickly determining the infection level and subsequent disease risk of the population before seedling transplanting.

[0007] The dot-ELISA immunoassay method uses RRSV-specific antibodies to detect viral antigens in rice samples, thereby determining whether the rice is infected with RRSV. Compared to nucleic acid detection, this method simplifies the process to some extent, but it still has shortcomings: the detection relies on RRSV-specific antibodies, and the preparation, procurement, and storage of antibodies are costly; the detection results are mainly used to determine whether the sample carries the virus, and it is difficult to directly obtain the actual virus carriage rate in the field population; when used for large-scale field testing, a large number of samples still need to be processed one by one, resulting in a high overall workload and detection cost; its sensitivity and quantitative ability are limited compared to RT-qPCR, making it difficult to meet the needs of fine assessment of samples with low virus accumulation and population infection levels.

[0008] Conventional pooled RT-PCR or RT-qPCR methods involve mixing multiple rice leaves for molecular detection to reduce the number of samples tested, increase throughput, and lower costs. However, most pooled tests can only determine the presence of the virus in the pooled sample, making it difficult to directly predict the actual viral load in the population. Furthermore, the pool size lacks optimization: too few samples result in limited efficiency improvement, while too many samples can dilute the viral signal due to the large number of non-virus-carrying samples, making it difficult to establish a stable and directly interpretable correlation between pooled RT-qPCR detection values ​​and the actual viral load in the population. Finally, there is a lack of technical solutions that utilize predictive models to convert pooled molecular detection data into population viral load rates.

[0009] In summary, existing technologies still lack a rapid assessment method for RRSV population carrying rate that can balance detection efficiency, detection cost, and prediction accuracy. In particular, there is a lack of a technical solution that can accurately estimate the RRSV carrying level of seedling populations or the entire field based on a single pooled RT-qPCR test result, combined with a quantitative prediction model, and further provide a basis for subsequent field disease risk assessment.

[0010] In view of the above-mentioned shortcomings of the existing technology, the following technical problems exist:

[0011] After RRSV infection, symptoms usually take 30 days or longer to appear. Infected seedlings often do not have obvious identifiable symptoms. If only field symptoms are relied upon for judgment, it is difficult to identify the risk of virus transmission in a population before transplanting or in the early stages of disease outbreak. This may lead to the transplanting of infected seedlings and cause serious yield losses later. There is a problem that simple symptom observation is delayed and it is difficult to achieve early warning in the seedling stage.

[0012] Existing single-plant PCR or RT-qPCR detection methods are time-consuming, consume a lot of reagents, and are labor-intensive when the number of field samples is large. They are difficult to meet the needs of large-scale field monitoring and rapid assessment before seedling transplanting. Methods such as dot-ELISA rely on specific antibodies, which are also costly and have limited application. They suffer from the problems of large workload, high cost, and low efficiency of traditional single-plant testing.

[0013] Conventional pooled RT-PCR or RT-qPCR detection can only reflect the presence of the virus in the pooled sample, and cannot directly estimate the actual proportion of infected plants in the field population. It is insufficient in the quantitative assessment of the population virus carrying rate. Conventional pooled detection can only determine whether the virus is present, but it is difficult to quantitatively predict the population virus carrying rate.

[0014] When the number of mixed samples is too small, the improvement in detection efficiency is limited. When the number of mixed samples is too large, the virus signal of the infected plant may be diluted by a large number of non-infected samples, affecting the detection sensitivity and prediction accuracy. There is a lack of optimal combination between virus signal and population virus carrying rate under different mixed sample numbers.

[0015] RT-qPCR can obtain the relative viral load (such as Ct value or 2). -ΔCt While the measured value can be directly converted into the field population virus carrying rate, a predictive model between the measured value and the actual virus carrying rate is needed to convert molecular detection data into population infection rate prediction results. There is a problem that RT-qPCR molecular detection data is difficult to convert into field population infection rate assessment results.

[0016] In view of this, we propose a rapid prediction method for the virus carrying rate of rice spur leaf dwarf disease populations to solve the existing problems. Summary of the Invention

[0017] The purpose of this invention is to provide a method for rapid prediction of the viral load of rice spur dwarf disease populations, in order to solve the problems mentioned in the background art.

[0018] To achieve the above objectives, the present invention provides the following technical solution: a method for rapid prediction of the virus carrier rate of rice spur dwarf virus (RRSV) populations, comprising the following steps: obtaining a mixed sample of leaves from the rice population to be tested, wherein the mixed sample is composed of leaf tissues from N rice plants, where N ≥ 2; performing RNA extraction, reverse transcription, and RT-qPCR detection on the mixed sample, detecting the target as the P10 fragment of RRSV and the endogenous reference gene of rice, and calculating the 2-1 of the mixed sample. -ΔCt Value; will 2 -ΔCt The values ​​are input into a pre-trained random forest prediction model, and the random forest prediction model outputs the predicted virus-carrying rate of the rice population to be tested.

[0019] The random forest prediction model was trained using the following method: leaf pool samples were obtained from multiple rice populations with known actual virus infection rates, and the RRSV P10 fragments of each pool sample were extracted. -ΔCt The values ​​are used as input features, and the actual virus-carrying rate of the corresponding mixed group determined by single-strain RT-PCR detection is used as the output label. The random forest algorithm is used for training to obtain the random forest prediction model.

[0020] Furthermore, N can take the values ​​10, 20, 30, or 50.

[0021] Furthermore, the primer sequences used to detect the RRSV P10 fragment are: F: TTCTCCACTGCGCTGTCTTA, R: ACGAGTGATGATGCCTCCAA; the primer sequences used to detect the rice endogenous reference gene OsActin are: F: CAGCACATTCCAGCAGAT, R: GGCTTAGCATTCTTGGGT.

[0022] Furthermore, the mixed sample is constructed as follows: 10 rice plants are used as a basic sampling unit, and 1, 2, 3 or 5 basic units are combined to form mixed samples of 10, 20, 30 or 50 plants respectively.

[0023] Furthermore, the actual virus-carrying rate was determined as follows: each rice plant in the mixed sample group was subjected to single-plant RT-PCR detection, and positive plants were identified using RRSV P8 fragment specific primers. The P8 fragment detection primer sequences were: F: CTGAATACAACCGATACCGT, R: GACCACTGTTACTGCCTTTA; Actual virus-carrying rate = (number of RRSV positive plants / total number of tested plants) × 100%.

[0024] Furthermore, when training the random forest prediction model, for each variance, a portion of the data is randomly selected as the training set, and the remaining data is used as the validation set; and the coefficient of determination R is used as the validation set. 2 The mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE), symmetric mean absolute percentage error (SMAPE), and mean absolute percentage error (MAPE) are used to evaluate the predictive performance of the model.

[0025] Furthermore, the steps also include: judging the transplanting risk of rice seedlings based on the predicted virus infection rate; when the predicted virus infection rate is less than 20%, it is judged to be suitable for transplanting; when the predicted virus infection rate reaches or exceeds 20%, it is judged to have a high risk of disease and should not be transplanted directly. Based on actual production, with 2 seedlings per clump, the probability of all 2 seedlings in a single clump becoming diseased is less than 4%, which is lower than the disease control threshold.

[0026] Furthermore, the method is used for sample testing during the rice seedling stage, tillering stage, heading stage, grain-filling stage, or maturity stage, and is applicable to rice field samples from different geographical regions.

[0027] Furthermore, the predicted RRSV carrier rate of the test field or seedling population was obtained through a single pooled RT-qPCR test.

[0028] Furthermore, the random forest prediction model was used to process the 2 obtained from pooled RT-qPCR detection. -ΔCt The nonlinear relationship between the value and the actual virus-carrying rate of the rice population.

[0029] Compared with the prior art, the beneficial effects of the present invention are:

[0030] 1. Traditional methods require testing at least 30 individual plants to predict the virus infection rate of the entire field. This invention uses a pooled sample of leaves from 30 rice plants to reduce the number of samples to 1 / 30 of the original, significantly increasing the detection throughput. It is especially suitable for large-scale field sample screening and rapid assessment before seedling transplanting, significantly reducing the workload and improving detection efficiency.

[0031] 2. This invention reduces the number of RNA extraction, reverse transcription, and RT-qPCR reactions by using pooled testing, significantly reducing reagent and material consumption and manual operation time. Single-plant testing requires 30 extraction reagents and 30 PCR reaction solutions, while a pooled sample of 30 plants requires only about 1, resulting in reagent cost savings of over 96%. Simultaneously, the reduced working hours for testing personnel lower labor costs, making large-scale field monitoring economically feasible and significantly reducing testing costs.

[0032] 3. This invention establishes a mixed sample 2 -ΔCtA random forest prediction model was used to correlate the viral load with the actual viral load, enabling a quantitative conversion from molecular detection data to population infection rates. The R-value of a pool of 30 strains combined with the random forest model was [data missing]. 2 The predicted viral load rate reached 0.862, with an MSE of only 71.57 and an RMSE of 8.45, indicating a high degree of agreement between the predicted and actual viral load rates. Compared to the actual disease incidence rates observed in the fields, the predicted viral load rate of the random forest model for seedling samples from five fields showed an error within 2 percentage points, demonstrating its ability to accurately and quantitatively assess the viral load level in the population and achieve accurate quantitative prediction of the viral load rate.

[0033] 4. Based on representative random sampling, this invention requires only one RT-qPCR test on a single 30-plant pooled sample to rapidly estimate the RRSV infection level of the seedling population or the entire field. Only one pooled sample test is needed for each of five fields to obtain predictive results that are essentially consistent with the later field disease incidence. This provides an efficient and practical quantitative basis for subsequent field disease risk assessment, overcoming the shortcomings of existing technologies that require extensive single-plant testing or rely on delayed symptom investigations, and rapidly predicting field infection levels through a single pooled RT-qPCR test.

[0034] 5. This invention does not rely on symptom observation and can identify the risk of virus transmission before seedling transplanting. Through pooled RT-qPCR detection during the seedling stage, the disease incidence rate in the field after transplanting was successfully predicted. Based on this, this invention sets a 20% transplanting risk reference threshold (taking 2 seedlings / clump in production as an example, ensuring that the complete disease incidence rate of each clump after transplanting is less than 4%): transplanting is recommended when the predicted virus transmission rate is less than 20%, and retesting or adjusting the seedling source is recommended when it reaches or exceeds 20%. This threshold combines disease incidence data and actual production conditions, and has clear guiding value, especially suitable for early warning during the seedling stage to guide production decisions.

[0035] 6. This invention compares mixed sample sizes of 10, 20, 30, and 50 plants, and three models—linear regression, decision tree, and random forest—using R... 2 The evaluation was based on a comprehensive assessment of multiple indicators, including MSE, RMSE, MAE, SMAPE, and MAPE. The R-squared of a 30-plant pool combined with a random forest model was used. 2 The highest R value was (0.862), while the lowest MSE (71.57) and RMSE (8.45) were significantly better than other combinations. The R values ​​for the 10-plant and 20-plant mixtures were... 2 R values ​​close to 0 or negative for a pool of 50 plants 2 The error is only 0.528, which significantly increases the overall error. This indicates that the combination of the 30-strain pool and the random forest model achieves the optimal balance between detection efficiency, viral signal strength, and prediction accuracy, and can handle 2... -ΔCt The nonlinear relationship between the value and the actual toxicity rate shows high prediction accuracy and good model stability.

[0036] 7. The method of this invention has been validated in field samples from different regions, at different rice growth stages, and with different infection levels. The predicted virus-carrying rate and the actual virus-carrying rate of the 30-plant mixed random forest model in samples from different regions maintained good consistency, and the model did not exhibit overfitting or regional shift. Therefore, this invention is applicable not only to the seedling stage but also to field monitoring throughout the entire rice growing season, possessing good versatility and stability, wide applicability, and good field adaptability.

[0037] 8. This invention can be used in the following practical production scenarios: rapid assessment of the risk of rice seedlings carrying viruses before transplanting to avoid transplanting virus-carrying seedlings into the field; early warning of RRSV in the field to guide the timing of brown planthopper control; regional disease monitoring to provide data support for plant protection departments; large-scale screening of RRSV-resistant germplasm resources to improve breeding efficiency; and has significant production application value. Attached Figure Description

[0038] Figure 1 This is an overall flowchart of a method for rapid prediction of the viral load of rice spur leaf dwarf disease population according to the present invention;

[0039] Figure 2 The model fitting results are shown for different sample sizes.

[0040] Figure 3 A comparison graph showing the model's predicted values ​​and actual values ​​under different sample sizes;

[0041] Figure 4 A comparison chart of predicted and actual virus-carrying rates of field samples from different regions. Detailed Implementation

[0042] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0043] Example 1

[0044] This application presents an overall flowchart of a rapid prediction method for the virus carriage rate of rice scab. By comparing different sample sizes of 10, 20, 30, and 50 plants, and different prediction models such as linear regression, decision tree, and random forest, the method using a 30-plant rice leaf sample and RT-qPCR detection of the RRSV P10 fragment was determined to be the most effective. -ΔCt The preferred technical solution is to use the value and random forest model to predict the RRSV population carrying rate.

[0045] This method reduces the workload of individual plant testing while converting pooled RT-qPCR results into predicted RRSV carrier rates for field populations. By performing a single RT-qPCR test on a pooled sample of 30 plants and combining it with an optimized prediction model, the RRSV infection level of the tested field or seedling population can be rapidly estimated, providing a basis for subsequent field disease risk assessment. This addresses the problems of low detection efficiency, high cost, delayed symptom observation, and difficulty in quantitatively predicting population carrier rates using pooled sample testing in existing technologies. Figure 1 As shown, the technical solution of this application includes a sampling module, a model construction module, a model validation module, and a result application module. Specifically, A is the sampling module, showing that a basic sampling unit of 10 plants is used, and mixed sample groups of 20, 30, and 50 plants are formed through combination; B is the model construction module, showing that a prediction model is constructed based on the mixed sample RT-qPCR detection values ​​and actual virus-carrying rate data; C is the model validation module, showing that the optimal mixed sample size and prediction model are selected through different evaluation indicators; and D is the result application module, showing that the field RRSV prediction virus-carrying rate is assessed using a 30-plant mixed sample, and 20% is used as the preferred transplanting risk assessment reference threshold.

[0046] Field samples were collected and basic unit divisions were performed using a sampling module. During the model construction phase, leaf samples were randomly collected from 500 rice seedlings in paddy fields with severe brown planthopper infestations. Each rice seedling was individually numbered, and an equal amount of leaf tissue was collected. Figure 1 As shown in Part A, the collected 500 rice seedling samples were divided into 50 basic sampling units, each consisting of 10 seedlings. Specifically: one basic unit corresponded to a pooled sample of 10 seedlings; two basic units formed a pooled sample of 20 seedlings; three basic units formed a pooled sample of 30 seedlings; and five basic units formed a pooled sample of 50 seedlings. By constructing samples of different pool sizes using this method, the impact of different pool sizes on the predictive performance of RRSV infection rate was compared.

[0047] Single-plant detection and calculation of the actual virus-carrying rate were performed using the model construction module. Individual rice leaf samples were ground, total RNA was extracted, and cDNA was obtained through reverse transcription. RT-PCR was performed using RRSV P8 fragment-specific primers to determine whether each rice plant was infected with RRSV. The RRSV P8 fragment detection primer sequences were: F: CTGAATACAACCGATACCGT; R: GACCACTGTTACTGCCTTTA. Simultaneously, RT-qPCR was performed using RRSV P10 fragment-specific primers to further determine RRSV infection status. The RRSV P10 fragment detection primer sequences were: F: TTCTCCACTGCGCTGTCTTA; R: ACGAGTGATGATGCCTCCAA. The actual virus-carrying rate was calculated based on the single-plant detection results using the formula: Actual virus-carrying rate = Number of RRSV-positive plants / Total number of detected plants × 100%.

[0048] RT-qPCR detection of different numbers of mixed samples was performed using a model building module. Rice leaf samples of different mixed sizes were thoroughly ground, total RNA was extracted, and cDNA was obtained through reverse transcription. RT-qPCR detection of each mixed sample was performed using RRSV P10 fragment-specific primers to obtain the corresponding Ct values. Normalization was performed using the endogenous reference gene OsActin. The OsActin detection primer sequences were: F: CAGCACATTCCAGCAGAT; R: GGCTTAGCATTCTTGGGT. The 2-1 of the mixed samples was calculated. -ΔCt The value serves as a molecular indicator reflecting the relative content of RRSV virus in a mixed sample.

[0049] A model for predicting RRSV carrier rate was established using a model building module. The RRSV P10 fragment in mixed samples of 10, 20, 30, and 50 strains was used as the basis for the model prediction. -ΔCt Using the values ​​as input variables and the actual virus carrying rate of the corresponding mixed sample group as the output variable, a prediction model for the RRSV population virus carrying rate was established. For each type of mixed sample size, 50 sets of data were obtained. Of these, 30 sets were randomly selected as the training set for model construction; the remaining 20 sets were used as the validation set for model validation and prediction performance evaluation. Based on the training set data, linear regression models, decision tree models, and random forest models were constructed respectively. After the models were established, the validation set samples were... -ΔCt Input the value into the corresponding model to obtain the predicted infection rate, and compare it with the actual infection rate.

[0050] The optimal sample size and prediction model were selected through the model validation module. The coefficient of determination R0 was used. 2The predictive performance of different sample sizes and models is evaluated using metrics such as mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), symmetric mean absolute percentage error (SMAPE), and mean absolute percentage error (MAPE). 2 Higher values ​​indicate a better fit between the model's predicted values ​​and the actual virus-carrying rate; lower values ​​for MSE, RMSE, MAE, SMAPE, and MAPE indicate smaller model prediction errors. Based on the results of the example, the 30-strain pool combined with the random forest model has a high R-value. 2 With lower MSE and RMSE values, the overall prediction performance is better than the pooled sampling schemes of 10, 20, and 50 plants. Therefore, the pooled sampling of 30 rice leaves combined with a random forest model is determined to be the preferred scheme for predicting the RRSV population virus carrying rate in this invention.

[0051] The results application module was used to predict the application of this method to field samples. A 30-plant mixed random forest prediction model was applied to field sample testing, including field seedling samples, samples from different regions, and samples from different rice growth stages. The model's predicted virus-carrying rate was compared with the actual virus-carrying rate obtained from single-plant RT-PCR detection or the subsequent field disease incidence rate to verify the applicability and stability of this method in actual field samples.

[0052] The results application module is used to assess transplanting risk based on the predicted virus-carrying rate. Based on the RRSV predicted virus-carrying rate output by the model, the transplanting risk of seedling populations or test plots can be further assessed. Combining field validation results and production management needs, in a preferred embodiment, a predicted virus-carrying rate of 20% can be used as a reference threshold for seedling transplanting risk assessment: when the predicted virus-carrying rate is below 20%, it indicates a relatively low RRSV infection risk for the seedling population or plot, and can be used as a reference for suitable transplanting; when the predicted virus-carrying rate reaches or exceeds 20%, it indicates a high RRSV disease risk for the seedling population or plot, and direct transplanting is not recommended. Measures such as re-testing, replacing seedling sources, strengthening pest and disease monitoring, or adjusting transplanting arrangements can be taken, based on the brown planthopper occurrence, field symptoms, and production management needs.

[0053] It should be noted that the 20% threshold is set after comprehensively considering the correlation between the model's predicted virus-carrying rate and subsequent field disease incidence (taking rice plants with at least two seedlings per clump as an example, the probability of the clump being completely infected is less than 4%), as well as the production risk that RRSV infection during the seedling stage has a sustained and serious impact on subsequent heading, grain filling, and yield formation. To reduce the possibility of large-scale field disease caused by transplanting infected seedlings, this invention preferably sets the predicted virus-carrying rate of 20% as the reference threshold for transplanting risk warning. In other application scenarios, this threshold can also be appropriately adjusted according to rice variety, planting area, brown planthopper population density, cultivation season, and production management requirements.

[0054] The core principle of this application is that after RRSV infects rice, the virus P10 fragment can be detected by RT-qPCR in rice samples, and its 2 -ΔCt The value can reflect the relative virus accumulation in the pooled sample. Since the virus accumulation in the pooled sample is related to the proportion of virus-carrying plants in that pool, a 23 value can be established to reflect the relative virus accumulation in the pooled sample. -ΔCt A predictive model between the value and the actual viral load rate transforms molecular detection results into population viral load prediction results.

[0055] Because a non-linear relationship may exist between pooled RT-qPCR detection values ​​and actual virus carriage rates, simple linear regression models struggle to achieve stable prediction results across all pool sizes. This application compares different models such as linear regression, decision trees, and random forests, finding that the random forest model better handles the non-linear relationship between pooled RT-qPCR detection data and actual virus carriage rates. Furthermore, by comparing pool sizes of 10, 20, 30, and 50 plants, this application finds that a 30-plant pool achieves a good balance between detection efficiency, virus signal intensity, sample representativeness, and prediction accuracy. Therefore, a pool of 30 rice leaves combined with a random forest model is preferred for predicting RRSV population carriage rates.

[0056] Example 2

[0057] This embodiment describes the construction of an RRSV virus carrier rate prediction model based on rice leaf mixed sample RT-qPCR data and the screening of the optimal mixed sample quantity.

[0058] To establish a model for rapid prediction of RRSV infection rates in rice seedlings during field sampling and pooling, leaf samples were randomly collected from 500 rice seedlings in paddy fields with severe RRSV infestations. Each seedling was individually numbered, and an equal amount of leaf tissue was collected from each seedling and brought back to the laboratory for later use. Subsequently, individual leaf samples from each of the 500 seedlings were ground, total RNA was extracted, and cDNA was obtained through reverse transcription. RT-PCR was then performed using RRSV P8 fragment-specific primers to determine whether each rice seedling was infected with RRSV. The primer sequences for P8 fragment detection were: F: CTGAATACAACCGATACCGT; R: GACCACTGTTACTGCCTTTA. Simultaneously, RT-qPCR was performed on the individual samples using RRSV P10 fragment-specific primers to further confirm RRSV infection status. The primer sequences for P10 fragment detection are: F: TTCTCCACTGCGCTGTCTTA; R: ACGAGTGATGATGCCTCCAA. Based on the results of single-plant RT-PCR and / or RT-qPCR detection, the virus-carrying status of each rice seedling sample was determined, serving as the basis for subsequent calculations of the actual virus-carrying rate in each pooled group.

[0059] like Figure 1 As shown in Part A, the 500 rice seedling samples were divided into 50 basic sampling units, each consisting of 10 seedlings. Based on this, sample groups of different mixing sizes were constructed: 1 basic unit constituted a 10-seedling mixing group; 2 basic units constituted a 20-seedling mixing group; 3 basic units constituted a 30-seedling mixing group; and 5 basic units constituted a 50-seedling mixing group. Each type of mixing group had 50 groups. In each mixing group, an equal amount of tissue was taken from the leaves of each rice seedling and mixed. After mixing, the samples were thoroughly ground, total RNA was extracted, and cDNA was obtained through reverse transcription.

[0060] Subsequently, RT-qPCR was performed on each pooled sample using RRSV P10 fragment-specific primers to obtain the Ct value for each pooled sample, which was then corrected using the endogenous reference gene OsActin. The OsActin detection primer sequences were: F: CAGCACATTCCAGCAGAT; R: GGCTTAGCATTCTTGGGT. The corresponding 2-1 values ​​for each pooled sample were calculated. -ΔCt Value. 2 -ΔCt The value serves as a molecular indicator reflecting the relative accumulation level of RRSV virus in mixed samples.

[0061] The actual virus-carrying rate of each pooled sample group is calculated based on the test results of individual plants within that group. The calculation formula is as follows: Actual virus-carrying rate = Number of RRSV-positive plants in the pooled sample group / Total number of plants in the pooled sample group × 100%.

[0062] In the construction of the RRSV infection rate prediction model, the RRSV P10 fragment in mixed samples of leaves from 10, 20, 30, and 50 rice plants was used as the basis for prediction. -ΔCt Using the values ​​as input variables and the actual virus carrying rate of the corresponding mixed sample group as the output variable, a prediction model for the virus carrying rate of the RRSV population was established. For each type of mixed sample, 50 sets of data were obtained. Among them, 30 sets of data were randomly selected as the training set for model construction; the remaining 20 sets of data were used as the validation set for model validation and prediction performance evaluation.

[0063] Based on the training set data mentioned above, linear regression, decision tree, and random forest models were constructed respectively. After the models were built, the validation set of 2... -ΔCt Input the values ​​into the corresponding model to obtain the predicted infection rate, and compare the predicted infection rate with the actual infection rate to evaluate the accuracy of the model prediction.

[0064] In the selection of optimal pool size and prediction model, to determine the optimal pool size and prediction model suitable for predicting the virus carrier rate of RRSV populations, the fitting and prediction effects of linear regression model, decision tree model, and random forest model were compared under pool conditions of 10, 20, 30, and 50 plants, respectively. The model fitting results are shown below. Figure 2 As shown, the model prediction results are as follows: Figure 3 As shown in Table 1, the model performance evaluation metrics are as follows.

[0065] Table 1. Predictive performance analysis results of different models

[0066]

[0067] The fitting performance of the three models was compared under different sample sizes. Figure 2 As shown, the fitting relationship between the relative expression level of RRSV P10 and the actual viral load rate differed significantly under different sample sizes. Under the conditions of 10 and 20 samples, similar relative expression levels of RRSV P10 in some samples corresponded to significantly different actual viral load rates, leading to unstable model fitting results. Under the condition of 50 samples, although the three models could reflect certain trends, deviations still existed between some samples and the fitted curves. In contrast, under the condition of 30 samples, the relative expression level of RRSV P10 and the actual viral load rate showed a more stable correspondence, and all three models demonstrated good fitting results.

[0068] like Figure 3 As shown, there are significant differences in the predictive performance of RRSV carrier rates depending on the pool size and the different prediction models. Under pool sizes of 10 and 20 strains, the predicted carrier rates for some samples deviated considerably from the actual rates. Under a pool size of 50 strains, significant differences remained between the predicted and actual values ​​for some samples. In contrast, the prediction performance of all three models was generally better under a pool size of 30 strains, with the random forest model showing the closest prediction to the actual values. These results indicate that a pool size of 30 strains can achieve a good balance between detection efficiency, sample representativeness, and prediction accuracy.

[0069] To further quantify and evaluate the predictive performance of different sample sizes and models, the coefficient of determination (R²), mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), symmetric mean absolute percentage error (SMAPE), and mean absolute percentage error (MAPE) were used for comparison. A higher R² indicates a better fit between the predicted and actual values; lower MSE, RMSE, MAE, SMAPE, and MAPE indicate smaller model prediction errors. Table 1 shows that different sample sizes can be used to establish RRSV toxicity prediction models, but the model prediction performance varies significantly with different sample sizes.

[0070] Under the condition of a mixed sample of 10 plants, the R² values ​​of the linear regression model, decision tree model, and random forest model were -0.144, -0.643, and -0.061, respectively, all of which were below 0. This indicates that the model prediction performance was poor under this mixed sample size, and the predicted values ​​had a low degree of fit with the actual virus-carrying rate.

[0071] Under a pooled sample of 20 plants, the R values ​​for the three models were... 2 The values ​​increased to 0.181, 0.353, and 0.459 respectively, indicating that the 20-strain pool can reflect the changing trend of RRSV infection rate to some extent, but the overall prediction accuracy is still limited.

[0072] Under a pooled sample of 30 plants, all three models showed good predictive performance. Among them, the linear regression model had the highest R-value. 2 The R-squared value of the decision tree model is 0.814. 2 The R-value of the random forest model is 0.847. 2 The R-value was 0.862, which was significantly higher than the R-values ​​of the corresponding models under the mixed sample conditions of 10, 20, and 50 plants. 2 The values ​​are as follows. In particular, when the 30-tree pool is combined with the random forest model, the MSE is 71.57 and the RMSE is 8.45, both of which are the lowest values ​​among all model combinations, indicating that the combination has a good fitting effect and a low overall prediction error.

[0073] Under a pooled sample of 50 trees, the predictive performance of all three models decreased compared to the pooled sample of 30 trees. Specifically, the R-value of the random forest model... 2 The accuracy was 0.528, MSE was 228.05, and RMSE was 15.10, all significantly worse than the 30-strain pool random forest model. This result indicates that when the pool size is too large, the viral signal in the pooled samples may be diluted, thus affecting the model's predictive accuracy.

[0074] Comparing the predictive performance of different sample sizes and models, the 30-plant pooled sample outperformed the 10, 20, and 50-plant pooled samples overall. Under the 30-plant pooled sample condition, the random forest model exhibited the highest R² value and lower MSE and RMSE, indicating its ability to better fit the relationship between pooled RT-qPCR detection values ​​and the actual virus-carrying rate. Although the MAE, SMAPE, and MAPE of the decision tree model were slightly lower than those of the random forest model under the 30-plant pooled sample condition, the random forest model performed better in terms of overall fit and overall error control, and also demonstrated better model stability. Therefore, the combination of a 30-plant rice leaf pooled sample and the random forest model can be considered the preferred RRSV population virus-carrying rate prediction scheme in this application.

[0075] The results of this embodiment demonstrate that, by optimizing the pooled sample size and prediction model, this application can achieve a relatively accurate prediction of the RRSV carrier rate in rice populations in the field while reducing the workload of individual plant testing. This method can convert pooled RT-qPCR detection results into population carrier rate prediction results, and is suitable for rapid field monitoring of RRSV, pre-transplant risk assessment of seedlings, and large-scale sample screening.

[0076] Example 3

[0077] This example compares the predicted virus-carrying rate in field seedling samples with the disease incidence rate in the later field.

[0078] To verify whether the method described in this application can reflect the risk of disease in the field after transplanting during the seedling stage, leaf samples were randomly collected from rice seedlings in paddy fields with severe infestations of infected brown planthoppers in Hongqi Village, Sanya. Five pooled samples of 30 rice seedlings each were prepared, corresponding to five different field samples. Each seedling in each pooled sample was individually numbered, and an equal amount of leaf tissue was collected. One portion of the samples was used for pooled RT-qPCR detection and virus carriage rate prediction, while the other portion was used for individual plant virus carriage detection and subsequent field disease surveys.

[0079] In the pooled RT-qPCR detection and virus carriage rate prediction, equal amounts of leaf tissue were taken from 30 rice seedlings in each group and mixed. RT-qPCR detection of the RRSV P10 fragment was performed according to the method described in Example 2, and the corresponding 2... -ΔCt Value. And 2 -ΔCt The values ​​were input into the linear regression model, decision tree model, and random forest model established in Example 2, respectively, to obtain the predicted RRSV virus carrying rate for each group of 30 mixed samples. This step shows that during the seedling stage, only one RT-qPCR test is needed on one 30-plant mixed sample set for each field to obtain the predicted RRSV virus carrying level of that field or batch of seedlings.

[0080] In the subsequent field disease incidence and actual virus carriage rate detection, the above 5 groups of seedling samples were observed in the field, and the actual disease incidence of each group of samples was calculated based on the typical symptoms of RRSV. Typical RRSV symptoms include significantly stunted plants, dark green leaves with serrated leaf margins. Simultaneously, RNA was extracted and reverse transcribed from individual rice leaves in each group of samples, and PCR detection was performed using RRSV P8 fragment-specific primers to determine whether each rice plant was virus-carrying, and the actual virus carriage rate of each group of samples was calculated. The actual disease incidence and actual virus carriage rate were calculated according to the following formulas: Actual disease incidence = Number of plants with typical RRSV symptoms / Total number of tested plants × 100%; Actual virus carriage rate = Number of RRSV-positive plants / Total number of tested plants × 100%.

[0081] In the comparison of prediction results, the model-predicted virus-carrying rate of 5 groups of 30 mixed samples was compared with the actual disease incidence rate obtained from the subsequent field survey, and the predicted virus-carrying rate of the 4th group of samples was compared with the actual virus-carrying rate obtained from single-plant molecular detection. The results are shown in Table 2.

[0082] Table 2. Predicted RRSV infection rate in rice seedlings and disease incidence rate in rice plants at the heading stage.

[0083]

[0084] Table 2 shows that in the five field samples from Hongqi Village, the RRSV infection rate predicted by the linear regression model was 17.91%–22.43%, the RRSV infection rate predicted by the decision tree model was 20.00%–30.00%, and the RRSV infection rate predicted by the random forest model was 18.68%–29.25%. Later field surveys during the heading stage showed that the RRSV incidence rate in the five fields was 20.83%–27.08%, with a total incidence rate of 22.92%.

[0085] Among them, the actual RRSV infection rate of seedlings obtained by single-plant molecular detection in the fourth group of samples was 20.00%. In this group of samples, the RRSV infection rates predicted by the linear regression model, decision tree model and random forest model were 19.85%, 20.00% and 21.27%, respectively, all of which were close to the actual infection rates.

[0086] Further analysis of the later-stage field incidence rates reveals that the predictions from the random forest model are generally close to the actual incidence rates. For example, in the second and third groups of samples, the random forest model predicted virus carriage rates of 24.73% and 29.25%, respectively, corresponding to actual field incidence rates of 25.00% and 27.08%, respectively.

[0087] The results indicate that the predictive model established based on RT-qPCR data from a pooled sample of 30 seedlings can effectively reflect the RRSV disease level in the corresponding field after transplanting. In other words, a single RT-qPCR test of 30 seedlings at the seedling stage can provide a relatively accurate prediction of the potential RRSV disease level in the entire field.

[0088] This embodiment demonstrates that the RT-qPCR detection method for mixed leaf samples of 30 rice seedlings established in this application can not only predict the actual virus-carrying rate in the sense of molecular detection of seedling populations, but also reflect the disease incidence in the corresponding fields after seedling transplantation.

[0089] This method can quickly estimate the risk of subsequent RRSV disease in the field by a single pooled RT-qPCR test when the seedlings have not yet shown obvious symptoms. It provides a basis for production decisions such as whether to continue transplanting seedlings, whether to conduct supplementary retesting, whether to adjust the source of seedlings, and whether to strengthen the control of brown planthoppers.

[0090] Example 4

[0091] This embodiment validates the predictive model for RRSV infection rate in rice samples from different regions and at different developmental stages.

[0092] To verify the applicability of the method of this invention in different regions, field environments, and rice growth stages, field sample collection and actual virus-carrying rate testing were conducted in Sanya, Hainan; Wuhan, Hubei; and Guangzhou, Guangdong from 2023 to 2024. Sampling locations included seven fields in Baihe Village, Baichao Village, Gongbei Village, Hongqi Village, and Shengli Village in Sanya City, Hainan Province, as well as Wuhan City in Hubei Province and Guangzhou City in Guangdong Province. The sampled rice growth stages included tillering, heading, grain-filling, and maturity.

[0093] Rice samples collected from the above-mentioned different regions were subjected to single-plant PCR and RT-qPCR tests to determine whether each rice plant was infected with RRSV, and the actual RRSV infection rate of each field was calculated. The results are shown in Table 3.

[0094] Table 3. Actual RRSV infection rate of rice samples from different regions as detected by RT-PCR.

[0095]

[0096] As shown in Table 3, the actual RRSV infection rate varied among rice samples from different regions and at different growth stages, ranging from 13.33% to 96.67%. These samples covered different regions, fields, developmental stages, and RRSV infection levels, and can be used to verify the applicability and stability of the mixed-sample detection and infection rate prediction method of this invention in actual field samples.

[0097] In the 30-plant pooled RT-qPCR detection and prediction of virus carriage rate calculation, based on the actual RRSV carriage rate of each field, rice leaf samples collected from the same field were mixed in equal quantities according to 30 plants per pooled unit, and RRSV P10 fragment RT-qPCR detection was performed according to the method described in Example 2. The corresponding 2 -ΔCt Value. 2 -ΔCtThe values ​​are input into the linear regression model, decision tree model, and random forest model established in Example 2, respectively, to obtain the predicted RRSV infection rate of samples from different fields. When the same field contains multiple 30-species pooled samples, the predicted infection rate of each pooled sample is calculated separately and can be used to comprehensively determine the overall RRSV infection level of the field.

[0098] In comparing the predicted and actual virus-carrying rates, the model-predicted virus-carrying rates of rice samples from different regions and at different growth stages were compared with the actual virus-carrying rates obtained from single-plant RT-PCR detection. The results are as follows: Figure 4 As shown, the prediction model established based on RT-qPCR detection data of mixed leaf samples from 30 rice plants can predict the RRSV population carrying rate in field samples from different regions. Among the different models, the RRSV carrying rate predicted by the random forest model is generally close to the actual carrying rate, indicating that the model can reflect the RRSV infection level in field samples from different regions and at different developmental stages relatively well.

[0099] Combined with Table 3 and Figure 4 The results show that the samples used in this embodiment cover different regions, including Sanya in Hainan, Wuhan in Hubei, and Guangzhou in Guangdong, and include different rice growth stages such as tillering, heading, grain filling, and maturity. The actual virus-carrying rate ranges from 13.33% to 96.67%. In these field samples with significant differences, the 30-plant pool combined with the random forest model can still obtain prediction results that are relatively consistent with the actual virus-carrying rate, indicating that the method of this application has good field applicability and stability.

[0100] This embodiment demonstrates that the RT-qPCR detection method and random forest prediction model for mixed samples of 30 rice leaves established in this application are applicable not only to seedling samples, but also to samples from different regions, different field environments, and different rice growth stages.

[0101] Compared with traditional single-plant RT-PCR detection methods, the method proposed in this application can significantly reduce the number of samples and workload, and can rapidly obtain the predicted RRSV carrier rate of rice populations in the field through pooled RT-qPCR detection results. This method can be used for assessing the risk of virus carriers before seedling transplanting and predicting the risk of disease in the field afterwards, as well as for monitoring RRSV infection levels in different regions and large-scale sample screening, and has good practical application value.

[0102] The above specific embodiments are merely several preferred embodiments of the present invention. Based on the technical solutions of the present invention and the relevant teachings of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. A method for rapid prediction of the viral load rate of rice spur dwarf disease populations, characterized in that, The steps include: Obtain a mixed sample of leaves from the rice population to be tested. The mixed sample is composed of leaf tissues from N rice plants, where N ≥ 2. RNA extraction, reverse transcription, and RT-qPCR were performed on the mixed samples to detect the P10 fragment of rice scab virus RRSV and the endogenous reference gene of rice. The 2-1 of the mixed samples was calculated. -ΔCt Value; will 2 -ΔCt The values ​​are input into a pre-trained random forest prediction model, and the random forest prediction model outputs the predicted virus-carrying rate of the rice population to be tested. The random forest prediction model was trained using the following method: leaf pool samples were obtained from multiple rice populations with known actual virus infection rates, and the RRSV P10 fragments of each pool sample were extracted. -ΔCt The values ​​are used as input features, and the actual virus-carrying rate of the corresponding mixed group determined by single-strain RT-PCR detection is used as the output label. The random forest algorithm is used for training to obtain the random forest prediction model.

2. The method for rapid prediction of the viral load rate of rice spur dwarf disease according to claim 1, characterized in that: The value of N is 10, 20, 30 or 50.

3. The method for rapid prediction of the virus carriage rate of rice spur dwarf disease according to claim 1, characterized in that, The primer sequences used to detect the RRSV P10 fragment are: F: TTCTCCACTGCGCTGTCTTA, R: ACGAGTGATGATGCCTCCAA; the primer sequences used to detect the rice endogenous reference gene OsActin are: F: CAGCACATTCCAGCAGAT, R: GGCTTAGCATTCTTGGGT.

4. The method for rapid prediction of the virus carriage rate of rice spur dwarf disease according to claim 1, characterized in that, The method for constructing mixed samples is as follows: using 10 rice plants as a basic sampling unit, and combining 1, 2, 3 or 5 basic units to form mixed samples of 10, 20, 30 or 50 plants respectively.

5. The method for rapid prediction of the virus-carrying rate of rice spur dwarf disease according to claim 1, characterized in that, The method for determining the actual virus-carrying rate is as follows: each rice plant in the mixed sample group is subjected to single-plant RT-PCR detection, and RRSV positive plants are identified using RRSVP8 fragment specific primers. The primer sequences for P8 fragment detection are: F: CTGAATACAACCGATACCGT, R: GACCACTGTTACTGCCTTTA. Actual virus carrying rate = (Number of RRSV positive plants / Total number of tested plants) × 100%.

6. The method for rapid prediction of the viral load rate of rice spur dwarf disease according to claim 1, characterized in that: When training a random forest prediction model, for each variance, a portion of the data is randomly selected as the training set, and the remaining data is used as the validation set; and the coefficient of determination R is used as the validation set. 2 The mean square error (MSE), root mean square error (RMSE), mean absolute error (MAE), symmetric mean absolute percentage error (SMAPE), and mean absolute percentage error (MAPE) are used to evaluate the predictive performance of the model.

7. The method for rapid prediction of the viral load rate of rice spur dwarf disease according to claim 1, characterized in that, The steps also include: judging the risk of transplanting rice seedlings based on the predicted virus infection rate; when the predicted virus infection rate is less than 20%, it is judged to be suitable for transplanting; when the predicted virus infection rate reaches or exceeds 20%, it is judged to have a high risk of disease and should not be transplanted directly.

8. The method for rapid prediction of the viral load rate of rice spur dwarf disease according to claim 1, characterized in that: The method is used for the detection of rice samples during the seedling, tillering, heading, grain-filling, or maturity stages, and is applicable to rice field samples from different geographical regions.

9. The method for rapid prediction of the viral load rate of rice spur dwarf disease according to claim 1, characterized in that: The predicted RRSV carrier rate of the test field or seedling population can be obtained by a single pooled RT-qPCR test.

10. A method for rapid prediction of the viral load rate of rice spur dwarf disease according to claim 1, characterized in that: Random forest prediction models were used to process pooled RT-qPCR assays. -ΔCt The nonlinear relationship between the value and the actual virus-carrying rate of the rice population.