Traditional Chinese medicine planting soil quality monitoring method and system based on big data analysis

By generating soil optimization schemes through big data analysis and predictive models, the problem of refined management of soil quality monitoring in Chinese medicinal herb cultivation has been solved, thereby improving the quality and yield of Chinese medicinal herbs.

CN120707324BActive Publication Date: 2025-12-12GUANGXI ZHUANG AUTONOMOUS REGION INST OF PROD QUALITY INSPECTION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510914340.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-03
Publication Date
2025-12-12
Estimated Expiration
2045-07-03

AI Technical Summary

Technical Problem

Traditional soil quality monitoring methods cannot meet the needs of refined management in Chinese medicinal herb cultivation, especially since Chinese medicinal herbs have long growth cycles and different soil requirements at different growth stages.

Method used

By using big data analytics, we collect and preprocess data related to the cultivation of Chinese medicinal herbs, define the scope of key data, generate multiple data input combinations, and use predictive models and genetic algorithms to generate soil optimization schemes to ensure that Chinese medicinal herbs are in optimal condition at each growth stage.

Benefits of technology

It has improved the quality and yield of Chinese medicinal herbs, enabled precise management of soil conditions, and met the multi-stage needs of Chinese medicinal herb cultivation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707324B_ABST
    Figure CN120707324B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of traditional Chinese medicine planting, and discloses a traditional Chinese medicine planting soil quality monitoring method and system based on big data analysis. The method comprises the following steps: collecting related data of traditional Chinese medicine planting, the related data comprising historical soil data, planting history data and traditional Chinese medicine growth data, pre-processing the related data, taking the pre-processed related data as learning data to train a prediction model; obtaining stage related data of a first growth stage of traditional Chinese medicine, and demarcating a key data range based on first data in the stage related data; generating multiple data input combinations based on the key data range, and generating a soil optimization scheme based on the multiple data input combinations and the prediction model; and judging whether the soil optimization schemes of all growth stages of traditional Chinese medicine have been generated, and if not, generating the soil optimization schemes of all growth stages of traditional Chinese medicine by using the above method. The growth soil optimization scheme improves the quality and yield of traditional Chinese medicine.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of traditional Chinese medicine planting, and in particular to a traditional Chinese medicine planting soil quality monitoring method and system based on big data analysis. BACKGROUND

[0002] With the continuous development of traditional Chinese medicine planting industry, the influence of soil quality on the quality and yield of traditional Chinese medicine is increasingly valued. Traditional soil quality monitoring methods mainly rely on manual sampling and laboratory analysis, which is not only time-consuming and laborious, but also cannot monitor the changes of soil quality in real time.

[0003] A similar prior art is Chinese patent application No. CN119204843A, which provides a soil environment quality monitoring system and method based on multi-source data, comprising: collecting data of several pollutants, including heavy metals, pesticide residues and organic matter pollution; preprocessing the collected data to form a soil pollution data set; calculating the single pollution index of each pollutant, grading the soil environment quality according to the calculated comprehensive pollution index through the Nemerow comprehensive pollution index; using the random forest algorithm to assign weights to obtain the land use intensity factor of the region; based on the multilayer perceptron model, input factory discharge data, pollutant treatment data and soil pollution coefficient to obtain the industrial activity factor of the region; calculate the urbanization rate of the region to obtain the human activity impact factor; use the Naive Bayes algorithm for different levels of hierarchical management, and according to the Bayes theorem, assign the management level according to the result with the maximum probability.

[0004] A similar prior art is Chinese patent application No. CN109102421A, which provides an intelligent and reliable monitoring system for farmland soil quality, comprising: a parameter acquisition module for acquiring soil quality parameters reflecting farmland environmental conditions and sending the acquired soil quality parameters to a preprocessing module; the preprocessing module is configured to preprocess the received soil quality parameters and send them to a data management module for storage; the data management module is configured to manage the stored data; the comparison module is configured to compare the soil quality parameters with the set safety threshold and output the comparison result; the alarm module is configured to receive the comparison result and output alarm information to the set user terminal when the soil quality parameter is greater than the set safety threshold.

[0005] However, due to the long growth cycle of traditional Chinese medicine, different growth stages have different requirements for soil conditions, and the above two documents use a single soil quality evaluation method, which is difficult to meet the fine management needs of traditional Chinese medicine planting. Therefore, the application provides a traditional Chinese medicine planting soil quality monitoring method and system based on big data analysis. SUMMARY

[0006] The application provides a traditional Chinese medicine planting soil quality monitoring method and system based on big data analysis. By monitoring the soil data of each growth stage of traditional Chinese medicine during growth, a corresponding soil optimization scheme is generated, so that the soil state is always suitable for the growth of traditional Chinese medicine, achieving the purpose of improving the quality and yield of traditional Chinese medicine.

[0007] In a first aspect, the application provides a traditional Chinese medicine planting soil quality monitoring method based on big data analysis, which comprises:

[0008] Step S1, collect related data of traditional Chinese medicine planting, the related data including historical soil data, planting history data and traditional Chinese medicine growth data, pre-process the related data, and use the pre-processed related data as learning data to train a prediction model;

[0009] Step S2, obtain stage-related data of the first growth stage of traditional Chinese medicine, and divide a key data range based on first data in the stage-related data;

[0010] Step S3, generate multiple data input combinations based on the key data range, and generate a soil optimization scheme based on the multiple data input combinations and the prediction model;

[0011] Step S4, determine whether a soil optimization scheme for all growth stages of traditional Chinese medicine has been generated, if yes, end the step, otherwise, repeat steps S2-S3 until a soil optimization scheme for all growth stages is generated.

[0012] In combination with the first aspect, in a first implementation manner of the first aspect of the application, the key data range is divided based on the first data in the stage-related data, which comprises:

[0013] Step S21, preliminarily divide the value range of each data parameter in the first data into several data intervals, and based on each group of sample data in the first data, correspond each sample data to the data interval combination to which it belongs;

[0014] Step S22, count the number of sample data contained in each data interval, and mark the data interval combination with a sample data number greater than zero as an effective data interval;

[0015] Step S23, obtain a key interval based on the effective data interval, and combine the key interval and the effective data interval to generate a key data range.

[0016] In combination with the first aspect, in a second implementation manner of the first aspect of the application, the key interval is obtained based on the effective data interval, which comprises:

[0017] Two first valid intervals are obtained, the two first valid intervals satisfy two data intervals belonging to the same data interval in any one data parameter dimension and belonging to different data intervals in other data parameter dimensions, in the direction of the coincident data dimension, a data interval between the two first valid intervals is obtained as a first data interval, the first data interval is taken as a first key interval, and a second key interval is obtained based on the first key interval and the valid data interval.

[0018] In a third implementation manner of the first aspect, the second key interval is obtained based on the first key interval and the valid data interval, including:

[0019] Two second data intervals satisfying the first condition and the second condition are obtained, the first condition is that one first key interval and one valid data interval are included, and the second condition is that the data intervals are the same in any one data parameter dimension and different in other data parameter dimensions, a data interval between the two second data intervals in the data parameter direction with the same data interval is taken as a second key interval.

[0020] In a fourth implementation manner of the first aspect, the plurality of data input combinations are generated based on the key data range, including:

[0021] Each data parameter in the key data range is evenly divided into a plurality of small data intervals based on a preset interval, the small data intervals of the data parameters are combined into a multi-dimensional data grid, a center value of each small data interval is taken as an initial data sampling point, and an additional data sampling point is randomly generated in each data interval, excellent soil data corresponding to a growth stage with good growth is selected from historical soil data, and the initial data sampling point, the additional data sampling point and the excellent soil data are taken as the plurality of data input combinations.

[0022] In a fifth implementation manner of the first aspect, the soil optimization scheme is generated based on the plurality of data input combinations and the prediction model, including:

[0023] The plurality of data input combinations are input into the prediction model, and a corresponding prediction result is output by the prediction model, a first prediction result with a prediction accuracy greater than a preset second threshold is obtained based on a prediction accuracy corresponding to the prediction result, a plurality of second prediction results are obtained from the first prediction result, a plurality of first data input combinations corresponding to the second prediction results are obtained, a plurality of new data input combinations are generated based on the plurality of first data input combinations using a genetic algorithm, and the soil optimization scheme is generated based on the new data input combinations and the prediction model.

[0024] In a sixth implementation manner of the first aspect, the soil optimization scheme is generated based on the new data input combinations and the prediction model, including:

[0025] The plurality of new data combinations are input into the prediction model to obtain corresponding prediction results, a third prediction result with a prediction accuracy greater than a second threshold is obtained based on all the prediction results, an optimal prediction result is obtained from the second prediction result and the third prediction result, and an optimal input data combination corresponding to the optimal prediction result is obtained based on the optimal prediction result; a data difference is obtained based on the optimal input data combination and the real-time soil data, and a soil optimization scheme is generated based on the data difference.

[0026] In combination with the first aspect, in a seventh implementation manner of the first aspect of the present application, the training of the prediction model comprises:

[0027] The first data combination in the learning data is obtained, and in the process of training the prediction model, the first data in the first data combination is randomly selected and randomly deleted from the learning data until the training of the prediction model is completed.

[0028] In a second aspect, the present application provides a traditional Chinese medicine planting soil quality monitoring system based on big data analysis, which comprises:

[0029] The model training module is configured to collect relevant data of traditional Chinese medicine planting, the relevant data comprising historical soil data, planting history data and traditional Chinese medicine growth data, pre-process the relevant data, and train a prediction model using the pre-processed relevant data as learning data.

[0030] The data processing module is configured to obtain stage-related data of a first growth stage of traditional Chinese medicine, and to determine a key data range based on first data in the stage-related data.

[0031] The soil optimization module is configured to generate a plurality of data input combinations based on the key data range, to generate a soil optimization scheme based on the plurality of data input combinations and the prediction model, and to determine whether the soil optimization scheme for all growth stages of traditional Chinese medicine has been generated; if yes, the step is ended; otherwise, the data processing module and the soil optimization module are repeatedly executed until the soil optimization scheme for all growth stages is generated.

[0032] Compared with the prior art, the present application has at least the following advantages:

[0033] In the technical scheme provided in the application, the relevant data of traditional Chinese medicine planting are collected, and the data are preprocessed to ensure the integrity and consistency of the data, thereby improving the data quality and providing a solid foundation for subsequent model training; the search range is narrowed by delimiting the key data range, thereby facilitating the subsequent faster finding of optimal soil data; the genetic algorithm and other optimization algorithms are used to generate soil optimization schemes based on multiple data input combinations, which helps to quickly find the optimal soil condition combination and improves the reliability and practicality of the soil optimization scheme; the soil optimization scheme of each growth stage is generated, so that the traditional Chinese medicine is in the optimal state at each growth stage, thereby achieving the purpose of finally improving the quality and yield of the traditional Chinese medicine. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor based on these drawings.

[0035] Figure 1 is an embodiment schematic diagram of the traditional Chinese medicine planting soil quality monitoring method based on big data analysis in the embodiment of the present application;

[0036] Figure 2 is a schematic diagram of the data division interval in the embodiment of the present application;

[0037] Figure 3 is a schematic diagram of the key data range divided into multiple small data intervals in the embodiment of the present application;

[0038] Figure 4 is an embodiment schematic diagram of the traditional Chinese medicine planting soil quality monitoring system based on big data analysis in the embodiment of the present application. DETAILED DESCRIPTION

[0039] The embodiments of the present application provide a traditional Chinese medicine planting soil quality monitoring method and system based on big data analysis. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0040] For ease of understanding, the specific process of the embodiments of the present application is described below. Please refer to Figure 1 One embodiment of the traditional Chinese medicine planting soil quality monitoring method based on big data analysis in the embodiments of the present application includes:

[0041] Step S1, collect related data of traditional Chinese medicine planting, the related data including historical soil data, planting history data and traditional Chinese medicine growth data, pre-process the related data, and use the pre-processed related data as learning data to train a prediction model.

[0042] Specifically, the historical soil data includes various soil parameter data, such as PH value, organic matter content, and nitrogen, phosphorus and potassium content, etc., the planting history data includes fertilization record data, irrigation record data, etc., the traditional Chinese medicine growth data includes traditional Chinese medicine growth stage, corresponding traditional Chinese medicine growth condition, final traditional Chinese medicinal material yield, effective component content, etc., the pre-processing of the related data includes numerical conversion and data normalization processing, etc., the related data is used as learning data to train a prediction model, and the prediction model is used to predict the mapping relationship between soil parameters and traditional Chinese medicine growth conditions, such as making the prediction model learn that under what combination of soil data and planting history data can make the traditional Chinese medicine grow well in each stage, and finally make the quality of the traditional Chinese medicine good and the yield high.

[0043] Step S2, obtain stage-related data of a first growth stage of traditional Chinese medicine, and delimit a key data range based on first data in the stage-related data.

[0044] Specifically, in order to improve the quality and yield of traditional Chinese medicine, the optimal soil optimization scheme for each growth stage of traditional Chinese medicine is obtained. First, the stage-related data of the first growth stage of traditional Chinese medicine is obtained. Generally, the first growth stage is the seeding period, so the historical stage-related data of the seeding period is obtained, including historical soil data and growth record data. The first data refers to the historical soil data in the stage-related data. The key data range is determined based on the first data, that is, the range where the optimal soil data is selected based on the historical data set. The method of determining the key data range will be explained in detail later. By narrowing the search range, it is convenient to find the optimal soil data more quickly. Based on the optimal soil data, the optimal soil optimization scheme is selected, so that the traditional Chinese medicine is in the optimal state at each growth stage, and finally the purpose of improving the quality and yield of traditional Chinese medicine is achieved.

[0045] Step S3, generating a plurality of data input combinations based on the key data range, and generating a soil optimization scheme based on the plurality of data input combinations and the prediction model.

[0046] Specifically, a plurality of data input combinations are generated within the key data range, and the specific generation method will be explained in detail later. The generated plurality of data input combinations are respectively input into the prediction model to obtain the growth of traditional Chinese medicine at the corresponding growth stage. The first data input combination with the best growth of traditional Chinese medicine is obtained. The soil optimization scheme is generated based on the difference between the real-time soil data and the first data input combination.

[0047] Step S4, determining whether the soil optimization schemes of all growth stages of traditional Chinese medicine have been generated. If yes, end this step, otherwise, repeat steps S2-S3 until the soil optimization schemes of all growth stages are generated.

[0048] Specifically, steps S2-S3 are repeated to generate soil optimization schemes for all growth stages.

[0049] Through the cooperation between the above steps, the soil optimization schemes for each growth stage are generated, so that the traditional Chinese medicine is in the optimal state at each growth stage, and the purpose of finally improving the quality and yield of traditional Chinese medicine is achieved.

[0050] In a specific embodiment, the key data range is determined based on the first data in the stage-related data, specifically including the following steps:

[0051] Step S21, the value range of each data parameter in the first data is preliminarily divided into a plurality of data intervals. Based on each group of sample data in the first data, each sample data is corresponded to the data interval combination to which it belongs.

[0052] Step S22, count the number of sample data contained in each data interval, and mark the data interval combination with a sample data number greater than zero as an effective data interval.

[0053] Step S23, obtain a key interval based on the effective data interval, and generate a key data range by combining the key interval and the effective data interval.

[0054] Specifically, the first data refers to historical soil data, and the soil data usually includes multiple data parameters such as pH value, organic matter content, and nitrogen, phosphorus, and potassium content. In order to clearly explain the embodiment, it is assumed that the soil data has two data parameters, pH value and organic matter content. The value range of the data parameter is divided into several data intervals. Taking the pH value as an example, it is assumed that the value range of the pH value is 4.5-6.5, and the value range of the organic matter content is 1.6%-2.4%. The pH value is divided into four intervals of 4.5-5.0, 5.0-5.5, 5.5-6.0, and 6.0-6.5 according to an interval of 0.5. The organic matter is divided into four intervals of 1.6%-1.8%, 1.8%-2.0%, 2.0%-2.2%, and 2.2%-2.4% according to an interval of 0.2%. Each data parameter is divided into multiple different data intervals. Each data parameter is corresponded to the data interval to which it belongs. For example, a group of soil data has a pH value of 6.1 and an organic matter content of 2.3%. This group of soil data belongs to the interval of pH: 6.0-6.5 and organic matter: 2.2%-2.4%. The number of samples in each data interval combination is counted. The sample number refers to the number of all soil data combinations in the data interval. The data interval combination with a sample number greater than or equal to zero is marked as an effective data interval. For example, Figure 2 As shown in the figure, A11, A14, A42, and A43 are effective data intervals, Figure 2 Each small black dot in the figure represents a group of soil data. There may be soil data in the effective data interval that can improve the quality and yield of traditional Chinese medicine. Therefore, the effective data interval needs to be focused on. In addition to the effective data interval, there may also be better soil data combinations in the intervals adjacent or similar to the effective data interval that can improve the quality and yield of traditional Chinese medicine. Therefore, the key interval is obtained based on the effective data interval. The process of obtaining the key interval will be explained in detail later. The key data range is generated by combining the key interval and the effective data interval, which facilitates the selection of reliable soil data based on the key data range.

[0055] The above method narrows the search range by obtaining the key data range, and can more quickly find relatively optimal soil data. Multiple model parameters are generated based on the key data range in the subsequent process to obtain the optimal soil data.

[0056] In one specific embodiment, obtaining the key interval based on the effective data interval specifically includes the following steps:

[0057] Two first valid intervals are obtained. The two first valid intervals are two data intervals that are the same data intervals in any data parameter dimension and different data intervals in other data parameter dimensions. In the direction of the overlapping data dimension, the data interval between the two first valid intervals is obtained as the first data interval. The first data interval is used as the first key interval. The second key interval is obtained based on the first key interval and the valid data interval.

[0058] Specifically, such as Figure 2 As shown, in terms of the data parameter dimension of organic matter content, A11 and A14 are two valid data intervals with the same data interval in this dimension, but different data intervals in terms of the data parameter dimension of pH value. Therefore, A11 and A14 are two first valid intervals. In terms of the data parameter dimension of organic matter content, the data intervals A12 and A13 between A11 and A14 are obtained, and A12 and A13 are taken as the first data intervals. The first data intervals are the first key intervals. Using the above method to obtain the first key intervals, all the first key intervals are obtained. Soil data exists in both of the first valid intervals, so the data intervals between the first valid intervals may also contain soil data. It is possible that the coverage is not comprehensive due to the small amount of historical soil data. Therefore, based on the above method, the key intervals are obtained to expand the range of the key intervals and improve the predictive continuity of the prediction model in the data edge area.

[0059] After obtaining the first key interval, the second key interval is obtained based on the first key interval and the effective data interval. The specific method for obtaining the second key interval will be explained in detail later.

[0060] In one specific embodiment, obtaining the second key interval based on the first key interval and the effective data interval specifically includes the following steps:

[0061] Obtain two second data intervals that simultaneously satisfy the first and second conditions. The first condition is that it includes a first key interval and a valid data interval. The second condition is that the data intervals are the same in any one data parameter dimension and different in other data parameter dimensions. The data interval between the two second data intervals in the same data parameter direction is taken as the second key interval.

[0062] Specifically, two second data intervals satisfying both the first condition and the second condition are obtained, the first condition is that the two second data intervals include a first key interval and a valid data interval, and the second condition is that the two second data intervals are the same in any data parameter dimension and different in other data parameter dimensions, such as Figure 2 As shown in the figure, A42 is a valid data interval, A12 is a first key interval, and A42 and A12 belong to the same data interval in the pH value dimension, so A42 and A12 are two second data intervals satisfying both the first condition and the second condition. In the pH value dimension, the data intervals A32 and A22 between A42 and A12 are obtained as second key intervals.

[0063] Repeat this step to obtain all second key intervals.

[0064] That is, all second key intervals are obtained using the above method.

[0065] As shown in the figure, A33 and A23 are also second key intervals between the two second data intervals A43 and A13. Figure 2

[0066] All valid data intervals, all first key intervals, and all second key intervals are combined to generate key intervals.

[0067] Through the above steps, reliable key intervals are extracted from the chaotic valid data intervals, which facilitates subsequent acquisition of optimal soil data suitable for each growth stage of traditional Chinese medicine in the key intervals. Based on the optimal soil data, a soil optimization scheme is generated, and based on the soil optimization scheme, precise agricultural measures such as fertilization, irrigation, or body modification measures are taken at each stage of traditional Chinese medicine growth, improving the quality and yield of traditional Chinese medicine.

[0068] In a specific embodiment, a plurality of data input combinations are generated based on the key data range, specifically including the following steps:

[0069] Based on the preset interval, each data parameter in the key data range is evenly divided into a plurality of small data intervals, the small data intervals of each data parameter are combined into a multi-dimensional data grid, the center value of each small data interval is taken as an initial data sampling point, and additional data sampling points are randomly generated in each data interval. Good soil data with good growth conditions for the corresponding growth stage is selected from historical soil data, and the initial data sampling points, additional data sampling points, and good soil data are used as a plurality of data input combinations.

[0070] Specifically, the value range of each data parameter is obtained based on the key data range, and the value range of each data parameter is divided into a plurality of small data intervals. Figure 3 ​For example, the key data range consists of A42, A43, A32, A33, A22, A23, A11, A12, A13 and A14, each data parameter in the key data range is divided into multiple small data intervals based on a preset interval, such as Figure 3 A42 is divided into 4 smaller small data intervals, the center value of each small data interval is taken as an initial data sampling point, the center value refers to the data parameter combination corresponding to the center point of the small data interval, in order to make up for the shortage of grid points, additional data sampling points are also randomly generated in each data interval, such as Figure 3 The positions of the hollow circles in the above figure are randomly generated additional data sampling points, and the soil data for growth condition quantification is also selected from historical soil data, if the soil data corresponding to each small black dot in A14 is excellent soil data, the soil data corresponding to each small black dot in A14 is also added to the data input combination, through the above steps, multiple different data input combinations are obtained, which cover all possible soil combinations, and potential optimal solutions are avoided.

[0071] In a specific embodiment, a soil optimization scheme is generated based on multiple data input combinations and a prediction model, specifically including the following steps:

[0072] The multiple data input combinations are input into the prediction model, and the prediction model outputs corresponding prediction results, based on the prediction accuracy corresponding to the prediction results, a first prediction result with a prediction accuracy greater than a preset second threshold is obtained, a plurality of second prediction results are obtained from the first prediction result, a plurality of first data input combinations corresponding to the second prediction results are obtained, a plurality of new data input combinations are generated based on the plurality of first data input combinations using a genetic algorithm, and a soil optimization scheme is generated based on the new data input combinations and the prediction model.

[0073] Specifically, in order to obtain the soil optimization scheme, first, the generated multiple data input combinations are input into the prediction model to obtain prediction results, the multiple data input combinations include multiple different soil data combinations, and the prediction results refer to the growth conditions of traditional Chinese medicine corresponding to the growth stage. Different data input combinations correspond to different prediction accuracies based on the data input combination. For example, if the data input combination belongs to the actual historical data with high data density, the prediction accuracy is relatively high, and vice versa. The prediction accuracy is relatively low. The prediction accuracy corresponding to the output prediction result is obtained. The first prediction result with a prediction accuracy greater than a preset second threshold is obtained. The first prediction result with high accuracy is screened out. The first prediction result is obtained. The first prediction result is obtained. The second prediction result refers to the top several prediction results with good growth conditions. Based on the second prediction result, a plurality of first data input combinations corresponding to the second prediction result are obtained. The first data input combination refers to the soil data combination corresponding to the second prediction result. A genetic algorithm is used to generate a plurality of new data input combinations based on the plurality of first data input combinations. That is, the high-quality soil data is recombined to generate a plurality of new data combinations. The generated new data combinations can be better than the first data input combinations and can make the traditional Chinese medicine grow better. Based on the new data input combination and the prediction model, a soil optimization scheme is generated. The specific process of generating the soil optimization scheme will be explained later.

[0074] In a specific embodiment, the soil optimization scheme is generated based on the new data input combination and the prediction model, and specifically includes the following steps:

[0075] The plurality of new data combinations are input into the prediction model to obtain corresponding prediction results. The prediction accuracy greater than the second threshold is obtained based on all the prediction results. The third prediction result is obtained. The optimal prediction result is obtained from the second prediction result and the third prediction result. The optimal input data combination corresponding to the optimal prediction result is obtained based on the optimal input data combination and the real-time soil data. The data difference is obtained based on the data difference. The soil optimization scheme is generated based on the data difference.

[0076] Specifically, a third prediction result of good growth condition is obtained, a second prediction result is also a prediction result of good growth condition, an optimal prediction result of the best growth condition is obtained from the second prediction result and the third prediction result, an optimal input data combination corresponding to the optimal prediction result is obtained based on the optimal prediction result, the optimal input data combination refers to an optimal soil data combination that can make the traditional Chinese medicine in each growth stage, and a soil optimization scheme is generated based on a data difference between the optimal input data combination and collected real-time soil data. For example, the optimal input data combination is that the pH value is 6.5 and the organic matter content is 2.2%, the pH value of the real-time soil data is 6.8 and the organic matter content is 1.9%, and the difference between the optimal input data combination and the real-time soil data generates the soil optimization scheme. For example, in the preparation stage of planting traditional Chinese medicine, organic fertilizer is applied to increase the organic matter content of the soil, and lime, slaked lime or limestone powder is used to increase the alkalinity of the soil.

[0077] In an embodiment, the prediction model is trained, specifically including the following steps:

[0078] A first data combination in the learning data is obtained, in the process of training the prediction model, a first data in the first data combination is randomly selected, and the first data is randomly removed from the learning data until the training of the prediction model is completed.

[0079] Specifically, in order to improve the accuracy of the prediction model and avoid overfitting of the prediction model, a first data combination in the learning data is obtained, the first data combination refers to learning data with high frequency in the learning data, and a first data refers to part of the data randomly obtained from the first data combination. In the process of training the prediction model, if the frequency of some learning data is high, it may cause overfitting of the prediction model. Therefore, in order to avoid overfitting of the prediction model, the first data combination is randomly removed from the learning data in the process of training the prediction model.

[0080] In an embodiment, the prediction model is trained, and further includes the following steps:

[0081] A second data combination in the learning data is obtained, and supplementary data is generated based on the second data combination, and the supplementary data is added to the learning data in the process of training the prediction model.

[0082] Specifically, in order to improve the accuracy of the prediction model, a second data combination in the learning data is obtained, the second data combination refers to learning data with low frequency in the learning data, supplementary data is generated based on the second combination data, the supplementary data refers to data similar to the second combination data, and the supplementary data is added to the learning data to increase the learning data with low frequency and improve the prediction accuracy of the prediction model.

[0083] The method for monitoring the quality of the soil for planting traditional Chinese medicine based on big data analysis in the embodiments of the application is described above, and the system for monitoring the quality of the soil for planting traditional Chinese medicine based on big data analysis in the embodiments of the application is described below. Please refer to Figure 4 An embodiment of the system for monitoring the quality of the soil for planting traditional Chinese medicine based on big data analysis in the embodiments of the application includes:

[0084] The model training module is configured to collect relevant data of traditional Chinese medicine planting, the relevant data including historical soil data, planting history data and traditional Chinese medicine growth data, pre-process the relevant data, and use the pre-processed relevant data as learning data to train a prediction model.

[0085] The data processing module is configured to obtain stage-related data of a first growth stage of traditional Chinese medicine, and determine a key data range based on first data in the stage-related data.

[0086] The soil optimization module is configured to generate a plurality of data input combinations based on the key data range, generate a soil optimization scheme based on the plurality of data input combinations and the prediction model, determine whether soil optimization schemes for all growth stages of traditional Chinese medicine have been generated, end the step if yes, and otherwise, repeatedly execute the data processing module and the soil optimization module until soil optimization schemes for all growth stages are generated.

[0087] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, the system and the unit described above can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein.

[0088] The integrated unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0089] The above-described embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for monitoring soil quality for planting traditional Chinese medicine based on big data analysis, characterized in that, The method comprises: Step S1, collecting relevant data of Chinese medicine planting, the relevant data including historical soil data, planting history data and Chinese medicine growth data, preprocessing the relevant data, and taking the preprocessed relevant data as learning data to train a prediction model; Step S2, obtaining stage-related data of the first growth stage of Chinese medicine, and determining a key data range based on first data in the stage-related data; Wherein, based on the first data in the stage-related data, the key data range is determined, which comprises: step S21, the value range of each data parameter in the first data is preliminarily divided into several data intervals, and each sample data is corresponded to the data interval group based on each group of sample data in the first data; Step S22, the number of sample data contained in each data interval is counted, and the data interval combination with the number of sample data greater than zero is marked as an effective data interval; Step S23, obtaining a key interval based on the effective data interval, and combining the key interval and the effective data interval to generate a key data range; Wherein, based on the effective data interval, the key interval is obtained, which comprises: obtaining two first effective intervals, the two first effective intervals satisfy that the two data intervals belonging to the same data interval in any one data parameter dimension belong to different data intervals in other data parameter dimensions, in the direction of the overlapping data dimension, the data interval between the two first effective intervals is obtained as a first data interval, the first data interval is taken as a first key interval, and a second key interval is obtained based on the first key interval and the effective data interval; Wherein, based on the first key interval and the effective data interval, the second key interval is obtained, which comprises: obtaining two second data intervals that satisfy the first condition and the second condition at the same time, the first condition is that it includes a first key interval and an effective data interval, and the second condition is that the data intervals are the same in any one data parameter dimension and different in other data parameter dimensions, the data interval between the two second data intervals in the data parameter direction with the same data interval is taken as a second key interval; Step S3, generating a plurality of data input combinations based on the key data range, and generating a soil optimization scheme based on the plurality of data input combinations and the prediction model; Wherein, based on the key data range, the plurality of data input combinations are generated, which comprises: dividing each data parameter in the key data range into a plurality of small data intervals based on a preset interval, combining the small data intervals of each data parameter into a multi-dimensional data grid, taking the center value of each small data interval as an initial data sampling point, and randomly generating additional data sampling points in each data interval, selecting good soil data with good growth conditions in the corresponding growth stage from the historical soil data, and taking the initial data sampling points, the additional data sampling points and the good soil data as the plurality of data input combinations; Step S4, determining whether the soil optimization scheme of all growth stages of Chinese medicine has been generated, if yes, ending the step, otherwise, repeating steps S2-S3 until the soil optimization scheme of all growth stages is generated.

2. The method of claim 1, wherein, Based on the plurality of data input combinations and the prediction model, the soil optimization scheme is generated, which comprises: The plurality of data input combinations are input into the prediction model, and the prediction model outputs corresponding prediction results. Based on the prediction accuracy corresponding to the prediction results, a first prediction result with a prediction accuracy greater than a preset second threshold is obtained. A plurality of second prediction results are obtained from the first prediction results. Based on the second prediction results, a plurality of first data input combinations are obtained. A genetic algorithm is used to generate a plurality of new data input combinations based on the plurality of first data input combinations. A soil optimization scheme is generated based on the new data input combinations and the prediction model.

3. The method of claim 2, wherein, The soil optimization scheme is generated based on the new data input combinations and the prediction model, comprising: The plurality of new data combinations are input into the prediction model to obtain corresponding prediction results. Based on all the prediction results, a third prediction result with a prediction accuracy greater than a second threshold is obtained. From the second prediction result and the third prediction result, an optimal prediction result is obtained. Based on the optimal prediction result, an optimal input data combination is obtained. Based on the optimal input data combination and real-time soil data, a data difference is obtained. Based on the data difference, a soil optimization scheme is generated.

4. The method of claim 1, wherein, The prediction model is trained, comprising: The first data combination in the learning data is obtained. In the process of training the prediction model, the first data in the first data combination is randomly selected. The first data is randomly deleted from the learning data until the prediction model is trained.

5. A traditional Chinese medicine planting soil quality monitoring system based on big data analysis, used to realize the traditional Chinese medicine planting soil quality monitoring method based on big data analysis according to any one of claims 1-4, characterized in that, The system comprises: The model training module is used to collect relevant data of traditional Chinese medicine planting, including historical soil data, planting history data and traditional Chinese medicine growth data. The relevant data is preprocessed, and the preprocessed relevant data is used as learning data to train the prediction model. The data processing module is configured to obtain stage-related data of a first growth stage of the traditional Chinese medicine, and to demarcate a key data range based on first data in the stage-related data. Demarcating the key data range based on the first data in the stage-related data includes: step S21, preliminarily dividing a value range of each data parameter in the first data into a plurality of data intervals, and based on each group of sample data in the first data, combining each sample data to a data interval group to which the sample data belongs; step S22, counting a number of sample data contained in each data interval, and marking a data interval group in which the number of sample data is greater than zero as an effective data interval; and step S23, obtaining a key interval based on the effective data interval, and combining the key interval and the effective data interval to generate the key data range. Obtaining the key interval based on the effective data interval includes: obtaining two first effective intervals, the two first effective intervals satisfying that, in any one data parameter dimension, two data intervals belonging to the same data interval in other data parameter dimensions, and in the direction of the coinciding data dimension, obtaining a data interval between the two first effective intervals as a first data interval, taking the first data interval as a first key interval, and obtaining a second key interval based on the first key interval and the effective data interval. Obtaining the second key interval based on the first key interval and the effective data interval includes: obtaining two second data intervals that satisfy a first condition and a second condition at the same time, the first condition being that one first key interval and one effective data interval are included, and the second condition being that the data intervals are the same in any one data parameter dimension and different in other data parameter dimensions, and taking a data interval between the two second data intervals in the data parameter direction in which the data intervals are the same as a second key interval. The soil optimization module is configured to generate a plurality of data input combinations based on the key data range, and to generate a soil optimization scheme based on the plurality of data input combinations and a prediction model. Generating the plurality of data input combinations based on the key data range includes: dividing each data parameter in the key data range into a plurality of small data intervals based on a preset interval, combining the small data intervals of each data parameter into a multi-dimensional data grid, taking a center value of each small data interval as an initial data sampling point, randomly generating additional data sampling points in each data interval, selecting excellent soil data corresponding to a good growth condition of the growth stage from historical soil data, taking the initial data sampling points, the additional data sampling points, and the excellent soil data as the plurality of data input combinations, determining whether soil optimization schemes for all growth stages of the traditional Chinese medicine have been generated, and if yes, ending the step, otherwise, repeatedly executing the data processing module and the soil optimization module until the soil optimization schemes for all growth stages are generated.

Citation Information

Patent Citations

  • An intelligent and reliable monitoring system for farmland soil quality

    CN109102421A

  • Soil environment quality monitoring system and method based on multi-source data

    CN119204843A

  • Nonlinear combination prediction method and system for cultivation water quality

    CN108038571A

  • Method and device for optimizing hyper-parameters in learning model algorithm and storage medium

    CN119047603A