Training method, prediction method and system of prediction model for processing degree of traditional Chinese medicinal materials

By training a prediction model for the degree of preparation of Chinese medicinal materials, using sample characteristic peak data to achieve accurate prediction of the degree of preparation, the problem of the inability to accurately judge the degree of preparation of Chinese medicinal materials in the prior art is solved, and the quality and efficacy of Chinese medicinal materials are ensured.

CN120072111APending Publication Date: 2025-05-30SHANGHAI INST OF PHARMA IND CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510125970.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art cannot accurately judge the degree of preparation of Chinese medicinal materials, resulting in the quality and efficacy of Chinese medicinal materials being unable to meet the requirements.

Method used

By obtaining the preparation degree data and mass spectrometry data of Chinese medicinal materials in the sample, extracting the sample characteristic peak data, and training a prediction model of the preparation degree of Chinese medicinal materials to achieve accurate prediction of the preparation degree.

Benefits of technology

Accurate prediction of the degree of preparation of Chinese medicinal materials has been achieved, quality control standards have been formed, and processing time has been guided to control, ensuring the quality and efficacy of Chinese medicinal materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072111A_ABST
    Figure CN120072111A_ABST
Patent Text Reader

Abstract

The invention provides a training method, a prediction method and a system of a prediction model for the processing degree of traditional Chinese medicinal materials. The training method comprises the following steps: acquiring a plurality of groups of sample processing degree data and sample mass spectrum data of sample traditional Chinese medicinal materials; extracting sample characteristic peak data from the sample mass spectrum data; and taking each group of sample characteristic peak data of the sample traditional Chinese medicinal materials as input, taking corresponding each group of sample processing degree data as output, and training to obtain a target prediction model of the processing degree of the traditional Chinese medicinal materials. The sample characteristic peak data is extracted through the sample mass spectrum data of the sample traditional Chinese medicinal materials, and then the target prediction model of the processing degree of the traditional Chinese medicinal materials is obtained through training, so that the effectiveness and reliability of the target prediction model are ensured. And taking the target characteristic peak data extracted from the target mass spectrum data of the target traditional Chinese medicinal material as the input of the target prediction model to output the target processing degree data of the target traditional Chinese medicinal material, so that the prediction of the processing degree of the traditional Chinese medicinal material is effectively and accurately realized, and the quality and efficacy of the traditional Chinese medicinal material are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of machine learning technology, and particularly to a training method, a prediction method, and a system for a prediction model of the processing degree of traditional Chinese medicinal materials. Background Art

[0002] Traditional Chinese medicinal materials have a long history of processing and numerous methods. The processing process involves complex biological transformations and chemical changes. After processing, the characteristic components or medicinal components of traditional Chinese medicine decoction pieces will change, and these changes directly affect the quality of traditional Chinese medicine decoction pieces. The 2020 edition of the Chinese Pharmacopoeia includes a total of 616 varieties of traditional Chinese medicinal materials and traditional Chinese medicine decoction pieces. Most traditional Chinese medicinal materials need to be processed, but there are no specific and clear processing process parameters. The existing processing methods do not clarify the detailed parameters of the processing time and frequency, resulting in uneven processing degrees and qualities of traditional Chinese medicinal materials sold on the market currently, and the medicinal effects cannot reach a stable consistency. This phenomenon not only affects the clinical efficacy of traditional Chinese medicine but also may pose potential risks to the health of patients.

[0003] The methods for evaluating the quality of traditional Chinese medicine are mainly divided into traditional empirical identification and modern analytical identification. With the development of modern analytical technologies, the methods and means for evaluating the quality of traditional Chinese medicine have become increasingly perfect, evolving from single-component detection to multi-component detection. However, due to the complex components and high integrity of traditional Chinese medicine, there are still certain objections as to whether this "component-only theory" can represent the quality of traditional Chinese medicine, and there are problems such as long experimental periods, complex operations, and high costs in practical applications.

[0004] Currently, traditional empirical identification is still an important way to evaluate the quality of traditional Chinese medicine. Experienced traditional Chinese medicine workers often judge the quality based on the shape, color, smell, and taste of traditional Chinese medicinal materials and decoction pieces. However, this subjective empirical judgment by humans often has uncertainties due to individual differences. In recent years, some scholars, based on traditional empirical identification methods, have quantified the color and smell information during the processing of traditional Chinese medicinal materials with the help of instruments such as electronic noses, electronic tongues, colorimeters, and color difference meters, and explored the dynamic change laws of color and smell information during the processing process, providing a new method for the quality evaluation of traditional Chinese medicinal materials.

[0005] The above methods for judging the processing degree have explained certain chemical and physical changes during the processing of traditional Chinese medicinal materials and can distinguish traditional Chinese medicinal material samples with known processing degrees. However, the traditional Chinese medicinal materials circulating on the market may be dried decoction piece products, resulting in certain losses of color and smell, making the results measured by instruments such as electronic noses, colorimeters, and color difference meters deviate greatly from the actual results. In addition, as the processing time increases, these appearance, smell, and marker contents tend to be stable, and it is impossible to determine the processing degree of traditional Chinese medicinal materials using physical and chemical methods, making it difficult to form quality control.

[0006] Therefore, the existing methods for judging the processing degree cannot accurately judge the processing degree of traditional Chinese medicinal materials, resulting in the quality and medicinal effects of traditional Chinese medicinal materials not meeting the requirements. Summary of the Invention

[0007] The technical problem to be solved by the present disclosure is to overcome the defect that the method for judging the processing degree of traditional Chinese medicinal materials in the prior art cannot accurately judge the processing degree of traditional Chinese medicinal materials, resulting in the quality and efficacy of traditional Chinese medicinal materials not meeting the requirements, etc., and to provide a training method, a prediction method and a system for a prediction model of the processing degree of traditional Chinese medicinal materials.

[0008] The present disclosure solves the above technical problem through the following technical solutions:

[0009] The present disclosure provides a training method for a prediction model of the processing degree of traditional Chinese medicinal materials, and the training method includes:

[0010] Obtaining a plurality of groups of sample processing degree data and sample mass spectrometry data of sample traditional Chinese medicinal materials;

[0011] Extracting sample characteristic peak data from the sample mass spectrometry data;

[0012] Wherein, the sample characteristic peak data is used to characterize a plurality of compounds in the sample traditional Chinese medicinal materials;

[0013] Using each group of the sample characteristic peak data of the sample traditional Chinese medicinal materials as an input and the corresponding group of the sample processing degree data as an output, training to obtain a target prediction model for the processing degree of traditional Chinese medicinal materials.

[0014] Preferably, the sample traditional Chinese medicinal materials include Rehmannia glutinosa;

[0015] And / or,

[0016] The sample processing degree data includes processing duration;

[0017] And / or,

[0018] The step of obtaining the sample processing degree data and the sample mass spectrometry data of the sample traditional Chinese medicinal materials includes:

[0019] Detecting a test solution of the sample traditional Chinese medicinal materials by using ultra-high performance liquid chromatography-quadrupole time-of-flight mass spectrometry technology to obtain the sample mass spectrometry data of the sample traditional Chinese medicinal materials;

[0020] And / or,

[0021] The step of extracting the sample characteristic peak data from the sample mass spectrometry data includes:

[0022] Performing alignment processing on the sample mass spectrometry data;

[0023] Extracting the sample characteristic peak data from the sample mass spectrometry data after alignment processing.

[0024] Preferably, the step of using each group of the sample characteristic peak data of the sample Chinese medicinal materials as input and the corresponding each group of the sample processing degree data as output to train a target prediction model for the processing degree of Chinese medicinal materials includes:

[0025] Obtain the correlation between each group of the sample processing degree data and each characteristic peak in the corresponding sample characteristic peak data;

[0026] Wherein, each characteristic peak is used to characterize a compound in the sample Chinese medicinal materials;

[0027] Based on the correlation, select several of the characteristic peaks to construct a characteristic data set;

[0028] Use each group of the characteristic data sets as input and the corresponding each group of the sample processing degree data as output to train a first target prediction model for the processing degree of Chinese medicinal materials;

[0029] And / or,

[0030] Preferably, the step of using each group of the sample characteristic peak data of the sample Chinese medicinal materials as input and the corresponding each group of the sample processing degree data as output to train a target prediction model for the processing degree of Chinese medicinal materials includes:

[0031] Obtain the projection variable importance value corresponding to each characteristic peak in each group of the sample characteristic peak data;

[0032] Wherein, each characteristic peak is used to characterize a compound in the sample Chinese medicinal materials;

[0033] Based on the projection variable importance value, select several of the characteristic peaks from the sample characteristic peak data to construct a characteristic data set;

[0034] Use each group of the characteristic data sets as input and the corresponding each group of the sample processing degree data as output to train a second target prediction model for the processing degree of Chinese medicinal materials.

[0035] Preferably, the step of using each group of the characteristic data sets as input and the corresponding each group of the sample processing degree data as output to train a first target prediction model / second target prediction model for the processing degree of Chinese medicinal materials includes:

[0036] Use each group of the characteristic data sets as input and the corresponding each group of the sample processing degree data as output to train an initial prediction model for the processing degree of Chinese medicinal materials;

[0037] Evaluate the initial prediction model to obtain an evaluation result;

[0038] Based on the evaluation results, the initial prediction model is adjusted to obtain the first target prediction model / the second target prediction model.

[0039] The present disclosure also provides a method for predicting the processing degree of traditional Chinese medicinal materials, and the prediction method includes:

[0040] Obtain the target mass spectrometry data of the target traditional Chinese medicinal materials;

[0041] Extract the target characteristic peak data from the target mass spectrometry data;

[0042] Wherein, the target characteristic peak data is used to characterize several compounds in the target traditional Chinese medicinal materials;

[0043] Input the target characteristic peak data into the target prediction model for the processing degree of traditional Chinese medicinal materials to output the target processing degree data of the target traditional Chinese medicinal materials;

[0044] Wherein, the target prediction model is obtained based on the training method of the prediction model for the processing degree of traditional Chinese medicinal materials described above.

[0045] The present disclosure also provides a training system for a prediction model of the processing degree of traditional Chinese medicinal materials, and the training system includes:

[0046] A sample data acquisition module, configured to acquire several groups of sample processing degree data and sample mass spectrometry data of sample traditional Chinese medicinal materials;

[0047] A sample peak extraction module, configured to extract sample characteristic peak data from the sample mass spectrometry data;

[0048] Wherein, the sample characteristic peak data is used to characterize several compounds in the sample traditional Chinese medicinal materials;

[0049] A model training module, configured to use each group of the sample characteristic peak data of the sample traditional Chinese medicinal materials as input and the corresponding each group of the sample processing degree data as output to train and obtain a target prediction model for the processing degree of traditional Chinese medicinal materials.

[0050] Preferably, the sample traditional Chinese medicinal materials include Rehmannia glutinosa;

[0051] And / or

[0052] The sample processing degree data includes the processing duration;

[0053] And / or

[0054] The sample data acquisition module is further configured to detect the test solution of the sample traditional Chinese medicinal materials by using ultra-high performance liquid chromatography - quadrupole time-of-flight mass spectrometry technology to obtain the sample mass spectrometry data of the sample traditional Chinese medicinal materials;

[0055] And / or

[0056] The sample peak extraction module includes:

[0057] An alignment processing unit for performing alignment processing on the sample mass spectrometry data;

[0058] A sample peak extraction unit for extracting the sample characteristic peak data from the aligned sample mass spectrometry data.

[0059] Preferably, the model training module includes:

[0060] A correlation acquisition unit for acquiring the correlation between each characteristic peak in each group of the sample processing degree data and the corresponding sample characteristic peak data;

[0061] wherein each of the characteristic peaks is used to characterize one of the compounds in the sample Chinese medicinal materials;

[0062] A first data set construction unit for constructing a characteristic data set by selecting a number of the characteristic peaks based on the correlation;

[0063] A first model training unit for training a first target prediction model of the processing degree of Chinese medicinal materials by taking each group of the characteristic data sets as input and the corresponding each group of the sample processing degree data as output;

[0064] and / or,

[0065] The model training module includes:

[0066] An importance value acquisition unit for acquiring the projection variable importance value corresponding to each characteristic peak in each group of the sample characteristic peak data;

[0067] wherein each of the characteristic peaks is used to characterize one of the compounds in the sample Chinese medicinal materials;

[0068] A second data set construction unit for constructing a characteristic data set by selecting a number of the characteristic peaks from the sample characteristic peak data based on the projection variable importance value;

[0069] A second model training unit for training a second target prediction model of the processing degree of Chinese medicinal materials by taking each group of the characteristic data sets as input and the corresponding each group of the sample processing degree data as output.

[0070] Preferably, the first model training unit / the second model training unit includes:

[0071] An initial model acquisition subunit for training an initial prediction model of the processing degree of Chinese medicinal materials by taking each group of the characteristic data sets as input and the corresponding each group of the sample processing degree data as output.

[0072] An evaluation result acquisition subunit, configured to evaluate the initial prediction model to obtain an evaluation result;

[0073] A target model acquisition subunit, configured to adjust the initial prediction model based on the evaluation result to obtain the first target prediction model / the second target prediction model.

[0074] The present disclosure also provides a prediction system for the processing degree of traditional Chinese medicine materials, the prediction system includes:

[0075] A target data acquisition module, configured to acquire target mass spectrometry data of target traditional Chinese medicine materials;

[0076] A target peak extraction module, configured to extract target characteristic peak data from the target mass spectrometry data;

[0077] Wherein, the target characteristic peak data is used to characterize several compounds in the target traditional Chinese medicine materials;

[0078] A model output module, configured to input the target characteristic peak data into a target prediction model for the processing degree of traditional Chinese medicine materials to output target processing degree data of the target traditional Chinese medicine materials;

[0079] Wherein, the target prediction model is obtained based on the training system of the prediction model for the processing degree of traditional Chinese medicine materials described above.

[0080] The present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and configured to run on the processor. When the processor executes the computer program, it implements the training method of the prediction model for the processing degree of traditional Chinese medicine materials described above, or implements the prediction method for the processing degree of traditional Chinese medicine materials described above.

[0081] The present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the training method of the prediction model for the processing degree of traditional Chinese medicine materials described above, or implements the prediction method for the processing degree of traditional Chinese medicine materials described above.

[0082] The present disclosure also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the training method of the prediction model for the processing degree of traditional Chinese medicine materials as described above, or implements the prediction method for the processing degree of traditional Chinese medicine materials as described above.

[0083] On the basis of conforming to common knowledge in the art, the above preferred conditions can be combined arbitrarily to obtain various preferred examples of the present disclosure.

[0084] The positive and progressive effects of the present disclosure are as follows:

[0085] The present disclosure extracts sample characteristic peak data from the sample mass spectrometry data of traditional Chinese medicine materials, and then trains a target prediction model for the processing degree of traditional Chinese medicine materials, ensuring the effectiveness and reliability of the target prediction model. Using the target characteristic peak data extracted from the target mass spectrometry data of the target traditional Chinese medicine materials as the input of the target prediction model to output the target processing degree data of the target traditional Chinese medicine materials, the prediction of the processing degree of traditional Chinese medicine materials is effectively and accurately realized, a quality control standard can be formed to guide the control of the processing time in the production process, and the quality and efficacy of traditional Chinese medicine materials can be guaranteed. Description of the Drawings

[0086] Figure 1 It is a flowchart of the training method of the prediction model for the processing degree of traditional Chinese medicine materials in Embodiment 1 of the present disclosure;

[0087] Figure 2 It is a flowchart of the training method of the prediction model for the processing degree of traditional Chinese medicine materials in Embodiment 2 of the present disclosure;

[0088] Figure 3 It is a flowchart of preparing cooked rehmannia root by the method of stewing with wine in Embodiment 2 of the present disclosure;

[0089] Figure 4 It is a flowchart of preparing cooked rehmannia root by the steaming method in Embodiment 2 of the present disclosure;

[0090] Figure 5 It is a steaming time graph of the cooked rehmannia root prepared by the method of stewing with wine and the steaming method in Embodiment 2 of the present disclosure;

[0091] Figure 6 It is a flowchart of step S133 in the training method of the prediction model for the processing degree of traditional Chinese medicine materials in Embodiment 2 of the present disclosure;

[0092] Figure 7 It is a flowchart of step S136 in the training method of the prediction model for the processing degree of traditional Chinese medicine materials in Embodiment 2 of the present disclosure;

[0093] Figure 8 It is a flowchart of the prediction method for the processing degree of traditional Chinese medicine materials in Embodiment 3 of the present disclosure;

[0094] Figure 9 It is a specific example graph of the prediction method for the processing degree of traditional Chinese medicine materials in Embodiment 3 of the present disclosure;

[0095] Figure 10 It is a module schematic diagram of the training system of the prediction model for the processing degree of traditional Chinese medicine materials in Embodiment 4 of the present disclosure;

[0096] Figure 11 It is a module schematic diagram of the training system of the prediction model for the processing degree of traditional Chinese medicine materials in Embodiment 5 of the present disclosure;

[0097] Figure 12 Schematic diagram of the modules of the prediction system for the processing degree of Chinese medicinal materials in Embodiment 6 of the present disclosure;

[0098] Figure 13 Schematic diagram of the structure of the electronic device in Embodiment 7 of the present disclosure. Detailed implementation manners

[0099] The present disclosure will be further described below by way of embodiments, but the present disclosure is not limited to the scope of the described embodiments.

[0100] In the embodiments of the present disclosure, prefix words such as "first" and "second" are only used to distinguish different described objects, and do not limit the position, order, priority, quantity or content of the described objects. The use of ordinal words and other prefix words for distinguishing described objects in the embodiments of the present disclosure does not constitute a limitation on the described objects. For the statements of the described objects, refer to the descriptions in the context of the embodiments, and no redundant limitation should be constituted due to the use of such prefix words. In addition, in the description of this embodiment, unless otherwise specified, the meaning of "a plurality" is two or more.

[0101] In the embodiments of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of the user's personal information and other processing all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0102] Embodiment 1

[0103] This embodiment provides a training method for a prediction model of the processing degree of Chinese medicinal materials, as Figure 1 shown, the training method includes:

[0104] S11. Obtain the sample processing degree data and sample mass spectrometry data of the sample Chinese medicinal materials;

[0105] S12. Extract the sample characteristic peak data from the sample mass spectrometry data;

[0106] Among them, the sample characteristic peak data is used to characterize several compounds in the sample Chinese medicinal materials;

[0107] S13. Use each group of sample characteristic peak data of the sample Chinese medicinal materials as input, and the corresponding group of sample processing degree data as output, and train to obtain the target prediction model for the processing degree of Chinese medicinal materials.

[0108] Specifically, the sample Chinese medicinal materials include Rehmannia glutinosa; the sample processing degree data includes the processing duration;

[0109] In this embodiment, the sample characteristic peak data is extracted from the sample mass spectrometry data of the sample Chinese medicinal materials, and then the target prediction model for the processing degree of Chinese medicinal materials is trained, ensuring the effectiveness and reliability of the target prediction model.

[0110] Example 2

[0111] This example provides a training method for a prediction model of the processing degree of traditional Chinese medicinal materials, which is a further improvement on Example 1.

[0112] In an implementable solution, as Figure 2 shown, step S11 includes:

[0113] S111. Detect the test solution of the sample traditional Chinese medicinal materials by using ultra-high performance liquid chromatography-quadrupole time-of-flight mass spectrometry technology to obtain the sample mass spectrometry data of the sample traditional Chinese medicinal materials.

[0114] Specifically, obtain the sample processing degree data and sample mass spectrometry data of the sample traditional Chinese medicinal materials. The sample traditional Chinese medicinal materials are, for example, Rehmanniae Radix Praeparata. First, Rehmanniae Radix Praeparata can be prepared by the method of stewing with wine or steaming. The sample processing degree data of this Rehmanniae Radix Praeparata is determined in advance according to the preparation process. Then, obtain the test solution of Rehmanniae Radix Praeparata. Finally, use the ultra-high performance liquid chromatography-quadrupole time-of-flight mass spectrometry (Ultra Performance Liquid Chromatography-Quadrupole Time-of-Flight Mass Spectrometry, UPLC-Q-TOF-MS) coupling technology to detect the test solution to obtain the mass spectrometry data corresponding to the test solution, that is, the sample mass spectrometry data of the sample traditional Chinese medicinal materials.

[0115] As Figure 3 shown, the process of preparing Rehmanniae Radix Praeparata by the method of stewing with wine is as follows: Fresh Rehmanniae Radix, weigh, quickly wash with clear water 3 times, place on a tray, that is, perform purification on Fresh Rehmanniae Radix, then dry at 55 °C for 45 h, let cool, weigh, add yellow rice wine according to the ratio of 100 g Rehmanniae Radix: 35 ml yellow rice wine, moisten for 24 h, turn several times during the period to make Rehmanniae Radix fully absorb the yellow rice wine. After moistening, place Rehmanniae Radix on a tray, put it into a steamer, steam for 12 h. While collecting the steaming juice, place Rehmanniae Radix in an oven at 50 °C and dry to 80% dryness, keep a sample, and obtain the first-steamed Rehmanniae Radix Praeparata; Mix the steaming juice back into the 80%-dry Rehmanniae Radix, repeat the first steaming operation 2 times to obtain the second-steamed Rehmanniae Radix Praeparata and the third-steamed Rehmanniae Radix Praeparata respectively; After the third steaming, mix the steamed Rehmanniae Radix Praeparata into the steaming juice and continue to steam for 8 h, collect the steaming juice, dry at 50 °C to 80% dryness, keep a sample, and obtain the fourth-steamed Rehmanniae Radix Praeparata; Mix the steaming juice back into the 80%-dry Rehmanniae Radix, repeat the fourth steaming operation 2 times to obtain the fifth-steamed Rehmanniae Radix Praeparata and the sixth-steamed Rehmanniae Radix Praeparata respectively; After the sixth steaming, mix the steamed Rehmanniae Radix Praeparata into the steaming juice and continue to steam for 6 h, collect the steaming juice, dry at 50 °C to 80% dryness, keep a sample, and obtain the seventh-steamed Rehmanniae Radix Praeparata; Mix the steaming juice back into the 80%-dry Rehmanniae Radix, repeat the seventh steaming operation 2 times to obtain the eighth-steamed Rehmanniae Radix Praeparata and the ninth-steamed Rehmanniae Radix Praeparata respectively.

[0116] As Figure 4As shown in the figure, the process of preparing Rehmanniae Radix Praeparata by steaming method is as follows: replace "100 g Rehmanniae Radix: 35 ml yellow rice wine for moistening raw Rehmanniae Radix" with "100 g Rehmanniae Radix: 35 ml clear water for moistening raw Rehmanniae Radix", and the remaining preparation methods are the same as those of Rehmanniae Radix Praeparata prepared by the wine stewing method, to obtain Rehmanniae Radix Praeparata prepared by clear steaming method.

[0117] 4 batches of Rehmanniae Radix Praeparata prepared by the wine stewing method are made per steaming, and 4 batches of Rehmanniae Radix Praeparata prepared by the steaming method are made per steaming. There are a total of 72 batches of samples for the first to ninth steaming. The steaming time diagrams of Rehmanniae Radix Praeparata prepared by the wine stewing method and the steaming method are as Figure 5 shown. The first-steamed Rehmanniae Radix Praeparata is 0 h - 12 h, the second-steamed Rehmanniae Radix Praeparata is 12 h - 24 h, the third-steamed Rehmanniae Radix Praeparata is 24 h - 36 h, the fourth-steamed Rehmanniae Radix Praeparata is 36 h - 44 h, the fifth-steamed Rehmanniae Radix Praeparata is 44 h - 52 h, the sixth-steamed Rehmanniae Radix Praeparata is 52 h - 60 h, the seventh-steamed Rehmanniae Radix Praeparata is 60 h - 66 h, the eighth-steamed Rehmanniae Radix Praeparata is 66 h - 72 h, and the ninth-steamed Rehmanniae Radix Praeparata is 72 h - 78 h.

[0118] Among them, the raw Rehmanniae Radix can be the raw Rehmanniae Radix with batch numbers DH002, DH003, DH004, DH005, DH006 purchased by oneself, and the place of origin is Nankaiyi Village, Huagong Town, Mengzhou, Henan. The batch number and place of origin are just an example and can be set or adjusted according to the actual situation.

[0119] Obtain the test solution of Rehmanniae Radix Praeparata, and use ultra-high performance liquid chromatography - quadrupole time-of-flight mass spectrometry to detect the test solution to obtain the mass spectrometry data corresponding to the test solution. The following experimental reagents are required in these processes: acetonitrile (LC-MS), methanol (AR), water (Watsons distilled water), formic acid (LC-MS), and the following experimental instruments are required: QTOF-MS system (Waters Xevo G2, Waters Corporation, USA), Waters ACUITY HSS T3 1.8 μm chromatographic column (2.1×50 mm); MS204TS analytical balance (Shanghai Mettler Toledo Instruments Co., Ltd.); KQ-250DE type numerical control ultrasonic cleaner (Kunshan Ultrasonic Instruments Co., Ltd.); 100 ml volumetric flask; 1.5 ml centrifuge tube. The specific process is as follows: Cut each of the 72 batches of Rehmanniae Radix Praeparata slices into small pieces about 5 mm in size, dry them under reduced pressure at 80 °C for 24 hours, grind them into coarse powder, take 0.3 g and put it into a 100 ml volumetric flask, add 95 ml of 75% methanol, extract ultrasonically for 90 min, after the extraction is completed, let it cool and make up the volume, centrifuge at a high speed of 12000 rpm for 10 min, and take the supernatant as the test solution. This operation is repeated twice to obtain a total of 144 samples.

[0120] Use ultra-high performance liquid chromatography - quadrupole time-of-flight mass spectrometry to detect 144 test solutions to obtain the mass spectrometry data corresponding to the test solutions.

[0121] The detection conditions include ultra-high performance liquid chromatography conditions and Xevo G2-XS Q-Tof-MS mass spectrometry conditions. The ultra-high performance liquid chromatography conditions are as follows: The chromatographic column is Waters ACUITY HSS T3 chromatographic column, mobile phase A is 0.1% formic acid in water, mobile phase B is acetonitrile, column temperature is 35°C, sample tray temperature is controlled at 10°C, injection volume is 10 μl, flow rate is 0.3 ml / min, and gradient elution is performed. The gradient elution program is as follows: 0 - 2 min, 1% B; 2 - 4 min, 1% - 9% B; 4 - 10 min; 9% - 29% B; 10 - 12 min, 29% - 48% B; 12 - 27 min, 48% - 100% B; 27 - 33 min, 100% B; 30 - 33.1 min, 100% - 1% B; 33.1 - 34 min, 1% B.

[0122] The Xevo G2-XS Q-Tof-MS mass spectrometry conditions are as follows: Electrospray ionization source (ESI), negative ion mode (scan mode is full scan; MS parameters: mass range m / z 50 - 1500 Da, scan time 0.3 s, nebulizing gas and cone gas are nitrogen, capillary voltage 3 kV (ESI+) and 2 kV (ESI-), cone sampling voltage 40 V, cone gas flow rate 50 L / h, ion source temperature 100°C, collision energy 30 - 50 eV, desolvation temperature 250°C, desolvation gas flow rate 600 L / h; accurate mass calibration is performed using leucine enkephalin ([M - H] - 554.2615) solution, with a mass of 1 μg / L, flow rate 5 μl / min, and calibration frequency 5 s.

[0123] The above values are just an example and can be set or adjusted according to the actual situation.

[0124] In this solution, by using ultra-high performance liquid chromatography - quadrupole time-of-flight mass spectrometry technology to detect the test solution, the sample mass spectrometry data of the Chinese herbal medicine in the sample is obtained, ensuring the accuracy and reliability of the sample mass spectrometry data.

[0125] In an implementable solution, step S12 includes:

[0126] S121. Perform alignment processing on the sample mass spectrometry data;

[0127] S122. Extract sample characteristic peak data from the aligned sample mass spectrometry data.

[0128] Specifically, the preset software is used to align the sample mass spectrometry data. The preset software is Progenesis QI software (a mass spectrometry data processing software for bioanalysis). The RAW (a digital image file format) file obtained by UPLC-Q-TOF-MS detection is imported into Progenesis QI software, and alignment processing is performed with the specified data. Then, peak extraction is carried out, and all peak data are exported to an Excel (a spreadsheet software) table and saved as a CSV (a text file format) file. All characteristic peaks and their corresponding peak normalization intensity values are selected as sample characteristic peak data.

[0129] The purpose of alignment processing is to eliminate peak shape offsets caused by instrument drift, sample preparation, etc., improve the accuracy and reliability of mass spectrometry data, and ensure that all samples are in the same position in the sample mass spectrometry data, that is, the mass spectrometry graph. Alignment can be automatically completed by the algorithm built in Progenesis QI software, or alignment parameters can be set manually, that is, specify the data. Peak extraction is performed on the aligned mass spectrometry graph. Progenesis QI software can identify and extract all characteristic peaks to form sample characteristic peak data. Characteristic peaks represent compounds or metabolites in the sample.

[0130] In the process of obtaining sample characteristic peak data, some attribute columns can also be removed according to actual needs to reduce the data dimension.

[0131] A group of sample characteristic peak data can also be randomly selected as unknown samples to verify the evaluation model.

[0132] In this solution, after aligning the sample mass spectrometry data and then performing peak extraction, sample characteristic peak data are obtained, reducing the error of the sample characteristic peak data and improving the reliability and accuracy of the sample characteristic peak data.

[0133] In an implementable solution, step S13 includes:

[0134] S131. Obtain the correlation between the processed degree data of each group of samples and each characteristic peak in the corresponding sample characteristic peak data;

[0135] Among them, each characteristic peak is used to characterize a compound in the sample Chinese herbal medicine;

[0136] S132. Based on the correlation, select several characteristic peaks to construct a characteristic data set;

[0137] S133. Use each group of characteristic data sets as input and the corresponding processed degree data of each group of samples as output to train the first target prediction model for the processed degree of Chinese herbal medicine.

[0138] Specifically, during the model training process, the sample processing degree data serves as the target variable, and the sample characteristic peak data serves as the feature dataset. The Mutual Information Classification (MIC) is used as the scoring function. By calculating the correlation between each feature in the feature dataset and the target variable, the SelectKBest method is adopted to select the top K features with the highest correlation with the target variable to construct the feature dataset, and then the first target prediction model is trained. The target prediction model includes this first target prediction model. Among them, the range of K is 50 - 200.

[0139] The formula corresponding to the scoring function is as follows:

[0140]

[0141] Among them, X and Y are two random variables. X represents the feature dataset, Y represents the target variable, x represents each feature in the feature dataset, and y represents each sample in the target variable; I(X; Y) represents the mutual information between the random variables X and Y; p(x, y) is the joint probability distribution of X and Y occurring simultaneously; p(x) and p(y) are the marginal probability distributions of X and Y respectively; log represents the logarithmic function, usually the natural logarithm or the logarithm with base 2.

[0142] In this solution, by obtaining the correlation between each characteristic peak in the sample processing degree data and the sample characteristic peak data, and then selecting several characteristic peaks according to the correlation to construct the feature dataset, redundant features and noise features are removed, the data dimension is reduced, and the training efficiency and prediction performance of the model are effectively improved.

[0143] In an implementable solution, step S13 includes:

[0144] S134. Obtain the projection variable importance value corresponding to each characteristic peak in each group of sample characteristic peak data;

[0145] Among them, each characteristic peak is used to characterize a compound in the sample Chinese medicinal material;

[0146] S135. Based on the projection variable importance value, select several characteristic peaks from the sample characteristic peak data to construct the feature dataset;

[0147] S136. Use each group of feature datasets as the input and the corresponding group of sample processing degree data as the output to train the second target prediction model for the processing degree of Chinese medicinal materials.

[0148] Specifically, the sample characteristic peak data extracted from the sample mass spectrometry data has an extremely large amount of data characteristic quantities, far more than the sample size, belonging to typical high-dimensional small-sample data. Directly applying high-dimensional small-sample data to a machine learning model may lead to problems such as low computational efficiency and overfitting. Therefore, after obtaining the sample characteristic peak data, the EZinfo (a bioinformatics data analysis and visualization tool) module in the Progenesis QI plugin is used to select partial least squares discriminant analysis (PLS-DA), and the characteristic peaks with variable importance in the projection (VIP) values greater than the preset projection variable importance threshold are selected and marked and exported to a CSV file. A characteristic data set is constructed with the selected characteristic peaks, and then a second target prediction model is trained. The target prediction model includes this second target prediction model. The preset projection variable importance threshold is, for example, 1, and it can be set or adjusted according to the actual situation.

[0149] A group of sample processing degree data and the characteristic data set can also be randomly selected as unknown samples for validating and evaluating the model.

[0150] In this solution, by obtaining the variable importance value in the projection corresponding to each characteristic peak in the sample characteristic peak data, and then constructing a characteristic data set according to the variable importance value in the projection, that is, screening out important characteristics from the sample characteristic peak data according to the variable importance value in the projection to construct a characteristic data set, the data dimension is effectively reduced, the computational efficiency in the model training process is improved, and the overfitting risk of the model is reduced.

[0151] In an implementable solution, as Figure 6 shown, step S133 includes:

[0152] S1331. Using each group of characteristic data sets as input and the corresponding each group of sample processing degree data as output, train an initial prediction model for the processing degree of traditional Chinese medicine.

[0153] S1332. Evaluate the initial prediction model to obtain an evaluation result.

[0154] S1333. Based on the evaluation result, adjust the initial prediction model to obtain a first target prediction model.

[0155] Specifically, using each group of characteristic data sets as input and the corresponding each group of sample processing degree data as output, an initial prediction model for the processing degree of traditional Chinese medicine is trained using the random forest algorithm.

[0156] The dataset includes sample processing degree data and a feature dataset. The dataset is divided according to the ratio of 80% training set and 20% test set to ensure sufficient sample data in the model training stage and an adequate independent test set for subsequent model evaluation. The values are just an example and can be set or adjusted according to the actual situation.

[0157] The initial prediction model is obtained by training the model using the Random Forest algorithm. The Random Forest algorithm improves the accuracy and stability of classification by integrating multiple decision trees. When instantiating the RandomForestClassifier, appropriate parameters are set. For example, the number of decision trees in the Random Forest, the parameter n_estimators, is set to 50 - 200 or 5 - 30 to ensure that the model has sufficient complexity and generalization ability; the randomness parameter random_state in the Random Forest construction process is set to 42 to ensure the reproducibility of the experiment. The maximum depth parameter max_depth of the Random Forest can also be set to 3 - 20. A feature filter can also be used to obtain a specific feature subset to train the RandomForestClassifier so that it learns the mapping relationship from features to target variables.

[0158] After obtaining the initial prediction model, the Cross-Validation method is used for evaluation. By dividing the dataset into different subsets for training and testing multiple times, the performance of the model on different subsets is evaluated to obtain more stable performance evaluation results, so as to more accurately understand the stability and generalization ability of the model. Preset metrics can also be used to evaluate the model. The preset metrics include Accuracy, Confusion Matrix, and Classification Report on the training set and test set.

[0159] In this solution, by evaluating the initial prediction model and then adjusting the initial prediction model, the generalization ability of the first target prediction model is ensured, and the performance of the first target prediction model is improved.

[0160] In an implementable solution, as Figure 7 shown, step S136 includes:

[0161] S1361. Using each group of feature datasets as input and the corresponding each group of sample processing degree data as output, train to obtain the initial prediction model of the Chinese herbal medicine processing degree;

[0162] S1362. Evaluate the initial prediction model to obtain the evaluation result;

[0163] S1363. Adjust the initial prediction model based on the evaluation results to obtain the second target prediction model.

[0164] Specifically, the evaluation method of the second target prediction model is similar to that of the first target prediction model, which will not be elaborated here.

[0165] In this solution, by evaluating the initial prediction model and then adjusting the initial prediction model, the generalization ability of the second target prediction model is ensured, and the performance of the second target prediction model is improved.

[0166] In this embodiment, the sample characteristic peak data is extracted from the sample mass spectrometry data of the sample Chinese herbal medicine, and then the target prediction model for the processing degree of the Chinese herbal medicine is trained, ensuring the effectiveness and reliability of the target prediction model.

[0167] Embodiment 3

[0168] This embodiment provides a prediction method for the processing degree of Chinese herbal medicine. As Figure 8 shown, the prediction method includes:

[0169] S21. Obtain the target mass spectrometry data of the target Chinese herbal medicine;

[0170] S22. Extract the target characteristic peak data from the target mass spectrometry data;

[0171] Among them, the target characteristic peak data is used to characterize several compounds in the target Chinese herbal medicine;

[0172] S23. Input the target characteristic peak data into the target prediction model for the processing degree of the Chinese herbal medicine to output the target processing degree data of the target Chinese herbal medicine;

[0173] Among them, the target prediction model is obtained based on the training method of the prediction model for the processing degree of the Chinese herbal medicine in Embodiment 1 or Embodiment 2.

[0174] Specifically, the target Chinese medicinal material can be, for example, Rehmanniae Radix Praeparata, and 10 batches of commercially available Rehmanniae Radix Praeparata and 3 batches of self-made Rehmanniae Radix Praeparata can be used. The commercially available Rehmanniae Radix Praeparata includes Rehmanniae Radix Praeparata with batch number 240306 from a certain Chinese herbal medicine slice manufacturer in Anhui, Rehmanniae Radix Praeparata with batch number 2305001 from a certain Chinese herbal medicine slice manufacturer in Jiangxi, Rehmanniae Radix Praeparata with batch number C22404065 from a certain Chinese herbal medicine slice manufacturer in Guangdong, Rehmanniae Radix Praeparata with batch number 240400399 from a certain Chinese herbal medicine slice manufacturer in Anhui, Rehmanniae Radix Praeparata with batch number 240100249 from a certain Chinese herbal medicine slice manufacturer in Anhui, Rehmanniae Radix Praeparata with batch number 240901 from a certain Chinese herbal medicine slice manufacturer in Guangdong, Rehmanniae Radix Praeparata with batch number 292240601 from a certain Chinese herbal medicine slice manufacturer in Hebei, Rehmanniae Radix Praeparata with batch number 202309011143 from a certain Chinese herbal medicine slice manufacturer in Hebei, Rehmanniae Radix Praeparata with batch number 220110 from a certain Chinese herbal medicine slice manufacturer in Shanghai, and Rehmanniae Radix Praeparata with batch number 230307 from a certain Chinese herbal medicine slice manufacturer in Shanghai. The batch numbers of the self-made Rehmanniae Radix Praeparata include DH004-SR, DH005-SR, and DH006-SR. The batch numbers and manufacturers can be selected according to the actual situation, and this is just an example here.

[0175] The self-made Rehmanniae Radix Praeparata can be prepared according to the Henan Province Chinese Herbal Medicine Slicing and Processing Specification Standard (November 12, 2021). The specific process is as follows: Take raw Rehmanniae Radix, mix it evenly with yellow rice wine and Amomi Fructus powder, place it in a suitable steaming container, seal it, heat it with strong fire, steam it over water for about 48 hours until it is completely black inside and outside and the center is black, take it out, air it until it is about 80% dry, slice it, and dry it in the sun to obtain it. For every 100 kg of raw Rehmanniae Radix, use 50 kg of yellow rice wine and 0.9 kg of Amomi Fructus powder.

[0176] The process of obtaining the target mass spectrometry data of the target Chinese medicinal material is similar to the process of obtaining the sample mass spectrometry data of the sample Chinese medicinal material. The detection techniques and experimental reagents used are not elaborated here. The specific process is as follows: Cut 10 batches of commercially available Rehmanniae Radix Praeparata slices and 3 batches of self-made Rehmanniae Radix Praeparata slices into small pieces about 5 mm each. After drying under reduced pressure at 80 °C for 24 hours, grind them into coarse powder. Take 0.3 g and put it into a 100 ml volumetric flask, add 95 ml of 75% methanol, and ultrasonically extract for 90 min. After the extraction is completed, let it cool, make up the volume, centrifuge at 12000 rpm for 10 min, and take the supernatant as the test solution. This process is carried out in two batches in parallel, and a total of 26 samples are obtained. Use ultra-high performance liquid chromatography-quadrupole tandem time-of-flight mass spectrometry to detect the 26 test solutions to obtain the target mass spectrometry data of the commercially available and self-made Rehmanniae Radix Praeparata.

[0177] The target prediction model includes the first target prediction model and the second target prediction model.

[0178] When using the first target prediction model to predict the processing degree of the target Chinese medicinal material, the detection data of commercially available and self-made Rehmanniae Radix Praeparata samples, that is, the target mass spectrometry data, are imported into the Progenesis QI software, and peak alignment is performed uniformly with the training data of the first target prediction model. All peak data are extracted and saved as a CSV file in an Excel sheet. All peaks and their corresponding peak normalization intensity values are selected as the sample data to be predicted. The sample data to be predicted are input into the first target prediction model for prediction. The predicted processing degree of the self-made Rehmanniae Radix Praeparata is 36 - 44 h. The processing degrees of Rehmanniae Radix Praeparata with batch numbers 240306, C22404065, 240400399, 240100249, 240901, 202309011143, 220110, and 230307 predicted by the first target prediction model are all 0 - 12 h. The processing degrees of Rehmanniae Radix Praeparata with batch numbers 2305001 and 292240601 predicted by the first target prediction model are 12 - 24 h.

[0179] When using the second target prediction model to predict the processing degree of the target Chinese medicinal material, the detection data of commercially available and self-made Rehmanniae Radix Praeparata samples are imported into the Progenesis QI software, and peak alignment is performed uniformly with the training data of the second target prediction model, and peak extraction is carried out. The partial least squares discriminant analysis is selected using the EZinfo module in the Progenesis QI plug-in. The characteristic peaks with VIP values greater than the preset importance threshold of projection variables are selected and marked and exported to a CSV file. The selected characteristic peaks and their corresponding peak normalization intensity values are used as the sample data to be predicted. The sample data to be predicted are input into the second target prediction model for prediction. The predicted processing degree of the self-made Rehmanniae Radix Praeparata is 36 - 44 h. The processing degrees of Rehmanniae Radix Praeparata with batch numbers C22404065, 240400399, 240100249, 240901, 202309011143, 220110, and 230307 predicted by the second target prediction model are all 0 - 12 h. The processing degrees of Rehmanniae Radix Praeparata with batch numbers 240306, 2305001, and 292240601 predicted by the second target prediction model are 12 - 24 h. The preset importance threshold of projection variables is, for example, 1, and it can be set or adjusted according to the actual situation.

[0180] When using two models to predict the processing degree of the target Chinese medicinal material, based on the target processing degree data corresponding to the first target prediction model and the target processing degree data corresponding to the second target prediction model, the target processing degree data corresponding to the target Chinese medicinal material are obtained. For example, the larger range in the target processing degree data corresponding to the two models is used as the target processing degree data corresponding to the target Chinese medicinal material, or the union of the two target processing degree data corresponding to the two models is used as the target processing degree data corresponding to the target Chinese medicinal material.

[0181] The predicted processing degrees of the self-made Rehmanniae Radix Praeparata by the two models are consistent. For the remaining 10 batches of Rehmanniae Radix Praeparata, the predicted processing degrees of 9 batches are consistent. For the Rehmanniae Radix Praeparata with batch number 240306, the predicted processing time by the first target prediction model is 0 - 12 h, and the predicted processing time by the second target prediction model is 12 - 24 h. It can be obtained that the processing degree of the Rehmanniae Radix Praeparata with batch number 240306 is between 0 - 24 h, and it has the relevant characteristics of both the first-steamed Rehmanniae Radix Praeparata and the second-steamed Rehmanniae Radix Praeparata. That is, the target prediction model has good generalization ability for predicting the processing degrees of different Rehmanniae Radix Praeparata.

[0182] The working principle of the prediction method for the processing degree of traditional Chinese medicinal materials in this embodiment is described below with specific examples, as Figure 9 shown:

[0183] First, a target prediction model is obtained according to the training method of the prediction model for the processing degree of traditional Chinese medicinal materials. The specific process is as follows: The sample traditional Chinese medicinal materials have known sample processing degree data. The sample traditional Chinese medicinal materials (such as Rehmanniae Radix Praeparata) are cut into pieces, dried under reduced pressure at 80 °C for 24 hours, ground into coarse powder, and then ultrasonically extracted to obtain a test solution. The test solution is detected by UPLC-Q-TOF-MS technology to obtain sample mass spectrometry data; Progenesis QI software is used to extract features from the sample mass spectrometry data to obtain sample characteristic peak data; a characteristic data set is constructed according to the sample characteristic peak data; Progenesis QI software is used to screen the sample characteristic peak data and then construct a characteristic data set; the data set includes sample processing degree data and the characteristic data set, and the data set is divided into a validation set, a training set, and a test set; model training and model evaluation are carried out according to the training set and the test set; if the evaluation result of the model evaluation fails, the parameters of the model are adjusted and model training is continued; if the evaluation result of the model evaluation passes, the model is saved, and the model accuracy is verified according to the validation set; different target prediction models are obtained according to different ways of constructing the characteristic data set. If the characteristic data set is constructed by the mutual information classification and SelectKBest method, the first target prediction model (i.e., model 1) is obtained; if the characteristic data set is constructed by the partial least squares discriminant analysis method, the second target prediction model (i.e., model 2) is obtained.

[0184] After obtaining the target prediction model, the target processing degree data of the target Chinese medicinal material is predicted according to the prediction method of the processing degree of Chinese medicinal materials. The specific process is as follows: The target processing degree data of the target Chinese medicinal material is unknown data. The target Chinese medicinal material (prepared rehmannia root) is cut into pieces, dried under reduced pressure at 80 °C for 24 hours, ground into coarse powder, and then ultrasonically extracted to obtain a test solution. The test solution is detected by UPLC-Q-TOF-MS technology to obtain target mass spectrometry data; Progenesis QI software is used to extract features from the target mass spectrometry data to obtain target characteristic peak data; a characteristic data set is constructed according to the target characteristic peak data; Progenesis QI software is used to screen the sample characteristic peak data, and then a characteristic data set is constructed; according to different ways of constructing the characteristic data set, different target prediction models are used to predict the characteristic data set to output the prediction result, that is, the target processing degree data. If the characteristic data set is constructed by using mutual information classification and SelectKBest method, the first target prediction model (i.e., model 1) is used for prediction; if the characteristic data set is constructed by using partial least squares discriminant analysis method, the second target prediction model (i.e., model 2) is used for prediction.

[0185] In this embodiment, the target characteristic peak data extracted from the target mass spectrometry data of the target Chinese medicinal material is used as the input of the target prediction model to output the target processing degree data of the target Chinese medicinal material, effectively and accurately realizing the prediction of the processing degree of Chinese medicinal materials, capable of forming a quality control standard, guiding the control of the processing time in the production process, and ensuring the quality and efficacy of Chinese medicinal materials.

[0186] Example 4

[0187] This embodiment provides a training system for a prediction model of the processing degree of Chinese medicinal materials, as Figure 10 shown. This training system includes:

[0188] A sample data acquisition module 1, configured to acquire a plurality of groups of sample processing degree data and sample mass spectrometry data of sample Chinese medicinal materials;

[0189] A sample peak extraction module 2, configured to extract sample characteristic peak data from the sample mass spectrometry data;

[0190] Among them, the sample characteristic peak data is used to characterize a plurality of compounds in the sample Chinese medicinal material;

[0191] A model training module 3, configured to use each group of sample characteristic peak data of the sample Chinese medicinal material as the input and the corresponding group of sample processing degree data as the output to train a target prediction model for the processing degree of Chinese medicinal materials.

[0192] In this embodiment, the sample characteristic peak data is extracted from the sample mass spectrometry data of the traditional Chinese medicine materials, and then the target prediction model for the processing degree of the traditional Chinese medicine materials is trained, ensuring the effectiveness and reliability of the target prediction model.

[0193] Example 5

[0194] This embodiment provides a training system for a prediction model of the processing degree of traditional Chinese medicine materials, which is a further improvement of Example 4.

[0195] In an implementable solution, the sample traditional Chinese medicine materials include Rehmannia glutinosa.

[0196] In an implementable solution, the sample processing degree data includes the processing duration.

[0197] In an implementable solution, the sample data acquisition module 1 is further configured to detect the test solution of the sample traditional Chinese medicine materials by using ultra-high performance liquid chromatography-quadrupole tandem time-of-flight mass spectrometry technology to obtain the sample mass spectrometry data of the sample traditional Chinese medicine materials.

[0198] In an implementable solution, the sample peak extraction module 2 includes:

[0199] An alignment processing unit 21 for performing alignment processing on the sample mass spectrometry data;

[0200] A sample peak extraction unit 22 for extracting sample characteristic peak data from the aligned sample mass spectrometry data.

[0201] In an implementable solution, the model training module 3 includes:

[0202] A correlation acquisition unit 31 for obtaining the correlation between each characteristic peak in each group of sample processing degree data and the corresponding sample characteristic peak data;

[0203] wherein each characteristic peak is used to represent a compound in the sample traditional Chinese medicine materials;

[0204] A first data set construction unit 32 for constructing a characteristic data set by selecting a number of characteristic peaks based on the correlation;

[0205] A first model training unit 33 for training a first target prediction model for the processing degree of traditional Chinese medicine materials by using each group of characteristic data sets as input and the corresponding each group of sample processing degree data as output.

[0206] In an implementable solution, the model training module 3 includes:

[0207] An importance value acquisition unit 34 for obtaining the projection variable importance value corresponding to each characteristic peak in each group of sample characteristic peak data;

[0208] Among them, each characteristic peak is used to characterize a compound in the sample of traditional Chinese medicine materials;

[0209] A second data set construction unit 35, configured to select a number of characteristic peaks from the sample characteristic peak data based on the projection variable importance value to construct a characteristic data set;

[0210] A second model training unit 36, configured to use each group of characteristic data sets as inputs and the corresponding each group of sample processing degrees data as outputs to train and obtain a second target prediction model for the processing degree of traditional Chinese medicine materials.

[0211] In an implementable solution, the first model training unit 33 includes:

[0212] An initial model acquisition subunit 331, configured to use each group of characteristic data sets as inputs and the corresponding each group of sample processing degrees data as outputs to train and obtain an initial prediction model for the processing degree of traditional Chinese medicine materials;

[0213] An evaluation result acquisition subunit 332, configured to evaluate the initial prediction model to obtain an evaluation result;

[0214] A target model acquisition subunit 333, configured to adjust the initial prediction model based on the evaluation result to obtain a first target prediction model / a second target prediction model.

[0215] In an implementable solution, the second model training unit 36 includes:

[0216] An initial model acquisition subunit 331, configured to use each group of characteristic data sets as inputs and the corresponding each group of sample processing degrees data as outputs to train and obtain an initial prediction model for the processing degree of traditional Chinese medicine materials;

[0217] An evaluation result acquisition subunit 332, configured to evaluate the initial prediction model to obtain an evaluation result;

[0218] A target model acquisition subunit 333, configured to adjust the initial prediction model based on the evaluation result to obtain a second target prediction model.

[0219] In this embodiment, the sample characteristic peak data is extracted from the sample mass spectrometry data of the sample of traditional Chinese medicine materials, and then the target prediction model for the processing degree of traditional Chinese medicine materials is trained, ensuring the effectiveness and reliability of the target prediction model.

[0220] Embodiment 6

[0221] This embodiment provides a prediction system for the processing degree of traditional Chinese medicine materials, as Figure 12 shown, the prediction system includes:

[0222] A target data acquisition module 4, configured to acquire target mass spectrometry data of target traditional Chinese medicine materials;

[0223] A target peak extraction module 5 for extracting target characteristic peak data from target mass spectrometry data;

[0224] Among them, the target characteristic peak data is used to characterize several compounds in the target Chinese medicinal materials;

[0225] A model output module 6 for inputting the target characteristic peak data into a target prediction model of the processing degree of Chinese medicinal materials to output target processing degree data of the target Chinese medicinal materials;

[0226] Among them, the target prediction model is obtained based on the training system of the prediction model of the processing degree of Chinese medicinal materials in Embodiment 4 or Embodiment 5.

[0227] In this embodiment, the target characteristic peak data extracted from the target mass spectrometry data of the target Chinese medicinal materials is used as the input of the target prediction model to output the target processing degree data of the target Chinese medicinal materials, effectively and accurately realizing the prediction of the processing degree of Chinese medicinal materials, being able to form a quality control standard, guiding the control of the processing time in the production process, and ensuring the quality and efficacy of Chinese medicinal materials.

[0228] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The system embodiment described above is only illustrative. The units described as separate components may or may not be physically separated. The components as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the present disclosure solution.

[0229] Embodiment 7

[0230] Figure 13 As shown in the structural schematic diagram of an electronic device shown in an exemplary embodiment of the present disclosure, the electronic device includes a memory, a processor, and a computer program stored on the memory and used to run on the processor. When the processor executes the computer program, it implements the training method of the prediction model of the processing degree of Chinese medicinal materials or the prediction method of the processing degree of Chinese medicinal materials described in any of the above embodiments. Figure 13 The electronic device 90 shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.

[0231] As Figure 13 shown, the electronic device 90 can be presented in the form of a general computing device. For example, it can be a server device. The components of the electronic device 90 may include but are not limited to: at least one of the above processors 91, at least one of the above memories 92, and a bus 93 connecting different system components (including the memory 92 and the processor 91).

[0232] The bus 93 includes a data bus, an address bus, and a control bus.

[0233] The memory 92 may include volatile memory, such as random access memory (RAM) 921 and / or cache memory 922, and may further include read-only memory (ROM) 923.

[0234] The memory 92 may also include a program tool 925 (or utility) having a set (at least one) of program modules 924. Such program modules 924 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.

[0235] The processor 91 executes various functional applications and data processing by running computer programs stored in the memory 92, such as the training method of the prediction model for the degree of traditional Chinese medicine processing or the prediction method for the degree of traditional Chinese medicine processing provided in any of the above embodiments.

[0236] The electronic device 90 can also communicate with one or more external devices 94 (such as a keyboard, a pointing device, etc.). Such communication can be carried out through the input / output (I / O) interface 95. And, the electronic device 90 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 96. As shown in the figure, the network adapter 96 communicates with other modules of the electronic device 90 through the bus 93. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 90, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID (disk array) systems, tape drives, and data backup storage systems, etc.

[0237] It should be noted that although several units / modules or sub-units / modules of the electronic device are mentioned in the above detailed description, this division is merely exemplary and not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more units / modules described above can be embodied in one unit / modules. Conversely, the features and functions of one unit / modules described above can be further divided and embodied by multiple units / modules.

[0238] Embodiment 8

[0239] The embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the training method of the prediction model for the degree of traditional Chinese medicine processing or the prediction method for the degree of traditional Chinese medicine processing provided in any of the above embodiments.

[0240] Among them, more specific examples of the readable storage medium may include, but are not limited to: portable disks, hard disks, random access memories, read-only memories, erasable programmable read-only memories, optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0241] Embodiment 9

[0242] The embodiments of the present disclosure further provide a computer program product, including a computer program, which when executed by a processor implements the training method of the prediction model for the degree of traditional Chinese medicine processing or the prediction method for the degree of traditional Chinese medicine processing described in any one of the above.

[0243] Among them, the program code for executing the computer program product of the present disclosure can be written in any combination of one or more programming languages, and the program code can be executed entirely on the user device, partially on the user device, executed as an independent software package, partially on the user device and partially on a remote device, or entirely on a remote device.

[0244] Although the specific embodiments of the present disclosure have been described above, those skilled in the art should understand that this is only an example, and the protection scope of the present disclosure is defined by the appended claims. Without departing from the principles and essence of the present disclosure, those skilled in the art can make various changes or modifications to these embodiments, but these changes and modifications all fall within the protection scope of the present disclosure.

Claims

1. A training method for a prediction model of the processing degree of Chinese medicinal materials, characterized in that: The training method comprises: Obtaining several groups of sample processing degree data and sample mass spectrum data of sample Chinese medicinal materials; Extracting sample characteristic peak data from the sample mass spectrum data; Wherein, the sample characteristic peak data is used to characterize several compounds in the Chinese medicinal materials of the sample; The sample characteristic peak data of each group of the sample Chinese medicinal materials are taken as input, and the corresponding sample processing degree data of each group are taken as output, so as to train and obtain a target prediction model of the processing degree of Chinese medicinal materials.

2. The method for training a prediction model for the processing degree of Chinese medicinal materials according to claim 1, characterized in that: The Chinese medicinal materials in the sample include Rehmannia root; and / or, The sample processing degree data includes the processing time; and / or, The step of obtaining sample processing degree data and sample mass spectrum data of the sample Chinese medicinal material comprises: Using ultra-high performance liquid chromatography-quadrupole tandem time-of-flight mass spectrometry to detect the test solution of the sample Chinese medicinal material to obtain the sample mass spectrum data of the sample Chinese medicinal material; and / or, The step of extracting sample characteristic peak data from the sample mass spectrum data comprises: Performing alignment processing on the sample mass spectrum data; The sample characteristic peak data is extracted from the sample mass spectrum data after the alignment process.

3. The training method of the prediction model of the processing degree of Chinese medicinal materials according to claim 1 or 2, characterized in that: The step of taking the sample characteristic peak data of each group of the sample Chinese medicinal materials as input and the corresponding sample processing degree data of each group as output, and training a target prediction model for the processing degree of Chinese medicinal materials comprises: Obtaining the correlation between each group of the sample processing degree data and each characteristic peak in the corresponding sample characteristic peak data; Wherein, each of the characteristic peaks is used to characterize one of the compounds in the Chinese medicinal material of the sample; Based on the correlation, selecting a number of the characteristic peaks to construct a characteristic data set; Taking each group of the feature data sets as input and the corresponding group of the sample processing degree data as output, a first target prediction model of the processing degree of Chinese medicinal materials is trained; and / or, The step of taking the sample characteristic peak data of each group of the sample Chinese medicinal materials as input and the corresponding sample processing degree data of each group as output, and training a target prediction model for the processing degree of Chinese medicinal materials comprises: Obtaining the projection variable importance value corresponding to each characteristic peak in each group of the sample characteristic peak data; Wherein, each of the characteristic peaks is used to characterize one of the compounds in the Chinese medicinal material of the sample; Based on the projection variable importance value, selecting a number of the characteristic peaks from the sample characteristic peak data to construct a characteristic data set; Each group of the feature data sets is used as input, and each group of the corresponding sample processing degree data is used as output, so as to train and obtain a second target prediction model for the processing degree of Chinese medicinal materials.

4. The method for training a prediction model of the processing degree of Chinese medicinal materials as claimed in claim 3, characterized in that: The step of taking each group of the feature data sets as input and each group of the corresponding sample processing degree data as output, and training to obtain the first target prediction model / second target prediction model of the processing degree of Chinese medicinal materials comprises: Taking each group of the feature data sets as input and the corresponding group of the sample processing degree data as output, an initial prediction model for the processing degree of Chinese medicinal materials is trained; Evaluating the initial prediction model to obtain an evaluation result; Based on the evaluation result, the initial prediction model is adjusted to obtain the first target prediction model / the second target prediction model.

5. A method for predicting the degree of processing of Chinese herbal medicines, characterized in that: The prediction method comprises: Acquire target mass spectrometry data of target Chinese medicinal materials; Extracting target characteristic peak data from the target mass spectrum data; Wherein, the target characteristic peak data is used to characterize several compounds in the target Chinese medicinal material; Inputting the target characteristic peak data into a target prediction model of the processing degree of Chinese medicinal materials to output the target processing degree data of the target Chinese medicinal materials; Wherein, the target prediction model is obtained based on the training method of the prediction model of the degree of processing of traditional Chinese medicine according to any one of claims 1-4.

6. A training system for a prediction model of the processing degree of Chinese herbal medicines, characterized in that: The training system comprises: A sample data acquisition module is used to acquire sample processing degree data and sample mass spectrometry data of sample Chinese medicinal materials; A sample peak extraction module, used to extract sample characteristic peak data from the sample mass spectrum data; Wherein, the sample characteristic peak data is used to characterize several compounds in the Chinese medicinal materials of the sample; The model training module is used to take the sample characteristic peak data of each group of the sample Chinese medicinal materials as input and the corresponding sample processing degree data of each group as output, and train to obtain a target prediction model for the processing degree of Chinese medicinal materials.

7. A system for predicting the degree of processing of Chinese herbal medicines, characterized in that: The prediction system comprises: A target data acquisition module, used to acquire target mass spectrometry data of target Chinese medicinal materials; A target peak extraction module, used to extract target characteristic peak data from the target mass spectrum data; Wherein, the target characteristic peak data is used to characterize several compounds in the target Chinese medicinal material; A model output module, used for inputting the target characteristic peak data into the target prediction model of the processing degree of the Chinese medicinal materials to output the target processing degree data of the target Chinese medicinal materials; Wherein, the target prediction model is obtained based on the training system of the prediction model of the degree of processing of traditional Chinese medicines as described in claim 6.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and used to run on the processor, characterized in that: When the processor executes the computer program, it implements the training method of the prediction model of the processing degree of traditional Chinese medicine according to any one of claims 1 to 4, or implements the prediction method of the processing degree of traditional Chinese medicine according to claim 5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the training method of the prediction model of the processing degree of traditional Chinese medicine according to any one of claims 1 to 4, or implements the prediction method of the processing degree of traditional Chinese medicine according to claim 5.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the training method of the prediction model of the processing degree of traditional Chinese medicine as described in any one of claims 1 to 4, or implements the prediction method of the processing degree of traditional Chinese medicine as described in claim 5.