Key compound identification method and device based on Daqu grade
By constructing a Daqu grade prediction model and using the SHAP algorithm to identify key compounds, the problem of low recognition efficiency and accuracy of Daqu key compounds in the existing technology is solved, and efficient and accurate compound recognition is achieved to ensure the stability and consistency of Daqu quality.
Patent Information
- Application Number
- CN202510364899.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-07-08
AI Technical Summary
In the prior art, the identification method of Daqu key compounds is low in efficiency and accuracy, relies on artificial experience and is highly subjective.
The key compound recognition method based on the Daqu level is adopted. By obtaining the metabolic compound data of the Daqu sample, the sample data set is constructed and divided into training sets and test sets. The Daqu level prediction model is trained using the neural network model, and the feature contribution module is embedded. The SHAP algorithm is used to calculate the contribution degree of each metabolic compound and identify the key compounds.
It improves the identification efficiency and accuracy of key compounds, and can quickly and accurately identify critical compounds in the fermentation process, ensure the stability and consistency of Daqu quality, reduce detection costs and time, optimize resource allocation, and quickly locate abnormal changes.
Smart Images

Figure CN120280037A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brewing, and particularly to a method and device for identifying key compounds based on the grades of Daqu. Background Art
[0002] Daqu is a saccharifying and fermenting agent used in traditional Chinese brewing technology, mainly for brewing Chinese spirits (especially high-proof spirits such as strong aroma type, sauce aroma type, and light aroma type). It uses grains such as wheat, barley, and peas as raw materials, and through natural inoculation of microorganisms in the environment (such as molds, yeasts, bacteria, etc.), it is made into "qu blocks" in the shape of blocks or bricks through processes such as cultivation, fermentation, and drying. Daqu is not only a key raw material for brewing but also directly affects the flavor, aroma, and quality of Chinese spirits.
[0003] The key compounds in Daqu refer to the compounds contained in Daqu during the brewing process that have a decisive impact on the flavor, taste, and aroma type of Chinese spirits. Key compounds may include various compounds such as alcohols, esters, acids, and aromatic compounds. These compounds interact with each other during the fermentation process and jointly affect and shape the unique flavor and taste of Chinese spirits. By accurately identifying and controlling the content and proportion of key compounds in Daqu, it is beneficial to ensure the stability and consistency of Chinese spirits, thereby improving the overall quality of the product. It can also help adjust parameters in the process, such as oxygen concentration, temperature, humidity, fermentation time, etc., to optimize the generation of key compounds and further enhance the quality and flavor of Chinese spirits.
[0004] The existing methods for determining key compounds in Daqu usually include sensory evaluation and physical and chemical detection, that is, matching the descriptive words obtained from sensory evaluation with the compound list obtained from physical and chemical detection to find the compounds related to specific flavor characteristics to identify key compounds. This method relies on manual experience, has strong subjectivity, and poor efficiency and accuracy. Summary of the Invention
[0005] The present invention aims to solve the problem of low efficiency and accuracy in the existing methods for determining key compounds in Daqu, and proposes a method and device for identifying key compounds based on the grades of Daqu.
[0006] The technical solutions adopted by the present invention to solve the above technical problems are as follows:
[0007] In a first aspect, the present invention provides a method for identifying key compounds based on the grades of Daqu, the method comprising:
[0008] Obtaining the Daqu grades and metabolic compound data of a plurality of Daqu samples, and constructing a sample data set based on the Daqu grades and metabolic compound data;
[0009] Divide the sample data set into a training set and a test set, use the training set to train a neural network model, and use the test set to evaluate the training effect of the neural network model. After the neural network model is trained, a Daqu grade prediction model is obtained;
[0010] Embed a feature contribution module in the Daqu grade prediction model;
[0011] For each Daqu sample in the test set, input the corresponding metabolic compound data into the Daqu grade prediction model to obtain the Daqu grade prediction result of each Daqu sample. Based on the Daqu grade prediction result and the feature contribution module, obtain the key compound identification result of the Daqu.
[0012] Furthermore, based on the Daqu grade prediction result and the feature contribution module, obtain the key compound identification result of the Daqu, which specifically includes:
[0013] For each Daqu sample in the test set, the feature contribution module calculates the SHAP value of each metabolic compound data for the Daqu grade prediction result through the SHAP algorithm;
[0014] Calculate the average SHAP value of each metabolic compound data, and determine the contribution degree of each metabolic compound data according to the average SHAP value;
[0015] Sort the contribution degrees from large to small, and take the compounds corresponding to the top K metabolic compound data with the largest contribution degrees as the key compounds.
[0016] Furthermore, the calculation formula of the SHAP value is as follows:
[0017]
[0018] where, φ kj represents the SHAP value of the j-th metabolic compound data of the k-th Daqu sample in the test set for the Daqu grade prediction result, N represents the set of all metabolic compound data, |N| represents the total number of metabolic compound data, S represents any subset of metabolic compound data that does not contain the j-th metabolic compound data, |S| represents the size of the metabolic compound data subset S, v(S) represents the Daqu grade prediction result corresponding to the metabolic compound data subset S, and v(S∪{j}) represents the Daqu grade prediction result corresponding to the metabolic compound data subset S after adding the j-th metabolic compound data;
[0019] The calculation formula of the average SHAP value is as follows:
[0020]
[0021] where, φ jrepresents the average SHAP value of the j-th metabolite data, and K represents the number of Daqu samples in the test set.
[0022] Further, the metabolite data includes content data of one or more compounds among esters, alcohols, aldehydes, acids, ketones, pyrazines, furans, and aromatics. The esters at least include ethyl hexadecanoate, ethyl caproate, ethyl oleate, and ethyl nonanoate. The alcohols at least include phenethyl alcohol, 4-octanol, R23-butanediol, S23-butanediol, and pentanol. The pyrazines at least include trimethylpyrazine.
[0023] Further, the Daqu grades include first-grade Daqu, second-grade Daqu, and third-grade Daqu.
[0024] Further, the neural network model is one model or a combination of multiple models among MLP, LSTM, RNN, CNN, GRU, and Transformer.
[0025] Further, the method further includes:
[0026] After constructing the sample data set, remove the missing values and outliers in the sample data set, and perform normalization processing on the metabolite data.
[0027] Further, training the neural network model according to the sample data set includes:
[0028] Taking the metabolite data in the sample data set as input features, taking the corresponding Daqu grade as the true label, training the neural network model using the training set, comparing the prediction result of the neural network model with the true label during the training process, calculating the corresponding loss function, and optimizing the neural network model through backpropagation;
[0029] Using the test set to determine the accuracy of the neural network model. When the corresponding loss function is less than the loss function threshold and the accuracy is greater than the accuracy threshold, the training of the neural network model is completed.
[0030] Further, the loss function is as follows:
[0031]
[0032] where Loss represents the loss function, M represents the number of Daqu samples in the training set, C represents the number of Daqu grades, p ij represents the true label of the i-th Daqu sample in the j-th Daqu grade. If the i-th Daqu sample belongs to the j-th Daqu grade, then p ij = 1, otherwise p ij = 0, represents the probability that the neural network model predicts the i-th Daqu sample belongs to the j-th Daqu grade.
[0033] In a second aspect, the present invention provides a key compound recognition device based on the grade of Daqu, and the device includes:
[0034] An acquisition module, configured to acquire the Daqu grade and metabolic compound data of a plurality of Daqu samples, and construct a sample data set according to the Daqu grade and metabolic compound data;
[0035] A training module, configured to divide the sample data set into a training set and a test set, train a neural network model using the training set, and evaluate the training effect of the neural network model using the test set. After the neural network model is trained, a Daqu grade prediction model is obtained; and a feature contribution module is embedded in the Daqu grade prediction model;
[0036] An identification module, configured to input the corresponding metabolic compound data of each Daqu sample in the test set into the Daqu grade prediction model, obtain the Daqu grade prediction result of each Daqu sample, and obtain the key compound identification result of the Daqu based on the feature contribution module according to the Daqu grade prediction result.
[0037] The beneficial effect of the present invention is that: the key compound recognition method and device based on the grade of Daqu provided by the present invention use the metabolic compound data of Daqu as input parameters, use the Daqu grade prediction model to predict the Daqu grade, and then use the feature contribution module to determine the contribution degree of each metabolic compound data to the prediction result, so as to realize the recognition of key compounds, thereby improving the efficiency and accuracy of key compound recognition. Description of the Drawings
[0038] Figure 1 It is a schematic flowchart of the key compound recognition method based on the grade of Daqu provided in the embodiment;
[0039] Figure 2 It is a schematic diagram of the average SHAP value of the compounds corresponding to each metabolic compound data provided in the embodiment;
[0040] Figure 3 It is a schematic structural diagram of the key compound recognition device based on the grade of Daqu provided in the embodiment. Detailed Embodiments
[0041] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the present embodiment will be clearly and completely described below with reference to the accompanying drawings in the present embodiment.
[0042] In some of the processes described in the specification of the present invention and the above-mentioned drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel.
[0043] In order to improve the efficiency and accuracy of determining key compounds in Daqu, the technical solution of the present invention is proposed. In the present invention, the Daqu grades and metabolic compound data of a plurality of Daqu samples are obtained, and a sample data set is constructed according to the Daqu grades and metabolic compound data; the sample data set is divided into a training set and a test set, the neural network model is trained using the training set, and the training effect of the neural network model is evaluated using the test set. After the neural network model is trained, a Daqu grade prediction model is obtained; a feature contribution module is embedded in the Daqu grade prediction model; for each Daqu sample in the test set, the corresponding metabolic compound data is input into the Daqu grade prediction model to obtain the Daqu grade prediction result of each Daqu sample, and based on the Daqu grade prediction result and the feature contribution module, the key compound identification result of the Daqu is obtained.
[0044] Specifically, the Daqu grade prediction model in the present invention predicts the Daqu grade through quantitative analysis of the metabolic compound data, and by quantifying the contribution degree of each metabolic compound data to the Daqu grade prediction result, it is possible to quickly and accurately identify which metabolic compound data plays a crucial role in the fermentation process, realizing the identification of key compounds in Daqu, thereby improving the identification efficiency and accuracy. By identifying key compounds, it is possible to more specifically monitor the activity of key compounds, thereby adjusting the production process parameters, ensuring the stability and consistency of the Daqu quality, while reducing the monitoring and analysis of non-key compounds, reducing the detection cost and time cost, optimizing the resource allocation, and when there are problems with the Daqu quality, it is possible to quickly locate the abnormal changes of key compounds and shorten the problem troubleshooting time.
[0045] Based on this, the technical solutions in the present embodiment will be clearly and completely described below in conjunction with the drawings in the present embodiment. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0046] Figure 1 A flowchart showing a method for identifying key compounds based on Daqu grades is shown. Please refer to Figure 1 , and the method includes the following steps:
[0047] S1. Obtain the Daqu grade and metabolite compound data of multiple Daqu samples, and construct a sample data set based on the Daqu grade and metabolite compound data.
[0048] Among them, the metabolite compound data includes the content data of one or more compounds in esters, alcohols, aldehydes, acids, ketones, pyrazines, furans, and aromatics. The esters at least include ethyl hexadecanoate, ethyl hexanoate, ethyl oleate, and ethyl nonanoate. The alcohols at least include phenethyl alcohol, 4-octanol, R2,3-butanediol, S2,3-butanediol, and pentanol. The pyrazines at least include trimethylpyrazine. In practical applications, gas chromatography and mass spectrometry can be used to obtain the metabolite compound data in Daqu samples.
[0049] In this embodiment, a total of 25 metabolite compounds including phenethyl alcohol, 4-octanol, R2,3-butanediol, S2,3-butanediol, ethyl hexadecanoate, ethyl hexanoate, trimethylpyrazine, ethyl oleate, ethyl nonanoate, and pentanol are considered, and are represented by metabolite compound 1, metabolite compound 2,..., metabolite compound 25.
[0050] The Daqu grade characteristics can include first-grade Daqu, second-grade Daqu, and third-grade Daqu. The Daqu grade characteristics can also be other grading standards recognized by experts in the field. In this embodiment, two categories of first-grade Daqu and second-grade Daqu are considered, and are represented by 0 and 1 respectively.
[0051] After obtaining the metabolite compound data of each Daqu sample at different Daqu grades, construct a sample data set that can reflect the corresponding relationship between the Daqu grade and the metabolite compound data.
[0052] In this embodiment, after constructing the sample data set, it also includes data cleaning, that is, removing the missing values and outliers in the sample data set. Through data cleaning, the data quality and the accuracy of the model can be improved. This embodiment also includes normalizing the metabolite compound data. In practical applications, the sklearn.preprocessing.StandardScaler function can be used for normalization processing to standardize the data into a normal distribution with a mean of 0 and a standard deviation of 1. Through normalization processing, the influence of different magnitudes can be eliminated, and the data quality and the accuracy of the model can be further improved. In this embodiment, the pickle.dump function can be used to save the normalized data as a pkl format file for convenient subsequent reading.
[0053] S2. Divide the sample data set into a training set and a test set, use the training set to train a neural network model, and use the test set to evaluate the training effect of the neural network model. After the neural network model is trained, a Daqu grade prediction model is obtained.
[0054] In this embodiment, the neural network model is one model or a combination of multiple models among MLP, LSTM, RNN, CNN, GRU, and Transformer. These neural network models can mine the complex non-linear relationships between metabolic compound data, thereby further improving the accuracy of Daqu grade prediction.
[0055] For example, the neural network model is MLP, which is implemented based on the Python language using the PyTorch library. The specific model is as follows:
[0056]
[0057] In this embodiment, since 25 metabolic compounds are considered, the input dimension input_size = 25, and the other hyperparameters are: hidden_size = 64, output_size = 2.
[0058] In this embodiment, the training process of the neural network model specifically includes:
[0059] The sample dataset is divided into a training set and a test set according to a preset ratio; the metabolic compound data in the sample dataset is used as input features, and the corresponding Daqu grade is used as the true label. The neural network model is trained using the training set, and the training effect of the neural network model is evaluated using the test set. During the training process, the prediction result of the neural network model is compared with the true label, the corresponding loss function is calculated, and the neural network model is optimized through backpropagation; the accuracy of the neural network model is determined using the test set. When the corresponding loss function is less than the loss function threshold and the accuracy is greater than the accuracy threshold, the training of the neural network model is completed, and a Daqu grade prediction model is obtained.
[0060] In this embodiment, the loss function is implemented based on the torch.nn.CrossEntropyLoss function, and the loss function is as follows:
[0061]
[0062] Among them, Loss represents the loss function, M represents the number of Daqu samples in the training set, C represents the number of Daqu grades, p ij represents the true label of the i-th Daqu sample at the j-th Daqu grade. If the i-th Daqu sample belongs to the j-th Daqu grade, then p ij = 1, otherwise p ij = 0, represents the probability that the neural network model predicts the i-th Daqu sample belongs to the j-th Daqu grade.
[0063] S3. Embed a feature contribution module in the Daqu grade prediction model.
[0064] After the Daqu grade prediction model is obtained through training, a feature contribution module is embedded in the Daqu grade prediction model, and the feature contribution module is used to determine the contribution of each metabolic compound data to the Daqu grade prediction result.
[0065] S4. For each Daqu sample in the test set, the corresponding metabolic compound data is input into the Daqu grade prediction model to obtain the Daqu grade prediction result of each Daqu sample. According to the Daqu grade prediction result and based on the feature contribution module, the key compound identification result of Daqu is obtained.
[0066] In practical applications, the metabolic compound data of each Daqu sample in the test set are sequentially input into the Daqu grade prediction model. The Daqu grade prediction model predicts the Daqu grade by quantitatively analyzing the metabolic compound data, and obtains the Daqu grade prediction result of each Daqu sample. The Daqu grade prediction model can mine the complex nonlinear relationship between metabolic compound data and realize accurate prediction of Daqu grade.
[0067] After obtaining the Daqu grade prediction result, the feature contribution module is used to determine the contribution of each metabolite compound data to the Daqu grade prediction result, thereby identifying the key compounds. In this embodiment, it specifically includes:
[0068] For each Daqu sample in the test set, the feature contribution module calculates the SHAP value of each metabolic compound data to the Daqu grade prediction result through the SHAP algorithm; calculates the average SHAP value of each metabolic compound data, and determines the contribution of each metabolic compound data according to the average SHAP value; sorts the contribution from large to small, and takes the compounds corresponding to the metabolic compound data with the top K contribution as key compounds.
[0069] It can be understood that the SHAP (Shapley Additive Explanations) algorithm is based on the SHAP value in game theory, which regards the contribution of each feature value to the model output as a fair distribution. The SHAP value provides an intuitive way to understand the impact of features on prediction results. It can accurately reflect the contribution of each feature to a single prediction and has local accuracy.
[0070] This embodiment calculates the SHAP value of each metabolic compound data of each Daqu sample in the test set for the Daqu grade prediction result based on the SHAP algorithm, and then calculates the average SHAP value of each metabolic compound data for the Daqu grade prediction result. The larger the average SHAP value, the greater the contribution of the corresponding metabolic compound data to the Daqu grade prediction result. Then, according to the average SHAP value, the compounds corresponding to the metabolic compound data with the largest contribution are screened out as key compounds to achieve the identification of key compounds.
[0071] In this embodiment, the calculation formula of the SHAP value is as follows:
[0072]
[0073] Among them, φ kj represents the SHAP value of the Daqu grade prediction result of the j-th metabolite compound data in the k-th Daqu sample in the test set, N represents the set of all metabolite compound data, |N| represents the total number of metabolite compound data, S represents any metabolite compound data subset that does not contain the j-th metabolite compound data, |S| represents the size of the metabolite compound data subset S, v(S) represents the Daqu grade prediction result corresponding to the metabolite compound data subset S, and v(S∪{j}) represents the Daqu grade prediction result corresponding to the metabolite compound data subset S after adding the j-th metabolite compound data;
[0074] The average SHAP value is calculated as follows:
[0075]
[0076] Among them, φ j represents the average SHAP value of the j-th metabolite compound data, and K represents the number of Daqu samples in the test set.
[0077] In this example, the average SHAP value of each metabolite compound data corresponding to the compound is as follows: Figure 2 As shown, metabolite 19, metabolite 6, metabolite 20, metabolite 15, and metabolite 9 are the five metabolites with the highest average SHAP value contribution, and can be considered as the key compounds in this example.
[0078] In summary, the key compound identification method based on Daqu grade provided in this embodiment uses the Daqu grade prediction model to quantitatively analyze the Daqu metabolic compound data. The Daqu grade prediction model can mine the complex nonlinear relationship between the metabolic compound data, thereby realizing the accurate prediction of the Daqu grade. By quantifying the contribution of each metabolic compound data to the Daqu grade prediction result, it is possible to quickly and accurately identify which metabolic compound data plays a vital role in the fermentation process, realize the identification of key compounds, and thus improve the identification efficiency and accuracy. By identifying key compounds, the activity of key compounds can be monitored more specifically, thereby adjusting the production process parameters, ensuring the stability and consistency of Daqu quality, while reducing the monitoring and analysis of non-key compounds, reducing detection costs and time costs, optimizing resource allocation, and when there is a problem with Daqu quality, it is possible to quickly locate abnormal changes in key compounds and shorten the time for troubleshooting.
[0079] Based on the above technical solution, this embodiment further provides a key compound recognition device based on the Daqu grade. Please refer to Figure 3 , the device includes:
[0080] An acquisition module, configured to acquire the Daqu grade and metabolic compound data of multiple Daqu samples, and construct a sample data set according to the Daqu grade and metabolic compound data;
[0081] A training module, configured to divide the sample data set into a training set and a test set, train a neural network model using the training set, and evaluate the training effect of the neural network model using the test set. After the neural network model is trained, a Daqu grade prediction model is obtained; and a feature contribution module is embedded in the Daqu grade prediction model;
[0082] An identification module, configured to input the corresponding metabolic compound data of each Daqu sample in the test set into the Daqu grade prediction model, obtain the Daqu grade prediction result of each Daqu sample, and obtain the key compound identification result of the Daqu based on the Daqu grade prediction result and the feature contribution module.
[0083] It can be understood that since the key compound recognition device based on the Daqu grade described in this embodiment is a device for implementing the key compound recognition method based on the Daqu grade described in the embodiment, for the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For the relevant parts, please refer to the partial description of the method, and details are not described herein again.
Claims
1. A method for identifying key compounds based on the grade of Daqu, characterized in that The method includes: Obtaining the Daqu grades and metabolic compound data of multiple Daqu samples, and constructing a sample data set based on the Daqu grades and metabolic compound data; Dividing the sample data set into a training set and a test set, training a neural network model according to the training set, and evaluating the training effect of the neural network model using the test set. After the neural network model is trained, a Daqu grade prediction model is obtained; Embedding a feature contribution module into the Daqu grade prediction model; For each Daqu sample in the test set, inputting the corresponding metabolic compound data into the Daqu grade prediction model to obtain the Daqu grade prediction result of each Daqu sample, and obtaining the key compound identification result of the Daqu based on the Daqu grade prediction result and the feature contribution module.
2. The key compound identification method based on the grade of Daqu according to claim 1, wherein Obtaining the key compound identification result of the Daqu based on the Daqu grade prediction result and the feature contribution module, specifically including: For each Daqu sample in the test set, the feature contribution module calculates the SHAP value of each metabolic compound data for the Daqu grade prediction result through the SHAP algorithm; Calculating the average SHAP value of each metabolic compound data, and determining the contribution degree of each metabolic compound data according to the average SHAP value; Sorting the contribution degrees from large to small, and taking the compounds corresponding to the metabolic compound data with the top K contribution degrees as the key compounds.
3. The key compound identification method based on the grade of Daqu according to claim 2, wherein The calculation formula of the SHAP value is as follows: Among them, φ kj represents the SHAP value of the k-th metabolite compound data of the j-th Daqu sample in the test set on the prediction result of the Daqu grade. N represents the set of all metabolite compound data, |N| represents the total number of metabolite compound data, S represents an arbitrary subset of metabolite compound data that does not contain the j-th metabolite compound data, |S| represents the size of the metabolite compound data subset S, v(S) represents the prediction result of the Daqu grade corresponding to the metabolite compound data subset S, and v(S∪{j}) represents the prediction result of the Daqu grade corresponding to the metabolite compound data subset S after adding the j-th metabolite compound data; The calculation formula of the average SHAP value is as follows: Among them, φ j represents the average SHAP value of the j-th metabolic compound data, and K represents the number of Daqu samples in the test set.
4. The key compound identification method based on the grade of Daqu according to claim 1, wherein The metabolic compound data includes the content data of one or more compounds among esters, alcohols, aldehydes, acids, ketones, pyrazines, furans, and aromatics. The esters include at least ethyl hexadecanoate, ethyl caproate, ethyl oleate, and ethyl nonanoate. The alcohols include at least phenethyl alcohol, 4-octanol, R2,3-butanediol, S2,3-butanediol, and pentanol. The pyrazines include at least trimethylpyrazine.
5. The key compound identification method based on the grade of Daqu according to claim 1, characterized in that The Daqu grades include first-grade Daqu, second-grade Daqu, and third-grade Daqu.
6. The key compound identification method based on the grade of Daqu according to claim 1, characterized in that, The neural network model is one model or a combination of multiple models among MLP, LSTM, RNN, CNN, GRU, and Transformer.
7. The key compound identification method based on Daqu grade according to claim 1, characterized in that The method further includes: After constructing the sample data set, removing the missing values and outliers in the sample data set, and normalizing the metabolic compound data.
8. The key compound identification method based on the grade of Daqu according to claim 1, characterized in that, Training the neural network model according to the sample data set, including: Taking the metabolic compound data in the sample data set as input features, taking the corresponding Daqu grade as the true label, training the neural network model using the training set, comparing the prediction result of the neural network model with the true label during the training process, calculating the corresponding loss function, and optimizing the neural network model through backpropagation; Using the test set to determine the accuracy of the neural network model. When the corresponding loss function is less than the loss function threshold and the accuracy is greater than the accuracy threshold, the neural network model is trained.
9. The key compound identification method based on the grade of Daqu according to claim 8, wherein, The loss function is as follows: Among them, Loss represents the loss function, M represents the number of Daqu samples in the training set, C represents the number of Daqu grades, and p ij represents the true label of the i-th Daqu sample in the j-th Daqu grade. If the i-th Daqu sample belongs to the j-th Daqu grade, then p ij = 1; otherwise p ij = 0. represents the probability that the neural network model predicts the i-th Daqu sample belongs to the j-th Daqu grade.
10. A key compound recognition device based on the grade of Daqu, characterized in that, The device includes: An acquisition module, configured to obtain the Daqu grades and metabolic compound data of multiple Daqu samples, and construct a sample data set based on the Daqu grades and metabolic compound data; A training module, which is used to divide the sample data set into a training set and a test set, train a neural network model according to the training set, and evaluate the training effect of the neural network model by using the test set. After the neural network model is trained, a Daqu grade prediction model is obtained; and a feature contribution module is embedded in the Daqu grade prediction model; An identification module, which is used to input the corresponding metabolite data of each Daqu sample in the test set into the Daqu grade prediction model, obtain the Daqu grade prediction result of each Daqu sample, and obtain the key compound identification result of the Daqu based on the Daqu grade prediction result and the feature contribution module.