Physical and chemical prediction method and device for fermented grains based on multiple modes
Through multimodal neural network model trained by multimodal data coupling, the problem of low efficiency and accuracy of existing methods for determining mash physical and chemical data is solved, and efficient and accurate mash physical and chemical data prediction is achieved.
Patent Information
- Application Number
- CN202510194826.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-27
AI Technical Summary
The existing methods for determining mash and chemical data are low in efficiency and accuracy, rely on manual testing, cumbersome operations, strict detection conditions, and strong subjectivity.
A multimodal physical and chemical prediction method is adopted to obtain multimodal data (visual data, physical and chemical data, team data and hierarchical data) of multiple mash samples, a multimodal database is established, and a multimodal neural network model is trained through multimodal data coupling to build a multimodal physical and chemical prediction model to predict physical and chemical data to be predicted.
It improves the efficiency and accuracy of the determination of malted physical and chemical data, avoids the subjectivity and cumbersome operation of manual detection, enhances the model's understanding and reasoning ability, and improves the model's accuracy and generalization ability.
Smart Images

Figure CN120048392A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brewing, and particularly to a multimodal-based physical and chemical prediction method and device for fermented grains. Background Art
[0002] Fermented grains are the core matrix in the fermentation process of Luzhou-flavor liquor. The physical and chemical data of fermented grains mainly include moisture, acidity, sugar content, and starch content. These data directly affect the normal progress of the fermentation process, the formation of liquor quality, and the accuracy of brewing operations. Therefore, determining the physical and chemical data of fermented grains is crucial for the brewing process and also of great significance for ensuring the quality and yield of liquor.
[0003] In the prior art, the physical and chemical data of fermented grains are usually determined by corresponding detection methods. For example, the moisture of fermented grains is detected by the drying method, the acidity is detected by the NaOH standard solution neutralization titration method, the sugar content is determined by the Fehling reagent method, and the starch is detected by the hydrochloric acid hydrolysis standard glucose solution back-titration method. These methods rely on manual detection, with cumbersome operation steps, strict detection condition limitations, high requirements for the experience and skills of detection personnel, and also have the problem of strong subjectivity, resulting in low efficiency and accuracy, and it is difficult to meet the requirements of modern production for high efficiency, precision, and real-time monitoring. Summary of the Invention
[0004] The present invention aims to solve the problem of low efficiency and accuracy in the existing method for determining the physical and chemical data of fermented grains, and proposes a multimodal-based physical and chemical prediction method and device for fermented grains.
[0005] The technical solution adopted by the present invention to solve the above technical problems is as follows:
[0006] In a first aspect, the present invention provides a multimodal-based physical and chemical prediction method for fermented grains, the method comprising:
[0007] Obtain the multimodal data of multiple fermented grain samples, and establish a multimodal database according to the multimodal data, the multimodal data including visual data, physical and chemical data, team data, and hierarchical data;
[0008] Perform multimodal data coupling on the visual data, team data, and hierarchical data in the multimodal database to obtain multimodal coupled data, and train a multimodal neural network model according to the multimodal coupled data and its corresponding physical and chemical data to obtain a multimodal physical and chemical prediction model;
[0009] Perform multimodal data coupling on the visual data, team data, and hierarchical data of the fermented grains to be predicted to obtain the multimodal coupled data of the fermented grains to be predicted, and input the multimodal coupled data of the fermented grains to be predicted into the multimodal physical and chemical prediction model to obtain the physical and chemical prediction result of the fermented grains to be predicted.
[0010] Further, the visual data is a color photo, and the shooting angles, spatial resolutions, and light source white balances of the color photos corresponding to each fermented grains sample and the fermented grains to be predicted are the same; the color photo is a square color photo with a pixel resolution of 64×64, 128×128, 256×256, or 512×512;
[0011] The physical and chemical data includes the moisture content, acidity, sugar content, and starch content of the fermented grains;
[0012] The hierarchical data is the upper-layer fermented grains, middle-layer fermented grains, and lower-layer fermented grains divided according to the position of the fermented grains, or the dry fermented grains and wet fermented grains divided according to the contact relationship between the fermented grains and the yellow water.
[0013] Further, the fermented grains to be predicted are the fermented grains to be predicted for out-cellar or the fermented grains to be predicted for in-cellar.
[0014] Further, the coupling of the multi-modal data specifically includes:
[0015] Use a convolutional neural network to extract high-dimensional visual features from the visual data, and represent the high-dimensional visual features as a one-dimensional visual tensor. Use one-hot vectors to represent the team data and hierarchical data respectively to obtain a one-dimensional team tensor and a one-dimensional hierarchical tensor;
[0016] Concatenate the one-dimensional visual tensor, one-dimensional team tensor, and one-dimensional hierarchical tensor to obtain a one-dimensional multi-modal data, and input the one-dimensional multi-modal data into a one-dimensional data processing neural network to obtain multi-modal coupled data.
[0017] Further, the multi-modal neural network model is a single multi-modal neural network that simultaneously predicts multiple physical and chemical data, or a composite multi-modal neural network composed of multiple multi-modal neural networks that independently predict a single physical and chemical data. The multi-modal neural network is a CNN, MLP, LSTM, GRU, 1DCNN, Transformer, or RNN.
[0018] Further, training the multi-modal neural network model specifically includes:
[0019] Divide the multi-modal data in the multi-modal database into a training set and a test set according to a preset ratio, use the training set to train the multi-modal neural network model, and use the test set to determine the prediction accuracy of the multi-modal neural network model. When the error function of the test set is less than the error function threshold, and the determination coefficient of the physical and chemical prediction results and the true physical and chemical results of the test set is greater than the determination coefficient, the training of the multi-modal neural network model is completed.
[0020] Further, the error function is a SmoothL1 loss function, and the expression is as follows:
[0021]
[0022] Where SmoothL1(x, y) represents the SmoothL1 loss function, x represents the output of the multimodal neural network model, y represents the target of the multimodal neural network model, and β represents the smoothing threshold.
[0023] Furthermore, the method further includes:
[0024] Updating and supplementing the multimodal data in the multimodal database according to a preset period, and synchronously fine-tuning and updating the multimodal physicochemical prediction model. The preset period is monthly, quarterly, or a period determined according to the brewing cycle of the fermented grains.
[0025] Furthermore, the method further includes:
[0026] After establishing the multimodal database, for the visual data in the multimodal database, data augmentation is performed by means of random rotation, random flipping, random cropping, color jittering, and brightness adjustment.
[0027] After establishing the multimodal database and before coupling the multimodal data of the fermented grains to be predicted, the visual data and physicochemical data in the multimodal database and the visual data and physicochemical data of the fermented grains to be predicted are respectively standardized by a standardization method.
[0028] After obtaining the physicochemical prediction result of the fermented grains to be predicted, an inverse operation of the standardization method is performed on the physicochemical prediction result.
[0029] In a second aspect, the present invention provides a multimodal-based fermented grains physicochemical prediction device, and the device includes:
[0030] An acquisition module, configured to acquire multimodal data of multiple fermented grains samples, and establish a multimodal database according to the multimodal data. The multimodal data includes visual data, physicochemical data, team data, and hierarchical data.
[0031] A training module, configured to couple the visual data, team data, and hierarchical data in the multimodal database to obtain multimodal coupled data, and train a multimodal neural network model according to the multimodal coupled data and its corresponding physicochemical data to obtain a multimodal physicochemical prediction model.
[0032] A prediction module, configured to couple the visual data, team data, and hierarchical data of the fermented grains to be predicted to obtain multimodal coupled data of the fermented grains to be predicted, and input the multimodal coupled data of the fermented grains to be predicted into the multimodal physicochemical prediction model to obtain a physicochemical prediction result of the fermented grains to be predicted.
[0033] The beneficial effects of the present invention are as follows: The multi-modal-based physicochemical prediction method and device for fermented grains provided by the present invention construct a multi-modal physicochemical prediction model for predicting the physicochemical properties of fermented grains by coupling visual data, team data, and hierarchical data. By inputting the multi-modal coupled data of the fermented grains to be predicted into the multi-modal physicochemical prediction model, the physicochemical data of the fermented grains to be predicted can be obtained, thereby improving the efficiency of determining the physicochemical properties of fermented grains. Moreover, the multi-modal physicochemical prediction model can combine text information in image recognition tasks, make full use of various types of data, and through multi-modal fusion, capture the associations and dependencies between different modalities, conduct more complete learning and understanding of the input content, thereby enhancing the understanding ability and reasoning ability of the model, improving the accuracy and generalization ability of the model, and further improving the accuracy of determining the physicochemical data of fermented grains. In addition, when predicting the physicochemical data of fermented grains, by coupling the team data and hierarchical data of fermented grains, the differences in physicochemical data caused by different teams and levels of fermented grains can be avoided, further improving the accuracy of determining the physicochemical data of fermented grains. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Schematic flowchart of the multi-modal-based physicochemical prediction method for fermented grains provided in the embodiment;
[0035] Figure 2 Schematic diagram of the visual data of the fermented grain sample provided in the embodiment;
[0036] Figure 3 Schematic diagram of the loss change during the training process of the multi-modal neural network model provided in the embodiment;
[0037] Figure 4 Scatter plot of moisture in the physicochemical prediction results of the fermented grain sample provided in the embodiment;
[0038] Figure 5 Comparative bar chart of moisture in the physicochemical prediction results of the fermented grain sample provided in the embodiment;
[0039] Figure 6 Schematic diagram of the structure of the multi-modal-based physicochemical prediction device for fermented grains provided in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the present embodiment will be clearly and completely described below with reference to the accompanying drawings in the present embodiment.
[0041] In some processes described in the specification of the present invention and the above-mentioned drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear herein or in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations can be executed in sequence or in parallel.
[0042] In order to improve the efficiency and accuracy of determining the physical and chemical data of fermented grains, the technical solution of the present invention is proposed. In the present invention, multi-modal data of a plurality of fermented grain samples are obtained, and a multi-modal database is established according to the multi-modal data. The multi-modal data includes visual data, physical and chemical data, team data, and hierarchical data; the visual data, team data, and hierarchical data in the multi-modal database are subjected to multi-modal data coupling to obtain multi-modal coupling data, and a multi-modal neural network model is trained according to the multi-modal coupling data and its corresponding physical and chemical data to obtain a multi-modal physical and chemical prediction model; the visual data, team data, and hierarchical data of the fermented grains to be predicted are subjected to multi-modal data coupling to obtain multi-modal coupling data of the fermented grains to be predicted, and the multi-modal coupling data of the fermented grains to be predicted is input into the multi-modal physical and chemical prediction model to obtain the physical and chemical prediction result of the fermented grains to be predicted.
[0043] Specifically, in the present invention, by coupling the visual data, team data, and hierarchical data of the fermented grain samples, the obtained multi-modal coupling data is used as the input, and the corresponding physical and chemical data is used as the label for training the multi-modal neural network model. After the training is completed, a multi-modal physical and chemical prediction model is obtained. Then, the multi-modal coupling data of the fermented grains to be predicted is input into the multi-modal physical and chemical prediction model, and the physical and chemical prediction result of the fermented grains to be predicted can be obtained. The present invention avoids the problems of cumbersome operation, high detection requirements, and subjectivity existing in the manual detection method, and improves the efficiency and accuracy of determining the physical and chemical data of fermented grains. And by coupling data from different modalities, the associations and dependencies between different modalities can be captured, thereby enhancing the understanding ability and reasoning ability of the model, improving the performance of the model, and further improving the accuracy of determining the physical and chemical data of fermented grains. By determining the physical and chemical data of fermented grains, the smooth progress of the brewing process, the improvement of wine quality, and the accuracy of brewing operations can be ensured.
[0044] Next, the technical solutions in this embodiment will be clearly and completely described with reference to the accompanying drawings in this embodiment. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0045] Figure 1 The flowchart of a multi-modal-based method for predicting the physical and chemical properties of fermented grains is shown. Please refer to Figure 1 and the method includes the following steps:
[0046] S1. Obtain the multi-modal data of multiple fermented grains samples, and establish a multi-modal database according to the multi-modal data. The multi-modal data includes visual data, physical and chemical data, team data, and hierarchical data.
[0047] Please refer to Figure 2 , the visual data of the fermented grains sample is a color photo of the fermented grains sample, and the shooting angles, spatial resolutions, and light source white balances corresponding to the color photos of each fermented grains sample are the same; the color photo is a square color photo with a pixel resolution of 64×64, 128×128, 256×256, or 512×512.
[0048] In practical applications, a color photo of the fermented grains sample can be obtained by using a camera. During shooting, keep the same shooting angle, spatial resolution, and light source white balance, and then crop the color photo of the fermented grains into a square.
[0049] The physical and chemical data of the fermented grains sample includes the moisture, acidity, sugar content, and starch content of the fermented grains sample. In practical applications, the moisture of the fermented grains sample can be detected by the drying method, the acidity of the fermented grains sample can be detected by the NaOH standard solution neutralization titration method, the sugar content of the fermented grains sample can be determined by the Fehling reagent method, and the starch content of the fermented grains sample can be detected by the hydrochloric acid hydrolysis standard glucose solution back-titration method. In practical applications, the physical and chemical data of the fermented grains sample can also be obtained through physical and chemical tests or near-infrared physical and chemical analyzers.
[0050] The team data is the information of the production team to which the fermented grains belong, represented in text or numbers. In this embodiment, the team data includes 6 teams.
[0051] The hierarchical data is the upper-layer fermented grains, middle-layer fermented grains, and lower-layer fermented grains divided according to the position of the fermented grains, or the dry fermented grains and wet fermented grains divided according to the contact relationship between the fermented grains and the yellow water. In this embodiment, the hierarchical data includes three levels: upper-layer fermented grains, middle-layer fermented grains, and lower-layer fermented grains.
[0052] In practical applications, the above physical and chemical data, team data, and hierarchical data may also include other physical and chemical data, team data, and hierarchical data recognized by experts in the field.
[0053] After obtaining the multi-modal data of each fermented grains sample, establish a multi-modal database, and save it to the multi-modal database according to the corresponding relationship between the fermented grains sample and the multi-modal data.
[0054] In this embodiment, it further includes: after establishing the multi-modal database, for the visual data in the multi-modal database, perform data augmentation by means of random rotation, random flipping, random cropping, color jittering, and brightness adjustment; and perform standardization processing on the visual data and physical and chemical data in the multi-modal database respectively by using the standardization method.
[0055] In this embodiment, the visual data is randomly cropped into a square color photo of 128×128 with a pixel resolution of 512×512 for data augmentation. In practical applications, based on Python, the three functions of transforms.RandomRotation, transforms.RandomHorizontalFlip, and transforms.RandomCrop in the torchvision library can be used to respectively implement random rotation, random flipping, and random cropping of the visual data. By performing data augmentation on the visual data, the number of training samples is increased, more samples can be provided for the model to learn, which helps to reduce the overfitting of the model to specific training data. Moreover, the data augmentation introduces randomness and generates more diverse training samples, enabling the model to learn more general features, so that the model has stronger generalization ability for new samples. In addition, it can also reduce the data acquisition cost and improve the training efficiency.
[0056] In this embodiment, by performing standardization processing on the visual data and physicochemical data, data of different magnitudes can be converted to a unified standard, enabling the algorithm to evaluate the importance of each feature on an equal footing, thereby improving the accuracy and stability of model prediction. The standardization methods include min-max standardization method, range standardization method, Z-score standardization method, etc.
[0057] S2. Couple the visual data, team data, and hierarchical data in the multimodal database to obtain multimodal coupled data, and train a multimodal neural network model based on the multimodal coupled data and its corresponding physicochemical data to obtain a multimodal physicochemical prediction model.
[0058] It can be understood that due to factors such as operation differences and different raw material batches in different teams, there may be differences in the physicochemical parameters of the fermented grains. There are significant differences in temperature, viscosity, sugar content, flavor substances, etc. among the fermented grains stacked at different levels. These differences may affect the fermentation effect of the fermented grains and the flavor of the final liquor product. Based on this, in this embodiment, by coupling the team data and hierarchical data, the differences in physicochemical data caused by different teams and levels of the fermented grains can be avoided, further improving the accuracy of determining the physicochemical data of the fermented grains.
[0059] In this embodiment, the multimodal data coupling specifically includes:
[0060] Extract high-dimensional visual features from visual data using a convolutional neural network, and represent the high-dimensional visual features as a one-dimensional visual tensor. Use one-hot vectors to represent the team data and hierarchical data respectively to obtain a one-dimensional team tensor and a one-dimensional hierarchical tensor. Concatenate the one-dimensional visual tensor, the one-dimensional team tensor, and the one-dimensional hierarchical tensor to obtain a one-dimensional multimodal data, and input the one-dimensional multimodal data into a one-dimensional data processing neural network to obtain multimodal coupled data.
[0061] Specifically, if the high-dimensional visual features extracted from visual data using a convolutional neural network are a three-dimensional tensor with a dimension of 6×6×3, then representing the high-dimensional visual features as a one-dimensional visual tensor can directly save each element as a one-dimensional tensor with 108 elements. In this embodiment, there are 6 teams in total, which are represented using one-hot vectors. Then the first team can be represented as [1,0,0,0,0,0]. Therefore, the one-dimensional team tensor is a one-dimensional tensor with 6 elements. In this embodiment, three levels, namely the upper layer of fermented grains, the middle layer of fermented grains, and the lower layer of fermented grains, are considered. Therefore, using one-hot vectors to represent, the upper layer of fermented grains can be represented as [1,0,0]. Concatenate the one-dimensional visual tensor, the one-dimensional team tensor, and the one-dimensional hierarchical tensor to obtain a one-dimensional multimodal data, which is a one-dimensional tensor with 117 elements. Input the one-dimensional multimodal data into a one-dimensional data processing neural network to obtain multimodal coupled data. If the output dimension of the one-dimensional data processing neural network is 32, then the finally obtained multimodal coupled data is a one-dimensional tensor with 32 elements.
[0062] After obtaining the multimodal coupled data of each fermented grains sample, use the multimodal coupled data as the input and the corresponding physicochemical data as the output to train the multimodal neural network model, which specifically includes:
[0063] Divide the multimodal data in the multimodal database into a training set and a test set according to a preset ratio. The visual data, team data, and hierarchical data in the training set and the test set are coupled into corresponding multimodal coupled data. Use the training set to train the multimodal neural network model, and use the test set to determine the prediction accuracy of the multimodal neural network model. When the error function of the test set is less than the error function threshold, and the determination coefficient between the physicochemical prediction result and the true physicochemical result of the test set is greater than the determination coefficient, the training of the multimodal neural network model is completed.
[0064] In this embodiment, the error function can use the loss function, or other error functions recognized by experts in the field, such as the root mean square error and the mean absolute error. When the test set R 2 is greater than the preset R 2 threshold, it is considered that the model training is completed. The R 2 threshold can be set according to the actual situation. For example, the R 2 threshold is 0.8.
[0065] Among them, the loss function can be the SmoothL1 loss function. During the training process, the changes in the loss functions of the training set and the test set are as Figure 3 shown, and the expression is as follows:
[0066]
[0067] Among them, SmoothL1(x, y) represents the SmoothL1 loss function, x represents the output of the multimodal neural network model, y represents the target of the multimodal neural network model, and β represents the smoothing threshold.
[0068] In practical applications, the multimodal neural network model can be a single multimodal neural network that simultaneously predicts multiple physical and chemical data, or a composite multimodal neural network composed of multiple multimodal neural networks that independently predict a single physical and chemical data. The multimodal neural network can be CNN, MLP, LSTM, GRU, 1DCNN, Transformer, or RNN.
[0069] In this embodiment, the multimodal neural network is a combined network of CNN and MLP. The multimodal data coupling and the construction of the multimodal neural network are realized by programming based on the Python language using the PyTorch library, as follows:
[0070]
[0071]
[0072]
[0073] Figure 4 Shows the comparison result between the moisture prediction result of the test set data corresponding to a certain work team and the true moisture content of the fermented grains. It can be seen that R 2 is greater than 0.9, greater than the R 2 threshold of 0.8 in this embodiment. It can be considered that the model training is completed. Figure 5 The comparison between the moisture prediction result of the test set data corresponding to a certain work team and the true moisture content of the fermented grains is further shown in the form of a bar chart. It can be seen that the model prediction effect is good.
[0074] S3. Perform multimodal data coupling on the visual data, work team data, and hierarchical data of the fermented grains to be predicted to obtain the multimodal coupled data of the fermented grains to be predicted. Input the multimodal coupled data of the fermented grains to be predicted into the multimodal physical and chemical prediction model to obtain the physical and chemical prediction result of the fermented grains to be predicted.
[0075] In this embodiment, the fermented grains to be predicted can be the fermented grains to be taken out of the cellar or the fermented grains to be put into the cellar. When it is necessary to obtain the physical and chemical data of the fermented grains to be predicted, first obtain the visual data, team data, and hierarchical data of the fermented grains to be predicted, and perform standardization processing on the visual data and physical and chemical data of the fermented grains to be predicted using the same standardization method as the multi-modal database. Then perform multi-modal data coupling on the hierarchical data of the fermented grains to be predicted and the visually and physically standardized data to obtain the multi-modal coupled data of the fermented grains to be predicted. The method of multi-modal data coupling is the same as that of the fermented grain samples and will not be elaborated here. Finally, input the multi-modal coupled data of the fermented grains to be predicted into the trained multi-modal physical and chemical prediction model to obtain the physical and chemical prediction results of the fermented grains to be predicted, that is, the physical and chemical data of the fermented grains to be predicted.
[0076] After obtaining the physical and chemical prediction results of the fermented grains to be predicted, perform the inverse operation of the standardization method on the physical and chemical prediction results to restore the original magnitude of the physical and chemical data of the fermented grains.
[0077] In this embodiment, it further includes: updating and supplementing multi-modal data in the multi-modal database according to a preset cycle, and synchronously performing fine-tuning updates on the multi-modal physical and chemical prediction model. The preset cycle is monthly, quarterly, or a cycle determined according to the brewing cycle of the fermented grains.
[0078] In practical applications, the preset cycle can also be other update frequencies recognized by experts in the field. By regularly updating the sample database, the latest state of the data can be ensured, reducing prediction accuracy problems caused by stale or incorrect data, thereby further improving the accuracy of the physical and chemical prediction of the fermented grains.
[0079] In summary, the multi-modal-based physical and chemical prediction method for fermented grains provided in this embodiment constructs a multi-modal physical and chemical prediction model for predicting the physical and chemical properties of fermented grains by coupling visual data, team data, and hierarchical data. Inputting the multi-modal coupled data of the fermented grains to be predicted into the multi-modal physical and chemical prediction model can obtain the physical and chemical data of the fermented grains to be predicted, thereby improving the efficiency of determining the physical and chemical properties of the fermented grains. Moreover, the multi-modal physical and chemical prediction model can combine text information in image recognition tasks, make full use of various types of data, and through multi-modal fusion, capture the associations and dependencies between different modalities, perform more complete learning and understanding of the input content, thereby enhancing the model's understanding and reasoning abilities, improving the accuracy and generalization ability of the model, and further improving the accuracy of determining the physical and chemical data of the fermented grains. By determining the physical and chemical data of the fermented grains, the smooth progress of the brewing process, the quality of the wine, and the accuracy of brewing operations can be ensured.
[0080] Based on the above technical solution, this embodiment further proposes a multi-modal-based physical and chemical prediction device for fermented grains. Please refer to Figure 6 The device includes:
[0081] An acquisition module, configured to acquire multimodal data of multiple fermented grains samples, and establish a multimodal database according to the multimodal data, where the multimodal data includes visual data, physical and chemical data, team data, and hierarchical data;
[0082] A training module, configured to couple the visual data, team data, and hierarchical data in the multimodal database to obtain multimodal coupled data, and train a multimodal neural network model according to the multimodal coupled data and its corresponding physical and chemical data to obtain a multimodal physical and chemical prediction model;
[0083] A prediction module, configured to couple the visual data, team data, and hierarchical data of the fermented grains to be predicted to obtain multimodal coupled data of the fermented grains to be predicted, and input the multimodal coupled data of the fermented grains to be predicted into the multimodal physical and chemical prediction model to obtain a physical and chemical prediction result of the fermented grains to be predicted.
[0084] It can be understood that since the multimodal-based fermented grains physical and chemical prediction device described in this embodiment is a device for implementing the multimodal-based fermented grains physical and chemical prediction method described in the embodiment, for the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For related parts, refer to the partial description of the method, and details are not described herein again.
Claims
1. A multimodal prediction method for physicochemical properties of mash, characterized in that: The method comprises: Acquire multimodal data of a plurality of mash samples, and establish a multimodal database according to the multimodal data, wherein the multimodal data includes visual data, physical and chemical data, team data, and hierarchical data; Performing multimodal data coupling on the visual data, team data and hierarchical data in the multimodal database to obtain multimodal coupling data, training a multimodal neural network model according to the multimodal coupling data and its corresponding physical and chemical data to obtain a multimodal physical and chemical prediction model; The visual data, team data and hierarchical data of the mash to be predicted are multimodally coupled to obtain the multimodal coupled data of the mash to be predicted, and the multimodal coupled data of the mash to be predicted are input into the multimodal physical and chemical prediction model to obtain the physical and chemical prediction result of the mash to be predicted.
2. The multimodal prediction method for physicochemical properties of fermented grains according to claim 1, characterized in that: The visual data is a color photo, and the shooting angle, spatial resolution and light source white balance corresponding to the color photo of each fermented grain sample and the fermented grain to be predicted are the same; the color photo is a square color photo with a pixel resolution of 64×64, 128×128, 256×256 or 512×512; The physical and chemical data include the moisture, acidity, sugar and starch content of the mash; The hierarchical data are upper layer mash, middle layer mash and lower layer mash divided according to the position of mash, or dry mash and wet mash divided according to the contact relationship between mash and yellow water.
3. The multimodal based physicochemical prediction method for mash according to claim 1, characterized in that: The mash to be predicted is mash to be predicted to be taken out of the cellar or mash to be predicted to be put into the cellar.
4. The multimodal based physicochemical prediction method for mash according to claim 1, characterized in that: The multimodal data coupling specifically includes: A convolutional neural network is used to extract high-dimensional visual features from visual data, and the high-dimensional visual features are represented as a one-dimensional visual tensor. One-hot vectors are used to represent the class data and the level data, respectively, to obtain a one-dimensional class tensor and a one-dimensional level tensor. The one-dimensional visual tensor, the one-dimensional team tensor and the one-dimensional hierarchical tensor are spliced to obtain one-dimensional multimodal data, and the one-dimensional multimodal data is input into a one-dimensional data processing neural network to obtain multimodal coupling data.
5. The multimodal based physicochemical prediction method for mash according to claim 4, characterized in that: The multimodal neural network model is a single multimodal neural network that simultaneously predicts multiple physical and chemical data, or a composite multimodal neural network composed of multiple multimodal neural networks that independently predict single physical and chemical data. The multimodal neural network is CNN, MLP, LSTM, GRU, 1DCNN, Transformer or RNN.
6. The multimodal based physicochemical prediction method for mash according to claim 1, characterized in that: Training a multimodal neural network model, including: The multimodal data in the multimodal database are divided into a training set and a test set according to a preset ratio, the multimodal neural network model is trained using the training set, and the prediction accuracy of the multimodal neural network model is determined using the test set. When the error function of the test set is less than the error function threshold and the determination coefficient between the physicochemical prediction results and the actual physicochemical results of the test set is greater than the determination coefficient threshold, the training of the multimodal neural network model is completed.
7. The multimodal based physicochemical prediction method for mash according to claim 6, characterized in that: The error function is the SmoothL1 loss function, and the expression is as follows: Where SmoothL1(x,y) represents the SmoothL1 loss function, x represents the output of the multimodal neural network model, y represents the target of the multimodal neural network model, and β represents the smoothing threshold.
8. The multimodal based physicochemical prediction method for mash according to claim 1, characterized in that: The method further comprises: The multimodal database is updated and supplemented with multimodal data according to a preset period, and the multimodal physicochemical prediction model is fine-tuned and updated synchronously. The preset period is monthly, quarterly, or a period determined according to the brewing cycle of the mash.
9. The method for predicting the physicochemical properties of fermented grains based on multimodality according to claim 1, characterized in that: The method further comprises: After establishing the multimodal database, data augmentation is performed on the visual data in the multimodal database by using random rotation, random flipping, random cropping, color jittering and brightness adjustment. After establishing the multimodal database and before coupling the multimodal data of the mash to be predicted, a standardization method is used to standardize the visual data and the physicochemical data in the multimodal database and the visual data and the physicochemical data of the mash to be predicted respectively; After obtaining the physicochemical prediction results of the mash to be predicted, the inverse operation of the standardization method is performed on the physicochemical prediction results.
10. A multi-modal mash physicochemical prediction device, characterized in that: The device comprises: An acquisition module, used for acquiring multimodal data of a plurality of mash samples, and establishing a multimodal database according to the multimodal data, wherein the multimodal data includes visual data, physical and chemical data, team data and hierarchical data; A training module, used for performing multimodal data coupling on the visual data, team data and hierarchical data in the multimodal database to obtain multimodal coupling data, training a multimodal neural network model according to the multimodal coupling data and its corresponding physical and chemical data to obtain a multimodal physical and chemical prediction model; The prediction module is used to perform multimodal data coupling on the visual data, team data and hierarchical data of the mash to be predicted to obtain the multimodal coupled data of the mash to be predicted, and input the multimodal coupled data of the mash to be predicted into the multimodal physical and chemical prediction model to obtain the physical and chemical prediction results of the mash to be predicted.