Multi-mode-based Luzhou-flavor liquor bag koji making grade identification method and device
By establishing a multimodal database and building a deep learning model, automatic recognition of baggage music levels is achieved, the problem of poor efficiency and accuracy of existing methods is solved, and the accuracy and efficiency of recognition is improved.
Patent Information
- Application Number
- CN202510331440.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-06-17
AI Technical Summary
The existing methods for determining the level of bag tufts have poor efficiency and accuracy, the sensory evaluation method is highly subjective, and the equipment threshold for physical and chemical index detection method is high and the efficiency is low.
Using a multimodal recognition method, by establishing a multimodal database, a bag curve cross-section recognition model and a grade recognition model are constructed, and combined with object detection and deep learning models, automatic recognition of bag curve grade is realized.
Reliance on manual experience is reduced, the accuracy and efficiency of grade recognition is improved, the identification threshold is reduced, and batch recognition of baggage music grades can be performed.
Smart Images

Figure CN120164038A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of brewing, and particularly relates to a method and device for identifying the grades of Baobaoqu for Luzhou-flavor liquor based on multi-modalities. Background Art
[0002] Baobaoqu is the core saccharifying and fermenting agent for Luzhou-flavor liquor. Because the middle part of the qu embryo bulges, it is called "Baobaoqu". The complex bacterial system and enzyme system of Baobaoqu can promote esterification, aroma generation and the accumulation of flavor substances. The compound aroma of Luzhou-flavor liquor is closely related to the rich flavor substances in Baobaoqu. The key saccharifying and fermenting ability of Baobaoqu supports the characteristics of mellow body and strong aroma of the liquor body.
[0003] The types of microorganisms, enzyme activities and metabolites contained in Baobaoqu of different grades are different. By accurately determining the grade of Baobaoqu, stable-quality qu blocks can be screened out, batch differences can be reduced, and it is convenient to accurately match fermentation conditions (such as temperature, humidity), optimize and control the Daqu fermentation process, and ensure the stable quality of the liquor body.
[0004] In the traditional production process, the methods for determining the grade of Baobaoqu mainly include the sensory evaluation method and the physical and chemical index detection method. The sensory evaluation method is mainly for the qu graders to classify by observing the appearance of the qu block (such as color, mycelium distribution), smelling the aroma (such as soy sauce aroma, cellar aroma characteristics) and feeling the touch (humidity, hardness). However, this method has high requirements for the qu graders, strong subjectivity, and there may be deviations in the grade judgment of the same qu block by different qu graders, with low robustness; and it takes a long time to train an experienced qu grader; the physical and chemical index detection method is mainly to determine the physical and chemical parameters of the qu block (such as saccharifying power, liquefying power, esterifying power, fermenting power, acidity, etc.) for classification. However, this method has high detection equipment and technical thresholds, the detection method is cumbersome, and it can only detect single Baobaoqu one by one, with low efficiency. Summary of the Invention
[0005] The present invention aims to solve the problems of poor efficiency and accuracy in the existing methods for determining the grade of Baobaoqu, and proposes a method and device for identifying the grades of Baobaoqu for Luzhou-flavor liquor based on multi-modalities.
[0006] The technical solution adopted by the present invention to solve the above technical problems is as follows:
[0007] In a first aspect, the present invention provides a method for identifying the grades of Baobaoqu for Luzhou-flavor liquor based on multi-modalities, and the method includes:
[0008] Establish a multi-modal database for Baobaoqu of Luzhou-flavor liquor, where the multi-modal database includes the overall image data and cross-section annotation data of Baobaoqu, as well as the cross-section visual data, production time data and grade data of single Baobaoqu;
[0009] Construct a bag curved cross-section recognition model based on the overall image data and cross-section annotation data, and construct a bag curved grade recognition model based on the cross-section visual data, production time data, and grade data;
[0010] Obtain the overall image data and production time data of the bag curved to be recognized, and input the overall image data of the bag curved to be recognized into the bag curved cross-section recognition model to obtain the cross-section annotation data corresponding to the bag curved to be recognized;
[0011] According to the cross-section annotation data corresponding to the bag curved to be recognized, crop the overall image data of the bag curved to be recognized to obtain the cross-section visual data of all individual bag curves corresponding to the bag curved to be recognized, and input the cross-section visual data and production time data of each individual bag curve corresponding to the bag curved to be recognized into the bag curved grade recognition model respectively to obtain the grade recognition results of each individual bag curve.
[0012] Further, the overall image data is a color photo taken by a high-resolution camera and includes multiple bag curves;
[0013] The cross-section annotation data is, based on the overall image data, the border coordinate data of the borders marking each individual bag curve cross-section;
[0014] The cross-section visual data is the cross-section image of an individual bag curve.
[0015] Further, the production time data is the production month data or production quarter data of the bag curve.
[0016] Further, the grade data at least includes first-class curve, second-class curve, and third-class curve.
[0017] Further, the bag curved cross-section recognition model is an object detection neural network model constructed based on Faster-RCNN, RetinaNet, YOLO, EfficientDet, or SSD.
[0018] Further, the bag curved grade recognition model includes a cross-section feature extraction module and a cross-section grade recognition module;
[0019] The cross-section feature extraction module is used to extract high-dimensional cross-section features from the cross-section visual data of an individual bag curve, reconstruct the high-dimensional cross-section features into a one-dimensional cross-section feature tensor, and splice and combine the one-dimensional cross-section feature tensor with the production time data to obtain a combined cross-section feature tensor. The cross-section grade recognition module is used to obtain the grade recognition result of an individual bag curve according to the combined cross-section feature tensor.
[0020] Further, the cross-section feature extraction module is a convolutional neural network or an attention mechanism model. The convolutional neural network is ResNet, EfficientNet or ConvNeXt, and the attention mechanism model is Self-Attention or Transformer. The cross-section level recognition module is an MLP network.
[0021] Further, the method further includes:
[0022] Determine the number of target-level baguettes in the baguettes to be recognized according to the level recognition results of each individual baguette in the baguettes to be recognized, and calculate the proportion of the target-level baguettes in the baguettes to be recognized according to the number of target-level baguettes and the total number of baguettes.
[0023] Further, the method further includes:
[0024] Update and expand the multi-modal database according to a preset period, and after updating and expanding the multi-modal database, update the baguette cross-section recognition model and the baguette level recognition model.
[0025] In a second aspect, the present invention provides a multi-modal-based recognition device for the levels of Nongxiangxing baijiu baguettes. The device includes:
[0026] A database establishment module for establishing a multi-modal database of Nongxiangxing baijiu baguettes. The multi-modal database includes the overall image data and cross-section annotation data of the baguettes, as well as the cross-section visual data, production time data and level data of each individual baguette;
[0027] A model establishment module for constructing a baguette cross-section recognition model according to the overall image data and cross-section annotation data, and constructing a baguette level recognition model according to the cross-section visual data, production time data and level data;
[0028] A level recognition module for obtaining the overall image data and production time data of the baguettes to be recognized, inputting the overall image data of the baguettes to be recognized into the baguette cross-section recognition model to obtain the corresponding cross-section annotation data of the baguettes to be recognized; according to the corresponding cross-section annotation data of the baguettes to be recognized, cropping the overall image data of the baguettes to be recognized to obtain the cross-section visual data of all individual baguettes corresponding to the baguettes to be recognized, and inputting the cross-section visual data and production time data of each individual baguette corresponding to the baguettes to be recognized into the baguette level recognition model respectively to obtain the level recognition results of each individual baguette.
[0029] The beneficial effects of the present invention are as follows: The method and device for identifying the grades of Baobaoqu for Luzhou-flavor liquor based on multimodality provided by the present invention combine object detection and deep learning models to achieve automatic identification of the overall image of Baobaoqu to the grades of individual Baobaoqu. Compared with the sensory evaluation method and the physical and chemical index detection method, it reduces the dependence on manual experience, avoids subjectivity, improves the accuracy of grade identification, and does not require professional detection equipment and detection methods, reducing the threshold of grade identification. At the same time, the present invention can perform batch identification of Baobaoqu grades from the overall image containing multiple Baobaoqu, improving the efficiency of grade identification. In addition, the present invention combines the cross-sectional visual data of Baobaoqu with the production time data for grade identification. The cross-sectional visual data can reflect the microscopic characteristics such as the hypha density and porosity of the koji block, and the production time data can reflect the growth environment characteristics of the koji block. The combination of the two can cover the multi-dimensional differences between Baobaoqu grades and reduce the misjudgment risk of a single data source, thereby improving the accuracy of Baobaoqu grade identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Schematic flow chart of the method for identifying the grades of Baobaoqu for Luzhou-flavor liquor based on multimodality provided in the embodiment;
[0031] Figure 2 Schematic diagram of the overall image data of Baobaoqu provided in the embodiment;
[0032] Figure 3 Schematic diagram of the cross-sectional annotation data of Baobaoqu provided in the embodiment;
[0033] Figure 4 Schematic diagram of the principle of the Baobaoqu grade identification model provided in the embodiment;
[0034] Figure 5 Schematic diagram of the grade identification results of each individual Baobaoqu provided in the embodiment;
[0035] Figure 6 Schematic diagram of the structure of the device for identifying the grades of Baobaoqu for Luzhou-flavor liquor based on multimodality provided in the embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0036] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments will be clearly and completely described below with reference to the accompanying drawings in the embodiments.
[0037] In some processes described in the specification of the present invention and the above-mentioned drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel.
[0038] In order to improve the efficiency and accuracy of the recognition of the grade of Baobaoqu, the technical solution of the present invention is proposed. In the present invention, a multi-modal database of Luzhou-flavor Baijiu Baobaoqu is established. The multi-modal database includes the overall image data and cross-section annotation data of Baobaoqu, as well as the cross-section visual data, production time data, and grade data of a single Baobaoqu; a Baobaoqu cross-section recognition model is constructed according to the overall image data and cross-section annotation data, and a Baobaoqu grade recognition model is constructed according to the cross-section visual data, production time data, and grade data; the overall image data and production time data of the Baobaoqu to be recognized are obtained, and the overall image data of the Baobaoqu to be recognized is input into the Baobaoqu cross-section recognition model to obtain the corresponding cross-section annotation data of the Baobaoqu to be recognized; according to the corresponding cross-section annotation data of the Baobaoqu to be recognized, the overall image data of the Baobaoqu to be recognized is cropped to obtain the cross-section visual data of all single Baobaoqu corresponding to the Baobaoqu to be recognized, and the cross-section visual data and production time data of each single Baobaoqu corresponding to the Baobaoqu to be recognized are respectively input into the Baobaoqu grade recognition model to obtain the grade recognition results of each single Baobaoqu.
[0039] Specifically, the present invention respectively constructs a Baobaoqu cross-section recognition model and a Baobaoqu grade recognition model, uses the Baobaoqu cross-section recognition model to perform Baobaoqu target detection on the overall image data containing multiple Baobaoqu images, and crops to obtain the cross-section visual data of a single Baobaoqu, and uses the Baobaoqu grade recognition model to respectively perform Baobaoqu grade recognition on the cross-section visual data of each Baobaoqu and its corresponding production time data. Through the above process, batch recognition of the grade of Baobaoqu can be performed from the overall image containing multiple Baobaoqu, thereby improving the recognition efficiency. In addition, by combining the cross-section visual data of Baobaoqu with the production time data for grade recognition, the multi-dimensional differences between Baobaoqu grades can be covered, and the risk of misjudgment of a single data source can be reduced, thereby improving the accuracy of Baobaoqu grade recognition.
[0040] Next, the technical solutions in the present embodiment will be clearly and completely described in conjunction with the accompanying drawings in the present embodiment. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0041] Figure 1 The flowchart of a method for recognizing the grade of Luzhou-flavor Baijiu Baobaoqu based on multi-modal is shown. Please refer toFigure 1 , the method includes the following steps:
[0042] Step 1: Establish a multi-modal database for the wrapped starter cakes of Luzhou-flavor liquor. The multi-modal database includes the overall image data and cross-section annotation data of the wrapped starter cakes, as well as the cross-section visual data, production time data, and grade data of individual wrapped starter cakes.
[0043] Among them, the overall image data is a color photo including multiple wrapped starter cakes taken by a high-resolution camera. In this embodiment, a Sony ILCE-7M4 camera and a SEL2470GM2 lens are used to take the cross-section images of the wrapped starter cakes under standard lighting conditions. All shootings are carried out in a controlled environment to ensure the unity of lighting, angle, and distance. As Figure 2 shown, each color photo contains 10 wrapped starter cakes, with a total of 20 cross-sections, and is stored at a high resolution of 3504×2336 pixels to capture the subtle differences in cross-section textures.
[0044] The cross-section annotation data is, based on the overall image data, the border coordinate data of each individual wrapped starter cake cross-section marked with a border. As Figure 3 shown, in this embodiment, data annotation is manually performed using the LabelStudio tool. During the annotation process, according to the actual contour of the wrapped starter cake, accurate border annotation is carried out on the cross-section area, and the border is saved in the multi-modal database in the form of coordinate data.
[0045] The cross-section visual data is the cross-section image of an individual wrapped starter cake. In practical applications, the cross-section annotation image can be cropped to obtain the cross-section visual data of each wrapped starter cake.
[0046] The production time data is the production month data or production quarter data of the wrapped starter cake. In this embodiment, the production time data is represented by the production quarter, that is, a year is divided into four quarters and represented by a one-hot vector. For example, [1,0,0,0] represents the first quarter, and [0,0,0,1] represents the fourth quarter.
[0047] The grade data is the grade judged according to enterprise specifications and by engineers with professional qualifications, and at least includes first-grade starter cakes, second-grade starter cakes, and third-grade starter cakes.
[0048] Step 2: Construct a wrapped starter cake cross-section recognition model based on the overall image data and cross-section annotation data, and construct a wrapped starter cake grade recognition model based on the cross-section visual data, production time data, and grade data.
[0049] In practical applications, the overall image data in the multi-modal database is used as input features, and the corresponding cross-section annotation data is used as the ground truth labels to train a bag curvature cross-section recognition model. Among them, the bag curvature cross-section recognition model is an object detection neural network model for object detection. The object detection neural network model can be Faster-RCNN, RetinaNet, YOLO, EfficientDet or SSD.
[0050] In this embodiment, the Faster-RCNN neural network is selected as the bag curvature cross-section recognition model. Its loss function mainly includes the ROIHead loss and the RPN (Region Proposal Network) loss. Among them, the ROIHead loss includes the classification loss and the bounding box regression loss, and the RPN (Region Proposal Network) loss includes the objectness loss and the anchor box regression loss. The above loss function is embedded in the definition of the Faster-RCNN neural network. The Faster-RCNN network is defined using the torchvision.models.detection.fasterrcnn_resnet50_fpn function. The above-mentioned losses are automatically calculated by the command loss_dict = model(images,targets), and then these loss terms are added together through losses = sum(loss for loss in loss_dict.values()) to obtain the final total loss, which is used for backpropagation and model parameter update. Among them, images is the processed image input to the neural network, target is a tensor with the shape of [N,4], N is the number of detected objects, and 4 means using 4 values to represent the rectangular bounding box. In this embodiment, [xmin, ymin, xmax, ymax] is used to represent it.
[0051] In practical applications, the cross-section visual data and production time data of a single bag curvature are used as input features, and the corresponding grade data is used as the ground truth labels to train a bag curvature grade recognition model.
[0052] In this embodiment, the bag curvature grade recognition model includes a cross-section feature extraction module and a cross-section grade recognition module. The cross-section feature extraction module is used to extract high-dimensional cross-section features from the cross-section visual data of a single bag curvature, reconstruct the high-dimensional cross-section features into a one-dimensional cross-section feature tensor, and splice and combine the one-dimensional cross-section feature tensor with the production time data to obtain a combined cross-section feature tensor. The cross-section grade recognition module is used to obtain the grade recognition result of a single bag curvature according to the combined cross-section feature tensor.
[0053] Among them, the cross-section feature extraction module is a convolutional neural network or an attention mechanism model. The convolutional neural network is ResNet, EfficientNet, or ConvNeXt, and the attention mechanism model is Self-Attention or Transformer. The cross-section level recognition module is an MLP network.
[0054] In this embodiment, the ResNet18 network is used as the cross-section feature extraction module. The cross-section data of a single wrapped curve is input into the ResNet18 network, and the output is high-dimensional cross-section features. Then, the high-dimensional cross-section features are reconstructed into a one-dimensional cross-section feature tensor. The one-dimensional cross-section feature tensor and the production time data of the wrapped curve represented by a one-hot vector are concatenated and combined to obtain a combined cross-section feature tensor. Then, the combined cross-section feature tensor is input into the cross-section level recognition module represented by the MLP network, and the output is the level of the wrapped curve corresponding to the cross-section. The architecture of the wrapped curve level recognition model is as Figure 4 shown.
[0055] Step 3: Obtain the overall image data and production time data of the wrapped curve to be recognized, and input the overall image data of the wrapped curve to be recognized into the wrapped curve cross-section recognition model to obtain the cross-section annotation data corresponding to the wrapped curve to be recognized.
[0056] When batch-level recognition of wrapped curves is required, the overall image of the wrapped curve to be recognized is captured using the same method as the overall image data in the multi-modal database, that is, keeping the same shooting parameters and background, and then the corresponding overall image data is input into the wrapped curve cross-section recognition model. The wrapped curve cross-section recognition model performs object detection on the overall image data to obtain the cross-section annotation data of the wrapped curve to be recognized.
[0057] Step 4: According to the cross-section annotation data corresponding to the wrapped curve to be recognized, crop the overall image data of the wrapped curve to be recognized to obtain the cross-section visual data of all single wrapped curves corresponding to the wrapped curve to be recognized. Input the cross-section visual data and production time data of each single wrapped curve corresponding to the wrapped curve to be recognized into the wrapped curve level recognition model respectively to obtain the level recognition results of each single wrapped curve.
[0058] After obtaining the cross-section annotation data of the wrapped curve to be recognized, crop the overall image data of the wrapped curve to be recognized one by one according to the cross-section annotation data of the wrapped curve to be recognized to obtain the cross-section visual data of each single wrapped curve in the overall image data of the wrapped curve to be recognized, and input them into the wrapped curve level recognition model respectively to obtain the level recognition results of each single wrapped curve. Please refer to Figure 5 , in this level recognition result, there are 4 first-level curves (8 cross-sections), 6 second-level curves (12 cross-sections), and 0 third-level curves.
[0059] In this embodiment, the method further includes: determining the number of target-level baguqu according to the recognition results of each individual baguqu in the baguqu to be recognized, and calculating the proportion of the target-level baguqu in the baguqu to be recognized according to the number of target-level baguqu and the total number of baguqu.
[0060] For example, in Figure 5 , there are 4 first-level baguqu, 6 second-level baguqu, and 0 third-level baguqu. Then the proportion of the first-level baguqu is 40%, the proportion of the second-level baguqu is 60%, and the proportion of the third-level baguqu is 0%. By calculating the proportion of the target-level baguqu, the staff can intuitively understand the grade composition of the baguqu to be recognized, so as to carry out the next step of processing.
[0061] In this embodiment, the method further includes: updating and expanding the multi-modal database according to a preset period, and after updating and expanding the multi-modal database, updating the baguqu cross-section recognition model and the baguqu grade recognition model.
[0062] Among them, the preset period can be set according to the actual situation, and this embodiment does not make restrictions. By regularly updating the multi-modal database and the model, the data quality and accuracy can be improved, thereby improving the performance and accuracy of the model, and further improving the accuracy of baguqu grade recognition.
[0063] In summary, the multi-modal-based Luzhou-flavor liquor baguqu grade recognition device provided in this embodiment combines object detection and deep learning models to achieve automatic recognition of the overall image of baguqu to the grade of individual baguqu. Compared with the sensory evaluation method and the physical and chemical index detection method, it reduces the dependence on manual experience, avoids subjectivity, improves the accuracy of grade recognition, and does not require professional detection equipment and detection methods, reducing the grade recognition threshold. At the same time, the present invention can perform batch recognition of baguqu grades from the overall image containing multiple baguqu, improving the grade recognition efficiency. In addition, the present invention combines the cross-section visual data and the production time data of the baguqu for grade recognition. The cross-section visual data can reflect the microscopic features such as the hypha density and porosity of the koji block, and the production time data can reflect the growth environment characteristics of the koji block. The combination of the two can cover the multi-dimensional differences between baguqu grades and reduce the misjudgment risk of a single data source, thereby improving the accuracy of baguqu grade recognition.
[0064] Based on the above technical solutions, this embodiment also proposes a multi-modal-based Luzhou-flavor liquor baguqu grade recognition device. Please refer to Figure 6 , the device includes:
[0065] A database establishment module, configured to establish a multi-modal database of Luzhou-flavor liquor baguqu, where the multi-modal database includes the overall image data and cross-section annotation data of the baguqu, as well as the cross-section visual data, production time data, and grade data of individual baguqu;
[0066] A model building module, configured to build a recognition model for the cross-section of the bag-shaped starter according to the overall image data and cross-section annotation data, and build a recognition model for the grade of the bag-shaped starter according to the cross-section visual data, production time data and grade data;
[0067] A grade recognition module, configured to obtain the overall image data and production time data of the bag-shaped starter to be recognized, input the overall image data of the bag-shaped starter to be recognized into the recognition model for the cross-section of the bag-shaped starter, and obtain the corresponding cross-section annotation data of the bag-shaped starter to be recognized; according to the cross-section annotation data corresponding to the bag-shaped starter to be recognized, crop the overall image data of the bag-shaped starter to be recognized to obtain the cross-section visual data of all single bag-shaped starters corresponding to the bag-shaped starter to be recognized, and input the cross-section visual data and production time data of each single bag-shaped starter corresponding to the bag-shaped starter to be recognized into the recognition model for the grade of the bag-shaped starter respectively, and obtain the grade recognition results of each single bag-shaped starter.
[0068] It can be understood that since the device for recognizing the grade of the bag-shaped starter of Luzhou-flavor liquor based on multi-modal in this embodiment is a device for implementing the method for recognizing the grade of the bag-shaped starter of Luzhou-flavor liquor based on multi-modal in the embodiment, for the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method, and details are not described herein again.
Claims
1. A multi-modal method for identifying the grade of Luzhou-flavor liquor Baobaoqu, characterized in that: The method comprises: Establishing a multimodal database of Luzhou-flavor liquor Baobaoqu, the multimodal database includes overall image data and cross-sectional annotation data of Baobaoqu, as well as cross-sectional visual data, production time data, and grade data of a single Baobaoqu; Constructing a Baobaoqu cross-section recognition model based on the overall image data and the cross-section annotation data, and constructing a Baobaoqu grade recognition model based on the cross-section visual data, the production time data and the grade data; Obtaining the overall image data and production time data of the bag song to be identified, inputting the overall image data of the bag song to be identified into the bag song cross-section recognition model, and obtaining the cross-section annotation data corresponding to the bag song to be identified; According to the cross-sectional annotation data corresponding to the bag song to be identified, the overall image data of the bag song to be identified is cropped to obtain the cross-sectional visual data of all individual bag songs corresponding to the bag song to be identified, and the cross-sectional visual data and production time data of each individual bag song corresponding to the bag song to be identified are respectively input into the bag song grade recognition model to obtain the grade recognition result of each individual bag song.
2. The multi-modal method for identifying the grade of Luzhou-flavor liquor Baobaoqu according to claim 1 is characterized in that: The overall image data is a color photo of a plurality of bags taken by a high-resolution camera; The cross-section annotation data is frame coordinate data of each single bag curve cross section marked with a frame based on the overall image data; The cross-sectional visual data is a cross-sectional image of a single bag of music.
3. The multi-modal-based Luzhou-flavor liquor bag-shaped koji grade identification method according to claim 1 is characterized in that: The production time data is the production month data or production quarter data of the Bao Bao song.
4. The method for identifying the grade of Luzhou-flavor liquor Baobaoqu based on multimodality according to claim 1 is characterized in that: The grade data at least includes a first-level song, a second-level song and a third-level song.
5. The multi-modal-based Luzhou-flavor liquor baguette grade identification method according to claim 1 is characterized in that: The bag curved section recognition model is a target detection neural network model built based on Faster-RCNN, RetinaNet, YOLO, EfficientDet or SSD.
6. The multi-modal-based Luzhou-flavor liquor Baobaoqu grade identification method according to claim 1 is characterized in that: The Baobaoqu grade recognition model includes a cross-section feature extraction module and a cross-section grade recognition module; The cross-sectional feature extraction module is used to extract high-dimensional cross-sectional features from the cross-sectional visual data of a single bag song, reconstruct the high-dimensional cross-sectional features into a one-dimensional cross-sectional feature tensor, and concatenate the one-dimensional cross-sectional feature tensor with the production time data to obtain a combined cross-sectional feature tensor. The cross-sectional grade identification module is used to obtain a grade identification result of a single bag song based on the combined cross-sectional feature tensor.
7. The multi-modal method for identifying the grade of Luzhou-flavor liquor Baobaoqu according to claim 6 is characterized in that: The cross-section feature extraction module is a convolutional neural network or an attention mechanism model, the convolutional neural network is ResNet, EfficientNet or ConvNeXt, the attention mechanism model is Self-Attention or Transformer, and the cross-section grade identification module is an MLP network.
8. The multi-modal-based Luzhou-flavor liquor bag-shaped koji grade identification method according to claim 1 is characterized in that: The method further comprises: The number of target-level bag songs is determined according to the level recognition results of each individual bag song in the bag songs to be recognized, and the proportion of the target-level bag songs in the bag songs to be recognized is calculated according to the number of target-level bag songs and the total number of bag songs.
9. The method for identifying the grade of Luzhou-flavor liquor Baobaoqu based on multimodality according to claim 1 is characterized in that: The method further comprises: The multimodal database is updated and expanded according to a preset period, and after the multimodal database is updated and expanded, the bag curve cross-section recognition model and the bag curve grade recognition model are updated.
10. A multi-modal Luzhou-flavor liquor bag-shaped koji grade recognition device, characterized in that: The device comprises: A database establishment module is used to establish a multimodal database of Luzhou-flavor liquor Baobaoqu, wherein the multimodal database includes overall image data and cross-sectional annotation data of Baobaoqu and cross-sectional visual data, production time data and grade data of a single Baobaoqu; A model building module, for building a bag-shaped curve cross-section recognition model according to the overall image data and the cross-section annotation data, and building a bag-shaped curve grade recognition model according to the cross-section visual data, the production time data and the grade data; The grade recognition module is used to obtain the overall image data and production time data of the bag song to be identified, input the overall image data of the bag song to be identified into the bag song cross-section recognition model, and obtain the cross-section annotation data corresponding to the bag song to be identified; according to the cross-section annotation data corresponding to the bag song to be identified, the overall image data of the bag song to be identified is cropped to obtain the cross-section visual data of all individual bag songs corresponding to the bag song to be identified, and the cross-section visual data and production time data of each individual bag song corresponding to the bag song to be identified are respectively input into the bag song grade recognition model to obtain the grade recognition result of each individual bag song.