Food sensory preference degree identification method, device and equipment based on 3D facial micro-expression and medium

Through the food sensory preference recognition method based on 3D facial micro-expression, and using teacher models and knowledge distillation strategies to train student models, the problems of low recognition accuracy and poor adaptability in the existing technology are solved, and higher recognition accuracy and reliability are achieved.

CN120198950AActive Publication Date: 2025-06-24INNOVATION CENTER OF YANGTZE RIVER DELTA ZHEJIANG UNIVERSITY

Patent Information

Application Number
CN202510668600.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing food sensory preference recognition methods have problems such as high cost, low accuracy, strong subjectivity and high technical complexity. Especially in complex environments and facial occlusion, the recognition accuracy of 2D facial recognition technology is low.

Method used

The food sensory preference recognition method based on 3D facial micro-expression is adopted. The teacher model is trained by using the first data set, and the student model is trained using the knowledge distillation strategy, combined with 3D facial micro-expression data for recognition, and converted into 2D facial micro-expression data to input into the student model to obtain recognition results.

Benefits of technology

It improves the accuracy and reliability of food sensory preference recognition, enhances light adaptability and head posture adaptability, and can provide stable recognition results in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198950A_ABST
    Figure CN120198950A_ABST
Patent Text Reader

Abstract

The invention discloses a food sensory preference degree recognition method, device and equipment based on 3D facial micro-expressions and a medium, and relates to the field of food sensory evaluation. According to the method, training of the teacher model is firstly carried out, under the condition that the precision of the teacher model is guaranteed, the teacher model is migrated to the student model by adopting a knowledge distillation strategy, training of the student model with higher generalization ability is carried out, and the high-precision student model can be obtained through training under a limited training set; according to the method, the 3D facial micro-expressions are applied in the recognition process, compared with 2D facial micro-expressions, the 3D facial micro-expression data have the advantages of high illumination adaptability, high head posture adaptability, facial geometric correlation information, facial depth information and the like, and the recognition precision of the food sensory preference degree is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of food sensory evaluation, and particularly to a method, device, equipment and medium for identifying food sensory preference based on 3D facial micro-expressions. Background Art

[0002] Emotion recognition technology, especially multi-modal data based on facial expressions, speech, electroencephalogram and physiological signals, has broad application prospects in modern intelligent sensory analysis. By capturing consumers' emotional responses in real time, emotion recognition technology can effectively evaluate consumers' feelings towards food, make up for the subjectivity and individual differences of traditional questionnaires, and provide more accurate data support.

[0003] Traditional questionnaire-based emotion assessment relies on the subjective feedback of subjects, and has problems of individual differences and low accuracy. Machine bionic sensors such as electronic noses and electronic tongues cannot fully reflect the true feelings of human consumers due to the limitations of the types and quantities of sensors, resulting in incomplete and inaccurate analysis results.

[0004] 2D camera technology is widely used and has relatively low costs, making it suitable for most basic emotion recognition applications. Existing 2D micro-expression data-based recognition methods are continuously optimized with the help of deep learning models (such as Multilayer Perceptron (MLP) and Convolutional Neural Network (CNN)). 2D facial recognition systems show high real-time performance in emotion recognition, can quickly feedback consumers' facial emotion responses, have mature algorithms and numerous open-source data sets. However, 2D technology relies too much on sufficient lighting conditions, and the accuracy of emotion recognition in low-light or strong-light environments is greatly reduced. Moreover, due to prominent problems such as facial occlusion (such as hand occlusion, object occlusion, etc.), the performance of 2D technology in complex scenarios is greatly limited, unable to accurately capture all emotion changes, and the natural expressions (baseline emotions) of individual facial expressions may affect the accuracy of emotion recognition.

[0005] Existing publicly available data sets for training deep learning models are limited and cannot be applied to the training of high-precision models.

[0006] Based on the above analysis of the existing technologies, the existing methods for identifying food sensory preference have the following disadvantages.

[0007] 1. Limitations of traditional questionnaire-based sensory evaluation methods. High cost: The price of a single test is expensive, increasing the R & D and evaluation burden on enterprises; long feedback cycle: Data collection and analysis take a long time, unable to meet the efficiency requirements of the fast-moving consumer goods industry; strong subjectivity: Evaluation results rely too much on individual subjective awareness, resulting in poor data consistency and reliability; large individual differences: Due to the differences in sensory experiences of each consumer, it is difficult to unify and standardize the results.

[0008] 2. Deficiencies of objective emotion recognition devices. Expensive equipment: The price of emotion recognition devices on the current market is too high, hindering the popularization and application of small and medium-sized enterprises; low accuracy rate: The accuracy of recognition results is not high, affecting the application effect; small sample size: The scale and diversity of existing data sets are insufficient, making it difficult to support the wide-ranging industry needs. Professional dependence: The operation of the device and data analysis require highly professional personnel support, which is not conducive to rapid promotion.

[0009] 3. Defects of existing 2D facial emotion recognition technologies. Poor recognition of 2D facial virtual 3D models, closed-source professional software, no commercial domestic software alternatives, no possibility of commercial customization, high usage costs, slow update of customized models, not meeting the requirements of fast-moving consumer goods industries such as food for high efficiency, low cost, and dataization; difficult integration of 2D recognition data with technologies such as AI and digital twins, making it difficult to promote the development of sensory evaluation towards intelligence and standardization. Summary of the Invention

[0010] The purpose of this application is to provide a method, device, equipment, and medium for identifying food sensory preference based on 3D facial micro-expressions to improve the accuracy of food sensory preference recognition.

[0011] To achieve the above purpose, the following solutions are provided in this application.

[0012] In the first aspect, this application provides a method for identifying food sensory preference based on 3D facial micro-expressions, including: Training a teacher model using a first data set to obtain a trained teacher model; Using the trained teacher model and a second data set, adopting a knowledge distillation strategy to train a student model to obtain a trained student model; Converting real-time collected 3D facial micro-expression data into 2D facial micro-expression data; Inputting the 2D facial micro-expression data into the trained student model to obtain a food sensory preference recognition result.

[0013] In a second aspect, the present application provides a food sensory preference recognition device based on 3D facial micro-expressions. This fuzzy action recognition device applies the above-mentioned food sensory preference recognition method based on 3D facial micro-expressions. The food sensory preference recognition device based on 3D facial micro-expressions includes: A teacher model training module, configured to train a teacher model using a first data set to obtain a trained teacher model; A student model training module, configured to train a student model using the trained teacher model and a second data set by adopting a knowledge distillation strategy to obtain a trained student model; A data conversion module, configured to convert real-time collected 3D facial micro-expression data into 2D facial micro-expression data; A food sensory preference recognition module, configured to input the 2D facial micro-expression data into the trained student model to obtain a food sensory preference recognition result.

[0014] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. The processor executes the computer program to implement the above-mentioned food sensory preference recognition method based on 3D facial micro-expressions.

[0015] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the above-mentioned food sensory preference recognition method based on 3D facial micro-expressions.

[0016] According to the specific embodiments provided by the present application, the following technical effects are disclosed by the present application: The present application provides a food sensory preference recognition method, device, equipment, and medium based on 3D facial micro-expressions. The present application trains a teacher model using a first data set to obtain a trained teacher model; uses the trained teacher model and a second data set, and adopts a knowledge distillation strategy to train a student model to obtain a trained student model; converts real-time collected 3D facial micro-expression data into 2D facial micro-expression data; inputs the 2D facial micro-expression data into the trained student model to obtain a food sensory preference recognition result. The present application first trains the teacher model. While ensuring the accuracy of the teacher model, it adopts a knowledge distillation strategy to transfer the teacher model to the student model and train a student model with stronger generalization ability, so that a high-precision student model can be trained under a limited training set. In the recognition process of the present application, 3D facial micro-expressions are applied. Compared with 2D facial micro-expressions, the 3D facial micro-expression data has advantages such as strong light adaptability, strong head pose adaptability, facial geometric correlation information, and facial depth information, further improving the food sensory preference recognition accuracy. Brief Description of the Drawings

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0018] Figure 1 It is a schematic flowchart of a method for identifying food sensory preference based on 3D facial micro-expressions provided by an embodiment of the present application.

[0019] Figure 2 It is a flowchart of student model training provided by an embodiment of the present application.

[0020] Figure 3 It is a schematic diagram of the principle of data dimensionality reduction provided by an embodiment of the present application.

[0021] Figure 4 It is a flowchart of the verification process of the food sensory preference identification method provided by an embodiment of the present application.

[0022] Figure 5 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed Description of the Embodiments

[0023] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0024] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the drawings and specific embodiments.

[0025] The current emotion recognition technology mainly focuses on 2D algorithms or image / video analysis and has not been specifically optimized for the food industry. The Dutch Noldus system has occupied nearly 90% of the market share in university research and food R & D fields in the past five years, and its commercial buyout price is as high as 64,000 yuan. However, 2D face recognition technology still has many limitations. For example, the recognition effect of 2D virtual 3D models is poor, the software is closed-source and lacks commercial customization support. At the same time, problems such as high price, slow model update speed, inability to extract original emotion data for model learning, inapplicability to food sensory emotion recognition, and poor protection of facial biometric information significantly limit its wide promotion and optimization potential in actual food sensory detection applications.

[0026] From the development trend of the technical route, 2D face recognition technology is suitable for simple emotion recognition tasks, has a relatively mature technical foundation, and low cost. However, in complex environments, especially under light changes and face occlusion, the recognition accuracy of 2D technology is relatively low. In contrast, 3D structured light face recognition technology, with its stronger robustness, adaptability, and ability to capture subtle expressions, can maintain high-precision emotion recognition in environments with low light, complex backgrounds, and head pose changes, and is especially suitable for fields such as consumer evaluation and food sensory preference surveys. With the further development of deep learning models, especially their application in food flavor evaluation, 3D face recognition technology will become an important tool for intelligent sensory analysis, helping to achieve the digital transformation of the industry and accurate market decision-making.

[0027] Advantages and disadvantages of the existing self-developed 3D face micro-expression recognition technology.

[0028] The advantages of 3D face micro-expression recognition technology include the following aspects.

[0029] Strong model robustness and high adaptability: 3D face recognition technology can maintain high accuracy in complex scenarios, has strong light adaptability, and can provide stable emotion recognition results under low light or strong light conditions.

[0030] Capturing facial depth information: By capturing the depth information of the face, 3D technology can effectively eliminate errors caused by individual differences in facial expression performance and improve the accuracy of emotion recognition.

[0031] Strong head pose adaptability: 3D technology has strong adaptability to head rotation or tilt and can accurately capture facial emotion features even when the consumer's head moves.

[0032] Higher recognition accuracy: 3D face recognition technology effectively compensates for the data loss of 2D technology in the case of face occlusion (such as the mouth and nose being blocked when raising the hand to taste) through facial geometric correlation information, increasing the effective continuous emotion observation value by 30%.

[0033] Micro-expression millisecond-level numerical identification: Based on deep learning micro-expression recognition algorithms, 3D technology can identify more subtle emotional changes, and the real-time response time is shortened to less than 100 milliseconds, with an accuracy rate of over 90% in dynamic expression recognition tasks.

[0034] The disadvantages of 3D facial micro-expression recognition technology include the following aspects.

[0035] High technical complexity and cost: Compared with 2D technology, 3D recognition requires more complex hardware support, such as depth cameras and stronger chip computing capabilities. Therefore, the cost is higher, and the requirements for the hardware development environment are also higher. The existing development environment is only for IOS.

[0036] Low application popularity: Currently, the popularity of 3D facial recognition technology is relatively low, and more market promotion and technical verification are needed.

[0037] Few 3D expression public datasets: Currently, there are few training sets that directly correspond 3D feature points to emotional intensity. It is necessary to reduce the dimension of the 2D database through data migration for fitting training, and a real 3D expression dataset needs to be established through long-term accumulation of user usage data to further improve the identification accuracy.

[0038] In an exemplary embodiment, a method for identifying food sensory preference based on 3D facial micro-expressions is provided, including the following steps 101 - step 104.

[0039] Step 101, training a teacher model using a first dataset to obtain a trained teacher model.

[0040] Step 102, using the trained teacher model and a second dataset, adopting a knowledge distillation strategy to train a student model to obtain a trained student model.

[0041] Step 103, converting the real-time collected 3D facial micro-expression data into 2D facial micro-expression data.

[0042] Step 104, inputting the 2D facial micro-expression data into the trained student model to obtain a food sensory preference recognition result.

[0043] In the above steps 101 - 104, facial videos of sensory evaluation personnel during the processes of smelling before tasting the samples, the moment of tasting contact, and aftertaste after swallowing are collected. 52 facial expression motion feature factors in the above videos are extracted through the ARKit face - tracking facial network function; a facial emotion recognition MLP model (a three - layer fully - connected network) is constructed and trained using the Emotion - Domestic Asian facial emotion dataset to obtain a trained model; the real - time input expression feature factors are processed by linear regression for dimensionality reduction to obtain the values of the feature factors in the two - dimensional state, and the values of the two - dimensional feature factors are input into the trained model for expression classification recognition to obtain the recognition intensity values of seven facial expressions; based on the recognized facial expression intensity values, combined with a preset evaluation algorithm, the food sensory preference evaluation result is calculated. By using a 3D facial network to extract spatio - temporal dynamic features that can reflect the tasting process of the subjects from the video, the present invention better characterizes the instantaneous changes of facial expressions compared with 2D micro - expression recognition, realizes the objectification and intelligentization of food sensory evaluation, and improves the accuracy and reliability of sensory evaluation preference. At the same time, in order to achieve high - precision facial emotion recognition under limited high - quality expression samples, while taking into account the coverage of large - scale datasets and the generalization ability of the model in complex environments, the present application designs a "teacher model - student model" cascaded knowledge transfer mechanism in the training framework. First, the teacher model is trained using the first dataset to obtain a trained teacher model; then, using the trained teacher model and the second dataset, and adopting a knowledge distillation strategy, the student model is trained to obtain a trained student model, ensuring the accuracy and generalization of the trained student model.

[0044] In another exemplary embodiment, the first dataset in the above step 101 is a high - credibility dataset, which has the characteristics of a small sample size but high label accuracy. In the above step 101, first, a publicly available high - credibility expression dataset (exemplarily, the JAFFE dataset is used in the embodiments of the present application) is selected as the basic data source for teacher model training. This dataset contains 7 basic facial emotion categories, with clear labels consistent with the expression postures.

[0045] Using a three - layer MLP neural network structure (input layer - hidden layer - output layer), a teacher model is constructed, with the input being 2D Blendshapes expression features (i.e., 2D facial micro - expression data) and the output being the emotion probability distribution in the form of Softmax.

[0046] The training uses the Cross Entropy Loss function and the Adam optimizer, and introduces an early stopping strategy for the validation set to prevent overfitting. The accuracy of the model can reach 97% after training, providing highly reliable guidance for the subsequent migration stage.

[0047] In another exemplary embodiment, the second data set in step 102 above is a low label consistency or weakly labeled data set (exemplarily, the Emotion-Domestic extended data set is used in the embodiments of the present application). This low label consistency or weakly labeled data set has the characteristics of a large sample size but low label accuracy. In step 102 above, after the teacher model is trained, its output is used as a "knowledge source" to guide the training of the student model, and the Knowledge Distillation strategy is used to guide the student model to learn more stably and generalize.

[0048] In the embodiments of the present application, the Knowledge Distillation strategy is used to output the teacher model as a SoftLabel, and the output of the student model is , and its loss function is shown in the following formula.

[0049] .

[0050] Among them, is the loss function used in the process of training the student model, is a factor for adjusting the guiding degree of the teacher model, is the cross entropy loss function, is the KL divergence loss function, is the recognition result of the food sensory preference degree corresponding to the 2D facial micro-expression data sample in the second data set output by the student model, is the recognition result of the food sensory preference degree corresponding to the 2D facial micro-expression data sample in the second data set output by the teacher model, is the food sensory preference degree label of the 2D facial micro-expression data sample in the second data set; α ∈ [0, 1]: adjusts the teacher's guiding degree.

[0051] This training method not only enhances the robustness of the student model to noise labels but also reduces the dependence on true labels.

[0052] In another exemplary embodiment, as Figure 2 shown, the specific implementation steps of step 102 above are: Load the Google open-source Mediapipe Face Landmarker model, detect and extract Asian facial emotion data samples in the Emotion-Domestic extended dataset, and construct a 2D Blendshapes training dataset for student model training; construct an MLP model (including a three-layer fully connected network), import the 2D Blendshapes training dataset into the MLP model, adjust the model parameters during model training, test the accuracy of the model's calculated emotion probability values and the confidence of emotion classification, and finally obtain an optimized emotion recognition model with an accuracy of 97%.

[0053] In another exemplary embodiment, as Figure 3 shown, in step 103 above, a linear dimensionality reduction mapping method is introduced to convert the real-time collected 3D Blendshapes (1220-dimensional high-dimensional facial dynamic features) into a 2D Blendshapes (52-dimensional) feature space compatible with the teacher model: Construct a linear regression mapping matrix that satisfies: .

[0054] Among them, is the 2D facial micro-expression data, is the 3D facial micro-expression data, is the linear regression mapping matrix.

[0055] Use the least squares method or LASSO (The Least Absolute Shrinkage and Selection Operator) regression for parameter estimation, and screen and retain dimensions through residual analysis.

[0056] This process ensures the consistency between the 3D feature space and the input domain of the teacher model, and supports cross-modal consistency learning.

[0057] The trained student model is cross-validated on the extended dataset and the performance of the independent test set is evaluated.

[0058] This application sets an adaptive loss weighting coefficient for samples with different noise levels (Label Noise) to reduce the adverse impact of uncertain labels on model training. At the same time, record the robustness indicators of the model for non-standard expression (head tilt, uneven lighting) scenarios to ensure generalization ability, and has the following advantages.

[0059] Soft labels improve sample effectiveness: The probability distribution output by the teacher model can be regarded as a "continuous label", which has stronger error correction ability when the label is biased.

[0060] Strong multi-source data fusion ability: It is compatible with high-quality static images and large-scale low-quality video frame data at the same time, enhancing the multi-modal expressiveness of the model.

[0061] Cross-modal feature compatibility design: Through the 3D→2D mapping mechanism, it is compatible with real facial capture data and 2D image datasets, improving the model's implementation ability.

[0062] Weak supervision and robust training mechanism: It realizes soft supervision training for some unlabeled samples, laying a foundation for subsequent large-scale industrial deployment.

[0063] In another exemplary embodiment, to verify the accuracy of the recognition method of the present application, as Figure 4 shown, the following verification process is provided.

[0064] Use the Live Link Face app on the iOS platform mobile device to collect 3D facial videos of the evaluator during the processes of smelling before tasting the sample, the moment of tasting contact, and aftertaste after swallowing, as well as information files including video sequences, frame capture time coding, etc. Detect and extract the 3D Blendshapes (facial expression motion feature factors) of the above three different evaluation stage intervals from the extracted videos and coding information files through the Face landmarker function suite in the ARKit open source program. Then, perform linear dimensionality reduction processing on the 3D Blendshapes through a linear regression model to obtain 2D Blendshapes with an error probability less than (1×10 -3 %).

[0065] Import the 2D Blendshapes of the evaluator in the three different evaluation stage intervals for facial expression classification recognition, and obtain the recognition intensity values of seven facial expressions (sadness, fear, disgust, anger, neutral, surprise, and happiness). The emotional intensity values are between 0 (none) and 1 (fully present).

[0066] Input the emotional intensity values into the linear regression PLSR (Partial Least Squares Regression) model of sample - emotional intensity for fitting, and calculate the evaluation result of food sensory preference.

[0067] Compare the food sensory preference evaluation result with the real feelings of the evaluator. The experimental results show that the recognition method of the present application can ensure high accuracy.

[0068] The recognition method of the present application can be applied as follows.

[0069] 1. Construction of a high-quality 3D facial emotion dataset.

[0070] Based on an exclusive dataset of Asian facial expressions and a 3D expression training set covering different regions of China, the platform has specifically optimized the regional adaptability of the model. By screening data related to Asian faces in large datasets such as BUPT and FER, and introducing the JAFFE high-precision expression dataset, combining the DeepFace classifier and DiFace technology, the data resolution and detail expressiveness have been improved, and a high-quality and highly adaptable 2D expression dataset has been constructed.

[0071] 2. High-precision emotion recognition model performance.

[0072] The platform uses the Mediapipe Face Mesh model to extract up to 478 feature points (2D blendshapes), which has a higher resolution ability than traditional models such as Dlib and Face++. At the same time, based on the most advanced JAFFE teacher model, through knowledge distillation technology, the high-precision knowledge of the teacher model is transferred to a relatively large-scale rough dataset (ordinary 2D data training set), enabling the 2D-trained MLP emotion recognition model to have both high precision and broad generalization ability, further improving the accuracy of emotion recognition. Through dimensionality reduction processing, 3D micro-expression feature vectors are extracted, and seven basic emotions (anger, disgust, fear, happiness, neutral, sadness, surprise) are classified into specific vector labels, and trained and cross-validated using the MLP deep learning model, effectively avoiding overfitting and achieving high-precision emotion recognition, with an emotion accuracy rate of 97%.

[0073] 3. Application of the OSC cloud platform and sensor digital twin scenarios in sensory preferences.

[0074] Through the OSC real-time data communication protocol, the platform can establish a low-latency data transmission channel between physical devices and digital twin virtual models, achieving high-efficiency, compatibility, and real-time data updates of real-time data. It supports the rapid transmission and integration of multi-type sensor data, providing support for the modeling of complex digital twin scenarios. In the intelligent sensory system, the OSC server can efficiently transmit 3D facial expression recognition, subjective questionnaire information, and other multi-modal physiological signal data, providing a real-time and comprehensive emotion decision data stream for food R & D. Combining sensors, real-time data streams, and artificial intelligence algorithms, this technology will be able to achieve comprehensive monitoring of the facial physical recognition sensor system in industrial-level production monitoring, sales scenarios, etc. and intelligent data optimization in the future.

[0075] Based on the same inventive concept, an embodiment of the present application further provides a device for recognizing food sensory preference based on 3D facial micro-expressions for implementing the method for recognizing food sensory preference based on 3D facial micro-expressions involved above. The solution provided by this device for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the device for recognizing food sensory preference based on 3D facial micro-expressions provided below can refer to the limitations on the method for recognizing food sensory preference based on 3D facial micro-expressions in the above text, and will not be elaborated here.

[0076] In an exemplary embodiment, a device for recognizing food sensory preference based on 3D facial micro-expressions is provided, including: A teacher model training module, configured to train a teacher model using a first data set to obtain a trained teacher model; A student model training module, configured to train a student model using the trained teacher model and a second data set by adopting a knowledge distillation strategy to obtain a trained student model; A data conversion module, configured to convert the 3D facial micro-expression data collected in real time into 2D facial micro-expression data; A food sensory preference recognition module, configured to input the 2D facial micro-expression data into the trained student model to obtain a food sensory preference recognition result.

[0077] In an exemplary embodiment, a computer device is provided. This computer device can be a server or a terminal, and its internal structure diagram can be as Figure 5 shown. This computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of this computer device is used to provide computing and control capabilities. The memory of this computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of this computer device is used for exchanging information between the processor and external devices. The communication interface of this computer device is used for communicating with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for recognizing food sensory preference based on 3D facial micro-expressions.

[0078] Those skilled in the art can understand, Figure 5The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0079] In an exemplary embodiment, a computer device is further provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above-mentioned embodiment of the method for identifying food sensory preference based on 3D facial micro-expressions are implemented.

[0080] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned embodiment of the method for identifying food sensory preference based on 3D facial micro-expressions are implemented.

[0081] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can include the processes of the above-mentioned method embodiments. Among them, any reference to a memory, database, or other medium used in the various embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0082] In each of the embodiments provided in this application, the database involved may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on a blockchain, etc., and is not limited thereto. In each of the embodiments provided in this application, the processor involved may be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, etc., and is not limited thereto.

[0083] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0084] Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.

Claims

1. A method for identifying food sensory preference based on 3D facial micro-expressions, characterized in that, It includes the following steps: Train a teacher model using a first data set to obtain a trained teacher model; Train a student model using the trained teacher model and a second data set by adopting a knowledge distillation strategy to obtain a trained student model; Convert the 3D facial micro-expression data collected in real time into 2D facial micro-expression data; Input the 2D facial micro-expression data into the trained student model to obtain a food sensory preference recognition result.

2. The method for identifying food sensory preference based on 3D facial micro-expressions according to claim 1, wherein The first data set includes: N 2D facial micro-expression data samples with food sensory preference labels, where N is less than a first preset threshold, and the accuracy of the food sensory preference labels of the N 2D facial micro-expression data samples in the first data set is greater than a first accuracy threshold; The second data set includes M 2D facial micro-expression data samples with food sensory preference labels, where M is greater than a second preset threshold, and the accuracy of the food sensory preference labels of the M 2D facial micro-expression data samples in the second data set is less than a second accuracy threshold. The second preset threshold is greater than the first preset threshold, and the first accuracy threshold is greater than the second accuracy threshold.

3. The method for identifying food sensory preference based on 3D facial micro-expressions according to claim 1 or 2, characterized in that, The first data set is the JAFFE data set, and the second data set is the Emotion-Domestic extended data set.

4. The method for identifying food sensory preference based on 3D facial micro-expressions according to claim 1, wherein The loss function adopted during the training of the student model is: ; Among them, is the loss function adopted during the training of the student model, is the factor for adjusting the guiding degree of the teacher model, is the cross-entropy loss function, is the KL divergence loss function, is the recognition result of the food sensory preference degree corresponding to the 2D facial micro-expression data sample in the second dataset output by the student model, is the recognition result of the food sensory preference degree corresponding to the 2D facial micro-expression data sample in the second dataset output by the teacher model, is the food sensory preference degree label of the 2D facial micro-expression data sample in the second dataset.

5. The method for identifying food sensory preference based on 3D facial micro-expressions according to claim 1, characterized in that Converting the 3D facial micro-expression data collected in real time into 2D facial micro-expression data specifically includes: Converting the 3D facial micro-expression data collected in real time into 2D facial micro-expression data by using a linear regression mapping matrix.

6. The method for identifying food sensory preference based on 3D facial micro-expressions according to claim 5, wherein The linear regression mapping matrix is obtained by solving a linear regression model by adopting the least squares method or the LASSO regression method; The linear regression model is as follows: ; Among them, is 2D facial micro-expression data, is 3D facial micro-expression data, is a linear regression mapping matrix.

7. The method for identifying food sensory preference based on 3D facial micro-expressions according to claim 1, characterized in that, Inputting the 2D facial micro-expression data into the trained student model to obtain a food sensory preference recognition result specifically includes: Input the 2D facial micro-expression data into the trained student model to obtain the recognition intensity value of each facial expression output by the trained student model; Input the recognition intensity value of each facial expression into a preference fitting model to obtain a food sensory preference recognition result output by the preference fitting model. The preference fitting model represents the relationship between the recognition intensity values of different facial expressions and food sensory preferences.

8. An apparatus for identifying the sensory preference of food based on 3D facial micro - expressions, characterized in that, The food sensory preference recognition device based on 3D facial micro-expressions is applied to the food sensory preference recognition method based on 3D facial micro-expressions according to any one of claims 1-7. The food sensory preference recognition device based on 3D facial micro-expressions includes: A teacher model training module for training a teacher model using a first data set to obtain a trained teacher model; A student model training module for training a student model using the trained teacher model and a second data set by adopting a knowledge distillation strategy to obtain a trained student model; A data conversion module for converting the 3D facial micro-expression data collected in real time into 2D facial micro-expression data; A food sensory preference recognition module for inputting the 2D facial micro-expression data into the trained student model to obtain a food sensory preference recognition result.

9. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the method for identifying food sensory preference based on 3D facial micro-expressions according to any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for identifying food sensory preference based on 3D facial micro-expressions according to any one of claims 1-7.

Citation Information

Patent Citations

  • Method and system for evaluating numb degree and favorite degree of zanthoxylum armatum oil based on facial expression recognition

    CN117765588A

  • Micro-expression recognition method, device and equipment based on artificial intelligence and storage medium

    CN119206813A

  • Fragrance preference degree detection system and method based on facial micro-expression

    CN119296153A

  • Model training method and communication apparatus

    WO2024187413A1

  • Facial expression recognition method and system based on multi-cue associative learning

    WO2024192844A1

Cited By

  • Method for preparing yellow rice wine based on ultrasonic aging-multistage membrane filtration-low-temperature sedimentation coupling and yellow rice wine prepared by method

    CN121022538A