A method, device, equipment and medium for identifying food sensory preference based on 3D facial micro-expressions

Through 3D facial micro-expression technology combined with knowledge distillation strategies, the high cost of traditional methods in food sensory preference recognition and insufficient accuracy of 2D facial recognition technology are solved, and high-precision and low-cost food sensory evaluation are achieved.

CN120198950BActive Publication Date: 2025-08-12INNOVATION CENTER OF YANGTZE RIVER DELTA ZHEJIANG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510668600.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-12
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The existing food sensory preference recognition methods have problems such as high cost, low efficiency, strong subjectivity, expensive and low accuracy of objective emotion recognition equipment, and insufficient accuracy of 2D facial emotion recognition technology in complex environments.

Method used

Using 3D facial micro-expression technology combined with knowledge distillation strategy, through teacher model training and student model migration, 3D facial micro-expression data is converted into 2D facial micro-expression data to recognize food sensory preferences, improving the accuracy and adaptability of recognition.

Benefits of technology

It realizes high-precision food sensory preference recognition in complex environments, improves the accuracy and adaptability of recognition, reduces equipment costs, and is suitable for the intelligent and standardized development of food sensory evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198950B_ABST
    Figure CN120198950B_ABST
Patent Text Reader

Abstract

This application discloses a method, device, equipment, and medium for identifying food sensory preferences based on 3D facial micro-expressions, relating to the field of food sensory evaluation. This application first trains a teacher model. While ensuring the accuracy of the teacher model, a knowledge distillation strategy is used to transfer the teacher model to a student model, training a student model with stronger generalization capabilities. This allows a high-precision student model to be trained using a limited training set. This application applies 3D facial micro-expressions during the recognition process. Compared to 2D facial micro-expressions, this 3D facial micro-expression data has advantages such as stronger adaptability to lighting and head posture, as well as facial geometry association information and facial depth information, further improving the accuracy of food sensory preference recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of food sensory evaluation, and in particular to a method, device, equipment and medium for identifying food sensory preference based on 3D facial micro-expressions. Background Art

[0002] Emotion recognition technology, particularly based on multimodal data from facial expressions, speech, brain waves, and physiological signals, has broad application prospects in modern intelligent sensory analysis. By capturing consumers' emotional responses in real time, emotion recognition technology can effectively assess their perceptions of food, addressing the subjectivity and individual differences inherent in traditional questionnaires and providing more accurate data support.

[0003] Traditional questionnaire-based emotion assessments rely on subjective responses from participants, which can lead to individual differences and low accuracy. Limited by the number and variety of sensors, robotic bionic sensors like electronic noses and tongues cannot fully reflect the true feelings of human consumers, resulting in incomplete and inaccurate analysis results.

[0004] 2D camera technology is widely used and relatively low-cost, making it suitable for most basic emotion recognition applications. Existing recognition methods based on 2D micro-expression data are continuously optimized with the help of deep learning models (such as multilayer perceptrons (MLPs) and convolutional neural networks (CNNs)). 2D facial recognition systems demonstrate high real-time performance in emotion recognition, enabling rapid feedback on consumers' facial emotional responses. These systems utilize mature algorithms and numerous open-source datasets. However, 2D technology relies heavily on sufficient lighting conditions, significantly reducing its accuracy in low-light or bright-light environments. Furthermore, due to the significant problem of facial occlusion (such as hands or objects), 2D technology is significantly limited in complex scenes, unable to accurately capture all emotional variations. Differences in the natural expression of individual facial expressions (baseline emotions) can also affect the accuracy of emotion recognition.

[0005] Existing public datasets for training deep learning models are limited and cannot be used for training high-precision models.

[0006] Based on the above analysis of the prior art, the existing food sensory preference identification methods have the following shortcomings.

[0007] 1. Limitations of traditional questionnaire-based sensory evaluation methods. High cost: A single test is expensive, increasing the R&D and evaluation burden on companies. Long feedback cycles: Data collection and analysis take a long time, failing to meet the efficiency demands of the fast-moving consumer goods industry. High subjectivity: Evaluation results are overly dependent on individual subjective perceptions, resulting in poor data consistency and reliability. Large individual differences: Due to the differences in sensory experiences among consumers, the results are difficult to standardize.

[0008] 2. Objective emotion recognition equipment is limited. Expensive equipment: Current emotion recognition equipment on the market is too expensive, hindering its widespread adoption and application by small and medium-sized enterprises. Low accuracy: Recognition results are not precise, impacting application effectiveness. Small sample sizes: Existing datasets are insufficiently large and diverse to support a wide range of industry needs. Dependence on specialized expertise: Equipment operation and data analysis require highly specialized personnel, hindering rapid adoption.

[0009] 3. Existing 2D facial emotion recognition technology has shortcomings. 2D facial virtual 3D models have poor recognition capabilities, professional software is closed-source, there are no commercial domestic alternatives, and commercial customization is impossible. This leads to high costs, slow updates of customized models, and a failure to meet the high-efficiency, low-cost, and data-driven needs of fast-moving consumer goods (FMCG) industries like food. Furthermore, 2D recognition data is difficult to integrate with technologies like AI and digital twins, hindering the development of intelligent and standardized sensory evaluation. Summary of the Invention

[0010] The purpose of this application is to provide a method, device, equipment and medium for identifying food sensory preference based on 3D facial micro-expressions to improve the accuracy of food sensory preference identification.

[0011] To achieve the above objectives, this application provides the following solutions.

[0012] In a first aspect, the present application provides a method for identifying food sensory preferences based on 3D facial micro-expressions, comprising:

[0013] Using the first data set to train the teacher model to obtain a trained teacher model;

[0014] Using the trained teacher model and the second dataset, the student model is trained using the knowledge distillation strategy to obtain a trained student model.

[0015] Convert the real-time collected 3D facial micro-expression data into 2D facial micro-expression data;

[0016] The 2D facial micro-expression data is input into the trained student model to obtain the food sensory preference recognition result.

[0017] In a second aspect, the present application provides a device for identifying food sensory preferences based on 3D facial micro-expressions. The fuzzy motion recognition device applies the above-mentioned method for identifying food sensory preferences based on 3D facial micro-expressions. The device for identifying food sensory preferences based on 3D facial micro-expressions comprises:

[0018] A teacher model training module is used to train the teacher model using the first data set to obtain a trained teacher model;

[0019] The student model training module is used to train the student model using the trained teacher model and the second dataset, adopting the knowledge distillation strategy, and obtain a trained student model;

[0020] A data conversion module, used to convert the real-time collected 3D facial micro-expression data into 2D facial micro-expression data;

[0021] The food sensory preference recognition module is used to input the 2D facial micro-expression data into the trained student model to obtain the food sensory preference recognition result.

[0022] In a third aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned method for identifying food sensory preferences based on 3D facial micro-expressions.

[0023] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned method for identifying food sensory preference based on 3D facial micro-expressions.

[0024] According to the specific embodiments provided in this application, this application discloses the following technical effects:

[0025] The present application provides a method, device, equipment and medium for identifying food sensory preference based on 3D facial micro-expressions. The present application uses a first data set to train a teacher model to obtain a trained teacher model; uses the trained teacher model and a second data set to adopt a knowledge distillation strategy to train a student model to obtain a trained student model; converts the 3D facial micro-expression data collected in real time into 2D facial micro-expression data; and inputs the 2D facial micro-expression data into the trained student model to obtain food sensory preference recognition results. The present application first trains a teacher model, and while ensuring the accuracy of the teacher model, adopts a knowledge distillation strategy to transfer the teacher model to the student model, and trains a student model with stronger generalization ability. A high-precision student model can be obtained by training with a limited training set. The present application applies 3D facial micro-expressions in the recognition process. Compared with 2D facial micro-expressions, the 3D facial micro-expression data has the advantages of strong illumination adaptability, strong head posture adaptability, facial geometry correlation information and facial depth information, etc., which further improves the accuracy of food sensory preference recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0027] Figure 1 A flowchart of a method for identifying food sensory preference based on 3D facial micro-expressions provided in one embodiment of the present application.

[0028] Figure 2 A flowchart of student model training provided in one embodiment of the present application.

[0029] Figure 3 A schematic diagram of the data dimensionality reduction principle provided in one embodiment of the present application.

[0030] Figure 4 This is a flowchart of the verification process of the food sensory preference identification method provided in one embodiment of the present application.

[0031] Figure 5 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0032] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0033] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0034] Current emotion recognition technology primarily relies on 2D algorithms or image and video analysis and has yet to be specifically optimized for the food industry. The Dutch Noldus system has captured nearly 90% of the market share in university research and food R&D over the past five years, with a commercial buyout price of up to 64,000 yuan. However, 2D facial recognition technology still has many limitations. For example, the recognition effect of 2D virtual 3D models is poor, the software is closed-source and lacks commercial customization support. Furthermore, it is expensive, has slow model updates, cannot extract raw emotion data for model learning, is unsuitable for food sensory emotion recognition, and has poor privacy protection for facial biometric information. These issues significantly limit its potential for widespread promotion and optimization in actual food sensory testing applications.

[0035] From the perspective of technology development trends, 2D facial recognition technology is suitable for simple emotion recognition tasks and has a relatively mature technical foundation and low cost. However, in complex environments, especially when lighting changes and facial occlusion occurs, the recognition accuracy of 2D technology is low. In contrast, 3D structured light facial recognition technology, with its greater robustness, adaptability, and ability to capture subtle expressions, can maintain high-precision emotion recognition in low-light, complex backgrounds, and environments with changing head postures. It is particularly suitable for areas such as consumer evaluation and food sensory preference surveys. With the further development of deep learning models, especially their application in food flavor evaluation, 3D facial recognition technology will become an important tool for intelligent sensory analysis, helping the industry's digital transformation and the realization of accurate market decisions.

[0036] The advantages and disadvantages of existing self-developed 3D facial micro-expression recognition technology.

[0037] The advantages of 3D facial micro-expression recognition technology include the following aspects.

[0038] The model is robust and adaptable: 3D facial recognition technology can maintain high accuracy in complex scenarios, has strong lighting adaptability, and can provide stable emotion recognition results in low light or strong light conditions.

[0039] Facial depth information capture: By capturing facial depth information, 3D technology can effectively eliminate errors caused by differences in individual facial expressions and improve the accuracy of emotion recognition.

[0040] Strong adaptability to head posture: 3D technology is highly adaptable to head rotation or tilt, and can accurately capture facial emotional characteristics when the consumer moves their head.

[0041] Higher recognition accuracy: 3D facial recognition technology uses facial geometry to effectively compensate for data loss in 2D technology when the face is obscured (such as when the mouth and nose are blocked when raising the hand to taste), thereby increasing the effective continuous emotion observation value by 30%.

[0042] Millisecond-level numerical identification of micro-expressions: Based on the deep learning micro-expression recognition algorithm, 3D technology can identify more subtle emotional changes, and the real-time response time is shortened to less than 100 milliseconds, achieving an accuracy rate of over 90% in dynamic expression recognition tasks.

[0043] The disadvantages of 3D facial micro-expression recognition technology include the following aspects.

[0044] High technical complexity and cost: Compared with 2D technology, 3D recognition requires more complex hardware support, such as depth cameras and stronger chip computing power, so the cost is higher and the hardware development environment is more demanding. The existing development environment is only IOS.

[0045] Low application popularity: Currently, the popularity of 3D facial recognition technology is relatively low, and more market promotion and technical verification are needed.

[0046] There are few public datasets of 3D expressions: Currently, there are few training sets that directly correspond 3D feature points to emotion intensity. It is necessary to reduce the dimension of the 2D database through data migration for fitting training. It takes a long period of user usage data accumulation to establish a real 3D expression dataset and improve the identification accuracy in one step.

[0047] In an exemplary embodiment, a method for identifying food sensory preference based on 3D facial micro-expressions is provided, comprising the following steps 101 to 104 .

[0048] Step 101: Use the first data set to train the teacher model to obtain a trained teacher model.

[0049] Step 102: Use the trained teacher model and the second data set to train the student model using a knowledge distillation strategy to obtain a trained student model.

[0050] Step 103: Convert the 3D facial micro-expression data collected in real time into 2D facial micro-expression data.

[0051] Step 104 : Input the 2D facial micro-expression data into the trained student model to obtain a food sensory preference recognition result.

[0052] In steps 101-104, facial videos of sensory evaluators are collected during their pre-tasting, tasting, and post-swallowing savoring process. ARKit face-tracking facial network functionality is used to extract facial motion feature factors from 52 faces in the videos. A multi-layered neural network (MLP) model for facial emotion recognition (MLP) is constructed and trained using the Emotion-Domestic Asian facial emotion dataset to obtain a trained model. The real-time input facial feature factors are subjected to dimensionality reduction using linear regression to obtain two-dimensional feature factor values. These two-dimensional feature factor values are then input into the trained model for expression classification and recognition, resulting in recognition strength values for seven facial expressions. Based on the recognized facial expression strength values and a pre-set evaluation algorithm, the sensory preference evaluation results of the food are calculated. This method utilizes a 3D facial network to extract and analyze spatiotemporal dynamic features from the video that reflect the tasting process of the subject. Compared to 2D micro-expression recognition, this method better characterizes instantaneous changes in facial expression, achieving objectivity and intelligentization in food sensory evaluation and improving the accuracy and reliability of sensory preference evaluation. At the same time, in order to achieve high-precision facial emotion recognition under limited high-quality expression samples, while taking into account the coverage of large-scale data sets and the generalization ability of the model in complex environments, this application designs a "teacher model-student model" cascade knowledge transfer mechanism in the training framework. First, the teacher model is trained using the first data set to obtain a trained teacher model; then, the student model is trained using the trained teacher model and the second data set, adopting the knowledge distillation strategy to obtain a trained student model, thereby ensuring the accuracy and generalization of the trained student model.

[0053] In another exemplary embodiment, the first dataset in step 101 is a high-confidence dataset, characterized by a small sample size and high label accuracy. In step 101, a publicly available high-confidence facial expression dataset (exemplarily, the JAFFE dataset is used in this embodiment) is selected as the basic data source for training the teacher model. This dataset contains seven basic facial emotion categories with clear labels consistent with the facial expressions.

[0054] A three-layer MLP neural network structure (input layer–hidden layer–output layer) is used to construct a teacher model. The input is 2D Blendshapes expression features (i.e., 2D facial micro-expression data), and the output is the emotion probability distribution in the form of Softmax.

[0055] Training used the cross-entropy loss function and the Adam optimizer, and introduced an early stopping strategy on the validation set to prevent overfitting. The model achieved an accuracy of 97% after training, providing highly reliable guidance for subsequent migration phases.

[0056] In another exemplary embodiment, the second dataset in step 102 is a low-label-consistency or weakly labeled dataset (exemplarily, the present embodiment uses the Emotion-Domestic extended dataset). This low-label-consistency or weakly labeled dataset has a large sample size but low label accuracy. In step 102, after the teacher model is trained, its output serves as a "knowledge source" to guide the student model training. A knowledge distillation strategy is employed to guide the student model's learning towards more stable generalization.

[0057] The embodiment of this application uses the knowledge distillation strategy to convert the output of the teacher model As soft label (SoftLabel), the student model output is , and its loss function is shown as follows.

[0058] .

[0059] in, is the loss function used in the process of training the student model, is a factor that adjusts the degree of guidance of the teacher model. is the cross entropy loss function, is the KL divergence loss function, The food sensory preference recognition results corresponding to the 2D facial micro-expression data samples in the second dataset output by the student model are: The food sensory preference recognition results corresponding to the 2D facial micro-expression data samples in the second dataset output by the teacher model, is the food sensory preference label of the 2D facial micro-expression data sample in the second dataset; α∈ [0, 1]: adjusts the degree of teacher guidance.

[0060] This training method not only enhances the robustness of the student model to noisy labels, but also reduces its dependence on real labels.

[0061] In another exemplary embodiment, Figure 2 As shown, the specific implementation steps of the above step 102 are:

[0062] The Mediapipe Face Landmarker Google open-source model was loaded to detect and extract facial emotion data samples of Asians from the Emotion-Domestic extended dataset. A 2D Blendshapes training dataset was constructed for student model training. An MLP model (containing a three-layer fully connected network) was built and the 2D Blendshapes training dataset was imported into the MLP model. Model training and parameter adjustment were performed, and the accuracy of the model's calculated emotion probability values and the confidence level of emotion classification were tested. Ultimately, an optimized emotion recognition model with an accuracy of 97% was obtained.

[0063] In another exemplary embodiment, Figure 3 As shown, in the above step 103, a linear dimensionality reduction mapping method is introduced to convert the real-time collected 3D Blendshapes (1220-dimensional high-dimensional facial dynamic features) into a 2DBlendshapes (52-dimensional) feature space compatible with the teacher model:

[0064] Construct a linear regression mapping matrix that satisfies: .

[0065] in, 2D facial micro-expression data, 3D facial micro-expression data, is the linear regression mapping matrix.

[0066] Parameter estimation was performed using the least squares method or LASSO (The Least Absolute Shrinkage and Selection Operator) regression, and the retained dimensions were screened through residual analysis.

[0067] This process ensures the consistency of the 3D feature space with the teacher model input domain, supporting cross-modal consistency learning.

[0068] The trained student model is cross-validated on the extended dataset and the performance of the independent test set is evaluated.

[0069] This application sets adaptive loss weighting coefficients for samples with different noise levels (label noise) to reduce the adverse effects of uncertain labels on model training. It also records the model's robustness to non-standard expressions (head tilt, uneven lighting) to ensure generalization, providing the following advantages.

[0070] Soft labels improve sample effectiveness: The probability distribution output by the teacher model can be regarded as a "continuous label", which has stronger error correction capabilities when there are deviations in the labels.

[0071] Strong multi-source data fusion capability: compatible with both high-quality static images and large-scale low-quality video frame data, improving the multimodal expressiveness of the model.

[0072] Cross-modal feature compatibility design: Through the 3D→2D mapping mechanism, it is compatible with real facial capture data and 2D image datasets, improving the model's implementation capabilities.

[0073] Weakly supervised and robust training mechanism: enables soft-supervised training of some unlabeled samples, laying the foundation for subsequent large-scale industrial deployment.

[0074] In another exemplary embodiment, in order to verify the accuracy of the recognition method of the present application, as shown in FIG. Figure 4 As shown, the following verification process is provided.

[0075] The Live Link Face app on iOS devices was used to collect 3D facial videos of the evaluators during the pre-tasting, contact, and post-tasting phases of the sample. These videos and encoded information files, including video sequences and frame time codes, were then used with the Face Landmarker feature in the ARKit open-source software to detect and extract 3D blendshapes (expression and motion features of the face) for the three different evaluation phases. The 3D blendshapes were then linearly reduced using a linear regression model to obtain a probability of error less than (1×10) from the 3D data. -3 %)2D Blendshapes.

[0076] The 2D Blendshapes of the evaluators at three different evaluation stages were imported for expression classification and recognition, and the recognition intensity values of seven facial expressions (sadness, fear, disgust, anger, neutrality, surprise and happiness) were obtained. The emotion intensity values were between 0 (absent) and 1 (completely present).

[0077] The emotion intensity value was input into the sample-emotion intensity linear regression PLSR (Partial Least Squares Regression) model to fit the food sensory preference evaluation results.

[0078] The sensory preference evaluation results of the food were compared with the evaluators' actual feelings. The experimental results showed that the recognition method of the present application can ensure high accuracy.

[0079] The identification method of the present application can be applied as follows.

[0080] 1. Construction of high-quality 3D facial emotion dataset.

[0081] The platform optimizes the model's regional adaptability based on a proprietary dataset of Asian facial expressions and a 3D expression training set covering different regions of China. By screening Asian-related data from large datasets such as BUPT and FER, and introducing the JAFFE high-precision expression dataset, combined with the DeepFace classifier and DiFace technology, the platform improves data resolution and detail expression, creating a high-quality, highly adaptable 2D expression dataset.

[0082] 2. High-precision emotion recognition model performance.

[0083] The platform utilizes the Mediapipe Face Mesh model to extract up to 478 feature points (2D blendshapes), achieving higher resolution than traditional models such as Dlib and Face++. Furthermore, based on the cutting-edge JAFFE teacher model, knowledge distillation techniques are used to transfer the teacher model's high-precision knowledge to a larger, coarse dataset (a standard 2D data training set). This enables the 2D-trained MLP emotion recognition model to achieve both high precision and broad generalization, further improving emotion recognition accuracy. Dimensionality reduction is used to extract 3D micro-expression feature vectors, and the seven basic emotions (anger, disgust, fear, happiness, neutrality, sadness, and surprise) are classified into specific vector labels. Using an MLP deep learning model for training and cross-validation, this effectively avoids overfitting and achieves highly accurate emotion recognition, with an accuracy rate of 97%.

[0084] 3. OSC cloud platform and sensor digital twin scenarios in sensory preference applications.

[0085] Through the OSC real-time data communication protocol, the platform can establish a low-latency data transmission channel between physical devices and digital twin virtual models, achieving real-time data efficiency, compatibility, and real-time data updates. It supports the rapid transmission and integration of multi-type sensor data, providing support for complex scene modeling of digital twins. In the intelligent sensory system, the OSC server can efficiently transmit 3D facial expression recognition, subjective questionnaire information and other multimodal physiological signal data, providing real-time and comprehensive emotional decision-making data streams for food research and development. Combining sensors, real-time data streams and artificial intelligence algorithms, this technology will be able to achieve comprehensive monitoring of facial physical recognition sensor systems and intelligent data optimization in industrial-grade production monitoring, sales scenarios, etc. in the future.

[0086] Based on the same inventive concept, embodiments of the present application also provide a device for identifying sensory preferences of foods based on 3D facial micro-expressions, for implementing the aforementioned method for identifying sensory preferences of foods based on 3D facial micro-expressions. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the device for identifying sensory preferences of foods based on 3D facial micro-expressions provided below can be found in the aforementioned limitations of the method for identifying sensory preferences of foods based on 3D facial micro-expressions, and will not be further elaborated here.

[0087] In an exemplary embodiment, a device for identifying food sensory preferences based on 3D facial micro-expressions is provided, comprising:

[0088] A teacher model training module is used to train the teacher model using the first data set to obtain a trained teacher model;

[0089] The student model training module is used to train the student model using the trained teacher model and the second dataset, adopting the knowledge distillation strategy, and obtain a trained student model;

[0090] A data conversion module, used to convert the real-time collected 3D facial micro-expression data into 2D facial micro-expression data;

[0091] The food sensory preference recognition module is used to input the 2D facial micro-expression data into the trained student model to obtain the food sensory preference recognition result.

[0092] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected via a system bus, and the communication interface is connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for identifying food sensory preferences based on 3D facial micro-expressions is implemented.

[0093] Those skilled in the art will understand that Figure 5The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0094] In an exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps in the above-mentioned embodiment of the method for identifying food sensory preference based on 3D facial micro-expressions are implemented.

[0095] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, which, when executed by a processor, implements the steps in the above-mentioned embodiment of the method for identifying food sensory preference based on 3D facial micro-expressions.

[0096] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0097] The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, and the like.

[0098] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0099] This document uses specific examples to illustrate the principles and implementation methods of this application. The description of the above examples is only intended to help understand the method and core concept of this application. At the same time, for those skilled in the art, based on the concept of this application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A method for identifying food sensory preferences based on 3D facial micro-expressions, characterized in that: The steps include: Using the first data set to train the teacher model to obtain a trained teacher model; Using the trained teacher model and the second data set, a knowledge distillation strategy is adopted to train the student model to obtain a trained student model; the first data set includes: N 2D facial micro-expression data samples with food sensory preference labels, N is less than a first preset threshold, and the accuracy of the food sensory preference labels of the N 2D facial micro-expression data samples in the first data set is greater than the first accuracy threshold; the second data set includes M 2D facial micro-expression data samples with food sensory preference labels, M is greater than a second preset threshold, the accuracy of the food sensory preference labels of the M 2D facial micro-expression data samples in the second data set is less than the second accuracy threshold, the second preset threshold is greater than the first preset threshold, and the first accuracy threshold is greater than the second accuracy threshold; Convert the real-time collected 3D facial micro-expression data into 2D facial micro-expression data; Inputting the 2D facial micro-expression data into the trained student model to obtain a recognition strength value of each facial expression output by the trained student model; The recognition strength value of each facial expression is input into the preference fitting model to obtain the food sensory preference recognition result output by the preference fitting model. The preference fitting model characterizes the relationship between the recognition strength values of different facial expressions and the food sensory preference.

2. The method for identifying food sensory preferences based on 3D facial micro-expressions according to claim 1, characterized in that: The first dataset is the JAFFE dataset, and the second dataset is the Emotion-Domestic extended dataset.

3. The method for identifying food sensory preferences based on 3D facial micro-expressions according to claim 1, characterized in that: The loss function used in the training of the student model is: ; in, is the loss function used in the process of training the student model, is a factor that adjusts the degree of guidance of the teacher model. is the cross entropy loss function, is the KL divergence loss function, The food sensory preference recognition results corresponding to the 2D facial micro-expression data samples in the second dataset output by the student model are: The food sensory preference recognition results corresponding to the 2D facial micro-expression data samples in the second dataset output by the teacher model, is the food sensory preference label of the 2D facial micro-expression data sample in the second dataset.

4. The method for identifying food sensory preferences based on 3D facial micro-expressions according to claim 1, wherein: Converting real-time collected 3D facial micro-expression data into 2D facial micro-expression data, specifically including: The linear regression mapping matrix is used to convert the real-time collected 3D facial micro-expression data into 2D facial micro-expression data.

5. The method for identifying food sensory preferences based on 3D facial micro-expressions according to claim 4, characterized in that: The linear regression mapping matrix is obtained by solving the linear regression model using the least squares method or the LASSO regression method; The linear regression model is: ; in, 2D facial micro-expression data, 3D facial micro-expression data, is the linear regression mapping matrix.

6. A food sensory preference recognition device based on 3D facial micro-expressions, characterized in that: The device for identifying food sensory preferences based on 3D facial micro-expressions is applied to the method for identifying food sensory preferences based on 3D facial micro-expressions according to any one of claims 1 to 5, and the device for identifying food sensory preferences based on 3D facial micro-expressions comprises: A teacher model training module is used to train the teacher model using the first data set to obtain a trained teacher model; The student model training module is used to train the student model using the trained teacher model and the second dataset, adopting the knowledge distillation strategy, and obtain a trained student model; A data conversion module, used to convert the real-time collected 3D facial micro-expression data into 2D facial micro-expression data; The food sensory preference recognition module is used to input the 2D facial micro-expression data into the trained student model to obtain the food sensory preference recognition result.

7. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for identifying food sensory preference based on 3D facial micro-expressions according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for identifying food sensory preference based on 3D facial micro-expressions according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Micro-expression recognition method, device and equipment based on artificial intelligence and storage medium

    CN119206813A

  • Fragrance preference degree detection system and method based on facial micro-expression

    CN119296153A