Milk preference degree prediction method and device based on micro-expression recognition and medium

Through the multi-scale optical flow method and cross-modal timing module combined with micro-expression features, the problem of difficult to capture consumer micro-expression in the prior art is solved, and high-precision and automated milk preference prediction are achieved.

CN120340096AActive Publication Date: 2025-07-18BEIJING FORESTRY UNIVERSITY
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510501052.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18
Estimated Expiration
2045-04-21

AI Technical Summary

Technical Problem

Existing consumer preference testing methods rely on explicit measurements, are susceptible to subjective biases, and are difficult to accurately capture the subtle expression changes of consumers when tasting similar foods.

Method used

The multi-scale optical flow method is used to extract the spatial and temporal features of the face, and a cross-modal timing module is built using the Vision Mamba encoder. Combining the micro-expression features and optical flow features, the automatic prediction of milk preferences is achieved through training models.

Benefits of technology

It improves the accuracy and automation of milk preference prediction, and can effectively capture consumers' unconscious emotional reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340096A_ABST
    Figure CN120340096A_ABST
Patent Text Reader

Abstract

The invention discloses a milk preference degree prediction method and device based on micro-expression recognition and a medium, and relates to the technical field of image processing and behavior analysis. The method comprises the following steps: decomposing face video data when a target tastes milk into an image sequence and preprocessing the image sequence; capturing facial micro-expression changes by adopting a multi-scale optical flow method, and extracting facial spatial-temporal features; constructing a cross-modal time sequence module, respectively processing the horizontal and vertical optical flow features by using an encoder, and fusing the processed horizontal and vertical optical flow features to obtain a fused feature; constructing and training a micro-expression recognition model; based on a cross-modal time sequence module, constructing a milk preference degree prediction model, and performing prediction in combination with the micro-expression features and the fusion features; training the milk preference degree prediction model; and preprocessing a video to be identified, inputting the preprocessed video into the trained milk preference degree prediction model, and finally outputting a preference degree score of a target on milk. According to the invention, high-automation and high-precision milk preference degree prediction can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of image processing and behavior analysis, and more particularly, to a method, device and medium for predicting milk preference based on micro-expression recognition. Background Art

[0002] Consumer preference testing plays a crucial role in product development and marketing strategies. Traditional consumer preference testing usually relies on explicit measurement methods such as questionnaires and rating scales. To ensure the reliability of consumer testing, these tests require a large amount of time to find consumers to participate. In addition, these methods are vulnerable to consumers' subjective biases and social expectations, resulting in the measurement results not accurately reflecting consumers' true preferences and limiting the testing accuracy of preferences.

[0003] In recent years, with the development of artificial intelligence technology, implicit measurement methods based on facial expression analysis have gradually become a research hotspot. Facial expression analysis can capture consumers' unconscious emotional reactions when tasting food, thus more accurately predicting their preferences. However, existing facial expression analysis technologies mainly target macro expressions (MaEs) and are difficult to capture the subtle expression changes of consumers when tasting similar foods. Micro-expressions (MEs) are facial expressions with extremely short durations (0.065 - 0.5 seconds) and are uncontrollable, which can more accurately reflect consumers' true emotions.

[0004] Therefore, it is of great practical significance to develop a highly automated and accurate method for automatically predicting milk preference based on micro-expression recognition. Summary of the Invention

[0005] To solve the above technical problems, the present invention provides a method, device and medium for predicting milk preference based on micro-expression recognition to achieve highly automated and accurate prediction of milk preference.

[0006] In a first aspect, the present invention provides a method for predicting milk preference based on micro-expression recognition, the method comprising:

[0007] Decompose the micro-expression video into a sequence of micro-expression images;

[0008] Decompose the facial video data when the target tastes milk into an image sequence;

[0009] Preprocess the image sequence to obtain a preprocessed image sequence; wherein, the preprocessing includes face detection, alignment and cropping;

[0010] Based on the preprocessed image sequence, the multi-scale optical flow method is used to capture facial micro-expression changes and extract facial spatio-temporal features, where the facial spatio-temporal features include horizontal optical flow features and vertical optical flow features between each frame and the starting frame;

[0011] Construct a cross-modal temporal module, use the encoder to process the horizontal and vertical optical flow features respectively, and fuse the processed horizontal and vertical optical flow features to obtain a fused feature, which can characterize the global temporal dependence relationship;

[0012] Construct and train a micro-expression recognition model, map the basic emotion categories to three categories: negative, positive, and surprised, and use the trained micro-expression model to extract micro-expression features from the preprocessed image sequence;

[0013] Based on the cross-modal temporal module, construct a milk preference prediction model, and combine micro-expression features and fused features for prediction;

[0014] Train the milk preference prediction model;

[0015] Preprocess the video to be recognized, input it into the trained milk preference prediction model, and finally output the target's preference score for milk.

[0016] Furthermore, preprocessing the image sequence to obtain a preprocessed image sequence includes:

[0017] Use the Dlib library to detect the facial features and multiple facial key points of each frame in the image sequence;

[0018] Use affine transformation to align multiple facial key points;

[0019] Taking the inner corner of the left eye of the facial feature in each frame image as the base point, crop out a valid facial area with the same size as the preprocessed image, and combine multiple preprocessed images in the original order to obtain a preprocessed image sequence.

[0020] Furthermore, based on the preprocessed image sequence, using the multi-scale optical flow method to capture facial micro-expression changes and extract facial spatio-temporal features, including:

[0021] Taking the first frame of the preprocessed image sequence as the starting frame;

[0022] Use the TV-L1 optical flow estimation algorithm to calculate the horizontal optical flow features and vertical optical flow features between each frame and the starting frame;

[0023] Taking the horizontal optical flow sequence and vertical optical flow sequence formed by the horizontal optical flow features and vertical optical flow features between each frame and the starting frame as the facial spatio-temporal features.

[0024] Further, the encoder is used to process the horizontal and vertical optical flow features respectively, and the processed horizontal and vertical optical flow features are fused to obtain fused features, including:

[0025] Use the Vision Mamba encoder to perform bidirectional temporal modeling on the horizontal optical flow feature and the vertical optical flow feature respectively;

[0026] Enhance the correlation between the horizontal optical flow feature and the vertical optical flow feature by sharing the state transition matrix;

[0027] Fuse the processed horizontal optical flow and vertical optical flow features to obtain fused features, thereby establishing the temporal relationship between local and global features.

[0028] Further, the steps for the milk preference prediction model to realize milk preference prediction include:

[0029] Adopt a cross-modal temporal module to process the optical flow features;

[0030] Fuse the micro-expression features with the processed optical flow features to obtain the first fused feature, and use the Vision Mamba encoder to perform temporal modeling on the first fused feature to obtain temporal features;

[0031] Input the temporal features into the fully connected layer, and output the consumer's preference score for milk through the fully connected layer.

[0032] Further, training the milk preference prediction model includes:

[0033] Use the video data corresponding to one of the targets as the test set, and the video data of the remaining targets as the training set;

[0034] Determine the best parameter combination of the model through cross-validation, and the best parameter combination includes multiple groups of best parameters;

[0035] Configure the milk preference prediction model with different best parameters to obtain multiple prediction models, and use all the video data to train the multiple prediction models, and save the prediction model with the best performance as the trained milk preference prediction model.

[0036] In a second aspect, the present invention provides a milk preference prediction device based on micro-expression recognition, and the device includes:

[0037] A video data decomposition module, configured to decompose the facial video data when the target tastes milk into an image sequence;

[0038] An image preprocessing module, configured to preprocess the image sequence to obtain a preprocessed image sequence; wherein, the preprocessing includes face detection, alignment and cropping;

[0039] A facial spatio-temporal feature extraction module, configured to capture facial micro-expression changes and extract facial spatio-temporal features based on the preprocessed image sequence by using a multi-scale optical flow method, where the facial spatio-temporal features include horizontal optical flow features and vertical optical flow features between each frame and the starting frame;

[0040] A feature fusion module, configured to construct a cross-modal temporal module, process the horizontal and vertical optical flow features respectively by using an encoder, and fuse the processed horizontal and vertical optical flow features to obtain a fusion feature, where the fusion feature can characterize the global temporal dependence relationship;

[0041] A micro-expression feature extraction module, configured to construct and train a micro-expression recognition model, map basic emotion categories to three categories: negative, positive, and surprised, and extract micro-expression features from the preprocessed image sequence by using the trained micro-expression model;

[0042] A prediction model construction module, configured to construct a milk preference prediction model based on the cross-modal temporal module, and make a prediction by combining the micro-expression features and the fusion features;

[0043] A prediction model training module, configured to train the milk preference prediction model;

[0044] A preference prediction module, configured to preprocess the video to be recognized, input it into the trained milk preference prediction model, and finally output the target's preference score for milk.

[0045] Further, the image preprocessing module is further configured to:

[0046] Use the Dlib library to detect the facial features and multiple facial key points of each frame of the image sequence;

[0047] Use an affine transformation to align multiple facial key points;

[0048] Taking the inner corner of the left eye of the facial feature in each frame of the image as the base point, crop out a valid facial area with the same size as the preprocessing image, and combine multiple preprocessing images in the original order to obtain a preprocessed image sequence.

[0049] Further, the facial spatio-temporal feature extraction module is further configured to:

[0050] Taking the first frame of the preprocessed image sequence as the starting frame;

[0051] Use the TV-L1 optical flow estimation algorithm to calculate the horizontal optical flow features and vertical optical flow features between each frame and the starting frame;

[0052] The horizontal optical flow sequence and vertical optical flow sequence formed by the horizontal optical flow features and vertical optical flow features between each frame and the starting frame are used as facial spatio-temporal features.

[0053] In a third aspect, the present invention provides a readable storage medium storing one or more programs, which can be executed by one or more processors to implement the method as described above.

[0054] The present invention has at least the following beneficial effects:

[0055] In view of the problem in the prior art that it is difficult to capture the subtle expression changes of the target, the present invention proposes an automatic prediction framework based on micro-expression recognition. First, facial spatio-temporal features are extracted by the multi-scale optical flow method to capture the micro-expression changes of the target when tasting milk; secondly, a cross-modal temporal module is constructed by using the Vision Mamba encoder to establish global temporal dependencies; finally, by fusing micro-expression features and optical flow features, automatic prediction of the target's milk preference degree is realized. This method can effectively capture the unconscious emotional reactions of consumers and significantly improve the accuracy and automation of prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 Shows a flowchart of a method for predicting milk preference degree based on micro-expression recognition according to an embodiment of the present invention;

[0057] Figure 2 Shows a flowchart of preprocessing an image sequence according to an embodiment of the present invention;

[0058] Figure 3 Shows a flowchart of facial spatio-temporal feature extraction according to an embodiment of the present invention;

[0059] Figure 4 Shows a flowchart of fusing the processed horizontal and vertical optical flow features according to an embodiment of the present invention;

[0060] Figure 5 Shows a structural diagram of a cross-modal temporal module according to an embodiment of the present invention;

[0061] Figure 6 Shows a flowchart of a milk preference degree prediction model for predicting milk preference degree according to an embodiment of the present invention;

[0062] Figure 7 Shows a structural diagram of a milk preference degree prediction framework based on micro-expression recognition according to an embodiment of the present invention;

[0063] Figure 8 Shows a structural diagram of a device for predicting milk preference degree based on micro-expression recognition according to an embodiment of the present invention. Specific Embodiments

[0064] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be described in detail below in conjunction with the accompanying drawings and specific embodiments. The embodiments of the present invention will be further described in detail below in conjunction with the accompanying drawings and specific examples, but shall not be construed as a limitation to the present invention. For the various steps described herein, if there is no necessity for a sequential relationship between them, the order in which they are described as examples herein shall not be regarded as a limitation. Those skilled in the art should know that they can be adjusted in order as long as the logic between them is not destroyed and the entire process cannot be realized.

[0065] An embodiment of the present invention provides a method for predicting milk preference based on micro-expression recognition, as Figure 1 shown. The method for predicting milk preference based on micro-expression recognition includes the following steps S100 - S700.

[0066] S100. Decompose the facial video data of the target when tasting milk into an image sequence.

[0067] It should be noted that the "target" described herein refers to a target that can taste milk, including but not limited to humans and other animals. For example, the method proposed in the present invention can determine the type of the target according to the object the milk is targeted at, such as consumers (humans) or animals, and can be used to evaluate the preference of consumers for milk or the preference of some animals for milk.

[0068] In this embodiment, the way to decompose the facial video data into an image sequence is frame-by-frame extraction, and the extraction method can use existing software, including but not limited to Adobe Premiere.

[0069] S200. Preprocess the image sequence to obtain a preprocessed image sequence; wherein, the preprocessing includes face detection, alignment, and cropping.

[0070] In some embodiments, as Figure 2 shown, preprocessing the image sequence to obtain a preprocessed image sequence includes the following steps:

[0071] S201. Use the Dlib library to detect the face features and multiple facial key points of each frame of the image sequence;

[0072] S202. Align multiple facial key points using affine transformation;

[0073] S203. Taking the inner canthus of the left eye of the face feature in each frame of the image as the base point, crop out an effective facial area with the same size as the preprocessed image. Combine multiple preprocessed images in the original order to obtain a preprocessed image sequence.

[0074] After being processed through the above steps, each of the preprocessed images included in the obtained preprocessed image sequence contains a valid facial region. For example, when the target is a consumer, the valid facial region is the valid face region.

[0075] S300. Based on the preprocessed image sequence, use the multi-scale optical flow method to capture facial micro-expression changes and extract facial spatio-temporal features, where the facial spatio-temporal features include the horizontal optical flow features and vertical optical flow features between each frame and the starting frame.

[0076] In some embodiments, as Figure 3 shown, based on the preprocessed image sequence, using the multi-scale optical flow method to capture facial micro-expression changes and extract facial spatio-temporal features includes the following steps:

[0077] S301. Use the first frame of the preprocessed image sequence as the starting frame;

[0078] S302. Use the TV-L1 optical flow estimation algorithm to calculate the horizontal optical flow features and vertical optical flow features between each frame and the starting frame;

[0079] S303. Use the horizontal optical flow sequence and vertical optical flow sequence formed by the horizontal optical flow features and vertical optical flow features between each frame and the starting frame as the facial spatio-temporal features.

[0080] S400. Construct a cross-modal temporal module, use the encoder to process the horizontal and vertical optical flow features respectively, and fuse the processed horizontal and vertical optical flow features to obtain a fused feature, where the fused feature can represent the global temporal dependence relationship.

[0081] In some embodiments, as Figure 4 shown, using the encoder to process the horizontal and vertical optical flow features respectively and fuse the processed horizontal and vertical optical flow features to obtain a fused feature includes the following steps:

[0082] S401. Use the Vision Mamba encoder to perform bidirectional temporal modeling on the horizontal optical flow features and vertical optical flow features respectively;

[0083] S402. Enhance the correlation between the horizontal optical flow features and vertical optical flow features through a shared state transition matrix;

[0084] S403. Fuse the processed horizontal optical flow and vertical optical flow features to obtain a fused feature, thereby establishing the temporal relationship between local and global features.

[0085] Exemplarily, as Figure 5As shown, the cross-modal temporal module designed in the present invention consists of two bidirectional Vision Mamba (VIM) encoders, which are respectively used to process horizontal optical flow and vertical optical flow features, and enhance the correlation between horizontal optical flow and vertical optical flow features by sharing a state transition matrix. By integrating a bidirectional temporal state model (SSM) to extend Mamba, allowing modeling of forward and backward temporal features, this module can capture the temporal dependencies before and after facial expressions. Specifically, the horizontal optical flow feature at time t The bidirectional SSM representation of can be expressed as follows:

[0086]

[0087] where respectively represent the forward and backward states of the horizontal optical flow feature at time t. represent the forward and backward outputs at time t. A, B, C are the state transition matrices of the forward SSM, A b , B b , C b are the state transition matrices of the backward SSM. The horizontal optical flow feature shares the state transition matrix A with the vertical optical flow feature, A b .

[0088] S500. Construct and train a micro-expression recognition model, map the basic emotion categories to three categories: negative, positive, and surprised, and use the trained micro-expression model to extract micro-expression features from the preprocessed image sequence;

[0089] In some embodiments, as Figure 6 shown, the steps for the milk preference prediction model to achieve milk preference prediction include:

[0090] S501. Process the optical flow features using a cross-modal temporal module;

[0091] S502. Fuse the micro-expression features with the processed optical flow features to obtain a first fused feature, and use a Vision Mamba encoder to perform temporal modeling on the first fused feature to obtain temporal features;

[0092] S503. Input the temporal features into a fully connected layer, and output the consumer's milk preference score through the fully connected layer.

[0093] S600. Based on the cross-modal temporal module, construct a milk preference prediction model, and perform prediction by combining micro-expression features and fused features.

[0094] S700. Train the milk preference prediction model.

[0095] Exemplarily, when training the milk preference prediction model, leave-one-consumer cross-validation is used for model training to determine the optimal parameters. All video data is used for training, and the weight model with the best performance is saved. The specific steps are as follows:

[0096] S701: Use the data of each consumer as the test set, and the data of the remaining consumers as the training set;

[0097] S702: Determine the optimal parameters of the model through cross-validation;

[0098] S703: Use all video data for training and save the weight model with the best performance.

[0099] S800: Preprocess the video to be recognized, input it into the trained milk preference prediction model, and finally output the target's preference score for milk.

[0100] In step S800, the method of preprocessing the video to be recognized can be implemented through steps S100 and S200, which will not be elaborated here.

[0101] The milk preference prediction method based on micro-expression recognition designed by the present invention can be implemented through the milk preference prediction framework based on micro-expression recognition. Its structure is as Figure 7 shown. This framework consists of a micro-expression recognition model and a milk preference prediction model. Both models are constructed based on the cross-modal temporal module. A bidirectional Vision Mamba encoder is used to fuse the micro-expression emotional features and the facial optical flow features of milk consumers to enhance the prediction accuracy of consumers' milk preferences.

[0102] The embodiment of the present invention also provides a milk preference prediction device based on micro-expression recognition, as Figure 8 shown. This device includes:

[0103] A video data decomposition module 801, configured to decompose the facial video data of the target when tasting milk into an image sequence;

[0104] An image preprocessing module 802, configured to preprocess the image sequence to obtain a preprocessed image sequence; wherein, the preprocessing includes face detection, alignment, and cropping;

[0105] A facial spatio-temporal feature extraction module 803, configured to capture facial micro-expression changes based on the preprocessed image sequence and extract facial spatio-temporal features. The facial spatio-temporal features include the horizontal optical flow features and vertical optical flow features between each frame and the starting frame;

[0106] A feature fusion module 804, configured to construct a cross-modal temporal module, process horizontal and vertical optical flow features respectively using an encoder, and fuse the processed horizontal and vertical optical flow features to obtain a fused feature, where the fused feature can characterize global temporal dependencies;

[0107] A micro-expression feature extraction module 805, configured to construct and train a micro-expression recognition model, map basic emotion categories into three categories: negative, positive, and surprised, and extract micro-expression features from the preprocessed image sequence using the trained micro-expression model;

[0108] A prediction model construction module 806, configured to construct a milk preference prediction model based on the cross-modal temporal module, and make a prediction by combining micro-expression features and fused features;

[0109] A prediction model training module 807, configured to train the milk preference prediction model;

[0110] A preference prediction module 808, configured to preprocess the video to be recognized, input it into the trained milk preference prediction model, and finally output the target's preference score for milk.

[0111] In some embodiments, the image preprocessing module is further configured to:

[0112] Use the Dlib library to detect the facial features and multiple facial key points of each frame of the image sequence;

[0113] Use affine transformation to align multiple facial key points;

[0114] Taking the inner corner of the left eye of the facial feature in each frame of the image as the base point, crop out a uniformly sized effective facial area as the preprocessed image, and combine multiple preprocessed images in the original order to obtain a preprocessed image sequence.

[0115] In some embodiments, the facial spatio-temporal feature extraction module is further configured to:

[0116] Taking the first frame of the preprocessed image sequence as the starting frame;

[0117] Use the TV-L1 optical flow estimation algorithm to calculate the horizontal and vertical optical flow features between each frame and the starting frame;

[0118] Taking the horizontal optical flow sequence and the vertical optical flow sequence formed by the horizontal and vertical optical flow features between each frame and the starting frame as the facial spatio-temporal features.

[0119] It should be noted that the structures of the various milk preference prediction devices based on micro-expression recognition described in this embodiment belong to the same technical concept as the previously described milk preference prediction method based on micro-expression recognition, and achieve the same beneficial effects through the same principle, which will not be elaborated here.

[0120] An embodiment of the present invention also provides a readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the method described in any of the above embodiments.

[0121] The above description is intended to be illustrative and not restrictive. For example, the above examples (or one or more of their solutions) can be used in combination with each other. For example, those of ordinary skill in the art can use other embodiments when reading the above description. Additionally, in the above specific embodiments, various features can be grouped together to simplify the present invention. This should not be construed as an intention that the features of an unclaimed invention are necessary for any claim. On the contrary, the subject matter of the present invention may be less than all the features of a particular embodiment of the invention. Thus, the following claims are incorporated herein as examples or embodiments into the specific embodiments, where each claim independently serves as a separate embodiment, and considering these embodiments, they can be combined with each other in various combinations or permutations. The scope of the present invention should be determined with reference to the appended claims and the full scope of the equivalent forms empowered by these claims.

Claims

1. A method for predicting milk preference based on micro-expression recognition, characterized in that, The method includes: Decompose the facial video data when the target tastes milk into an image sequence; Preprocess the image sequence to obtain a preprocessed image sequence; wherein, the preprocessing includes face detection, alignment, and cropping; Based on the preprocessed image sequence, adopt the multi-scale optical flow method to capture facial micro-expression changes and extract facial spatio-temporal features, where the facial spatio-temporal features include the horizontal optical flow features and vertical optical flow features between each frame and the starting frame; Construct a cross-modal temporal module, use the encoder to process the horizontal and vertical optical flow features respectively, and fuse the processed horizontal and vertical optical flow features to obtain a fused feature, which can represent the global temporal dependence relationship; Construct and train a micro-expression recognition model, map the basic emotion categories to three categories: negative, positive, and surprised, and use the trained micro-expression model to extract micro-expression features from the preprocessed image sequence; Based on the cross-modal temporal module, construct a milk preference prediction model, and combine the micro-expression features and the fused features for prediction; Train the milk preference prediction model; Preprocess the video to be recognized, input it into the trained milk preference prediction model, and finally output the target's milk preference score.

2. The method for predicting milk preference based on micro-expression recognition according to claim 1, wherein, Preprocessing the image sequence to obtain a preprocessed image sequence includes: Use the Dlib library to detect the face features and multiple facial key points in each frame of the image sequence; Align multiple facial key points using affine transformation; Taking the inner corner of the left eye of the face feature in each frame image as the base point, crop out an effective facial area with the same size as the preprocessed image, and combine multiple preprocessed images in the original order to obtain a preprocessed image sequence.

3. The method for predicting milk preference based on micro-expression recognition according to claim 1, characterized in that, Based on the preprocessed image sequence, adopt the multi-scale optical flow method to capture facial micro-expression changes and extract facial spatio-temporal features, including: Taking the first frame of the preprocessed image sequence as the starting frame; Use the TV-L1 optical flow estimation algorithm to calculate the horizontal optical flow features and vertical optical flow features between each frame and the starting frame; Taking the horizontal optical flow sequence and vertical optical flow sequence formed by the horizontal optical flow features and vertical optical flow features between each frame and the starting frame as the facial spatio-temporal features.

4. The method for predicting milk preference based on micro-expression recognition according to claim 3, characterized in that Using the encoder to process the horizontal and vertical optical flow features respectively, and fusing the processed horizontal and vertical optical flow features to obtain a fused feature, including: Use the Vision Mamba encoder to perform bidirectional temporal modeling on the horizontal optical flow features and vertical optical flow features respectively; Enhance the correlation between the horizontal optical flow features and vertical optical flow features through a shared state transition matrix; Fuse the processed horizontal optical flow and vertical optical flow features to obtain a fused feature, thereby establishing the temporal relationship between local and global features.

5. The method for predicting milk preference based on micro-expression recognition according to claim 1, wherein The steps for the milk preference prediction model to achieve milk preference prediction include: Use the cross-modal temporal module to process the optical flow features; Fuse the micro-expression features with the processed optical flow features to obtain a first fused feature, and use the Vision Mamba encoder to perform temporal modeling on the first fused feature to obtain temporal features; Input the temporal features into the fully connected layer, and output the consumer's milk preference score through the fully connected layer.

6. The method for predicting milk preference based on micro-expression recognition according to claim 1, wherein Training the milk preference prediction model includes: Use the video data corresponding to one of the targets as the test set, and the video data of the remaining targets as the training set; Determine the best parameter combination of the model through cross-validation, and the best parameter combination includes multiple groups of best parameters; Configure the milk preference prediction model with different best parameters to obtain multiple prediction models, train the multiple prediction models using all the video data, and save the prediction model with the best performance as the trained milk preference prediction model.

7. A milk preference prediction device based on micro-expression recognition, characterized in that, The device includes: A video data decomposition module configured to decompose the facial video data when the target tastes milk into an image sequence; An image preprocessing module configured to preprocess the image sequence to obtain a preprocessed image sequence; wherein, the preprocessing includes face detection, alignment, and cropping; A facial spatio-temporal feature extraction module configured to capture facial micro-expression changes using a multi-scale optical flow method based on the preprocessed image sequence, and extract facial spatio-temporal features, where the facial spatio-temporal features include horizontal optical flow features and vertical optical flow features between each frame and the starting frame; A feature fusion module configured to construct a cross-modal temporal module, process the horizontal and vertical optical flow features respectively using an encoder, and fuse the processed horizontal and vertical optical flow features to obtain a fused feature, where the fused feature can represent global temporal dependency; A micro-expression feature extraction module configured to construct and train a micro-expression recognition model, map basic emotion categories to three categories: negative, positive, and surprised, and extract micro-expression features from the preprocessed image sequence using the trained micro-expression model; A prediction model construction module configured to construct a milk preference prediction model based on the cross-modal temporal module, and make predictions by combining micro-expression features and fused features; A prediction model training module configured to train the milk preference prediction model; A preference prediction module configured to preprocess the video to be recognized, input it into the trained milk preference prediction model, and finally output the preference score of the target for milk.

8. The milk preference prediction device based on micro-expression recognition according to claim 7, wherein, The image preprocessing module is further configured to: Use the Dlib library to detect the face features and multiple facial key points of each frame image in the image sequence; Align multiple facial key points using an affine transformation; Using the inner corner of the left eye of the face feature in each frame image as a base point, crop out an effective facial area with the same size as the preprocessed image, and combine multiple preprocessed images in the original order to obtain a preprocessed image sequence.

9. The milk preference prediction device based on micro-expression recognition according to claim 7, characterized in that, The facial spatio-temporal feature extraction module is further configured to: Use the first frame of the preprocessed image sequence as the starting frame; Use the TV-L1 optical flow estimation algorithm to calculate the horizontal optical flow features and vertical optical flow features between each frame and the starting frame; Use the horizontal optical flow sequence and vertical optical flow sequence formed by the horizontal optical flow features and vertical optical flow features between each frame and the starting frame as the facial spatio-temporal features.

10. A non-transitory computer-readable storage medium storing instructions, which when executed by a processor, execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Micro-expression recognition method based on multi-feature fusion and double-flow network

    CN115359534A

  • Micro-expression recognition method based on Transform motion feature fusion

    CN118366202A

  • Micro-expression recognition pre-training method based on space-time double-flow mask reconstruction

    CN118644882A

  • Micro-expression recognition method based on convolutional neural network and optical flow features

    CN118675209A

  • Three-dimensional motion capture and intelligent analysis system and method based on monocular camera

    CN119169701A