Content recommendation method based on multi-modal emotion recognition, medium and electronic device

CN121567930APending Publication Date: 2026-02-24QINGDAO HAIER TECH +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511492429.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

[0004]本申请提供一种基于多模态情绪识别的内容推荐方法、介质及电子装置,用以解决用户通过手动操作或语音指令来调整电视节目内容的问题

Benefits of technology

[0068]This application provides a content recommendation method, medium, and electronic device based on multimodal emotion recognition. Upon responding to an emotion recommendation activation command, the method acquires multimodal data of the target user, including at least two of the following: facial image, facial video, and audio information. Based on this, it determines the user's first emotion category and its corresponding emotion confidence level. When the emotion confidence level is greater than or equal to a preset confidence level, it determines the recommended content corresponding to the first emotion category. This method, by combining multimodal data with emotion confidence level judgment, makes content recommendations more aligned with the user's current emotional state, improving the accuracy and applicability of recommendations, thereby solving the problem of users manually adjusting television program content through voice commands.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121567930A_ABST
    Figure CN121567930A_ABST
Patent Text Reader

Abstract

The invention discloses a content recommendation method based on multi-modal emotion recognition, a medium and an electronic device, and relates to the technical field of smart home / smart home, and the content recommendation method based on multi-modal emotion recognition comprises the steps: obtaining multi-modal data of a target user in response to an emotion recommendation starting instruction, the multi-modal data comprises at least two items of a face image, a face video and audio information; according to the multi-modal data, determining a first emotion category of the target user and an emotion confidence coefficient corresponding to the first emotion category; and when it is determined that the emotion confidence is greater than or equal to the preset confidence, determining recommended content corresponding to the first emotion category. According to the method, intelligent recommendation is performed by considering the emotional state of the user, so that the problem that the user needs to adjust the television program content through manual operation or voice instructions is effectively solved, and the watching experience of the user is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of smart home / intelligent home technology, and more specifically, to a content recommendation method, medium, and electronic device based on multimodal emotion recognition. Background Technology

[0002] With the fast pace of modern life, leisure time has become increasingly precious. Many users choose to relax and enjoy entertainment by watching television programs, making television programs one of the most important choices for daily leisure and entertainment.

[0003] Currently, when watching TV programs, users still need to manually operate or use voice commands to select the content they are interested in. This method cannot recommend suitable content based on the user's actual needs, which affects the user's viewing experience. Summary of the Invention

[0004] This application provides a content recommendation method, medium, and electronic device based on multimodal emotion recognition to solve the problem of users adjusting television program content through manual operation or voice commands.

[0005] Firstly, this application provides a content recommendation method based on multimodal emotion recognition, including:

[0006] In response to an instruction to enable emotion-based recommendations, multimodal data of the target user is acquired, wherein the multimodal data includes at least two of the following: facial images, facial videos, and audio information.

[0007] Based on the multimodal data, determine the first emotion category of the target user and the emotion confidence level corresponding to the first emotion category;

[0008] If the confidence level of the emotion is determined to be greater than or equal to a preset confidence level, recommended content corresponding to the first emotion category is determined.

[0009] Optionally, determining the first emotion category of the target user and the emotion confidence level corresponding to the first emotion category based on the multimodal data includes:

[0010] Based on the facial images and / or the facial videos, determine the facial expression confidence level and the second emotion category;

[0011] Based on the audio information, determine the confidence level of the first speech tone and the third emotion category;

[0012] If the second emotion category and the third emotion category are determined to be the same, then the first emotion category is determined to be the second emotion category, or the first emotion category is determined to be the third emotion category.

[0013] The confidence score of the facial expression is obtained by multiplying it by a first coefficient and then adding it to the confidence score of the first voice tone multiplied by a second coefficient.

[0014] Optionally, determining the facial expression confidence level and the second emotion category based on the facial image and / or the facial video includes:

[0015] The facial image and / or the facial video are input into the face detection model to obtain the facial image output by the face detection model;

[0016] The face image is preprocessed to obtain a processed face image;

[0017] Based on the processed facial image, the confidence level of the facial expression and the second emotion category are determined.

[0018] Optionally, determining the facial expression confidence level and the second emotion category based on the processed facial image includes:

[0019] The processed face image is input into the expression classification model to obtain multiple fourth emotion categories output by the expression classification model, as well as the probability value corresponding to each emotion category;

[0020] Determine the maximum probability value from the probability values ​​corresponding to each of the aforementioned emotion categories;

[0021] The maximum probability value is determined as the facial expression confidence level, and the fourth emotion category corresponding to the maximum probability value is determined as the second emotion category.

[0022] Optionally, determining the first speech tone confidence level and the third emotion category based on the audio information includes:

[0023] Determine the signal-to-noise ratio of the audio information;

[0024] Determine the second speech tone confidence level and the fifth emotion category of the audio information;

[0025] If the signal-to-noise ratio is determined to be greater than the preset signal-to-noise ratio, the second speech tone confidence is determined as the first speech tone confidence, and the fifth emotion category is determined as the third emotion category.

[0026] Optionally, determining the second speech tone confidence level and the fifth emotion category of the audio information includes:

[0027] The audio information is preprocessed using voice activation detection technology to obtain a voice signal;

[0028] The speech signal is subjected to feature processing to obtain the speech features corresponding to the speech signal;

[0029] The speech features are input into the speech recognition model to obtain the second speech tone confidence score and the fifth emotion category.

[0030] Optionally, the method further includes:

[0031] If the confidence level of the emotion is determined to be less than the preset confidence level, a prompt message is generated to remind the target user to actively correct the emotion category.

[0032] Optionally, the method further includes:

[0033] Determine the identity information of the target user, the identity information being used to indicate whether the target user is using smart home appliances for the first time;

[0034] When the identity information indicates that the target user is a first-time user of smart home appliances, a preset recommendation strategy is obtained, and new recommended content is determined based on the recommendation strategy.

[0035] If the identity information indicates that the target user is not a first-time user of smart home appliances, the step of determining recommended content based on the target user's emotional confidence level is performed.

[0036] Secondly, this application provides a content recommendation device based on multimodal emotion recognition, comprising:

[0037] The acquisition module is used to acquire multimodal data of the target user in response to the emotion recommendation activation command. The multimodal data includes at least two of the following: facial image, facial video, and audio information.

[0038] The determination module is used to determine the first emotion category of the target user and the emotion confidence level corresponding to the first emotion category based on the multimodal data.

[0039] The determining module is further configured to determine recommended content corresponding to the first emotion category when the emotion confidence level is determined to be greater than or equal to a preset confidence level.

[0040] Optionally, the determining module is further configured to determine facial expression confidence and a second emotion category based on the facial image and / or the facial video;

[0041] The determining module is further configured to determine the first speech tone confidence level and the third emotion category based on the audio information;

[0042] The determining module is specifically used to determine the first emotion category as the second emotion category, or to determine the first emotion category as the third emotion category, when the second emotion category and the third emotion category are the same.

[0043] The determining module is specifically used to multiply the facial expression confidence score by a first coefficient, and then add the first voice tone confidence score multiplied by a second coefficient to obtain the emotion confidence score.

[0044] Optionally, the device further includes: an input module;

[0045] The input module is used to input the facial image and / or the facial video into the face detection model to obtain the facial image output by the face detection model;

[0046] The device further includes: a processing module;

[0047] The processing module is used to preprocess the face image to obtain a processed face image;

[0048] The determining module is further configured to determine the facial expression confidence level and the second emotion category based on the processed facial image.

[0049] Optionally, the input module is further configured to input the processed face image into the expression classification model to obtain multiple fourth emotion categories output by the expression classification model, and the probability value corresponding to each emotion category;

[0050] The determining module is further configured to determine the maximum probability value from the probability values ​​corresponding to each of the emotion categories;

[0051] The determining module is specifically used to determine the maximum probability value as the facial expression confidence level, and to determine the fourth emotion category corresponding to the maximum probability value as the second emotion category.

[0052] Optionally, the determining module is further configured to determine the signal-to-noise ratio of the audio information;

[0053] The determining module is further configured to determine the second speech tone confidence level and the fifth emotion category of the audio information;

[0054] The determining module is specifically used to determine the second speech tone confidence as the first speech tone confidence when the signal-to-noise ratio is determined to be greater than the preset signal-to-noise ratio, and to determine the fifth emotion category as the third emotion category.

[0055] Optionally, the processing module is further configured to preprocess the audio information using voice activation detection technology to obtain a voice signal;

[0056] The processing module is further configured to perform feature processing on the speech signal to obtain the speech features corresponding to the speech signal;

[0057] The input module is specifically used to input the speech features into the speech recognition model to obtain the second speech tone confidence and the fifth emotion category.

[0058] Optionally, the apparatus further includes: a generation module;

[0059] The generation module is used to generate a prompt message when it is determined that the confidence level of the emotion is less than the preset confidence level. The prompt message is used to remind the target user to actively correct the emotion category.

[0060] Optionally, the determining module is further configured to determine the identity information of the target user, the identity information being used to indicate whether the target user is using smart home appliances for the first time;

[0061] The determining module is further configured to, when the identity information indicates that the target user is a first-time user of smart home appliances, obtain a preset recommendation strategy and determine new recommended content based on the recommendation strategy;

[0062] The determining module is further configured to perform a step of determining recommended content based on the target user's emotional confidence level when the identity information indicates that the target user is not a first-time user of smart home appliances.

[0063] Thirdly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0064] The memory stores computer-executed instructions;

[0065] The processor executes computer execution instructions stored in the memory to implement the content recommendation method based on multimodal emotion recognition as described in the first aspect and various possible implementations of the first aspect above.

[0066] Fourthly, this application provides a computer-readable storage medium storing computer-executable instructions thereon, which, when executed by a processor, are used to implement the content recommendation method based on multimodal emotion recognition as described in the first aspect and various possible implementations of the first aspect.

[0067] Fifthly, this application provides a program product, including a computer program, which, when executed by a processor, implements the content recommendation method based on multimodal emotion recognition as described above.

[0068] This application provides a content recommendation method, medium, and electronic device based on multimodal emotion recognition. Upon responding to an emotion recommendation activation command, the method acquires multimodal data of the target user, including at least two of the following: facial image, facial video, and audio information. Based on this, it determines the user's first emotion category and its corresponding emotion confidence level. When the emotion confidence level is greater than or equal to a preset confidence level, it determines the recommended content corresponding to the first emotion category. This method, by combining multimodal data with emotion confidence level judgment, makes content recommendations more aligned with the user's current emotional state, improving the accuracy and applicability of recommendations, thereby solving the problem of users manually adjusting television program content through voice commands. Attached Figure Description

[0069] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0070] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0071] Figure 1 This is a schematic diagram of the hardware environment for an interaction method of a smart device according to an embodiment of this application;

[0072] Figure 2 A flowchart illustrating a content recommendation method based on multimodal emotion recognition provided in this application. Figure 1 ;

[0073] Figure 3 A flowchart illustrating a content recommendation method based on multimodal emotion recognition provided in this application. Figure 2 ;

[0074] Figure 4 A flowchart illustrating a content recommendation method based on multimodal emotion recognition provided in this application. Figure 3 ;

[0075] Figure 5 A schematic diagram of the structure of a content recommendation device based on multimodal emotion recognition provided in this application;

[0076] Figure 6 This is a schematic diagram of the structure of a content recommendation device based on multimodal emotion recognition provided in this application. Detailed Implementation

[0077] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0078] It should be noted that the terms "first," "second," etc., used in this application's specification are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0079] According to one aspect of the embodiments of this application, a content recommendation method based on multimodal emotion recognition is provided. This content recommendation method based on multimodal emotion recognition is widely applicable to whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned content recommendation method based on multimodal emotion recognition can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.

[0080] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.

[0081] With the fast pace of modern life, people's leisure time has become increasingly limited. In this context, watching television programs has become a primary way for users to relax and entertain themselves, allowing them to efficiently unwind during their limited rest time.

[0082] Currently, when watching TV programs, users still need to manually operate or use voice commands to select the content they are interested in. This method cannot recommend suitable content based on the user's actual needs, which affects the user's viewing experience.

[0083] To address the aforementioned issues, this application provides a content recommendation method based on multimodal emotion recognition. In response to an emotion recommendation activation command, the method acquires multimodal data of the target user; based on the multimodal data, it determines the target user's first emotion category and the corresponding emotion confidence level; and if the emotion confidence level is greater than or equal to a preset confidence level, it determines recommended content corresponding to the first emotion category. This method achieves intelligent content recommendation without requiring users to actively perform button operations or issue voice commands, effectively solving the problem of users adjusting TV program content through manual operation or voice commands, thereby improving user convenience and experience.

[0084] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0085] Figure 2 A flowchart illustrating a content recommendation method based on multimodal emotion recognition provided in this application embodiment. Figure 1 ,like Figure 2As shown in this embodiment, the content recommendation method based on multimodal emotion recognition includes:

[0086] S101. In response to the emotion recommendation activation command, acquire multimodal data of the target user, including at least two of the following: facial image, facial video, and audio information.

[0087] Among them, the emotion recommendation activation command is used to instruct smart home appliances to activate the content recommendation function based on the target user's emotional state.

[0088] Facial images are used to characterize the static facial features of the target user, such as the curvature of the corners of the mouth, the state of the eyes, and the degree of frowning.

[0089] Facial videos are used to represent dynamic facial changes in users, such as the process of a target user's smile unfolding or the rhythm of a frown relaxing.

[0090] Audio information includes, but is not limited to: tone of voice, speaking speed, and volume.

[0091] The purpose of this step is to proactively collect multimodal data from target users by responding to sentiment-based recommendation activation commands.

[0092] Understandably, single-modal data cannot fully and accurately reflect a user's true emotions. For example, relying solely on facial images may fail to distinguish the subtle differences between a genuine smile and a fake one; relying solely on audio information may also lead to biased emotion judgments due to environmental noise interference. Therefore, by acquiring data from at least two modalities, the limitations of single-modal data can be reduced through complementary verification of data from different dimensions. This allows for more accurate subsequent emotion category recognition and emotion confidence calculation, ensuring a better match between the final recommended content and the user's emotions.

[0093] For example, the emotion recommendation activation command can be issued through a physical button on a smart home appliance, such as a dedicated emotion recommendation button on the remote control that the user presses to trigger it; it can also be a voice command issued by the user, such as saying a specific phrase like "activate emotion recommendation mode"; or it can be an instruction automatically generated by the smart home appliance's control system based on user habits, scenarios, and other factors, such as automatically activating it when the user is in a specific time period. This application does not impose any special restrictions on this.

[0094] S102. Based on the multimodal data, determine the target user's first emotion category and the corresponding emotion confidence level.

[0095] The first emotion category refers to the specific category that best matches the target user's current emotional state, determined through a series of algorithmic analyses based on collected multimodal data. The first emotion category includes, but is not limited to: happiness, sadness, anger, surprise, fear, disgust, and calmness.

[0096] Emotion confidence score is used to characterize the reliability of determining that a target user belongs to the first emotion category based on multimodal data. The higher the emotion confidence score, the higher the reliability of determining that the target user belongs to the first emotion category based on multimodal data; conversely, the lower the emotion confidence score, the lower the reliability of determining that the target user belongs to the first emotion category based on multimodal data.

[0097] Emotional confidence can be expressed as a decimal or as a percentage. This application does not impose any special restrictions on this.

[0098] The purpose of this step is to determine the user's current emotion category and the credibility of that emotion category.

[0099] Understandably, facial images, videos, and audio information can provide multi-dimensional clues about a target user's emotions, helping to improve the accuracy of emotion recognition. Therefore, by comprehensively considering facial images, videos, and audio information, it is possible to accurately capture changes in a user's emotions, thereby determining the user's current emotion category and the credibility of that category.

[0100] S103. If the confidence level of the emotion is greater than or equal to the preset confidence level, determine the recommended content corresponding to the first emotion category.

[0101] For example, the pre-set reliability can be expressed as a decimal or as a percentage. The pre-set reliability can be, for example, 0.8 or 80%. This application does not impose any special restrictions on this.

[0102] The purpose of determining whether the target user's emotional confidence level is greater than or equal to the preset confidence level is to determine whether the target user's first emotional category is trustworthy.

[0103] If the target user's sentiment confidence level is greater than or equal to the preset confidence level, it indicates that the target user's first sentiment category is trustworthy. Therefore, recommended content corresponding to the first sentiment category can be determined.

[0104] Understandably, emotion confidence level can characterize the accuracy of a smart home appliance's control system in judging emotions. Only when the emotion confidence level reaches a certain level can it be said that the smart home appliance's control system's judgment of emotions is relatively reliable. If the emotion confidence level is too low, it means that the smart home appliance's control system may be inaccurate in judging the user's emotions. In this case, recommending content based on the incorrectly judged emotion category will not only fail to meet the target user's needs but may also lead to a bad experience for the target user, causing the target user to distrust smart home appliances.

[0105] Therefore, when the emotion confidence level is greater than or equal to the preset confidence level, it indicates that the control system of smart home appliances accurately identifies the emotions of the target user. In this case, recommending content that corresponds to the user's emotions can improve the accuracy and effectiveness of the recommendations, better meet the user's emotions and needs, and provide the user with high-quality and personalized services.

[0106] Optionally, this application provides an implementation method for determining recommended content corresponding to a first emotion category, including:

[0107] The first step is to obtain the content recommendation database for the target users. The content recommendation database includes multiple emotion categories and a list of videos corresponding to each emotion category.

[0108] The video list includes at least one video to be played.

[0109] The purpose of this step is to obtain a video content resource library that is exclusive to the target user and contains a mapping between "emotion category" and "video list".

[0110] It is understandable that different users have different video preferences, and a general content recommendation database cannot adapt to individual differences; while a content recommendation database dedicated to a target user stores video resources that fit that user's preferences.

[0111] Therefore, by acquiring the content recommendation database of the target users, it is possible to avoid pushing content that the target users are not interested in, so that emotion-based recommendations can take into account the users' personal preferences at the same time, thereby significantly improving the recommendation effect and user experience.

[0112] The second step is to determine the target emotion category that corresponds to the first emotion category from multiple emotion categories.

[0113] The purpose of this step is to determine the target emotion category that corresponds to the user's current actual emotion from multiple emotion categories contained in the content recommendation database.

[0114] Understandably, content recommendation databases contain multiple emotion categories and corresponding video lists. If the target emotion category is not clearly defined, it is impossible to determine the range of content that matches the user's current emotion, which may result in recommended videos that do not match the user's emotion.

[0115] Therefore, by identifying the target emotion category, we can obtain the specific category that represents the target user's current emotion, thereby improving the personalization of subsequent content recommendations and providing users with content that better meets their needs.

[0116] The third step is to identify the video list corresponding to the target emotion category as recommended content.

[0117] The purpose of this step is to identify the video resources that best match the user's emotions from the content recommendation database.

[0118] For example, suppose the content recommendation database contains:

[0119] The list of things to be happy, including comedy and variety shows;

[0120] Sadness, and its corresponding lists, includes comedic videos and emotional videos;

[0121] Anger, and the corresponding list of anger, includes comedy videos and action videos;

[0122] Surprise and its corresponding list include suspense videos;

[0123] Fear and its corresponding lists include horror videos;

[0124] The list of things to dislike, and the corresponding lists of dislikes, includes variety shows;

[0125] The list of calm and tranquil items includes variety shows.

[0126] Given that the target user's primary emotion category is happiness, we can use the list of emotions associated with happiness as the recommended content.

[0127] This embodiment provides a content recommendation method based on multimodal emotion recognition. First, in response to an emotion recommendation activation command, multimodal data of the target user, including at least two of the following: facial image, facial video, and audio information, is acquired. Then, the user's first emotion category and its corresponding emotion confidence score are determined. Finally, when the emotion confidence score is greater than or equal to a preset confidence score, recommended content corresponding to the first emotion category is determined. This method combines multimodal data with emotion confidence score judgment, making content recommendations more aligned with the user's current emotional state, improving the accuracy and applicability of content recommendations. This solves the problem of users manually adjusting TV program content through voice commands, thereby enhancing the user experience.

[0128] Figure 3 A flowchart illustrating a content recommendation method based on multimodal emotion recognition provided in this application embodiment. Figure 2 .like Figure 3 As shown, in Figure 2 Based on the embodiments, a possible implementation of content recommendation based on multimodal emotion recognition is described in detail. The content recommendation method based on multimodal emotion recognition shown in this embodiment includes:

[0129] S201. In response to the emotion recommendation activation command, obtain multimodal data of the target user, including at least two of the following: facial image, facial video, and audio information.

[0130] The explanation of step S201 is the same as that in the above embodiments, and will not be repeated here.

[0131] S202. Determine the confidence level of facial expressions and the second emotion category based on facial images and / or facial videos.

[0132] Among them, facial expression confidence score is used to characterize the accuracy of judging the emotion category represented by a user's facial expressions based on facial images and / or facial videos. The higher the facial expression confidence score, the higher the accuracy of judging the emotion represented by a user's facial expressions based on facial images and / or facial videos; conversely, the lower the facial expression confidence score, the lower the accuracy of judging the emotion represented by a user's facial expressions based on facial images and / or facial videos.

[0133] Facial expression confidence can be expressed as a decimal or as a percentage. This application does not impose any special restrictions on this.

[0134] The second emotion category refers to the specific category that best matches the target user's current emotional state, determined through a series of algorithmic analyses based on captured facial images and / or videos. Second emotion categories include, but are not limited to: happiness, sadness, anger, surprise, fear, disgust, and calmness.

[0135] The purpose of this step is to determine the target user's emotional category and the accuracy of that emotional category determination based on the acquired facial images and / or facial videos.

[0136] This is understandable. First, facial images can reflect the static facial features of the target user, such as the curvature of the corners of the mouth, the state of the eyes, and the degree of frowning. Second, facial videos can reflect the dynamic changes in the user's face, such as the process of a smile unfolding and the rhythm of a frown relaxing.

[0137] Therefore, by comprehensively considering facial images and / or facial videos, it is possible to analyze the facial expressions of target users, thereby determining the user's current emotion category and the accuracy of that emotion category.

[0138] Optionally, this application provides a possible method for determining facial expression confidence and a second emotion category based on facial images and / or facial videos, including:

[0139] The first step is to input facial images and / or facial videos into the face detection model to obtain the facial images output by the face detection model.

[0140] The purpose of this step is to extract valid face images from facial images and / or facial videos containing facial information.

[0141] Understandably, facial images and / or videos may contain a significant amount of content irrelevant to emotion analysis, such as the surrounding scene and other people. Therefore, by inputting facial images and / or videos into a face detection model, the model can use pre-trained algorithms and rules to analyze the input images or videos, thereby identifying images containing only faces from the original facial image or video data.

[0142] The second step is to preprocess the face image to obtain the processed face image.

[0143] The preprocessing includes, but is not limited to: resizing, adjusting the cropped face image to a fixed pixel size (e.g., 224x224 pixels) to ensure that all images are consistent in size when input into subsequent models; grayscale conversion, converting the color face image to a grayscale image to reduce redundant data from the color channels; and normalization, scaling the pixel values ​​of the face image from the original 0-255 range to the range of 0-1 or -1 to 1, making the face image more suitable for computational needs.

[0144] Understandably, preprocessing facial images, including scaling to a fixed size, grayscale conversion, and normalization, can reduce unnecessary computational burden, thereby ensuring more accurate subsequent facial image-based analysis.

[0145] The third step is to determine the confidence level of facial expressions and the second emotion category based on the processed facial images.

[0146] Understandably, the processed facial images have undergone preprocessing to eliminate issues such as inconsistent sizes, pixel range clutter, and background interference. Therefore, by analyzing the processed facial images, it is possible to more accurately identify the emotion category that matches the user's facial expressions.

[0147] Optionally, this application provides a possible method for determining facial expression confidence and a second emotion category based on a processed facial image, including:

[0148] The first step is to input the processed face image into the expression classification model to obtain multiple fourth emotion categories output by the expression classification model, as well as the probability value corresponding to each emotion category.

[0149] The fourth emotion category includes, but is not limited to: happiness, sadness, anger, surprise, fear, disgust, and calmness.

[0150] The probability value corresponding to the emotion category represents how likely the expression classification model is to classify a face image as belonging to that emotion category. A higher probability value indicates a greater likelihood that the expression classification model considers the face image to belong to that emotion category. Conversely, a lower probability value indicates a lower likelihood that the expression classification model considers the face image to belong to that emotion category.

[0151] It is understandable that, given the complexity and diversity of facial expressions, the same facial image does not simply express only one emotion, but may simultaneously contain multiple emotional characteristics.

[0152] Therefore, by inputting the processed facial image into the expression classification model, the built-in algorithm of the expression classification model can be used to recognize the facial image, thereby obtaining multiple fourth emotion categories and the probability values ​​corresponding to each emotion category.

[0153] The second step is to determine the maximum probability value from the probability values ​​corresponding to each emotion category.

[0154] The purpose of this step is to determine the maximum probability value from multiple probability values.

[0155] Understandably, facial expression classification models can accurately identify processed facial images and output multiple emotion categories and their corresponding probability values. By selecting the highest probability value from these values, the most likely emotion category of the target user can be further determined.

[0156] The third step is to determine the maximum probability value as the facial expression confidence level, and to determine the fourth emotion category corresponding to the maximum probability value as the second emotion category.

[0157] Understandably, by determining the maximum probability value as the facial expression confidence level and the fourth emotion category corresponding to the maximum probability value as the second emotion category, we can obtain the emotion category that best matches the user's current emotion and obtain the facial expression confidence level corresponding to that emotion category.

[0158] S203. Based on the audio information, determine the confidence level of the first speech tone and the third emotion category.

[0159] The first speech tone confidence score is used to characterize the accuracy of judging the emotion category represented by the user's facial expression based on the audio information. The higher the first speech tone confidence score, the higher the accuracy of judging the emotion represented by the user's facial expression based on the audio information; conversely, the lower the first speech tone confidence score, the lower the accuracy of judging the emotion represented by the user's facial expression based on the audio information.

[0160] The confidence level of the first speech tone can be expressed as a decimal or as a percentage. This application does not impose any special restrictions on this.

[0161] The third emotion category refers to the specific category that best matches the target user's current emotional state, determined through a series of algorithmic analyses based on the collected audio information. The second emotion category includes, but is not limited to: happiness, sadness, anger, surprise, fear, disgust, and calmness.

[0162] The purpose of this step is to determine the target user's emotional category based on the acquired audio information, and to assess the accuracy of that emotional category determination.

[0163] Understandably, audio information contains clues about a user's emotions; characteristics such as tone, speed, and volume can directly reflect the target user's emotional state. For example, when angry, the voice may be high-pitched, fast-paced, and loud; when sad, the voice may be low-pitched, slow-paced, and soft.

[0164] Therefore, by considering audio information, it is possible to analyze the target user's speech and thus determine the user's current emotion category and the accuracy of that emotion category.

[0165] Optionally, this application provides a possible implementation method for determining a first speech tone confidence level and a third emotion category based on audio information, including:

[0166] The first step is to determine the signal-to-noise ratio of the audio information.

[0167] Signal-to-noise ratio (SNR) refers to the ratio of the power of the useful signal to the power of the noise signal in an audio signal. A higher SNR indicates that the useful signal is stronger relative to the noise signal, resulting in better audio quality; conversely, a lower SNR means that the noise signal interferes more with the useful signal, leading to poorer audio quality.

[0168] The purpose of this step is to assess the quality of the audio information.

[0169] Understandably, in real-life scenarios, the acquired audio information may contain various noise interferences, such as environmental noise and equipment noise, which can affect the judgment of the target user's emotions. Therefore, by determining the signal-to-noise ratio (SNR), we can determine the ratio of the power of the useful signal to the power of the noise signal in the audio information, and thus determine the quality of the audio information.

[0170] The second step is to determine the second speech tone confidence level and the fifth emotion category of the audio information.

[0171] The second speech tone confidence score is used to characterize the accuracy of judging the emotion category represented by a user's facial expression based on audio information. The higher the second speech tone confidence score, the higher the accuracy of judging the emotion represented by a user's facial expression based on audio information; conversely, the lower the second speech tone confidence score, the lower the accuracy of judging the emotion represented by a user's facial expression based on audio information.

[0172] The confidence level of the second speech tone can be expressed as a decimal or as a percentage. This application does not impose any special restrictions on this.

[0173] The fifth emotion category refers to the specific category that best matches the target user's current emotional state, determined through a series of algorithmic analyses based on collected audio information. The fifth emotion category includes, but is not limited to: happiness, sadness, anger, surprise, fear, disgust, and calmness.

[0174] The purpose of this step is to determine the target user's current emotional category and the accuracy of that emotional category determination based on the target user's audio information.

[0175] The third step is to determine the second speech tone confidence level as the first speech tone confidence level, and the fifth emotion category as the third emotion category, provided that the signal-to-noise ratio is greater than the preset signal-to-noise ratio.

[0176] The preset signal-to-noise ratio can be, for example, 20 dB or 25 dB. This application does not impose any special restrictions on this.

[0177] The purpose of this step, which determines whether the signal-to-noise ratio of the audio information is greater than the preset signal-to-noise ratio, is to determine whether the confidence level of the second speech tone and the fifth emotion category determined based on the audio information are reliable.

[0178] If the signal-to-noise ratio (SNR) of the audio information is greater than a preset SNR, it indicates that the second speech tone confidence level and the fifth emotion category determined based on the audio information are reliable. Therefore, the second speech tone confidence level can be determined as the first speech tone confidence level, and the fifth emotion category can be determined as the third emotion category.

[0179] For example, assuming a preset signal-to-noise ratio (SNR) of 20 dB, and the target user's emotion category is known to be "happy," with a corresponding speech pitch confidence level of 0.8, and the SNR of the target user's audio information is 26 dB, then based on this information, it can be determined that the target user's current emotion category is "happy," and the corresponding speech pitch confidence level is 0.8.

[0180] Optionally, this application provides a possible implementation method for determining the second speech tone confidence level and the fifth emotion category of audio information, including:

[0181] The first step is to use voice activation detection technology to preprocess the audio information and obtain the voice signal.

[0182] Among them, voice activation detection technology is a technology that can automatically distinguish between "effective voice segments (i.e., the target user's voice)" and "non-voice segments (such as environmental noise, background noise, etc.)" from audio information.

[0183] Speech signals refer to audio segments containing only valid human voices extracted from the original audio information after preprocessing with speech activation detection technology.

[0184] The purpose of this step is to extract audio segments from the original audio information that contain only the target user's voice.

[0185] The second step is to perform feature processing on the speech signal to obtain the speech features corresponding to the speech signal.

[0186] Among them, speech features include, but are not limited to: fundamental frequency, Mel frequency cepstral coefficient, energy, spectral centroid, and harmonic noise ratio.

[0187] The fundamental frequency refers to the basic frequency of sound, which directly corresponds to the pitch. For example, when happy, the pitch is higher and more varied; when sad, the pitch is lower and flatter.

[0188] Mel frequency cepstral coefficients are used to characterize the spectral shape of sound, are related to the perceptual characteristics of the human ear, and are key features for speech recognition and emotion recognition.

[0189] Energy refers to the intensity of sound. For example, the energy level is usually higher when one is angry.

[0190] The spectral centroid is used to distinguish speech differences under different emotional states. For example, the spectral centroid of excited speech is higher than that of calm speech.

[0191] The harmonic noise ratio (HNR) is used to assess the purity of a sound. A higher HNR indicates that the harmonic components in the sound are more prominent and the noise is less, resulting in a purer sound (such as a cappella singing in a quiet environment). Conversely, a lower HNR indicates more severe noise interference and a muddier sound (such as talking in a noisy environment).

[0192] The purpose of this step is to extract acoustic features that characterize pitch and emotion from audio segments containing only the target user's voice.

[0193] The third step is to input the speech features into the speech recognition model to obtain the second speech tone confidence score and the fifth emotion category.

[0194] The purpose of this step is to analyze the extracted speech features using a speech recognition model to determine the current emotion category of the target user and the accuracy of that emotion category determination.

[0195] Understandably, speech features contain multiple fragmented emotional characteristics, which cannot directly reflect the emotion category or the reliability of the judgment. Therefore, by using speech features as input into a speech recognition model, the built-in algorithms of the speech recognition model can integrate and analyze the speech features to obtain the target user's current emotion category and the accuracy of the emotion category judgment.

[0196] S204. If the second emotion category and the third emotion category are the same, determine the first emotion category as the second emotion category, or determine the first emotion category as the third emotion category.

[0197] The purpose of this step is to determine the user's final emotion category, provided that the emotion results identified based on facial images and / or facial videos are the same as the emotion results identified based on audio information.

[0198] Understandably, facial images / videos and audio information reflect a user's emotions from different dimensions. When the emotion results recognized based on facial images and / or facial videos are consistent with the emotion results recognized based on audio information, it indicates that the two emotion recognition results are consistent. Therefore, whether the second emotion category or the third emotion category is used as the target user's final emotion category, the target user's emotion can be accurately determined.

[0199] S205. Multiply the facial expression confidence score by the first coefficient, and then add the first voice tone confidence score multiplied by the second coefficient to obtain the emotion confidence score.

[0200] The first coefficient can be, for example, 0.6, and the second coefficient can be, for example, 0.4.

[0201] For example, suppose the first coefficient is 0.6 and the second coefficient is 0.4. Given that the confidence level for facial expression is 0.8 and the confidence level for the first tone of voice is 0.6, then based on the above information, the confidence level for emotion can be determined to be 0.72.

[0202] S206. If the confidence level of the emotion is determined to be less than the preset confidence level, a prompt message is generated to remind the target user to actively correct the emotion category.

[0203] The purpose of this step is to remind the target user to proactively adjust their emotion category.

[0204] Understandably, when the emotion confidence level is lower than the preset confidence level, it indicates that the smart home appliance control system's judgment of the user's emotion category based on multimodal data has low reliability. Therefore, by generating prompts, the target user can be guided to actively correct their judgment, thereby obtaining an emotion category that better reflects the user's actual emotions.

[0205] For example, suppose the preset confidence level is 0.8. The target user's sentiment confidence level is known to be 0.72. Based on this information, a prompt message can be generated to remind the target user to actively correct their sentiment category.

[0206] This embodiment provides a content recommendation method based on multimodal emotion recognition. First, in response to an emotion recommendation activation command, it acquires multimodal data of the target user, including at least two of the following: facial image, facial video, and audio information. Next, it determines facial expression confidence and a second emotion category based on the facial image and / or facial video, and determines a first voice tone confidence and a third emotion category based on the audio information. When the second and third emotion categories are determined to be the same, this same category is designated as the first emotion category. Subsequently, based on the facial expression confidence, the first voice tone confidence, and the corresponding coefficients, it calculates the target user's emotion confidence. If the emotion confidence is less than a preset confidence level, a prompt message is generated to remind the target user to actively correct the emotion category.

[0207] This method first identifies emotions through both facial and audio data, improving the accuracy of emotion category judgment. Second, it calculates the confidence score of emotions by weighting coefficients, making the confidence score result more consistent with the actual recognition situation. Then, it provides a prompt when the confidence score is unreliable, which not only ensures the reliability of subsequent content recommendations, but also further simplifies user operations and improves ease of use and user experience.

[0208] Figure 4 This application provides a schematic diagram of a content recommendation process based on multimodal emotion recognition, as illustrated in an embodiment of the present application. Figure 3 .like Figure 4 As shown, based on the above embodiments, content recommendation can also be performed based on the user's identity information. The content recommendation method based on multimodal emotion recognition shown in this embodiment includes:

[0209] S301. Determine the target user's identity information, which is used to indicate whether the target user is using smart home appliances for the first time.

[0210] The purpose of this step is to determine whether the target user is a first-time user of smart home appliances.

[0211] It's understandable that users in different usage states have significantly different needs for smart home appliances. For example, first-time users need basic services such as device guidance and function explanations to help them get started quickly, while users who are not first-time users need more personalized function recommendations.

[0212] Therefore, by identifying the target user's identity information, smart home appliances can differentiate between user types and provide targeted services. This not only improves the user experience for first-time users but also meets the personalized needs of repeat users, avoiding the poor user experience caused by standardized services.

[0213] Optionally, this application provides several possible implementation methods, including:

[0214] (1) First method: First, obtain the target user's account information, which includes the target user's first identification information; and obtain a preset user database, which includes multiple second identification information; then, when it is determined that the first information identifier is the same as any one of the multiple second identification information, determine the target user's identity information.

[0215] Understandably, the first identification information can be used to determine whether the target user is using smart home appliances for the first time. If the target user is using it for the first time, it means that the target user is in the registration process, and the corresponding first identification information has not yet been stored in the user database; if the target user is not using it for the first time, it means that they have completed registration, and the first identification information has been stored in the user database.

[0216] Therefore, by considering the storage status of the target user's primary identification information in the user database, it means that the user's identity information can be determined.

[0217] Therefore, the identity information of the target user can be determined from the user database based on the target user's primary identification information.

[0218] (2) The second method: First, perform feature extraction processing on the facial image of the target user to obtain the facial feature vector corresponding to the facial image; then, obtain a preset facial feature vector database, which includes multiple preset facial feature vectors; next, determine the feature similarity between the facial feature vector and the multiple preset facial feature vectors respectively; then, if the feature similarity between the facial feature vector and any preset facial feature vector is greater than the preset feature similarity, determine that the target user is not using smart home appliances for the first time; otherwise, if the feature similarity between the facial feature vector and any preset facial feature vector is less than or equal to the preset feature similarity, determine that the target user is using smart home appliances for the first time.

[0219] Understandably, facial feature vectors can be used to determine whether a target user is using smart home appliances for the first time. If the target user has used smart home appliances before, it means that the facial feature vector extracted from the target user's facial image has a feature similarity greater than the preset feature similarity of any preset facial feature vector stored in the facial feature vector database. If the target user is using it for the first time, it means that the facial feature vector extracted from the target user's facial image has a feature similarity less than or equal to the preset feature similarity of any preset facial feature vector in the facial feature vector database.

[0220] Therefore, by calculating the feature similarity between the target user's facial feature vector and a preset facial feature vector in the facial feature vector database, and comparing it with the preset feature similarity, the target user's identity information can be determined.

[0221] S302. If the identity information indicates that the target user is a first-time user of smart home appliances, obtain the preset recommendation strategy and determine new recommended content based on the recommendation strategy.

[0222] Understandably, once the target user is identified as a first-time user of smart home appliances, generating appropriate recommended content based on a preset recommendation strategy can provide targeted guidance and services to first-time users, helping them quickly become familiar with the device's functions, master basic operations, avoid usage obstacles due to unfamiliarity with the device, and improve the first-time user experience.

[0223] S303. If the identity information indicates that the target user is not a first-time user of smart home appliances, perform the step of determining the recommended content based on the target user's emotional confidence level.

[0224] Understandably, once it's determined that the target user is not a first-time user of smart home appliances, it means they are already familiar with the device's functions and have mastered basic operations, eliminating the need for introductory guidance. Therefore, content tailored to the target user's emotions can be generated to meet their personalized needs, further enhancing their user experience and device satisfaction.

[0225] This embodiment provides a content recommendation method based on multimodal emotion recognition. The method first determines the target user's identity information to ascertain whether the target user is a first-time user of smart home appliances. If the identity information indicates that the target user is a first-time user, a preset recommendation strategy is obtained and recommended content is determined accordingly. If the target user is not a first-time user, the step of determining the target user's emotion confidence level is executed. This method first employs a preset recommendation strategy for first-time users, quickly providing suitable content without complex calculations, lowering the barrier to entry for new users and improving the initial experience. Secondly, by introducing emotion confidence level judgment for non-first-time users, it can provide recommendations based on the user's real-time emotional state, making the recommended content more closely aligned with the user's personalized needs.

[0226] Figure 5 This is a schematic diagram of a content recommendation device based on multimodal emotion recognition provided in this application. Figure 5 As shown, this application provides a content recommendation device based on multimodal emotion recognition. The content recommendation device 400 based on multimodal emotion recognition includes:

[0227] The acquisition module 401 is used to acquire multimodal data of the target user in response to the emotion recommendation activation command. The multimodal data includes at least two of the following: facial image, facial video, and audio information.

[0228] The determination module 402 is used to determine the first emotion category of the target user and the emotion confidence level corresponding to the first emotion category based on multimodal data;

[0229] The determination module 402 is also used to determine the recommended content corresponding to the first emotion category when the determination of the emotion confidence is greater than or equal to the preset confidence.

[0230] Optionally, the determining module 402 is also configured to determine facial expression confidence and a second emotion category based on facial images and / or facial videos;

[0231] The determination module 402 is also used to determine the confidence level of the first speech tone and the third emotion category based on the audio information;

[0232] The determination module 402 is specifically used to determine the first emotion category as the second emotion category when the second emotion category and the third emotion category are the same, or to determine the first emotion category as the third emotion category.

[0233] The determination module 402 is specifically used to multiply the facial expression confidence score by a first coefficient, and then add the first voice tone confidence score multiplied by a second coefficient to obtain the emotion confidence score.

[0234] Optionally, the device may also include: an input module 403;

[0235] The input module 403 is used to input facial images and / or facial videos into the face detection model to obtain the facial image output by the face detection model;

[0236] The device also includes: a processing module 404;

[0237] Processing module 404 is used to preprocess the face image to obtain the processed face image;

[0238] The determination module 402 is also used to determine the facial expression confidence level and the second emotion category based on the processed face image.

[0239] Optionally, the input module 403 is also used to input the processed face image into the expression classification model to obtain multiple fourth emotion categories output by the expression classification model, as well as the probability value corresponding to each emotion category;

[0240] The determination module 402 is also used to determine the maximum probability value from the probability values ​​corresponding to each emotion category;

[0241] The determination module 402 is specifically used to determine the maximum probability value as the facial expression confidence level, and to determine the fourth emotion category corresponding to the maximum probability value as the second emotion category.

[0242] Optionally, the determining module 402 is also used to determine the signal-to-noise ratio of the audio information;

[0243] The determination module 402 is also used to determine the second speech tone confidence level and the fifth emotion category of the audio information;

[0244] The determination module 402 is specifically used to determine the second speech tone confidence as the first speech tone confidence when the signal-to-noise ratio is determined to be greater than the preset signal-to-noise ratio, and to determine the fifth emotion category as the third emotion category.

[0245] Optionally, the processing module 404 is also used to preprocess the audio information using voice activation detection technology to obtain a voice signal;

[0246] The processing module 404 is also used to perform feature processing on the speech signal to obtain the speech features corresponding to the speech signal;

[0247] The input module 403 is specifically used to input speech features into the speech recognition model to obtain the second speech tone confidence and the fifth emotion category.

[0248] Optionally, the apparatus also includes: a generation module 405;

[0249] The generation module 405 is used to generate a prompt message when the confidence level of the emotion is determined to be less than the preset confidence level. The prompt message is used to remind the target user to actively correct the emotion category.

[0250] Optionally, the determining module 402 is also used to determine the identity information of the target user, which is used to indicate whether the target user is using smart home appliances for the first time;

[0251] The determination module 402 is also used to obtain a preset recommendation strategy and determine new recommended content based on the recommendation strategy when the identity information indicates that the target user is a first-time user of smart home appliances;

[0252] The determination module 402 is further configured to perform the step of determining recommended content based on the target user's emotional confidence level when the identity information indicates that the target user is not a first-time user of smart home appliances.

[0253] The content recommendation device based on multimodal emotion recognition provided in this application embodiment has a similar implementation principle and technical effect to the implementation of each part of the aforementioned content recommendation method based on multimodal emotion recognition, and will not be described again here.

[0254] Figure 6 This is a schematic diagram of the structure of a content recommendation device based on multimodal emotion recognition provided in this application. Figure 6 As shown, this application provides a content recommendation device based on multimodal emotion recognition. The content recommendation device 500 based on multimodal emotion recognition includes: a receiver 501, a transmitter 502, a processor 503, and a memory 504.

[0255] Receiver 501 is used to receive instructions and data;

[0256] Transmitter 502 is used to send commands and data;

[0257] Memory 504 is used to store instructions executed by the computer;

[0258] Processor 503 is used to execute computer execution instructions stored in memory 504 to implement the various steps of the test method in the above embodiments. For details, please refer to the relevant descriptions in the foregoing test method embodiments.

[0259] Optionally, the memory 504 can be either standalone or integrated with the processor 503.

[0260] When the memory 504 is set up independently, the electronic device also includes a bus for connecting the memory 504 and the processor 503.

[0261] The implementation principle and technical effects of the electronic device provided in this embodiment can be found in the foregoing embodiments, and will not be repeated here.

[0262] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method of any of the foregoing embodiments.

[0263] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the method of any of the foregoing embodiments.

[0264] The integrated modules described above, implemented as software functional modules, can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods of the various embodiments of this application.

[0265] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor. The memory may include high-speed RAM, and may also include non-volatile memory (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk, or optical disc, etc.

[0266] The aforementioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0267] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. Both the processor and the storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic device or host device.

[0268] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0269] The above are merely preferred embodiments of this application. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A content recommendation method based on multimodal emotion recognition, characterized in that, The method includes: In response to an instruction to enable emotion-based recommendations, multimodal data of the target user is acquired, wherein the multimodal data includes at least two of the following: facial images, facial videos, and audio information. Based on the multimodal data, determine the first emotion category of the target user and the emotion confidence level corresponding to the first emotion category; If the confidence level of the emotion is determined to be greater than or equal to a preset confidence level, recommended content corresponding to the first emotion category is determined.

2. The method according to claim 1, characterized in that, The step of determining the first emotion category of the target user and the corresponding emotion confidence level based on the multimodal data includes: Based on the facial images and / or the facial videos, determine the facial expression confidence level and the second emotion category; Based on the audio information, determine the confidence level of the first speech tone and the third emotion category; If the second emotion category and the third emotion category are the same, the first emotion category is determined to be the second emotion category, or the first emotion category is determined to be the third emotion category. The confidence score of the facial expression is obtained by multiplying it by a first coefficient and then adding it to the confidence score of the first voice tone multiplied by a second coefficient.

3. The method according to claim 2, characterized in that, The step of determining facial expression confidence and a second emotion category based on the facial image and / or the facial video includes: The facial image and / or the facial video are input into the face detection model to obtain the facial image output by the face detection model; The face image is preprocessed to obtain a processed face image; Based on the processed facial image, the confidence level of the facial expression and the second emotion category are determined.

4. The method according to claim 3, characterized in that, The step of determining the facial expression confidence level and the second emotion category based on the processed facial image includes: The processed face image is input into the expression classification model to obtain multiple fourth emotion categories output by the expression classification model, as well as the probability value corresponding to each emotion category; Determine the maximum probability value from the probability values ​​corresponding to each of the aforementioned emotion categories; The maximum probability value is determined as the facial expression confidence level, and the fourth emotion category corresponding to the maximum probability value is determined as the second emotion category.

5. The method according to claim 2, characterized in that, The step of determining the first speech tone confidence level and the third emotion category based on the audio information includes: Determine the signal-to-noise ratio of the audio information; Determine the second speech tone confidence level and the fifth emotion category of the audio information; If the signal-to-noise ratio is determined to be greater than the preset signal-to-noise ratio, the second speech tone confidence is determined as the first speech tone confidence, and the fifth emotion category is determined as the third emotion category.

6. The method according to claim 5, characterized in that, The determination of the second speech tone confidence level and the fifth emotion category of the audio information includes: The audio information is preprocessed using voice activation detection technology to obtain a voice signal; The speech signal is subjected to feature processing to obtain the speech features corresponding to the speech signal; The speech features are input into the speech recognition model to obtain the second speech tone confidence score and the fifth emotion category.

7. The method according to claim 1, characterized in that, The method further includes: If the confidence level of the emotion is determined to be less than the preset confidence level, a prompt message is generated to remind the target user to actively correct the emotion category.

8. The method according to claim 1, characterized in that, The method further includes: Determine the identity information of the target user, the identity information being used to indicate whether the target user is using smart home appliances for the first time; When the identity information indicates that the target user is a first-time user of smart home appliances, a preset recommendation strategy is obtained, and new recommended content is determined based on the recommendation strategy. If the identity information indicates that the target user is not a first-time user of smart home appliances, the step of determining recommended content based on the target user's emotional confidence level is performed.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method of any one of claims 1 to 8.

10. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 8 through the computer program.