Media content recommendation method and device

By acquiring the user's image and sound information from multimedia devices and analyzing the user's reaction status to obtain evaluation information, the problem of inaccurate recommendations in the existing technology is solved, more accurate media content recommendations are achieved, and the user experience and the accuracy of evaluation information are improved.

CN113574525BActive Publication Date: 2025-09-09HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN201980094051.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-04-29
Publication Date
2025-09-09
Estimated Expiration
2039-04-29

AI Technical Summary

Technical Problem

When existing multimedia devices recommend media content, methods based on historical playback records cannot accurately recommend content of interest to users.

Method used

By acquiring the user's image or sound information and analyzing the user's reaction status to obtain evaluation information as the basis for recommendation, the image acquisition device may acquire images at irregular intervals or continuously within a preset time period, and the sound acquisition device may acquire sounds within a preset time period. The device may be turned on based on the user's authorization information, and the evaluation may be determined based on the facial expression and sound emotional information.

Benefits of technology

It improves the accuracy of media content recommendations, ensures that the recommended content is in line with user interests, and enhances user experience and the accuracy of evaluation information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113574525B_ABST
    Figure CN113574525B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a method and device for recommending media content. The method includes: obtaining user reaction status information during media content playback, the reaction status information including at least one of the following types of information: user image information obtained via an image capture device or user voice information obtained via a voice capture device; and obtaining user evaluation information of the media content based on the reaction status information, wherein the evaluation information serves as a basis for recommending other media content to the user. This embodiment can accurately determine whether the user is interested in the currently playing media content based on the user's image information or voice information, thereby recommending programs of interest to the user, thereby improving the accuracy of media content recommendations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of multimedia technology, and in particular to a method and device for recommending media content. Background Art

[0002] With the development of computer technology, the functions of intelligent multimedia playback devices are becoming more and more abundant, and there are more and more content providers serving intelligent multimedia playback devices, generating more and more playback content.

[0003] In the prior art, multimedia devices often recommend media content to users based on historical playback records. Specifically, a user logs into a multimedia device using an account. The device then retrieves the historical playback records corresponding to that account, retrieves similar media content based on these historical playback records, and recommends this similar media content to the user. After the user watches the media content, the device records the played media content in the historical playback records. The next time the user recommends media content, the device will use this new historical playback record to recommend similar media content.

[0004] However, recommending programs based on historical playback records cannot accurately recommend media content to users. Summary of the Invention

[0005] The embodiments of the present application provide a method and device for recommending media content to users in an accurate manner.

[0006] In a first aspect, an embodiment of the present application provides a media content recommendation method, which obtains user reaction status information when media content is played, wherein the reaction status information includes at least one of the following types of information: user image information obtained by an image acquisition device or user sound information obtained by a sound acquisition device; and secondly, based on the reaction status information, obtains user evaluation information on the media content, wherein the evaluation information serves as a basis for recommending other media content to the user.

[0007] In the above process, the user's image information or sound information can be used to accurately determine whether the user is interested in the currently playing media content. Secondly, the obtained evaluation information is used as the basis for recommending other media content to the user, so that the other media content recommended to the user can be the content that the user is interested in, thereby improving the accuracy of the recommended media content.

[0008] In a possible implementation, the image information includes one or more images captured by an image capture device at irregular intervals or continuously or based on a first image capture frequency within a preset time period, where the preset time period is a time period for playing preset content in the media content.

[0009] Among them, by setting a variety of image acquisition frequencies within a preset time period, the images acquired by the image acquisition device can be applied to various different scenarios, thereby improving the applicability of the image acquisition device.

[0010] In a possible implementation, the image information also includes one or more images captured by the image capture device at a second image capture frequency in time periods other than the preset time period; wherein the first image capture frequency is higher than the second image capture frequency.

[0011] By capturing images based on the second image capturing frequency in other time periods, wherein the first image capturing frequency is higher than the second image capturing frequency, processing efficiency can be improved and storage space of the multimedia device can be saved.

[0012] In one possible embodiment, the sound information includes one or more sound segments collected by the sound collection device within a preset time period, where the preset time period is the time period for playing preset content in the media content, and the frequency of collecting sound within the preset time period includes any one of the following: continuous collection, collection at a preset frequency, or collection at irregular intervals.

[0013] Among them, by setting various optional collection frequencies in a preset time period to collect one or more sound segments, the applicability of the sound collection device is improved.

[0014] In a possible implementation, before obtaining the user's reaction status information, first obtain the user's authorization information for instructing the device to turn on, specifically:

[0015] If the reaction status information includes image information, the image acquisition device is turned on according to the first authorization information of the user, where the first authorization information is used to instruct the turning on of the image acquisition device;

[0016] If the response status information includes sound information, turning on the sound collection device according to the user's second authorization information, where the second authorization information is used to instruct the turning on of the sound collection device;

[0017] If the reaction status information includes image information and sound information, the image acquisition device is activated according to the first authorization information and the sound acquisition device is activated according to the second authorization information.

[0018] In the above process, the image acquisition device and / or sound acquisition device is turned on according to the user's authorization information, so as to ensure that the user's reaction status information is obtained with the user's authorization, avoid infringing the user's privacy, and improve the user experience. In addition, the device that needs to be turned on is specifically determined based on the reflected information, thereby saving resources and avoiding waste.

[0019] In a possible implementation, obtaining user evaluation information of media content according to the response status information includes:

[0020] If the reaction state information includes image information, obtaining user evaluation information of the media content based on the user's facial expression information;

[0021] If the reaction status information includes voice information, obtaining the user's evaluation information of the media content according to the user's voice emotion information;

[0022] If the reaction state information includes image information and sound information, obtaining the user's evaluation information of the media content based on the user's facial expression information and sound emotion information;

[0023] The facial expression information is the facial emotion information of the user when watching the media content obtained based on the image information, and the voice emotion information is the voice emotion information of the user when watching the preset content obtained based on the voice information.

[0024] In the above process, the user's facial expression information and the user's voice emotion information can intuitively and accurately reflect the user's state when watching media content, thereby ensuring the reliability of the obtained evaluation information.

[0025] In one possible implementation, obtaining user evaluation information of media content based on user facial expression information includes:

[0026] If the user's facial expression information is obtained within a preset time period, then obtaining standard facial expression information corresponding to the preset time period, where the standard facial expression information is expression information predefined according to preset content;

[0027] If the user's facial expression information is consistent with the standard facial expression information, then the user's evaluation of the media content is determined to be a plus-point evaluation;

[0028] If the user's facial expression information is inconsistent with the standard facial expression information, it is determined that the user's evaluation of the media content is a deductible evaluation.

[0029] Among them, the implementation process of comparing the user's facial expression information with the standard facial expression information is relatively simple. Secondly, based on the comparison, the user's evaluation information is determined to be a plus-point evaluation or a minus-point evaluation. The plus-minus-point method can effectively improve the efficiency of obtaining evaluation information.

[0030] In one possible implementation, obtaining user evaluation information of media content based on user facial expression information includes:

[0031] If the user's facial expression information is obtained in other time periods, obtaining an evaluation mapping table, wherein the evaluation mapping table is used to indicate evaluation information corresponding to different facial expression information;

[0032] The user's evaluation information on the media content is obtained based on the user's facial expression information and the evaluation mapping table.

[0033] In the above process, when the user's facial expression information is obtained in other time periods, the user's evaluation information on the media content is obtained in real time through the user's facial expression information and the evaluation mapping table, so as to quickly and effectively determine the user's feedback effect on the program and ensure the accuracy of the evaluation information.

[0034] In one possible implementation, obtaining user evaluation information of media content based on user voice emotion information includes:

[0035] Obtaining standard sound emotion information corresponding to a preset time period, where the standard sound emotion information is sound information predefined according to preset content;

[0036] If the user's voice emotion information is consistent with the standard voice emotion information, the user's evaluation of the media content is determined to be a plus-point evaluation;

[0037] If the user's voice emotion information is inconsistent with the standard voice emotion information, the user's evaluation of the media content is determined to be a deductible evaluation.

[0038] Among them, the user's voice emotion information is compared with the standard voice emotion information to determine the user's evaluation information on the currently played media content, which can accurately obtain the user's feedback on the program content in the preset time period, thereby improving the authenticity and effectiveness of the user's evaluation information.

[0039] In a possible implementation, after obtaining the user's reaction status information, the method further includes:

[0040] Obtaining an identity identifier of at least one user based on the image information or the sound information;

[0041] After obtaining the user's evaluation information of the media content according to the response status information, the method further includes:

[0042] The evaluation information of each user on the media content is associated with the user's identity.

[0043] In the above process, by associating the evaluation information of the media content with the user's identity, when subsequently recommending media content to the user, the media content can be recommended based on the current associated information, thereby improving the utilization rate of the evaluation information and improving the accuracy of the recommendation.

[0044] In one possible implementation, obtaining user evaluation information of media content based on user facial expression information and user voice emotion information includes:

[0045] obtaining an identity identifier of at least one user based on the image information;

[0046] Obtaining target facial expression information that matches the voice emotion information, and obtaining a target identity corresponding to the target facial expression information from at least one user identity, wherein the target facial expression information and the emotion corresponding to the voice emotion information are consistent;

[0047] According to the sound emotion information and the standard sound emotion information corresponding to the preset time period, the evaluation information of the user corresponding to the target identity on the media content is obtained, and the standard sound emotion information is information predefined according to the preset content.

[0048] In the above process, facial expression information is matched with voice emotion information, and secondly, evaluation information of each user can be determined based on their facial expression information and voice emotion information, so as to improve the comprehensiveness and pertinence of user evaluation information acquisition.

[0049] In a possible implementation, before the media content is played, the method further includes:

[0050] Obtaining a correspondence between the user's identity identifier and the user's features based on the user's identity identifier and user features, where the user features include one of a face or a voice;

[0051] Obtaining at least one user's identity based on the image information or the sound information includes:

[0052] Obtaining at least one user's identity identifier based on a correspondence between the user's identity identifier and the face and the face included in the image information; or

[0053] At least one user's identity identifier is obtained according to the correspondence between the user's identity identifier and the sound and the sound included in the sound information.

[0054] Among them, by inputting the user's identity identifier and user characteristics in advance, it is possible to directly match when performing media content recommendation in the future without the need for real-time acquisition and matching, thereby improving operation efficiency.

[0055] In a possible implementation, after obtaining user evaluation information of the media content according to the reaction status information, the method further includes:

[0056] If the evaluation information is plus-point information, then updating the user's evaluation information on the media content to obtain updated evaluation information;

[0057] If the evaluation information is deduction information, it is determined whether to update the user's evaluation information on the media content according to the status information obtained after the status information.

[0058] When the user's evaluation of the media content is a negative evaluation, the status information after the current status information is comprehensively used to determine whether the user's evaluation information needs to be updated, thereby avoiding evaluation information errors caused by false detection and improving the accuracy of the evaluation information.

[0059] In a possible implementation, before the media content is played, the method further includes:

[0060] Identify the user's identity, and if the user's historical evaluation information is obtained based on the user's identity, determine the media content recommended to the user based on the historical evaluation information.

[0061] In the above implementation, if historical evaluation information can be obtained according to the user's identity, media content recommended to the user can be determined based on the historical evaluation information, which can effectively improve the efficiency of recommendation and the availability of historical evaluation information.

[0062] In a possible implementation, after the media content is played, the method further includes:

[0063] Determine an evaluation list based on the user's evaluation information, where the evaluation list includes media content that has been evaluated by the user;

[0064] Send the evaluation list to the media source platform and obtain the recommendation list returned by the media source platform, which includes media content to be recommended to the user.

[0065] In the above implementation, by sending the evaluation list determined according to the user's evaluation information to the media source platform, the media content currently to be recommended to the user is of interest to the user, thereby improving the accuracy of media content recommendation.

[0066] In a possible implementation, determining the evaluation list based on the user's evaluation information includes:

[0067] If a user to be recommended is identified, determining media content whose evaluation information associated with the user to be recommended satisfies a first preset condition; wherein the first preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking;

[0068] An evaluation list is determined based on the media content that meets the first preset condition.

[0069] In the above implementation, when there is only one user to be recommended, media programs that meet the first preset condition are recommended, wherein the first preset condition only considers the preferences of the current single user, thereby improving the efficiency and accuracy of media program recommendation.

[0070] In a possible implementation, determining the evaluation list based on the user's evaluation information includes:

[0071] If at least two users to be recommended are identified, determining media content whose evaluation information associated with each of the at least two users to be recommended satisfies a second preset condition, where the second preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking;

[0072] An evaluation list is determined based on the media contents that all meet the second preset condition.

[0073] When multiple users watch media content at the same time, media content is recommended based on the evaluation information of each user and the second preset condition. The second preset condition takes into account the common preferences of multiple users, so that media content suitable for multiple users can be recommended to improve user experience.

[0074] In a second aspect, an embodiment of the present application provides a media content recommendation module, comprising an input module and a processing module, wherein the input module is configured to obtain user reaction status information when media content is played, the reaction status information including at least one of the following types of information: user image information obtained by an image acquisition device or user voice information obtained by a voice acquisition device;

[0075] The processing module is used to obtain the user's evaluation information on the media content according to the response status information, wherein the evaluation information is used as a basis for recommending other media content to the user.

[0076] In a possible implementation, the image information includes one or more images captured by an image capture device at irregular intervals or continuously or based on a first image capture frequency within a preset time period, where the preset time period is a time period for playing preset content in the media content.

[0077] In a possible implementation, the image information also includes one or more images captured by the image capture device at a second image capture frequency in time periods other than the preset time period; wherein the first image capture frequency is higher than the second image capture frequency.

[0078] In one possible embodiment, the sound information includes one or more sound segments collected by the sound collection device within a preset time period, where the preset time period is the time period for playing preset content in the media content, and the frequency of collecting sound within the preset time period includes any one of the following: continuous collection, collection at a preset frequency, or collection at irregular intervals.

[0079] In a possible implementation, before obtaining the user's reaction status information, the processing module is further configured to:

[0080] If the reaction status information includes image information, the image acquisition device is turned on according to the first authorization information of the user, where the first authorization information is used to instruct the turning on of the image acquisition device;

[0081] If the response status information includes sound information, turning on the sound collection device according to the user's second authorization information, where the second authorization information is used to instruct the turning on of the sound collection device;

[0082] If the reaction status information includes image information and sound information, the image acquisition device is activated according to the first authorization information and the sound acquisition device is activated according to the second authorization information.

[0083] In a possible implementation, the processing module is specifically configured to:

[0084] If the reaction state information includes image information, obtaining user evaluation information of the media content based on the user's facial expression information;

[0085] If the reaction status information includes voice information, obtaining the user's evaluation information of the media content according to the user's voice emotion information;

[0086] If the reaction state information includes image information and sound information, obtaining the user's evaluation information of the media content based on the user's facial expression information and sound emotion information;

[0087] The facial expression information is the facial emotion information of the user when watching the media content obtained based on the image information, and the voice emotion information is the voice emotion information of the user when watching the preset content obtained based on the voice information.

[0088] In a possible implementation, the processing module is specifically configured to:

[0089] If the user's facial expression information is obtained within a preset time period, then obtaining standard facial expression information corresponding to the preset time period, where the standard facial expression information is expression information predefined according to preset content;

[0090] If the user's facial expression information is consistent with the standard facial expression information, then the user's evaluation of the media content is determined to be a plus-point evaluation;

[0091] If the user's facial expression information is inconsistent with the standard facial expression information, it is determined that the user's evaluation of the media content is a deductible evaluation.

[0092] In a possible implementation, the processing module is specifically configured to:

[0093] If the user's facial expression information is obtained in other time periods, obtaining an evaluation mapping table, wherein the evaluation mapping table is used to indicate evaluation information corresponding to different facial expression information;

[0094] The user's evaluation information on the media content is obtained based on the user's facial expression information and the evaluation mapping table.

[0095] In a possible implementation, the processing module is specifically configured to:

[0096] Obtaining standard sound emotion information corresponding to a preset time period, where the standard sound emotion information is sound information predefined according to preset content;

[0097] If the user's voice emotion information is consistent with the standard voice emotion information, the user's evaluation of the media content is determined to be a plus-point evaluation;

[0098] If the user's voice emotion information is inconsistent with the standard voice emotion information, the user's evaluation of the media content is determined to be a deductible evaluation.

[0099] In a possible implementation, after obtaining the user's reaction status information, the processing module is further configured to:

[0100] Obtaining an identity identifier of at least one user based on the image information or the sound information;

[0101] After obtaining the user's evaluation information on the media content according to the response status information, the evaluation information on the media content by each user is associated with the user's identity.

[0102] In a possible implementation, the processing module is specifically configured to:

[0103] obtaining an identity identifier of at least one user based on the image information;

[0104] Obtaining target facial expression information that matches the voice emotion information, and obtaining a target identity corresponding to the target facial expression information from at least one user identity, wherein the target facial expression information and the emotion corresponding to the voice emotion information are consistent;

[0105] According to the sound emotion information and the standard sound emotion information corresponding to the preset time period, the evaluation information of the user corresponding to the target identity on the media content is obtained, and the standard sound emotion information is information predefined according to the preset content.

[0106] In a possible implementation, before the media content is played, the processing module is further configured to:

[0107] Obtaining a correspondence between the user's identity identifier and the user's features based on the user's identity identifier and user features, where the user features include one of a face or a voice;

[0108] Obtaining at least one user's identity identifier based on a correspondence between the user's identity identifier and the face and the face included in the image information; or

[0109] At least one user's identity identifier is obtained according to the correspondence between the user's identity identifier and the sound and the sound included in the sound information.

[0110] In a possible implementation, the processing module is further configured to:

[0111] After obtaining the user's evaluation information on the media content according to the response status information, if the evaluation information is plus-point information, updating the user's evaluation information on the media content to obtain updated evaluation information;

[0112] If the evaluation information is deduction information, it is determined whether to update the user's evaluation information on the media content according to the status information obtained after the status information.

[0113] In a possible implementation, the processing module is further configured to:

[0114] Before playing the media content, the user's identity is identified. If the user's historical evaluation information is obtained based on the user's identity, the media content recommended to the user is determined based on the historical evaluation information.

[0115] In a possible implementation, the invention further includes: an output module;

[0116] The processing module is further configured to: after the media content is played, determine an evaluation list based on the user's evaluation information, wherein the evaluation list includes the media content evaluated by the user;

[0117] The output module is used to: send the evaluation list to the media source platform;

[0118] The input module is further used to obtain a recommendation list returned by the media source platform, where the recommendation list includes media content to be recommended to the user.

[0119] In a possible implementation, the processing module is specifically configured to:

[0120] If a user to be recommended is identified, determining media content whose evaluation information associated with the user to be recommended satisfies a first preset condition; wherein the first preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking;

[0121] An evaluation list is determined based on the media content that meets the first preset condition.

[0122] In a possible implementation, the processing module is specifically configured to:

[0123] If at least two users to be recommended are identified, determining media content whose evaluation information associated with each of the at least two users to be recommended satisfies a second preset condition, where the second preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking;

[0124] An evaluation list is determined based on the media contents that all meet the second preset condition.

[0125] In a third aspect, an embodiment of the present application provides a media content recommendation device, comprising: a processor and a memory;

[0126] The processor is configured to call the computer program stored in the memory to perform the following operations:

[0127] Acquiring user reaction status information when the media content is played, the reaction status information including at least one of the following types of information: image information of the user acquired by an image acquisition device or voice information of the user acquired by a voice acquisition device;

[0128] According to the reaction status information, the user's evaluation information on the media content is obtained, wherein the evaluation information serves as a basis for recommending other media content to the user.

[0129] In a possible implementation, the memory is further configured to store image information of the user acquired by the image acquisition device and / or sound information acquired by the sound acquisition device.

[0130] In one possible implementation, the image information includes one or more images captured by the image capture device at irregular intervals or continuously or based on a first image capture frequency within a preset time period, where the preset time period is a time period for playing preset content in the media content.

[0131] In a possible implementation, the image information also includes one or more images captured by the image capture device in other time periods outside the preset time period based on the second image capture frequency; wherein the first image capture frequency is higher than the second image capture frequency.

[0132] In one possible implementation, the sound information includes one or more sound segments collected by the sound collection device within a preset time period, where the preset time period is the time period for playing preset content in the media content, and the frequency of collecting sound within the preset time period includes any one of the following: continuous collection, collection at a preset frequency, or collection at irregular intervals.

[0133] In a possible implementation, the processor is further configured to:

[0134] Before obtaining the user's reaction status information, if the reaction status information includes image information, turning on the image acquisition device according to the user's first authorization information, where the first authorization information is used to instruct turning on the image acquisition device;

[0135] If the response status information includes sound information, turning on the sound collection device according to the user's second authorization information, where the second authorization information is used to instruct the turning on of the sound collection device;

[0136] If the reaction status information includes image information and sound information, the image acquisition device is activated according to the first authorization information and the sound acquisition device is activated according to the second authorization information.

[0137] In a possible implementation, the processor is specifically configured to:

[0138] If the reaction state information includes image information, obtaining user evaluation information of the media content based on the user's facial expression information;

[0139] If the reaction status information includes voice information, obtaining the user's evaluation information of the media content according to the user's voice emotion information;

[0140] If the reaction state information includes image information and sound information, obtaining the user's evaluation information of the media content based on the user's facial expression information and sound emotion information;

[0141] The facial expression information is the facial emotion information of the user when watching the media content obtained based on the image information, and the voice emotion information is the voice emotion information of the user when watching the preset content obtained based on the voice information.

[0142] In a possible implementation, the processor is specifically configured to:

[0143] If the user's facial expression information is obtained within a preset time period, then obtaining standard facial expression information corresponding to the preset time period, where the standard facial expression information is expression information predefined according to preset content;

[0144] If the user's facial expression information is consistent with the standard facial expression information, then the user's evaluation of the media content is determined to be a plus-point evaluation;

[0145] If the user's facial expression information is inconsistent with the standard facial expression information, it is determined that the user's evaluation of the media content is a deductible evaluation.

[0146] In a possible implementation, the processor is specifically configured to:

[0147] If the user's facial expression information is obtained in other time periods, obtaining an evaluation mapping table, wherein the evaluation mapping table is used to indicate evaluation information corresponding to different facial expression information;

[0148] The user's evaluation information on the media content is obtained based on the user's facial expression information and the evaluation mapping table.

[0149] In a possible implementation, the processor is specifically configured to:

[0150] Obtaining standard sound emotion information corresponding to a preset time period, where the standard sound emotion information is sound information predefined according to preset content;

[0151] If the user's voice emotion information is consistent with the standard voice emotion information, the user's evaluation of the media content is determined to be a plus-point evaluation;

[0152] If the user's voice emotion information is inconsistent with the standard voice emotion information, the user's evaluation of the media content is determined to be a deductible evaluation.

[0153] In a possible implementation, the processor is further configured to:

[0154] After obtaining the user's reaction status information, obtaining at least one user's identity based on the image information or the sound information;

[0155] After obtaining the user's evaluation information of the media content according to the response status information, the following is also included:

[0156] The evaluation information of each user on the media content is associated with the user's identity.

[0157] In a possible implementation, the processor is specifically configured to:

[0158] obtaining an identity identifier of at least one user based on the image information;

[0159] Obtaining target facial expression information that matches the voice emotion information, and obtaining a target identity corresponding to the target facial expression information from at least one user identity, wherein the target facial expression information and the emotion corresponding to the voice emotion information are consistent;

[0160] According to the sound emotion information and the standard sound emotion information corresponding to the preset time period, the evaluation information of the user corresponding to the target identity on the media content is obtained, and the standard sound emotion information is information predefined according to the preset content.

[0161] In a possible implementation, the processor is further configured to:

[0162] Before playing the media content, obtaining a correspondence between the user's identity identifier and the user's characteristics based on the identity identifier and user characteristics input by the user, where the user characteristics include one of a face or a voice;

[0163] Obtaining at least one user's identity based on the image information or the sound information includes:

[0164] Obtaining at least one user's identity identifier based on a correspondence between the user's identity identifier and the face and the face included in the image information; or

[0165] At least one user's identity identifier is obtained according to the correspondence between the user's identity identifier and the sound and the sound included in the sound information.

[0166] In a possible implementation, the processor is further configured to:

[0167] After obtaining the user's evaluation information on the media content according to the response status information, if the evaluation information is plus-point information, updating the user's evaluation information on the media content to obtain updated evaluation information;

[0168] If the evaluation information is deduction information, it is determined whether to update the user's evaluation information on the media content according to the status information obtained after the status information.

[0169] In a possible implementation, the processor is further configured to:

[0170] Before playing the media content, the user's identity is identified. If the user's historical evaluation information is obtained based on the user's identity, the media content recommended to the user is determined based on the historical evaluation information.

[0171] In a possible implementation, the device further includes: a communication module;

[0172] The processor is further configured to: after the media content is played, determine an evaluation list based on the user's evaluation information, the evaluation list including the media content evaluated by the user;

[0173] The communication module is used to send an evaluation list to a media source platform and obtain a recommendation list returned by the media source platform, where the recommendation list includes media content to be recommended to the user.

[0174] In a possible implementation, the processor is specifically configured to:

[0175] If a user to be recommended is identified, determining media content whose evaluation information associated with the user to be recommended satisfies a first preset condition; wherein the first preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking;

[0176] An evaluation list is determined based on the media content that meets the first preset condition.

[0177] In a possible implementation, the processor is specifically configured to:

[0178] If at least two users to be recommended are identified, determining media content whose evaluation information associated with each of the at least two users to be recommended satisfies a second preset condition, where the second preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking;

[0179] An evaluation list is determined based on the media contents that all meet the second preset condition.

[0180] In a fourth aspect, an embodiment of the present application provides a terminal device, comprising: a media content recommendation device, a camera and / or a microphone.

[0181] The media content recommendation device is configured to execute the method in the first aspect and any one of various possible implementations of the first aspect.

[0182] In a fifth aspect, an embodiment of the present application provides a storage medium, which is used to store a computer program. When the computer program is executed by a computer or a processor, it is used to implement the authentication method described in any one of the first aspects.

[0183] The media content recommendation method and device provided by the embodiment of the present application include: obtaining the corresponding relationship between the user's identity identifier and user characteristics based on the identity identifier and user characteristics input by the user, and the user characteristics include one of face or voice. When the media content is played, the user's reaction status information is obtained, and the reaction status information includes at least one of the following types of information: the user's image information obtained by an image acquisition device or the user's voice information obtained by a voice acquisition device. According to the reaction status information, the user's evaluation information on the media content is obtained. Each user's evaluation information on the media content is associated with the user's identity identifier. By obtaining the corresponding relationship between the user's identity identifier and the user characteristics, each user's evaluation information on the media content is associated with the corresponding identity identifier, so that personalized media content recommendations can be made for users with different identity identifiers in the future, so as to improve the accuracy of media content recommendations. The user's evaluation information on the media content is obtained through the user's facial expression information or the user's voice emotion information, and the evaluation information can be obtained in real time based on the user's feedback on the media content, thereby ensuring the authenticity and accuracy of the evaluation information. BRIEF DESCRIPTION OF THE DRAWINGS

[0184] Figure 1A Schematic diagram 1 of a media content recommendation system provided in one embodiment of the present application;

[0185] Figure 1B A media content recommendation system according to an embodiment of the present invention is shown in FIG. Figure 2 ;

[0186] Figure 2 Flowchart 1 of a media content recommendation method provided in one embodiment of the present application;

[0187] Figure 3 The process of the media content recommendation method provided in one embodiment of the present application Figure 2 ;

[0188] Figure 4 The process of the media content recommendation method provided in one embodiment of the present application Figure 3 ;

[0189] Figure 5 The process of the media content recommendation method provided in one embodiment of the present application Figure 4 ;

[0190] Figure 6 The process of the media content recommendation method provided in one embodiment of the present application Figure 5 ;

[0191] Figure 7 The process of the media content recommendation method provided in one embodiment of the present application Figure 6 ;

[0192] Figure 8 Flowchart 1 of a method for recommending media content according to an embodiment of the present application;

[0193] Figure 9 A flowchart of a media content recommendation method provided in an embodiment of the present application Figure 2 ;

[0194] Figure 10 Figure 1 is a signaling flow chart of a media content recommendation method provided in one embodiment of the present application;

[0195] Figure 11 The signaling process of the media content recommendation method provided in one embodiment of the present application Figure 2 ;

[0196] Figure 12 Schematic diagram 1 of the structure of a media content recommendation device provided in one embodiment of the present application;

[0197] Figure 13 A schematic diagram of the structure of a media content recommendation device provided in one embodiment of the present application Figure 2 ;

[0198] Figure 14 A schematic diagram of the hardware structure of a media content recommendation device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0199] The system architecture and business scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that, with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are equally applicable to similar technical problems.

[0200] Figure 1A The media content recommendation system provided in one embodiment of the present application is shown in FIG1. Figure 1A As shown, the recommendation system includes: a multimedia device 101 and a media source platform 102 , wherein the multimedia device 101 includes an image acquisition device 1011 and a sound acquisition device 1012 .

[0201] Specifically, the multimedia device 101 may include but is not limited to a digital television (DTV), a mobile device, a laptop computer, an external advertising device, a tablet device, a personal digital assistant (PDA), a smart terminal, a handheld device with wireless connection function, or a vehicle-mounted device and other portable devices.

[0202] In this embodiment, an image acquisition device 1011 is provided on the multimedia device 101, wherein the image acquisition device 1011 can be, for example, a camera, or a network camera device, etc., wherein the network camera device can include, for example: a lens, an image sensor, a microprocessor, an image processor, a memory, etc. The specific implementation method of the image acquisition device 1011 is not limited here.

[0203] Secondly, the multimedia device is also provided with a sound collection device 1021. Sound collection device 1021 may include, but is not limited to, a far-field microphone, a digital broadcast terminal, or a personal digital assistant, and primarily functions to collect sound. Sound collection device 1021 is equipped with a microphone capable of collecting sound signals from the surrounding environment. This embodiment does not specifically limit the use of sound collection device 1021.

[0204] The media source platform 102 is a platform for providing media content to the multimedia device 101. The media source platform can be, for example, a platform provided by different operators or different video providers, or a platform storing local videos, etc., without particular limitation. These media source platforms can provide multimedia files such as audio and video.

[0205] The multimedia device 101 and the media source platform 102 interact with each other, where the interaction can be carried out, for example, through a wired network, such as a coaxial cable, a twisted pair, or an optical fiber, or through a wireless network, such as a 2G network, a 3G network, a 4G network, a 5G network, a Wireless Fidelity (WIFI) network, etc. The specific type or form of interaction is not limited herein, as long as the function of interaction between the multimedia device 101 and the media source platform 102 is achieved.

[0206] In another possible implementation, the image acquisition device 1011 and the sound acquisition device 1021 described above can also be independent external devices. Figure 1B Make an introduction, Figure 1B A media content recommendation system according to an embodiment of the present invention is shown in FIG. Figure 2 ,like Figure 1B As shown, the recommendation system includes: a multimedia device 101 , a media source platform 102 , an image acquisition device 103 and a sound acquisition device 104 .

[0207] Among them, the image acquisition device 103 and the sound acquisition device 104 can be two independent devices, or they can be coupled into one device and externally connected to the multimedia device 101. The specific implementation method of the external connection can be, for example, a wired connection, such as connection through cables such as coaxial cables, twisted pairs and optical fibers, or a wireless connection, such as connection through Bluetooth, wireless network, etc., which is not limited in this embodiment.

[0208] Based on the problem that the existing technology recommends programs based on historical playback records, which leads to the problem that media content cannot be accurately recommended to users, the embodiment of the present application provides a media content recommendation method. Figure 2 The media content recommendation method provided in the embodiment of the present application is introduced in detail.

[0209] Figure 2 Flowchart 1 of the media content recommendation method provided in one embodiment of the present application. The execution subject of this embodiment can be, for example, the multimedia device in the above recommendation system. Figure 2 As shown, the method includes:

[0210] S201. When media content is played, obtain user reaction status information, where the reaction status information includes at least one of the following types of information: user image information obtained by an image acquisition device or user voice information obtained by a voice acquisition device.

[0211] Specifically, when a user plays media content through a multimedia device, the user's reaction status information can be obtained. The media content can be audio or video, and this embodiment does not impose any particular restrictions on the implementation of multimedia content.

[0212] While media content is playing, the multimedia device can control the operation of at least one of the image capture device and the sound capture device. For example, the image capture device can be controlled to capture the user's image information. Alternatively, the sound capture device can be controlled to capture the user's sound information. Alternatively, the image capture device and the sound capture device can be controlled to operate simultaneously to obtain the user's image information and sound information.

[0213] In this embodiment, the image information may be, for example, picture information in frames, or a segment of video information, etc. When the user is near the multimedia device, the image information includes the user's image, and the sound information includes the user's sound.

[0214] In one possible implementation, for example, the user's reaction status information can be obtained according to a preset period, or the user's reaction status information can be obtained in real time, that is, the user's reaction status information is obtained when a change in the user's action is detected or the user's voice information is monitored. Those skilled in the art can understand that the specific method of obtaining the user's reaction status information can be set according to actual needs, and this embodiment does not limit this.

[0215] S202: Obtain user evaluation information of the media content according to the response status information, wherein the evaluation information serves as a basis for recommending other media content to the user.

[0216] In this embodiment, the user's evaluation information of the media content may be, for example, an evaluation score, such as a numerical value between 1 and 100, indicating different levels of user satisfaction. Alternatively, the user's evaluation information of the media content may be, for example, a preset degree index, such as "not interested," "interested," or "very interested," as the user's evaluation information. The specific evaluation information may be selected based on actual needs and is not particularly limited herein.

[0217] The reaction status information includes image information or sound information, and the user's evaluation information on the media content can be obtained based on the image information or sound information.

[0218] Regarding image information, user feedback on the currently playing media content can be obtained based on the user's image information. For example, the image information can be used to determine whether the user's face is facing the multimedia device, or whether the user is sleeping. Based on the user's feedback on the media content, the user's interest in the media content can be determined, and the media content can be rated positively or negatively. For example, if the user is sleeping, the media content will be rated negatively, while if the user is paying attention to the media content, the media content will be rated positively.

[0219] Regarding audio information, user feedback on the currently playing media content can be obtained based on the user's audio information. For example, audio information can be used to determine whether the user is still watching the program, or whether the user has a corresponding audio reaction to the currently playing media content, such as whether the user laughs when the media content reaches a funny point, or whether the user cries when the media content reaches a tearful point. For example, if the user laughs at a funny point, the media content will be rated positively; if the user does not laugh at a funny point, the media content will be rated negatively.

[0220] In a specific implementation, the evaluation information may be obtained based solely on image information or solely on sound information. Alternatively, the evaluation information may be obtained by combining image and sound information. When combining image and sound information, the evaluations of the two may be superimposed, or a weighted approach may be used to obtain a comprehensive evaluation. This embodiment does not impose any particular limitation on the specific implementation method.

[0221] In an embodiment of the present application, the evaluation information is used as a basis for recommending other media content to the user. For example, highly rated media content is obtained based on the evaluation information, and related or similar programs of the highly rated media content are recommended to the user.

[0222] The media content recommendation method provided by an embodiment of the present application includes: obtaining user reaction status information when media content is played, the reaction status information including at least one of the following types of information: user image information obtained by an image capture device or user sound information obtained by a sound capture device; and obtaining user evaluation information of the media content based on the reaction status information, wherein the evaluation information serves as a basis for recommending other media content to the user. The user's image information or sound information can be used to accurately determine whether the user is interested in the currently playing media content. Secondly, the obtained evaluation information is used as a basis for recommending other media content to the user, so that the other media content recommended to the user is content of interest to the user, thereby improving the accuracy of the recommended media content.

[0223] Based on the above embodiments, the media content recommendation method provided by this application is further described in detail below in conjunction with specific embodiments. Figure 3 To explain, Figure 3 The process of the media content recommendation method provided in one embodiment of the present application Figure 2 ,like Figure 3 As shown, the method includes:

[0224] S301. Obtain a correspondence between the user's identity identifier and the user's features according to the identity identifier and user features input by the user, where the user features include one of a face or a voice.

[0225] Among them, the user's identity is used to indicate different user identities. In order to ensure that the evaluation information can be associated with the user's identity when obtaining the user's evaluation information, so as to make personalized content recommendations for each user later, the user's identity can be pre-stored.

[0226] Specifically, the user-entered identity identifier can be, for example, the user's name, user account number, or nickname. Those skilled in the art will appreciate that any identity identifier is sufficient as long as it can distinguish different users. This embodiment does not impose any particular limitation on the specific configuration of the identity identifier. When recommending programs to a user, using the user-entered identity identifier can make the user feel close and enhance the user experience.

[0227] The user feature includes one of a face or a voice. For example, a user inputs the identity of user 1 and records facial information, thereby obtaining the user feature of user 1's face. For another example, a user inputs the identity of Zhang San and records voice information, thereby obtaining the user feature of Zhang San's voice. The user feature can also include both a face and a voice.

[0228] Those skilled in the art will appreciate that, in another possible implementation, to simplify user operations, the user may not enter an identity identifier and user characteristics, but rather the multimedia device may obtain the user's identity identifier based on the user's image information and / or voice information. For example, with respect to image information, a face may be captured using technologies such as facial recognition, and then the identity identifier "User A" may be generated based on the face, and the face may be associated with "User A."

[0229] For voice information, voice recognition can be used to obtain the voice, and then the voice can be associated with "User B". Those skilled in the art will understand that when there is a user in the image information, a user identifier can be generated based on the face and voice, and the face and voice can be associated with "User C".

[0230] S302: When the media content is played, obtain user reaction status information, where the reaction status information includes at least one of the following types of information: user image information obtained by an image acquisition device or user voice information obtained by a voice acquisition device.

[0231] In the embodiment of the present application, different preset time periods are set for different media content, wherein the preset time period is the time period for playing preset content in the media program. The preset content in the media content may be, for example, funny content in the media content, tearful content in the media content, or other preset memes. Those skilled in the art will understand that the preset content in the media content may be content that can produce program effects, and its specific selection can be set according to the actual content of the media content, and this is not limited here.

[0232] The preset time period is a time period for playing a preset content in the media content, corresponding to the playback duration of the preset content. The media content may include at least one preset time period. In a specific implementation, the attribute information of the media content will include the time period of the preset content. The multimedia device can determine the preset time period based on the attribute information. The full duration of the media program may correspond to at least one preset time period, as well as other time periods in addition to the preset time periods.

[0233] In an embodiment of the present application, in order to improve processing efficiency and save storage space of multimedia devices, the collection frequency in the preset time period is different from that in other time periods. Secondly, the collection of image information and sound information may not be real-time, but may be collected according to a certain collection frequency.

[0234] In one optional implementation, the image information includes one or more images captured by the image capture device at irregular intervals within a preset time period, where the irregular intervals may be, for example, randomly generated time intervals or preset irregular intervals, and are not limited to this herein. Alternatively, the image information includes one or more images captured continuously and uninterruptedly by the image capture device during a preset time period. Specifically, the image capture device continuously and uninterruptedly captures the user's image information as long as it is in the on state. Alternatively, the image information includes one or more images captured by the image capture device during a preset time period based on a first image capture frequency, where the first image capture frequency is a frequency selected based on actual needs, and is not limited to this embodiment.

[0235] In another optional implementation, the image information also includes one or more images captured by the image acquisition device based on a second image acquisition frequency in time periods other than the preset time period, where the second image acquisition frequency is a frequency selected according to actual needs, and this implementation does not limit this.

[0236] Those skilled in the art can understand that in order to ensure that the user's response to the program effect corresponding to the preset content can be accurately obtained within the preset time period, the first image acquisition frequency is set to be higher than the second image acquisition frequency. For example, the currently playing comedy short film has a funny content from 2 minutes 15 seconds to 2 minutes 45 seconds, and a funny content from 3 minutes 35 seconds to 4 minutes 5 seconds. In this case, there are two preset time periods. The image acquisition frequency is set to once every 1 second in these two preset time periods, and the image acquisition frequency can be set to once every 5 seconds in other time periods.

[0237] Alternatively, the first image acquisition frequency may be set to be less than the second image acquisition frequency. Both settings may be made according to actual needs and are not limited here. As long as the first image acquisition frequency is not equal to the second image acquisition frequency, it will be sufficient.

[0238] In an optional embodiment, the sound information includes one or more sound segments collected by the sound collection device within a preset time period, wherein the frequency of collecting sound within the preset time period includes any one of the following: continuous collection, collection at a preset frequency, or collection at irregular intervals. The specific implementation method is similar to the collection of image information and will not be repeated here.

[0239] In an embodiment of the present application, in order to avoid infringing user privacy, the user's reaction status information can be obtained after obtaining the user's authorization to open the image acquisition device and / or sound acquisition device.

[0240] In an optional implementation, if the reflected status information includes only image information, the image acquisition device is turned on according to the first authorization information of the user, wherein the first authorization information is used to instruct the turning on of the image acquisition device;

[0241] Specifically, first authorization information of the user is obtained, where the first authorization information is used to instruct the image acquisition device to be turned on. Then, the image acquisition device is turned on according to the first authorization information.

[0242] For example, when a user turns on a multimedia device or before playing media content, a prompt message indicating that the image capture device needs to be turned on can be displayed to the user through the multimedia device. Then, a user operation input by the user is received to obtain the user's first authorization information, and then the permission to turn on the image capture device is obtained based on the first authorization information, thereby turning on the image capture device.

[0243] In another optional implementation, if the reaction status information includes sound information, the sound acquisition device is enabled based on the user's second authorization information, where the second authorization information indicates the activation of the sound acquisition device. Alternatively, if the reaction status information includes both image information and sound information, the image acquisition device is enabled based on the first authorization information, and the sound acquisition device is enabled based on the second authorization information, in a manner similar to the above-described activation of the image acquisition device.

[0244] For example, the permission to enable the image acquisition device is obtained based on the first authorization information, and the permission to enable the sound acquisition device is obtained based on the second authorization information, thereby enabling the image acquisition device and enabling the sound acquisition device N seconds before a preset time period, where N is an integer greater than or equal to 0.

[0245] By turning on the image acquisition device and / or the sound acquisition device according to the user's authorization information, it is possible to ensure that the user's reaction status information is obtained with the user's authorization, avoiding infringement of the user's privacy and improving the user experience.

[0246] S303: Identify at least one user's identity based on the image information or the sound information.

[0247] In this embodiment, the correspondence between the user's identity identifier and the user's features is obtained in advance. When identifying the user's identity identifier based on image information, at least one user's identity identifier is obtained based on the correspondence between the user's identity identifier and the face and the face included in the image information.

[0248] Specifically, the image information includes the face of at least one user, and face recognition is performed on at least one user in the image information. The specific implementation method of face recognition can refer to the existing technology and will not be repeated here. Secondly, the face is obtained based on face recognition, and the identity identifier of at least one user is obtained based on the correspondence between the face included in the image information and the user's identity identifier.

[0249] When identifying the user's identity based on the sound information, at least one user's identity is obtained based on the correspondence between the user's identity and the sound and the sound included in the sound information.

[0250] When identifying the identity of at least one user based on sound information, sound processing and feature analysis are performed on the sound contained in the sound information, such as frequency band analysis, timbre analysis, etc., so as to obtain the sound included in the sound information. Secondly, the identity of at least one user is obtained based on the correspondence between the sound and the user's identity.

[0251] If the correspondence between the user's identity identifier and user characteristics is not obtained in advance, that is, the user has not entered the corresponding settings, and the current user's face or voice appears for the first time, the identity identifier can be obtained from the face included in the image information or the voice included in the voice information, and the user's identity identifier can be set to user A, user B, user C, etc., so that users can be distinguished. Secondly, the correspondence between the user's identity identifier and user characteristics is established based on the face in the image information or the voice in the voice information.

[0252] In this embodiment, the operation of identifying the user's identity identifier only needs to be performed once, and the evaluation information corresponding to the user's identity identifier can be updated directly based on the user's identity identifier subsequently, without having to obtain the identity identifier again, thereby simplifying the operation.

[0253] S304: Obtain user evaluation information on the media content according to the response status information.

[0254] If the reaction status information includes image information, the user's evaluation information on the media content is obtained according to the user's facial expression information.

[0255] Specifically, the image information of at least one user in the image information is analyzed and processed to obtain the user's facial expression information, wherein the facial expression information is the facial emotional information of the user when watching media content obtained based on the image information. The facial expression information can be, for example, the user's overall facial expression, or it can be, for example, the state of part of the user's facial muscles, such as the state of the corners of the mouth, the state of the eyes, etc. There is no restriction on the facial expression information here.

[0256] The facial expression information of the user when watching the media content can reflect the user's feedback on the effect of the program when watching the media content. Therefore, by obtaining the facial expression information, the user's satisfaction with the currently played media content can be accurately obtained.

[0257] If the reaction status information includes voice information, the user's evaluation information on the media content is obtained according to the user's voice emotion information.

[0258] Specifically, the sound information of at least one user in the sound information is analyzed and processed to obtain the user's sound emotion information, wherein the sound emotion information may include, for example, the emotional state of the sound, or the decibel of the sound, etc., which is not limited here.

[0259] Among them, the voice emotion information of the user when watching media content can also reflect the user's feedback on the effect of the program when watching the media content. The implementation method is similar to that obtained based on facial expression information, which will not be repeated here.

[0260] If the reaction status information includes image information and sound information, the user's evaluation information on the media content is obtained according to the user's facial expression information and sound emotion information.

[0261] When combining facial expression information and voice emotion information to obtain evaluation information, the evaluations of the two can be superimposed, or a weighted approach can be used to obtain a comprehensive evaluation. This embodiment does not impose any special restrictions on the specific implementation method, and all three methods can achieve the acquisition of evaluation information.

[0262] S305: Associating each user's evaluation information on the media content with the user's identity.

[0263] Furthermore, in the above steps, the user's identity is identified. After determining at least one user's evaluation information on the media content, the evaluation information of each user on the media content can be associated with the user's identity, so that when media content is subsequently recommended to the user, the media content can be recommended based on the currently associated information.

[0264] The media content recommendation method provided by the embodiment of the present application includes: obtaining the correspondence between the user's identity identifier and user features based on the identity identifier and user features input by the user, and the user features include one of face or voice. When the media content is played, the user's reaction status information is obtained, and the reaction status information includes at least one of the following types of information: the user's image information obtained by an image acquisition device or the user's voice information obtained by a voice acquisition device. According to the reaction status information, the user's evaluation information on the media content is obtained. Each user's evaluation information on the media content is associated with the user's identity identifier. By obtaining the correspondence between the user's identity identifier and the user features, each user's evaluation information on the media content is associated with their respective corresponding identity identifiers, so that personalized media content recommendations can be made for users with different identity identifiers in the future, so as to improve the accuracy of media content recommendations, wherein the user's evaluation information on the media content is obtained through the user's facial expression information or the user's voice emotion information, and the evaluation information can be obtained in real time based on the user's feedback on the media content, thereby ensuring the authenticity and accuracy of the evaluation information.

[0265] On the basis of the above embodiments, the media content recommendation method provided by the present application can obtain the user's evaluation information on the media content based on the user's facial expression information or voice emotion information alone, and can also obtain the user's evaluation information on the media content based on the user's facial expression information and voice emotion information together. Figure 4 This paper introduces the implementation method of obtaining user evaluation information of media content based on user facial expression information.

[0266] Figure 4 The process of the media content recommendation method provided in one embodiment of the present application Figure 3 ,like Figure 4 As shown, the method includes:

[0267] S401. Determine whether the user's facial expression information is obtained within a preset time period. If so, execute S402; if not, execute S406.

[0268] In this embodiment, the media content includes a preset time period and other time periods except the preset time period, wherein the image acquisition frequencies in the preset time period and other time periods are different, so it is first determined whether the acquisition time period corresponding to the user's facial expression information is the preset time period.

[0269] Specifically, when obtaining the user's facial expression information based on image information, the obtained time node is associated with each facial expression information, and then the obtained time node is compared with the starting time point corresponding to the preset time period to determine whether the user's facial expression is obtained within the preset time period.

[0270] S402: Obtain standard facial expression information corresponding to a preset time period, where the standard facial expression information is expression information predefined according to preset content.

[0271] If the user's facial expression information is obtained within a preset time period, it indicates that the media content at this time has a corresponding program effect, so it is necessary to detect whether the user has feedback information on the corresponding media content within the preset time period.

[0272] Specifically, the standard facial expression information corresponding to the preset time period is the expression information pre-set according to the program content of the preset time period. For example, when the program content corresponding to the preset time period is a funny content, the corresponding standard facial expression information is "laughing". For example, when the program content corresponding to the third time period is a tearful content, the corresponding standard facial expression information is "crying".

[0273] Those skilled in the art will appreciate that the standard facial expression information corresponding to each preset time period is specifically set according to the program content, and this embodiment does not impose any limitation on this.

[0274] S403: Determine whether the user's facial expression information is consistent with standard facial expression information. If so, execute S404; if not, execute S405.

[0275] Furthermore, it is determined whether the user's facial expression information conforms to the standard facial expression information. For example, when the standard facial expression information is "smile", then smiling, laughing, and covering the mouth to laugh can all be considered to be consistent with the standard facial expression information. For example, when the standard facial expression information is "cry", then shedding tears, wiping the corners of the eyes, and turning the corners of the mouth downwards can all be determined to be consistent with the standard facial expression information.

[0276] Among them, the specific implementation method of the judgment can be, for example, to extract and analyze features based on the user's facial expression information, and then make a judgment based on the results of the feature analysis. For example, shape points can be extracted from the face contained in the user's image information, and the extracted shape points can be compared with the shape points of the preset standard facial expression information to make a judgment. This embodiment does not impose any special restrictions on the specific implementation method of the judgment.

[0277] S404: Determine that the user's evaluation of the media content is a plus-point evaluation.

[0278] If the user's facial expression information is consistent with the standard facial expression information, it can be determined that the user's feedback information on the current program content is positive feedback, thereby determining that the user's evaluation of the media content is a plus-point evaluation. If the image information is not the first image, the user's evaluation information on the media content can also be updated to obtain updated evaluation information.

[0279] Specifically, for example, when the evaluation information is a score, the score corresponding to the bonus evaluation can be directly added to the score of the program. For example, the evaluation information of the media content before the update is 88 points. When the media content plays to a funny point, the user's facial expression information obtained is a smile, and the score corresponding to the smile is, for example, 2 points. Then the updated evaluation information is 90 points.

[0280] Among them, the evaluation information can also be, for example, degree indicators such as interest and disinterest. Continuing to use the data of the above example, for example, the weight value of the degree indicator of interest can be increased, and then the updated evaluation information can be obtained based on the updated weight value of each degree indicator.

[0281] S405: Determine that the user's evaluation of the media content is a deductible evaluation.

[0282] If the user's facial expression information is inconsistent with the standard facial expression information, it can be determined that the user's feedback information on the current program content is negative feedback, and thus the user's evaluation of the media content is determined to be a deductible evaluation.

[0283] S406: Determine whether to update the user's evaluation information on the media content based on the status information obtained after the status information is obtained.

[0284] Specifically, when the evaluation of media content is a deduction, in order to avoid false detection or incorrect time nodes, for example, the user sheds tears when watching tearful content, but the user's facial expression information at the current time node is not detected to be crying, or when watching some funny content, because the user reacts slowly, he laughs a few seconds after the preset time period corresponding to the funny content, so for the deduction evaluation, it is necessary to continue to obtain the status information after the status information.

[0285] Based on the status information obtained after the status information, it is comprehensively determined whether the user's evaluation information needs to be updated. For example, if the current evaluation of the media content is a negative evaluation, it can be determined whether to update the evaluation information based on the user's facial expression information corresponding to a preset number of image information after the current user's facial expression information.

[0286] For example, assuming that among the 10 image information after the current image information, 6 user facial expression information corresponds to deduction evaluation, and 4 user facial expression information corresponds to addition evaluation, then it is determined to deduct points according to the deduction information corresponding to the current deduction evaluation.

[0287] For another example, the evaluation information can also be updated according to the weight values ​​of the plus and minus information of the 10 image information after the current user's facial expression information. Those skilled in the art can understand that when the user's evaluation of the media content is a minus evaluation, the method of updating the user's evaluation information of the media content can be selected according to needs. By determining whether to update the evaluation information based on the status information obtained after the status information, the accuracy of the user's evaluation information can be improved.

[0288] S407: Obtain an evaluation mapping table, wherein the evaluation mapping table is used to indicate evaluation information corresponding to different facial expression information.

[0289] If the user's facial expression information is not obtained within the preset time period, that is, it is obtained in a time period other than the preset time period, it indicates that the preset program effect does not exist at this time, and the evaluation information is directly obtained based on the user's facial expression information and the evaluation mapping table, where the evaluation mapping table is used to indicate the evaluation information corresponding to different facial states.

[0290] For example, the evaluation mapping table stores facial information such as yawning, face not facing the multimedia device, and distracted eyes, where each different facial information corresponds to its own evaluation information. For example, yawning will be deducted by 5 points, face not facing the multimedia device will be deducted by 10 points, or face not facing the multimedia device corresponds to a degree of lack of interest. The specific evaluation mapping table can be set according to actual needs, and there is no restriction on this here.

[0291] S408: Obtain user evaluation information on the media content based on the user's facial expression information and the evaluation mapping table.

[0292] According to the user's facial expression information and the evaluation mapping table, facial information matching the current user's facial expression information is searched in the evaluation mapping table, and then the first user's evaluation information on the media content is obtained according to the evaluation information corresponding to the matching facial information.

[0293] In this embodiment, the evaluation information corresponding to the facial information stored in the evaluation mapping table can also be divided into plus-point evaluation or minus-point evaluation. Secondly, the user's evaluation information on the media content is updated according to the plus-point evaluation or minus-point evaluation. The implementation method is similar to that described above and will not be repeated here.

[0294] The media content recommendation method provided by the embodiment of the present application includes: determining whether the user's facial expression information is obtained within a preset time period, and if so, obtaining standard facial expression information corresponding to the preset time period, wherein the standard facial expression information is expression information predefined according to preset content. Determining whether the user's facial expression information is consistent with the standard facial expression information, and if so, determining that the user's evaluation of the media content is a plus-point evaluation. If not, determining that the user's evaluation of the media content is a minus-point evaluation. Determining whether to update the user's evaluation information of the media content based on the status information obtained after the status information. If the user's facial expression information is not obtained within the preset time period, obtaining an evaluation mapping table, wherein the evaluation mapping table is used to indicate evaluation information corresponding to different facial expression information. Based on the user's facial expression information and the evaluation mapping table, obtain the user's evaluation information of the media content. By determining whether the user's evaluation of media content is a negative or positive evaluation based on the user's facial expression information and standard facial expression information within a preset time period, it is possible to quickly and effectively determine the user's feedback effect on the program. Secondly, when the user's evaluation of media content is a negative evaluation, the status information after the current status information is comprehensively determined to determine whether the user's evaluation information needs to be updated, thereby avoiding evaluation information errors caused by false detection and improving the accuracy of the evaluation information. Secondly, in other time periods, the user's evaluation information is obtained through the user's facial expression information and the evaluation mapping table, so that it is possible to determine in real time based on the user's facial expression whether the user is interested in the currently playing media content, thereby ensuring the authenticity of the user's evaluation information.

[0295] Based on the above embodiment, Figure 5 This paper introduces the implementation method of obtaining user evaluation information on media content based on user voice information.

[0296] Figure 5 The process of the media content recommendation method provided in one embodiment of the present application Figure 4 ,like Figure 5 As shown, the method includes:

[0297] S501. Obtain standard sound emotion information corresponding to a preset time period, where the standard sound emotion information is sound information predefined according to preset content.

[0298] Specifically, the program content within the preset time period has corresponding program effects, wherein the standard sound emotion information corresponding to the preset time period is the sound information pre-set according to the program content of the preset time period, wherein the standard sound emotion information is the sound information pre-defined according to the preset content, for example, it may include the emotional state of the sound, the decibel level of the sound, etc.

[0299] For example, if the program content in the preset time period corresponds to funny content, the emotional state of the corresponding standard sound emotional information is laughter, where the decibel of the sound can be 1 decibel, for example; or if the program content in the preset time period corresponds to tearful content, the emotional state of the corresponding standard sound emotional information is crying.

[0300] Those skilled in the art will appreciate that the standard sound emotion information corresponding to each preset time period is specifically set according to the program content, and this embodiment does not impose any limitation on this.

[0301] S502: Determine whether the user's voice emotion information is consistent with the standard voice emotion information. If so, execute S503; if not, execute S504.

[0302] The user's voice emotion information is compared with the standard voice emotion information. For example, the user's voice information can be feature analyzed to obtain the user's voice emotion information. The user's voice decibels can also be obtained, and then it is determined whether the user's voice emotion information is consistent with the standard voice emotion information.

[0303] For example, if the user's voice emotional information is detected as crying during tearful content, and the decibel level is greater than 1 decibel, it can be determined that the user's voice information is consistent with the standard voice emotional information. For example, if the user's voice emotional information is detected as less than 1 decibel during tearful content, or the emotional state is not crying, it can be determined that the user's voice emotional information is inconsistent with the standard voice emotional information.

[0304] S503: Determine that the user's evaluation of the media content is a plus-point evaluation.

[0305] S504: Determine that the user's evaluation of the media content is a deductible evaluation.

[0306] Among them, the implementation method of S503 and S504 is similar to that of S404 and S405. The specific content can be found in the introduction of the above embodiment and will not be repeated here.

[0307] The media content recommendation method provided by the embodiment of the present application includes: obtaining standard sound emotion information corresponding to a preset time period, where the standard sound emotion information is sound information predefined according to preset content. Determine whether the user's sound emotion information is consistent with the standard sound emotion information. If so, determine that the user's evaluation of the media content is a plus evaluation. If not, determine that the user's evaluation of the media content is a minus evaluation. By comparing the user's sound emotion information with the standard sound emotion information, the user's evaluation information for the currently playing media content is determined, and the user's feedback on the program content of the preset time period can be accurately obtained, thereby improving the authenticity and effectiveness of the user's evaluation information.

[0308] On the basis of the above embodiment, when the sound information and the image information correspond to the identity of the same user, the above Figure 4 Evaluation information for image information and Figure 5 The evaluation information of the sound information is superimposed or weighted to obtain the user's comprehensive score for the media content.

[0309] For example, if the identity identifiers corresponding to the current sound information and image information are both user 1, it means that the image information and sound information of user 1 have been obtained, and the evaluation information of user 1 on the currently played media content can be determined based on the sound information and image information of user 1.

[0310] In one possible implementation, for example, based on the image information of user 1, it is determined that user 1 responded with a smile to the program content with preset funny points, and based on the sound information of user 1, it is determined that user 1 laughed at the program content with preset funny points, then the score corresponding to the smile and the score corresponding to the laughter are added to the evaluation information of user 1 on the media content.

[0311] In another possible implementation, for example, within a preset time period, 20 images of the user's smile are obtained based on the user's image information, and the user's laughter is obtained based on the user's voice information, but the user's laughter is lower than the preset decibels, then the user's evaluation information can be updated based on the bonus weight corresponding to the degree of smile in each picture, and the bonus weight corresponding to the laughter lower than the preset decibels.

[0312] On the basis of the above embodiment, in order to reduce the processing capacity of the multimedia device and simplify the processing process, the sound recognition process can be weakened, that is, there is no need to establish an association between the sound and the user's identity, so as to obtain the user's evaluation information. Figure 6 Provide a detailed introduction.

[0313] Figure 6 The process of the media content recommendation method provided in one embodiment of the present application Figure 5 ,like Figure 6 As shown, the method includes:

[0314] S601: Obtain an identity identifier of at least one user according to image information.

[0315] The implementation of S601 is similar to that of S303 and will not be described in detail here.

[0316] S602. Obtain target facial expression information that matches the voice emotion information, and obtain a target identity identifier corresponding to the target facial expression information from at least one user's identity identifier, where the target facial expression information is consistent with the emotion corresponding to the voice emotion information.

[0317] In this embodiment, there is a matching relationship between the sound emotion information and the facial expression information. For example, if the currently detected sound emotion information is laughter, then the facial expression information of "laughing" is determined as the matching target facial expression information based on the facial expression information of at least one user.

[0318] If there are multiple users whose facial expressions are all "laughing", facial expression information matching the sound emotion information can be obtained based on the correspondence between the sound decibels in the sound emotion information and the degree of laughter. The method of obtaining the target facial expression information matching the sound emotion information can be selected according to actual needs, and is not limited here.

[0319] Among them, different user identities correspond to different user features, and the target facial expression information is processed to obtain the user features corresponding to the target facial expression information, such as a human face. Secondly, based on the user features corresponding to the identity of at least one user, the identity of the user that matches the user features of the target facial expression information is determined, thereby obtaining the target identity corresponding to the target facial expression.

[0320] The target facial expression information and the emotion corresponding to the voice emotion information are consistent. For example, when the target facial expression indicates a smile, the corresponding voice emotion information indicates a laugh, etc. The emotion can also be set to sadness, grief, joy, etc. This embodiment does not limit the specific corresponding emotion.

[0321] S603: Acquire evaluation information of the media content by the user corresponding to the target identity according to the voice emotion information and standard voice emotion information corresponding to the preset time period.

[0322] Determine whether the sound emotion information is consistent with the standard sound emotion information corresponding to the preset time period. If consistent, determine that the user's evaluation of the media content is a plus-point evaluation. If inconsistent, determine that the user's evaluation of the media content is a minus-point evaluation, thereby obtaining the evaluation information of the user corresponding to the target identity identifier on the media content. The specific implementation method is similar to the method of obtaining the user's evaluation information separately based on the sound emotion information introduced in the above embodiment, and will not be repeated here.

[0323] The media content recommendation method provided by the embodiment of the present application includes: obtaining the identity of at least one user based on image information. Obtaining target facial expression information that matches the sound emotion information, and obtaining the target identity corresponding to the target facial expression information from the identity of at least one user, the target facial expression information being consistent with the emotion corresponding to the sound emotion information. Based on the sound emotion information and the standard sound emotion information corresponding to the preset time period, obtaining the evaluation information of the user corresponding to the target identity on the media content. By matching the facial expression information with the sound emotion information, and then obtaining the identity corresponding to the target facial expression information, it is possible to determine the evaluation information of each user's identity based on their respective facial expression information and sound emotion information, so as to improve the comprehensiveness and pertinence of the acquisition of user evaluation information.

[0324] Based on the above embodiment, before playing the media content, the user's identity may be identified. If the user's historical evaluation information is obtained based on the user's identity, the media content recommended to the user may be determined based on the historical evaluation information.

[0325] Specifically, first, the user's identity identifier is obtained, and then it is detected whether the user's historical evaluation information can be obtained based on the user's identity identifier. If it can be detected, it means that the user corresponding to the identity identifier has watched media content on the multimedia device. At this time, the media content recommended to the user is determined based on the user's historical evaluation information. The recommended media content can, for example, be media content of the same type as the media content with evaluation information higher than a preset score in the user's historical evaluation information, or the recommended media content can also be content recommended by other users for the media content with evaluation information higher than a preset score in the user's historical evaluation information, etc. The specific implementation method can be set according to needs and is not limited here.

[0326] On the basis of the above embodiments, the media content recommendation method provided by the present application can obtain the user's evaluation information of the media content according to the status information, and can also make differentiated recommendations according to the number of users to be recommended. The following describes the implementation method of media content recommendation in conjunction with specific embodiments. Figure 7 Make an introduction.

[0327] Figure 7 The process of the media content recommendation method provided in one embodiment of the present application Figure 6 ,like Figure 7 As shown, the method includes:

[0328] S701: Obtain the number of identified users to be recommended.

[0329] In this embodiment, before recommending media content, it is necessary to first identify the user to be recommended and then determine the number of users to be recommended. For example, if only Zhang San is currently viewing media content, then media content that Zhang San is interested in can be recommended based on Zhang San's rating information. Alternatively, if Zhang San and Li Si are currently viewing media content together, then media content that is of mutual interest to Zhang San and Li Si can be recommended.

[0330] In one possible implementation, image information may be used for identification to obtain the number of users to be recommended, or the identifier or number of users to be recommended input by the user may be received before the media content starts playing. This embodiment does not limit the implementation of the number of users to be recommended.

[0331] S702. If a user to be recommended is identified, determine the media content whose evaluation information associated with the user to be recommended meets a first preset condition; wherein the first preset condition is specifically one of the following: the evaluation score is higher than a preset score or the evaluation ranking is higher than a preset ranking.

[0332] If a user to be recommended is identified, the media content that the user is interested in can be directly recommended. In this embodiment, each user is associated with evaluation information of different media content, where the association relationship can be shown in Table 1, for example:

[0333] Identity Media Content Evaluation information Zhang San fast and Furious 97 Zhang San Neptune 80 Li Si fast and Furious 34 Li Si Neptune 77

[0334] Specifically, determine the media content in which the evaluation information associated with the user to be recommended meets the first preset condition, where the first preset condition is a condition using the user's evaluation as a measurement indicator, specifically one of the following: the evaluation score is higher than the preset score or the evaluation ranking is higher than the preset ranking, or the first preset condition can also include conditions pre-entered by the user, such as the user setting not to recommend horror movies, etc. The specific setting method of the first preset condition can be set according to implementation requirements, and is not limited here.

[0335] The following is an example of the user's identity and evaluation information in Table 1. For example, if Zhang San is currently the user to be recommended, and the first preset condition is, for example, that the evaluation information is higher than 90 points, then based on the evaluation information of different media content associated with the user in Table 1, it can be determined that the media content whose evaluation meets the first preset condition is "Fast and Furious".

[0336] S703: Determine an evaluation list according to the media content that meets the first preset condition.

[0337] In one possible implementation, the program type of the media content that meets the first preset condition can be obtained, and then all media content of the same type can be obtained, and an evaluation list can be determined based on at least one media content ranked before a preset number. Alternatively, the evaluation list can be determined based on recommendation information from other users who have watched the media content that meets the first preset condition, thereby determining media content similar to the media content that meets the first preset condition.

[0338] S704. If at least two users to be recommended are identified, determine the media content whose evaluation information associated with each of the at least two users to be recommended meets a second preset condition; the second preset condition is specifically one of the following: the evaluation score is higher than a preset score or the evaluation ranking is higher than a preset ranking.

[0339] If at least two users to be recommended are identified, it is necessary to recommend media content that is of common interest to multiple users. Specifically, first obtain the evaluation information associated with each of the at least two users to be recommended, and determine the media content whose evaluation information associated with each user meets the second preset condition, where the second preset condition is a condition that uses the evaluations of multiple users as a measurement indicator, specifically one of the following: the evaluation score is higher than the preset score or the evaluation ranking is higher than the preset ranking. The second preset condition may also include conditions pre-entered by the user, etc., which are not limited here.

[0340] Take the user identity and evaluation information in Table 1 as an example for explanation. For example, currently Zhang San and Li Si are the users to be recommended, and it is necessary to recommend media content that both users are interested in. The second preset condition may be, for example, that the evaluation score is higher than 70 points. Then, based on the evaluation information of different media content associated with the two users in Table 1, it can be determined that the media content whose evaluation information meets the second preset condition is "Aquaman".

[0341] S705: Determine an evaluation list based on the media content that meets the second preset condition.

[0342] Secondly, an evaluation list is determined based on the media content whose evaluation information meets the second preset condition. The specific implementation method is similar to the above-mentioned determination of the evaluation list based on the first preset condition, and will not be repeated here.

[0343] S706: Send the evaluation list to the media source platform, and obtain a recommendation list returned by the media source platform, where the recommendation list includes media content to be recommended to the user.

[0344] Furthermore, the evaluation list is sent to the media source platform, and then the media source platform returns the content of the media content to the multimedia device, and finally the media content returned by the media source platform to be recommended to the user is obtained and the media content is recommended.

[0345] The media content recommendation method provided by the embodiment of the present application includes: obtaining the number of identified users to be recommended. If a user to be recommended is identified, the media content whose evaluation information associated with the user to be recommended meets the first preset condition is determined; wherein the first preset condition is specifically one of the following: the evaluation score is higher than the preset score or the evaluation ranking is higher than the preset ranking. Based on the media content whose evaluation meets the first preset condition, an evaluation list is determined. If at least two users to be recommended are identified, the media content whose evaluation information associated with each of the at least two users to be recommended meets the second preset condition is determined; the second preset condition is specifically one of the following: the evaluation score is higher than the preset score or the evaluation ranking is higher than the preset ranking. Based on the media content that meets the second preset condition, an evaluation list is determined. The evaluation list is sent to the media source platform, and a recommendation list returned by the media source platform is obtained, the recommendation list including the media content to be recommended to the user. Media content is recommended based on user evaluation information. When there is only a single user, program recommendations are made based only on the user's evaluation information and a first preset condition, thereby improving the accuracy of program recommendations. When multiple users watch media content at the same time, media content recommendations are made based on each user's evaluation information and a second preset condition. The second preset condition takes into account the common preferences of multiple users, so media content suitable for multiple users can be recommended to improve user experience.

[0346] The following is an example of a specific scenario to illustrate the media content recommendation method provided by this application. Figure 8 as well as Figure 9 Make an introduction, Figure 8 Flowchart 1 of a method for recommending media content provided in an embodiment of the present application. Figure 9 A flowchart of a media content recommendation method provided in an embodiment of the present application Figure 2 .

[0347] like Figure 8 As shown, assuming that an emotional drama is being played at this time, when the media content is playing, the camera records photos of the audience every 30 seconds, where 30 seconds is the first image acquisition frequency, and the far-field microphone is turned on in advance according to the preset "stalk", where the preset "stalk" is the program content corresponding to the preset time period.

[0348] During a preset time period corresponding to a preset "meme", the far-field microphone collects the user's voice and / or collects the user's image information at a second image acquisition frequency, such as recording photos of viewers every 5 seconds, and then inputs the collected photos and sounds into a model for detecting the viewer's status.

[0349] Among them, when the model for detecting the state of the moviegoer is implemented, the evaluation information can be obtained based on the viewer's photo only, the viewer's voice only, or the viewer's photo and voice combined. The specific operation steps of the three implementation methods can be referred to Figure 4 、 Figure 5 as well as Figure 6 Example of .

[0350] Secondly, the user's evaluation information of the media content is obtained based on the analysis results of the state model. For example, based on the analysis results such as husband A is not in front of the TV, husband A has his eyes closed, and husband A is facing the TV, it can be determined that husband A does not like to watch such videos. Based on the analysis results such as wife B wiping tears, watching the TV, and laughing, it can be determined that wife B likes to watch such videos.

[0351] Specifically, when the user starts watching a movie, the camera first detects the identity of the current viewer. If only wife B is watching the movie, similar emotional dramas will be recommended to wife B. If it is detected that husband A and wife B are watching the movie together, because the husband's previous evaluation of emotional dramas indicates that he is not interested in such films, it is necessary to recommend films that both of them like to improve the user experience.

[0352] The following combination Figure 9 The specific implementation process of the model for detecting the state of moviegoers is introduced, such as Figure 9 As shown, the multimedia device first enters the viewing interface, and then turns on the TV camera with the user's authorization, obtains the user's image information through the camera to identify the information of the current viewer, and determines the number of users to be recommended based on the identification results. The current viewer is, for example, husband A and / or husband B.

[0353] If the current moviegoer is only wife B, similar programs can be recommended based on the media content that wife B likes. If the current moviegoers are husband A and wife B, similar programs can be recommended based on the media content that both husband A and wife B like. For example, based on the evaluation information of husband A and wife B on multiple media contents, it is determined that husband A and wife B both like comedy variety shows, similar comedy variety shows will be recommended to the two.

[0354] Secondly, husband A and wife B start watching a movie, and the camera takes photos of husband A and wife B at regular intervals according to the first image acquisition frequency, thereby obtaining the viewing status of husband A and wife B, and judging in real time whether the time point of the current media content playback reaches the preset joke time point.

[0355] If the preset joke time point is not reached, that is, the time period of the current media content playback time node is in another time period, then it is judged whether each viewer is immersed in the viewing experience based on the viewing status of each viewer. The judgment method may be, for example, whether the viewer has left his seat, whether his face is facing the TV, whether his eyes are open, etc. If it is determined that the viewer is immersed in the viewing experience, it indicates that the viewing effect is good, and the viewer's evaluation of the media content is increased. If it is determined that the viewer is not immersed in the viewing experience, it indicates that the viewing effect is poor, and the viewer's evaluation of the media content is deducted.

[0356] If the time point of the preset joke is reached, the far-field microphone is first turned on. The far-field microphone can also be turned on before the time point of the preset joke. Secondly, the viewing status of each viewer is detected and judged separately. First, it is judged whether the viewer is detected to be laughing during the joke time period. If the viewer is not detected to be laughing, it can be determined that the effect of the joke is not achieved, and the viewer's evaluation of the media content will be deducted. Among them, when the viewer is not detected to be laughing, the evaluation information is directly updated without the need to judge whether laughter is detected, thereby simplifying the judgment process and improving the efficiency of the system.

[0357] Secondly, if it is detected that the viewer is laughing, it can be preliminarily determined that the joke content has produced a certain program effect. Then, it is determined whether laughter is detected during the joke time period. If no laughter is detected, it indicates that the viewer's feedback on the joke content is average. At this time, the score corresponding to the laughter in the image information can be added to the viewer's evaluation of the media content. If laughter is detected, it indicates that the user's feedback on the program effect produced by the joke content is very good. At this time, the laughter in the image information and the score corresponding to the laughter are added to the viewer's evaluation of the media content.

[0358] Among them, for different viewers, according to the time sequence of media content playback, the evaluation of the currently playing media content by each viewer is updated in real time during the preset time period corresponding to the preset jokes and other time periods outside the preset jokes, thereby obtaining the evaluation of the program by husband A and wife B respectively.

[0359] The media content recommendation method provided in this application determines each viewer's evaluation information on the media content based on the viewer's image information and / or sound information, and then recommends the media content based on the viewer's evaluation information, thereby improving the accuracy of media content recommendation.

[0360] Figure 10 This is a signaling flow chart of a media content recommendation method provided in an embodiment of the present application. Figure 11 The signaling process of the media content recommendation method provided in one embodiment of the present application Figure 2 .

[0361] like Figure 10 As shown, first, the multimedia device receives the instruction sent by the viewer to enter the viewing application, and then the multimedia device pops up an authorization page, where the authorization page is used to obtain the permission to open the image acquisition device and the sound acquisition device. Then, according to the viewer's authorization operation, the image acquisition device and / or the sound acquisition device is turned on, where the sound acquisition device can be turned on when the preset time period arrives to save resources.

[0362] The image acquisition device obtains the image information of the viewer and identifies the identity information of the viewer based on the image information. For example, the image recognition result of the image information can be compared with the pre-stored image information to obtain the identity information of the viewer. Secondly, a list of media content preferred by the viewer is obtained based on the identity information of the viewer. If there are multiple viewers at this time, the media content obtained is the media content that multiple viewers like.

[0363] The media content list is transmitted to the media source platform, where the media source platform obtains specific media content based on the list information. Then, the multimedia device receives the media content sent by the media source platform, recommends the media content as similar media content to the user, and plays the preferred media content at the same time.

[0364] like Figure 11 As shown, the viewer starts watching a movie, and the multimedia device turns on the image acquisition device to collect the viewer's picture information. Then, the multimedia device receives the viewer's viewing status picture sent by the image acquisition device, and identifies whether the viewer is immersed in the program content of the media content based on the viewing status picture, and adds or subtracts points from the viewer's evaluation of the media content based on the identification result.

[0365] Specifically, the media source platform pre-processes the media content, and sets a preset time period in the playback information of the media content. When the time node corresponding to the preset stalk is not reached, the operation of adding or subtracting points based on the viewer's viewing status picture is cyclically executed. For example, the operation can be cyclically executed according to the first image acquisition frequency until the start time node corresponding to the preset stalk is reached.

[0366] Secondly, the image acquisition device is controlled to start the high-speed continuous shooting mode N seconds before the preset stalk is about to arrive, where the frequency corresponding to the high-speed continuous shooting mode is the second image acquisition frequency, and the multimedia device controls the sound acquisition device to start and prepare to collect the sound information of the viewer. It should be understood that N is an integer greater than or equal to 0.

[0367] Specifically, within a preset time period corresponding to a preset meme, the image acquisition device acquires picture information of the viewer according to a second image acquisition frequency and sends it to the multimedia device. Next, the multimedia device identifies the viewer's reaction to the preset meme based on the picture information, and performs operations such as evaluation and scoring or evaluation and scoring based on the recognition result.

[0368] Secondly, the sound collection device collects the voices of the viewers within a preset time period corresponding to the preset meme, and sends the sound collection structure to the multimedia device. Then, the multimedia device identifies the viewers' reactions to the preset meme based on the sound information collected, and adds or subtracts points for the corresponding evaluation.

[0369] Among them, within the preset time period corresponding to the preset meme, the addition and subtraction of points based on the image information and sound information is also executed in a loop. For example, it can be executed in a loop according to the second collection frequency until the end time node corresponding to the preset meme is reached.

[0370] When the program content corresponding to the preset stalk ends, the image acquisition device is controlled to resume the low-speed shooting mode, and the sound acquisition device is controlled to be turned off, thereby avoiding waste of resources.

[0371] During the specific implementation process, there may be multiple preset memes in the media content. Specifically, according to the playback order of the current media content, the evaluation information of each user on the media content is updated in real time during the preset time period corresponding to the preset memes, as well as other time periods outside the preset memes, so as to obtain the evaluation information of each user on the media content at the end of the media content playback to ensure the authenticity and validity of the evaluation information.

[0372] Figure 12 This is a structural diagram of a media content recommendation device according to an embodiment of the present application. Figure 12 As shown, the device 120 includes: an input module 1201 and a processing module 1202 .

[0373] The input module 1201 is configured to obtain user reaction status information when media content is played, where the reaction status information includes at least one of the following types of information: user image information obtained by an image acquisition device or user voice information obtained by a voice acquisition device;

[0374] The processing module 1202 is configured to obtain the user's evaluation information on the media content according to the response status information, wherein the evaluation information is used as a basis for recommending other media content to the user.

[0375] In one possible design, the image information includes one or more images captured by the image capture device at irregular intervals or continuously or based on a first image capture frequency within a preset time period, and the preset time period is the time period for playing preset content in the media content.

[0376] In one possible design, the image information also includes one or more images captured by the image capture device based on the second image capture frequency in time periods other than the preset time period; wherein the first image capture frequency is higher than the second image capture frequency.

[0377] In one possible design, the sound information includes one or more sound segments collected by the sound collection device within a preset time period. The preset time period is the time period for playing preset content in the media content. The frequency of collecting sound within the preset time period includes any one of the following: continuous collection, collection at a preset frequency, or collection at irregular intervals.

[0378] In one possible design, before obtaining the user's reaction status information, the processing module 1202 is further configured to:

[0379] If the reaction status information includes image information, the image acquisition device is turned on according to the first authorization information of the user, where the first authorization information is used to instruct the turning on of the image acquisition device;

[0380] If the response status information includes sound information, turning on the sound collection device according to the user's second authorization information, where the second authorization information is used to instruct the turning on of the sound collection device;

[0381] If the reaction status information includes image information and sound information, the image acquisition device is activated according to the first authorization information and the sound acquisition device is activated according to the second authorization information.

[0382] In one possible design, the processing module 1202 is specifically configured to:

[0383] If the reaction state information includes image information, obtaining user evaluation information of the media content based on the user's facial expression information;

[0384] If the reaction status information includes voice information, obtaining the user's evaluation information of the media content according to the user's voice emotion information;

[0385] If the reaction state information includes image information and sound information, obtaining the user's evaluation information of the media content based on the user's facial expression information and sound emotion information;

[0386] The facial expression information is the facial emotion information of the user when watching the media content obtained based on the image information, and the voice emotion information is the voice emotion information of the user when watching the preset content obtained based on the voice information.

[0387] In one possible design, the processing module 1202 is specifically configured to:

[0388] If the user's facial expression information is obtained within a preset time period, then obtaining standard facial expression information corresponding to the preset time period, where the standard facial expression information is expression information predefined according to preset content;

[0389] If the user's facial expression information is consistent with the standard facial expression information, then the user's evaluation of the media content is determined to be a plus-point evaluation;

[0390] If the user's facial expression information is inconsistent with the standard facial expression information, it is determined that the user's evaluation of the media content is a deductible evaluation.

[0391] In one possible design, the processing module 1202 is specifically configured to:

[0392] If the user's facial expression information is obtained in other time periods, obtaining an evaluation mapping table, wherein the evaluation mapping table is used to indicate evaluation information corresponding to different facial expression information;

[0393] The user's evaluation information on the media content is obtained based on the user's facial expression information and the evaluation mapping table.

[0394] In one possible design, the processing module 1202 is specifically configured to:

[0395] Obtaining standard sound emotion information corresponding to a preset time period, where the standard sound emotion information is sound information predefined according to preset content;

[0396] If the user's voice emotion information is consistent with the standard voice emotion information, the user's evaluation of the media content is determined to be a plus-point evaluation;

[0397] If the user's voice emotion information is inconsistent with the standard voice emotion information, the user's evaluation of the media content is determined to be a deductible evaluation.

[0398] In one possible design, after obtaining the user's reaction status information, the processing module 1202 is further configured to:

[0399] Obtaining an identity identifier of at least one user based on the image information or the sound information;

[0400] After obtaining the user's evaluation information on the media content according to the response status information, the evaluation information on the media content by each user is associated with the user's identity.

[0401] In one possible design, the processing module 1202 is specifically configured to:

[0402] obtaining an identity identifier of at least one user based on the image information;

[0403] Obtaining target facial expression information that matches the voice emotion information, and obtaining a target identity corresponding to the target facial expression information from at least one user identity, wherein the target facial expression information and the emotion corresponding to the voice emotion information are consistent;

[0404] According to the sound emotion information and the standard sound emotion information corresponding to the preset time period, the evaluation information of the user corresponding to the target identity on the media content is obtained, and the standard sound emotion information is information predefined according to the preset content.

[0405] In one possible design, before the media content is played, the processing module 1202 is further configured to:

[0406] Obtaining a correspondence between the user's identity identifier and the user's features based on the user's identity identifier and user features, where the user features include one of a face or a voice;

[0407] Obtaining at least one user's identity identifier based on a correspondence between the user's identity identifier and the face and the face included in the image information; or

[0408] At least one user's identity identifier is obtained according to the correspondence between the user's identity identifier and the sound and the sound included in the sound information.

[0409] In one possible design, the processing module 1202 is further configured to:

[0410] After obtaining the user's evaluation information on the media content according to the response status information, if the evaluation information is plus-point information, updating the user's evaluation information on the media content to obtain updated evaluation information;

[0411] If the evaluation information is deduction information, it is determined whether to update the user's evaluation information on the media content according to the status information obtained after the status information.

[0412] In one possible design, the processing module 1202 is further configured to:

[0413] Before playing the media content, the user's identity is identified. If the user's historical evaluation information is obtained based on the user's identity, the media content recommended to the user is determined based on the historical evaluation information.

[0414] The device provided in this embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.

[0415] Figure 13 A schematic diagram of the structure of a media content recommendation device provided in one embodiment of the present application Figure 2 .like Figure 13 As shown, this embodiment Figure 12 Based on the embodiment, it further includes: an output module 1303.

[0416] In one possible design, the processing module 1302 is further configured to: after the media content is played, determine an evaluation list based on the user's evaluation information, where the evaluation list includes media content that has been evaluated by the user;

[0417] The output module 1303 is used to: send the evaluation list to the media source platform;

[0418] The input module 1301 is further used to obtain a recommendation list returned by the media source platform, where the recommendation list includes media content to be recommended to the user.

[0419] In one possible design, the processing module 1302 is specifically configured to:

[0420] If a user to be recommended is identified, determining media content whose evaluation information associated with the user to be recommended satisfies a first preset condition; wherein the first preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking;

[0421] An evaluation list is determined based on the media content that meets the first preset condition.

[0422] In one possible design, the processing module 1302 is specifically configured to:

[0423] If at least two users to be recommended are identified, determining media content whose evaluation information associated with each of the at least two users to be recommended satisfies a second preset condition, where the second preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking;

[0424] An evaluation list is determined based on the media contents that all meet the second preset condition.

[0425] The device provided in this embodiment can be used to execute the technical solution of the above method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.

[0426] Figure 14 This is a hardware structure diagram of a media content recommendation device provided in one embodiment of the present application. Figure 14 As shown:

[0427] The media content recommendation device 1401 can perform wireless communication via NFC-related protocols, for example, wireless communication with a media source platform, or communication with a third-party device such as an image acquisition device or a sound acquisition device.

[0428] Exemplarily, the media content recommendation device 1401 can be connected to the electronic device to be communicated through one or more communication networks (e.g., wired or wireless). Exemplarily, the communication network can be a local area network or a wide area network (WAN) such as the Internet. The communication network can be implemented using any known network communication protocol, which can be various wired or wireless communication protocols, such as Ethernet, universal serial bus (USB), FireWire, global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TD-SCDMA), long term evolution (LTE), Bluetooth, wireless fidelity (Wi-Fi), NFC, voice over Internet protocol (VoIP), or any other suitable communication protocol.

[0429] Exemplarily, the media content recommendation device 1401 can establish a connection with the image acquisition device via Wi-Fi or Bluetooth. In another exemplary embodiment, the media content recommendation device 1401 not only establishes a connection with the image acquisition device via Bluetooth, but also establishes a connection with the media source platform via a wide area network.

[0430] Among them, the media content recommendation device 1401 can be a mobile terminal (Mobile Terminal) or user equipment, such as a mobile phone, tablet computer, TV, external advertising equipment, vehicle-mounted processing device or a mobile computer, etc. The mobile computer can be, for example, a portable computer, a pocket computer or a handheld computer.

[0431] For example, Figure 1414 shows a schematic structural diagram of a media content recommendation device 1401. The media content recommendation device 1401 may include a processor 1410, an external memory interface 1420, an internal memory 1421, a universal serial bus (USB) interface 1430, a charging management module 1440, a power management module 1441, a battery 1442, antenna 1, antenna 2, a mobile communication module 1450, a wireless communication module 1460, an audio module 1470, a speaker 1470A, a receiver 1470B, a microphone 1470C, an earphone interface 1470D, a sensor 1480, a button 1490, a motor 1491, an indicator 1492, a camera 1493, and a display 1494.

[0432] It should be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the media content recommendation device 1401. In other embodiments of the present application, the media content recommendation device 1401 may include more or fewer components than illustrated, or may combine or separate certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0433] The processor 1410 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). The different processing units may be independent devices or integrated into one or more processors. In some embodiments, the media content recommendation device 1401 may also include one or more processors 1410. The controller may be the nerve center and command center of the media content recommendation device 1401. The controller may generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. The processor 1410 may also be provided with a memory for storing instructions and data.

[0434] In some embodiments, the memory in processor 1410 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 1410. If processor 1410 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated access, reduces the waiting time of processor 1410, and thus improves the processing efficiency of media content recommendation device 1401.

[0435] In some embodiments, the processor 1410 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0436] The USB interface 1430 is an interface that complies with USB standards and specifications, and may be a Mini USB interface, a MicroUSB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the media content recommendation device 1401, transfer data between the media content recommendation device 1401 and peripheral devices, and connect headphones to play audio.

[0437] It should be understood that the interface connection relationships between the modules illustrated in the embodiments of the present invention are merely illustrative and do not constitute a structural limitation on the media content recommendation device 1401. In other embodiments of the present application, the media content recommendation device 1401 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0438] The charging management module 1440 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 1440 can receive charging input from the wired charger via the USB interface 1430. In some wireless charging embodiments, the charging management module 1440 can receive wireless charging input via the wireless charging coil of the media content recommendation device 1401. While charging the battery 1442, the charging management module 1440 can also provide power to the media content recommendation device 1401 via the power management module 1441.

[0439] The power management module 1441 is used to connect the battery 1442, the charging management module 1440 and the processor 1410. The power management module 1441 receives input from the battery 1442 and / or the charging management module 1440, and provides power to the processor 1410, the internal memory 1421, the display 1494, the camera 1493, and the wireless communication module 1460. The power management module 1441 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 1441 can also be set in the processor 1410. In other embodiments, the power management module 1441 and the charging management module 1440 can also be set in the same device.

[0440] The wireless communication functionality of media content recommendation device 1401 can be implemented using antenna 1, antenna 2, mobile communication module 1450, wireless communication module 1460, a modem processor, and a baseband processor. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in media content recommendation device 1401 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antenna can be used in conjunction with a tuning switch.

[0441] The mobile communication module 1450 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc., applied to the media content recommendation device 1401. The mobile communication module 1450 may include at least one filter, a switch, a power amplifier, a low-noise amplifier, etc. The mobile communication module 1450 can receive electromagnetic waves from the antenna 1, filter, amplify, and perform other processing on the received electromagnetic waves, and transmit them to the modem processor for demodulation. The mobile communication module 1450 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 1450 can be set in the processor 1410. In some embodiments, at least some of the functional modules of the mobile communication module 1450 can be set in the same device as at least some of the modules of the processor 110.

[0442] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 1470A, the receiver 1470B, etc.) or displays an image or video through the display screen 1494. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 1410 and be set in the same device as the mobile communication module 1450 or other functional modules.

[0443] The wireless communication module 1460 can provide wireless communication solutions including wireless local area networks (WLAN), Bluetooth, global navigation satellite system (GNSS), frequency modulation (FM), NFC, infrared technology (IR), etc. applied to the media content recommendation device 14401. The wireless communication module 1460 can be one or more devices integrating at least one communication processing module. The wireless communication module 1460 receives electromagnetic waves via antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 1410. The wireless communication module 1460 can also receive the signal to be transmitted from the processor 1410, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through antenna 2.

[0444] In some embodiments, antenna 1 of the media content recommendation device 1401 is coupled to the mobile communication module 1450, and antenna 2 is coupled to the wireless communication module 1460, so that the media content recommendation device 1401 can communicate with the network and other devices via wireless communication technologies. The wireless communication technologies may include GSM, GPRS, CDMA, WCDMA, TD-SCDMA, LTE, GNSS, WLAN, NFC, FM, and / or IR technologies. The GNSS may include the global positioning system (GPS), the global navigation satellite system (GLONASS), the Beidou navigation satellite system (BDS), the quasi-zenith satellite system (QZSS), and / or the satellite-based augmentation system (SBAS).

[0445] Media content recommendation device 1401 can implement display functions through a GPU, display screen 1494, and an application processor. A GPU is a microprocessor for image processing that connects display screen 1494 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 1410 may include one or more GPUs that execute instructions to generate or modify display information.

[0446] The display screen 1494 is used to display images, videos, etc. The display screen 1494 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the media content recommendation device 1401 may include one or N display screens 1494, where N is a positive integer greater than 1.

[0447] The media content recommendation device 1401 can implement a shooting function through an ISP, one or more cameras 1493, a video codec, a GPU, one or more display screens 1494, and an application processor.

[0448] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications such as image recognition, face recognition, speech recognition, and text comprehension in the media content recommendation device 1401.

[0449] External memory interface 1420 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of media content recommendation device 1401. The external memory card communicates with processor 1410 via external memory interface 1420 to implement data storage. For example, data files such as music, photos, and videos can be stored on the external memory card.

[0450] The internal memory 1421 can be used to store one or more computer programs, which include instructions. The processor 1410 can execute the above instructions stored in the internal memory 1421, so that the media content recommendation device 1401 performs the voice switching method provided in some embodiments of the present application, as well as various functional applications and data processing. The internal memory 1421 may include a program storage area and a data storage area. The program storage area may store an operating system; the program storage area may also store one or more application programs (such as user characteristics, user voice information, etc.). The data storage area may store data created during the use of the media content recommendation device 1401 (such as user historical viewing records, etc.). In addition, the internal memory 1421 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc. In some embodiments, the processor 1410 can enable the media content recommendation device 1401 to execute the media content recommendation method provided in the embodiments of the present application, as well as various functional applications and data processing by running instructions stored in the internal memory 1421 and / or instructions stored in the memory set in the processor 1410.

[0451] The media content recommendation device 1401 can implement audio functions through an audio module 1470, a speaker 1470A, a receiver 1470B, a microphone 1470C, a headphone jack 1470D, and an application processor. For example, audio playback of media content, music playback, etc. Among them, the audio module 1470 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signals. The audio module 1470 can also be used to encode and decode audio signals. In some embodiments, the audio module 1470 can be set in the processor 1410, or some functional modules of the audio module 1470 can be set in the processor 1410.

[0452] Speaker 170A, also known as a "loudspeaker," is used to convert audio signals into sound signals. Media content recommendation device 1401 can listen to music or radio programs through speaker 1470A. Receiver 1470B, also known as a "handset," is used to convert audio signals into sound signals.

[0453] When the media content recommendation device 1401 obtains the user's voice information, it can be obtained through the receiver 1470B. Microphone 1470C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When obtaining the user's voice information, the user's voice signal can be obtained by microphone 1470C, wherein microphone 1470C can be, for example, a far-field microphone, and then the sound signal is input to microphone 1470C. The media content recommendation device 1401 can be provided with at least one microphone 1470C. In other embodiments, the media content recommendation device 1401 can be provided with two microphones 1470C, which can not only collect sound signals but also realize noise reduction functions. In other embodiments, the media content recommendation device 1401 can also be provided with three, four or more microphones 1470C to realize sound signal collection, noise reduction, and recognition of sound information, and realize directional recording functions, etc.

[0454] The headphone jack 1470D is used to connect a wired headphone and can be a USB interface 1430, a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0455] The sensor 1480 may include a pressure sensor, a gyro sensor, an air pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, and the like.

[0456] The pressure sensor is used to sense pressure signals and convert them into electrical signals. In some embodiments, the pressure sensor can be installed on the display screen. There are many types of pressure sensors, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor can include at least two parallel plates made of conductive material. When a force acts on the pressure sensor electrodes, the capacitance between them changes. The media content recommendation device 1401 determines the intensity of the pressure based on the change in capacitance. When a touch operation is applied to the display screen, the media content recommendation device 1401 detects the intensity of the touch operation using the pressure sensor. The media content recommendation device can also calculate the touch location based on the detection signal from the pressure sensor. In some embodiments, touch operations applied to the same touch location but with different touch operation intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a first pressure threshold is applied to a short message application icon, an instruction to view short messages is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the short message application icon, an instruction to create a new short message is executed.

[0457] The gyroscope sensor can be used to determine the motion posture of the media content recommendation device 1401 (for example, when it is a tablet or mobile phone). In some embodiments, the angular velocity of the media content recommendation device 1401 around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor. The gyroscope sensor can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor detects the angle of the shake of the media content recommendation device 1401, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shake of the media content recommendation device 1401 through reverse movement to achieve anti-shake. The gyroscope sensor can also be used for navigation, somatosensory game scenes, etc.

[0458] The accelerometer can detect the magnitude of acceleration of the media content recommendation device 1401 in all directions (generally three axes). When the media content recommendation device 1401 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the electronic device's posture, enabling applications such as switching between landscape and portrait modes and pedometers.

[0459] The distance sensor is used to measure distance. The media content recommendation device 1401 can measure distance using infrared or laser. In some embodiments, when shooting a scene, the media content recommendation device 1401 can use the distance sensor 180F to measure distance to achieve fast focusing.

[0460] The proximity light sensor may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The media content recommendation device 1401 emits infrared light through the light emitting diode. The media content recommendation device 1401 uses the photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the media content recommendation device 1401. When insufficient reflected light is detected, the media content recommendation device 1401 can determine that there is no object near the media content recommendation device 1401. The proximity light sensor 180G can also be used for automatic unlocking and locking of the screen.

[0461] The ambient light sensor is used to sense ambient light brightness. The media content recommendation device 1401 can adaptively adjust the brightness of the display screen 1494 based on the perceived ambient light brightness. The ambient light sensor can also be used to automatically adjust the white balance based on the user's line of sight when acquiring image information. The ambient light sensor can also be used in conjunction with the proximity light sensor to detect whether the media content recommendation device 1401 is operating normally, for example, to facilitate invalid detection.

[0462] The fingerprint sensor (also known as a fingerprint reader) is used to collect fingerprints. The media content recommendation device 1401 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, obtain user permissions, etc. In addition, for further information about fingerprint sensors, please refer to International Patent Application PCT / CN2017 / 082773 entitled "Method and Electronic Device for Processing Notifications", the entire contents of which are incorporated by reference into this application.

[0463] A touch sensor may also be referred to as a touch panel or touch-sensitive surface. The touch sensor may be provided on the display screen 1494. The touch sensor and the display screen 1494 form a touch screen, also referred to as a touch screen. The touch sensor 1480K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor may transmit the detected touch operations to the application processor to determine the type of touch event. Visual output related to the touch operations may be provided through the display screen 1494. In other embodiments, the touch sensor may also be provided on the surface of the media content recommendation device 1401, at a location different from that of the display screen 1494.

[0464] The bone conduction sensor 1480M can obtain vibration signals. In some embodiments, the bone conduction sensor 1480M can obtain vibration signals of the vibrating bones of the human body. The bone conduction sensor 1480M can also contact the human pulse to receive blood pressure pulse signals. In some embodiments, the bone conduction sensor 1480M can also be set in headphones to form bone conduction headphones. The audio module 1470 can parse out voice signals based on the vibration signals of the vibrating bones of the human body obtained by the bone conduction sensor 1480M to implement voice functions. The application processor can parse heart rate information based on the blood pressure pulse signals obtained by the bone conduction sensor 1480M to implement heart rate detection functions.

[0465] The buttons 1490 include a power button, a volume button, etc. The buttons 1490 may be mechanical buttons or touch buttons. The media content recommendation device 1401 may receive key inputs and generate key signal inputs related to user settings and function control of the media content recommendation device 1401.

[0466] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When loading and executing computer program instructions on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) mode to another website, computer, server or data center. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more media that can be integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).

Claims

1. A media content recommendation method, characterized in that: include: Acquiring user reaction status information during media content playback, the reaction status information including at least one of the following types of information: image information of the user acquired by an image acquisition device or voice information of the user acquired by a voice acquisition device; wherein the media content includes at least one preset time period, the preset time period being a time period for playing a preset content in the media content and corresponding to a playback duration of the preset content; obtaining, based on the reaction status information, evaluation information of the user on the media content, wherein the evaluation information serves as a basis for recommending other media content to the user; If the reaction status information includes the image information, obtaining the user's evaluation information of the media content according to the user's facial expression information; If the user's facial expression information is obtained within the preset time period, determining the user's evaluation information of the media content based on whether the user's facial expression information is consistent with standard facial expression information; If the facial expression information of the user is obtained in a time period other than the preset time period, obtaining the user's evaluation information of the media content according to the facial expression information of the user and an evaluation mapping table, wherein the evaluation mapping table is used to indicate the evaluation information corresponding to different facial expression information; The image information includes one or more images captured by the image capture device based on a first image capture frequency within a preset time period; The image information also includes one or more images captured by the image capture device at a second image capture frequency in other time periods outside the preset time period; wherein the first image capture frequency is higher than the second image capture frequency.

2. The method according to claim 1, characterized in that The sound information includes one or more sound segments collected by the sound collection device within a preset time period, where the preset time period is the time period for playing the preset content in the media content. The frequency of collecting sound within the preset time period includes any one of the following: continuous collection, collection at a preset frequency, or collection at irregular intervals.

3. The method according to any one of claims 1-2, characterized in that Before obtaining the user's reaction status information, the method further includes: If the reaction status information includes the image information, turning on the image acquisition device according to the first authorization information of the user, where the first authorization information is used to instruct turning on the image acquisition device; If the reaction status information includes the sound information, turning on the sound collection device according to the second authorization information of the user, where the second authorization information is used to instruct turning on the sound collection device; If the reaction status information includes the image information and the sound information, the image acquisition device is enabled according to the first authorization information and the sound acquisition device is enabled according to the second authorization information.

4. The method according to any one of claims 1 to 2, characterized in that The acquiring, based on the reaction status information, the user's evaluation information of the media content includes: If the reaction status information includes the voice information, obtaining the user's evaluation information on the media content according to the user's voice emotion information; If the reaction state information includes the image information and the sound information, obtaining the user's evaluation information on the media content according to the user's facial expression information and the sound emotion information; The facial expression information is facial emotion information of the user when viewing the media content obtained based on the image information, and the voice emotion information is voice emotion information of the user when viewing the preset content obtained based on the voice information.

5. The method according to claim 4, characterized in that The step of determining the user's evaluation of the media content based on whether the user's facial expression information is consistent with the standard facial expression information includes: Acquire standard facial expression information corresponding to the preset time period, where the standard facial expression information is expression information predefined according to the preset content; If the facial expression information of the user is consistent with the standard facial expression information, determining that the user's evaluation of the media content is a plus-point evaluation; If the facial expression information of the user is inconsistent with the standard facial expression information, it is determined that the user's evaluation of the media content is a deductible evaluation.

6. The method according to claim 4, characterized in that Before obtaining the user's evaluation information on the media content based on the user's facial expression information and the evaluation mapping table, the method includes: Get the evaluation mapping table.

7. The method according to claim 4, characterized in that The obtaining, based on the user's voice emotion information, the user's evaluation information of the media content includes: Acquiring standard sound emotion information corresponding to the preset time period, wherein the standard sound emotion information is sound information predefined according to the preset content; If the user's voice emotion information is consistent with the standard voice emotion information, determining that the user's evaluation of the media content is a plus-point evaluation; If the user's voice emotion information is inconsistent with the standard voice emotion information, it is determined that the user's evaluation of the media content is a deductible evaluation.

8. The method according to any one of claims 1-2 and 5-7, characterized in that: After obtaining the user's reaction status information, the method further includes: Obtaining an identity identifier of at least one user based on the image information or the sound information; After obtaining the user's evaluation information of the media content according to the reaction status information, the method further includes: The evaluation information of each user on the media content is associated with the identity identifier of the user.

9. The method according to claim 4, characterized in that The obtaining, based on the user's facial expression information and the user's voice emotion information, of the user's evaluation information of the media content includes: obtaining an identity identifier of at least one user according to the image information; Obtaining target facial expression information that matches the voice emotion information, and obtaining a target identity corresponding to the target facial expression information from the identity identifier of the at least one user, wherein the target facial expression information is consistent with the emotion corresponding to the voice emotion information; According to the sound emotion information and the standard sound emotion information corresponding to the preset time period, the evaluation information of the user corresponding to the target identity on the media content is obtained, and the standard sound emotion information is information predefined according to the preset content.

10. The method according to claim 9, characterized in that Before playing the media content, the method further includes: Obtaining a correspondence between the user's identity identifier and the user feature according to the identity identifier and user feature input by the user, wherein the user feature includes one of a face or a voice; The acquiring of at least one user's identity according to the image information or the sound information includes: Obtaining at least one user's identity identifier based on the correspondence between the user's identity identifier and the face and the face included in the image information; or The identity identifier of at least one user is obtained according to the correspondence between the identity identifier of the user and the sound and the sound included in the sound information.

11. The method according to any one of claims 1-2, 5-7, and 9, characterized in that: After obtaining the user's evaluation information of the media content according to the reaction status information, the method further includes: If the evaluation information is bonus information, updating the user's evaluation information on the media content to obtain updated evaluation information; If the evaluation information is deduction information, it is determined whether to update the user's evaluation information on the media content according to status information acquired after the status information.

12. The method according to any one of claims 9-10, characterized in that Before playing the media content, the method further includes: The user's identity identifier is identified, and if historical evaluation information of the user is obtained according to the user's identity identifier, the media content recommended to the user is determined according to the historical evaluation information.

13. The method according to any one of claims 1-2, 5-7, 9-10, characterized in that: After the media content is played, the method further includes: Determining an evaluation list according to the evaluation information of the user, wherein the evaluation list includes media content evaluated by the user; The evaluation list is sent to a media source platform, and a recommendation list returned by the media source platform is obtained, where the recommendation list includes media content to be recommended to the user.

14. The method according to claim 13, characterized in that Determining an evaluation list according to the user's evaluation information includes: If a user to be recommended is identified, determining media content whose evaluation information associated with the user to be recommended meets a first preset condition; wherein the first preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking; An evaluation list is determined according to the media content that meets the first preset condition.

15. The method according to claim 13, characterized in that Determining an evaluation list according to the user's evaluation information includes: If at least two users to be recommended are identified, determining media content whose evaluation information associated with each of the at least two users to be recommended meets a second preset condition, where the second preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking; An evaluation list is determined based on the media contents that all meet the second preset condition.

16. A media content recommendation device, characterized in that: include: An input module, configured to obtain user reaction status information when media content is played, the reaction status information including at least one of the following types of information: image information of the user acquired by an image acquisition device or voice information of the user acquired by a voice acquisition device; wherein the media content includes at least one preset time period, the preset time period being a time period for playing a preset content in the media content, corresponding to the playback duration of the preset content; a processing module, configured to obtain, based on the reaction status information, evaluation information of the user on the media content, wherein the evaluation information serves as a basis for recommending other media content to the user; If the reaction status information includes the image information, obtaining the user's evaluation information of the media content according to the user's facial expression information; If the user's facial expression information is obtained within the preset time period, determining the user's evaluation information of the media content based on whether the user's facial expression information is consistent with standard facial expression information; If the facial expression information of the user is obtained in a time period other than the preset time period, obtaining the user's evaluation information of the media content according to the facial expression information of the user and an evaluation mapping table, wherein the evaluation mapping table is used to indicate the evaluation information corresponding to different facial expression information; The image information includes one or more images captured by the image capture device at irregular intervals within a preset time period, or continuously and uninterruptedly, or based on a first image capture frequency; The image information also includes one or more images captured by the image capture device at a second image capture frequency in other time periods outside the preset time period; wherein the first image capture frequency is higher than the second image capture frequency.

17. The device according to claim 16, characterized in that The sound information includes one or more sound segments collected by the sound collection device within a preset time period, where the preset time period is the time period for playing the preset content in the media content. The frequency of collecting sound within the preset time period includes any one of the following: continuous collection, collection at a preset frequency, or collection at irregular intervals.

18. The device according to any one of claims 16-17, characterized in that Before obtaining the user's reaction status information, the processing module is further configured to: If the reaction status information includes the image information, turning on the image acquisition device according to the first authorization information of the user, where the first authorization information is used to instruct turning on the image acquisition device; If the reaction status information includes the sound information, turning on the sound collection device according to the second authorization information of the user, where the second authorization information is used to instruct turning on the sound collection device; If the reaction status information includes the image information and the sound information, the image acquisition device is enabled according to the first authorization information and the sound acquisition device is enabled according to the second authorization information.

19. The device according to any one of claims 16-17, characterized in that The processing module is specifically used for: If the reaction status information includes the voice information, obtaining the user's evaluation information on the media content according to the user's voice emotion information; If the reaction state information includes the image information and the sound information, obtaining the user's evaluation information on the media content according to the user's facial expression information and the sound emotion information; The facial expression information is facial emotion information of the user when viewing the media content obtained based on the image information, and the voice emotion information is voice emotion information of the user when viewing the preset content obtained based on the voice information.

20. The device according to claim 19, characterized in that The processing module is specifically used for: If the facial expression information of the user is obtained in the preset time period, then the standard facial expression information should be obtained, where the standard facial expression information is expression information predefined according to the preset content; If the facial expression information of the user is consistent with the standard facial expression information, determining that the user's evaluation of the media content is a plus-point evaluation; If the facial expression information of the user is inconsistent with the standard facial expression information, it is determined that the user's evaluation of the media content is a deductible evaluation.

21. The device according to claim 19, characterized in that The processing module is further configured to: Get the evaluation mapping table.

22. The device according to claim 19, characterized in that The processing module is specifically used for: Acquiring standard sound emotion information corresponding to the preset time period, wherein the standard sound emotion information is sound information predefined according to the preset content; If the user's voice emotion information is consistent with the standard voice emotion information, determining that the user's evaluation of the media content is a plus-point evaluation; If the user's voice emotion information is inconsistent with the standard voice emotion information, it is determined that the user's evaluation of the media content is a deductible evaluation.

23. The device according to any one of claims 16-17, 20-22, characterized in that After obtaining the user's reaction status information, the processing module is further configured to: Obtaining an identity identifier of at least one user based on the image information or the sound information; After obtaining the user's evaluation information on the media content according to the reaction status information, the evaluation information of each user on the media content is associated with the user's identity.

24. The device according to claim 19, characterized in that The processing module is specifically used for: obtaining an identity identifier of at least one user according to the image information; Obtaining target facial expression information that matches the voice emotion information, and obtaining a target identity corresponding to the target facial expression information from the identity identifier of the at least one user, wherein the target facial expression information is consistent with the emotion corresponding to the voice emotion information; According to the sound emotion information and the standard sound emotion information corresponding to the preset time period, the evaluation information of the user corresponding to the target identity on the media content is obtained, and the standard sound emotion information is information predefined according to the preset content.

25. The device according to claim 24, characterized in that Before the media content is played, the processing module is further configured to: Obtaining a correspondence between the user's identity identifier and the user feature according to the identity identifier and user feature input by the user, wherein the user feature includes one of a face or a voice; Obtaining at least one user's identity identifier based on the correspondence between the user's identity identifier and the face and the face included in the image information; or The identity identifier of at least one user is obtained according to the correspondence between the identity identifier of the user and the sound and the sound included in the sound information.

26. The device according to any one of claims 16-17, 20-22, 24-25, characterized in that The processing module is further configured to: After obtaining the user's evaluation information on the media content according to the reaction status information, if the evaluation information is bonus information, updating the user's evaluation information on the media content to obtain updated evaluation information; If the evaluation information is deduction information, it is determined whether to update the user's evaluation information on the media content according to status information acquired after the status information.

27. The device according to any one of claims 24-25, characterized in that The processing module is further configured to: Before the media content is played, the user's identity identifier is identified. If the user's historical evaluation information is obtained based on the user's identity identifier, the media content recommended to the user is determined based on the historical evaluation information.

28. The device according to any one of claims 16-17, 20-22, 24-25, characterized in that Also includes: Output module; The processing module is further configured to: after the media content is played, determine an evaluation list based on the user's evaluation information, wherein the evaluation list includes media content evaluated by the user; The output module is used to: send the evaluation list to the media source platform; The input module is further configured to obtain a recommendation list returned by the media source platform, wherein the recommendation list includes media content to be recommended to the user.

29. The device according to claim 28, characterized in that The processing module is specifically used for: If a user to be recommended is identified, determining media content whose evaluation information associated with the user to be recommended meets a first preset condition; wherein the first preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking; An evaluation list is determined according to the media content that meets the first preset condition.

30. The device according to claim 28, wherein The processing module is specifically used for: If at least two users to be recommended are identified, determining media content whose evaluation information associated with each of the at least two users to be recommended meets a second preset condition, where the second preset condition is specifically one of the following: an evaluation score higher than a preset score or an evaluation ranking higher than a preset ranking; An evaluation list is determined based on the media contents that all meet the second preset condition.

31. A terminal device, characterized in that: The device comprises a media content recommendation device according to any one of claims 16 to 30, a camera and / or a microphone.

Citation Information

Patent Citations

  • System and method for user-behavior based content recommendations

    CN108476259A