Service mode recommendation method and apparatus, electronic device, and storage medium

By acquiring user voice and images and combining them with historical evaluation information to adjust the service mode of smart home appliances, the problem that the service mode in the existing technology cannot meet user needs has been solved, thus improving the user experience.

CN114550243BActive Publication Date: 2026-05-05QINGDAO HAIER TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QINGDAO HAIER TECH
Filing Date
2022-02-10
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

The existing service models for smart home appliances fail to meet users' actual needs, resulting in a poor user experience.

Method used

By acquiring users' voice and/or images, the current target service mode is determined, and the service mode is adjusted in conjunction with evaluation information from historical usage to better match user needs.

Benefits of technology

It improved the compatibility of smart home appliance service models and optimized the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114550243B_ABST
    Figure CN114550243B_ABST
Patent Text Reader

Abstract

This invention provides a service mode recommendation method, apparatus, electronic device, and storage medium. The method includes: acquiring a user's first voice and / or first image; determining a current target service mode based on the acquired first voice and / or first image; adjusting the current target service mode based on historical usage; wherein the historical usage includes evaluation information of the previous target service mode within a set time period, the evaluation information being determined based on the user's second voice and / or second image; and using the evaluation information of the previous target service mode in the current service mode recommendation, thereby enabling the currently recommended target service mode to better meet the user's actual needs, improving the service mode recommendation function while optimizing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart home technology, and in particular to a service mode recommendation method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the rapid development of science and technology, smart home appliances are becoming increasingly popular. The widespread adoption of smart home appliances not only brings great convenience to people's daily lives, but also provides diversified choices and services.

[0003] Most current smart home appliances come with voice control or image control functions, which can provide intelligent services to users through voice or image control. However, they fail to consider whether the intelligent services provided meet the actual needs of users, or whether they can satisfy users. Once the services provided by smart home appliances deviate from the user's requirements, it is easy to lead to a poor user experience. Summary of the Invention

[0004] This invention provides a service mode recommendation method, apparatus, electronic device, and storage medium to address the shortcomings of existing technologies where recommended service modes fail to meet users' actual needs, resulting in a poor user experience.

[0005] This invention provides a service model recommendation method, comprising:

[0006] Acquire the user's first voice and / or first image;

[0007] Based on the acquired first voice and / or first image, determine the current target service mode;

[0008] Adjust the current target service mode based on historical usage;

[0009] The historical usage includes evaluation information for the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image.

[0010] According to a pattern recommendation method provided by the present invention, determining the current target service pattern based on the acquired first voice and / or first image includes:

[0011] Based on the acquired first voice and / or first image, determine the user's identity information and / or target service mode usage requirements;

[0012] Based on the identity information and / or the usage requirements of the target service mode, the current target service mode is determined.

[0013] According to a pattern recommendation method provided by the present invention, determining the user's identity information based on the acquired first voice and / or first image includes:

[0014] Voiceprint extraction is performed on the first speech to obtain the user's voiceprint features;

[0015] Based on the voiceprint features, the user's voice identity information is determined;

[0016] And / or, perform user identification on the first image to obtain the user's image identity information;

[0017] The user's identity information is determined based on the voice identity information and / or the image identity information.

[0018] According to a pattern recommendation method provided by the present invention, the evaluation information is determined based on the following steps:

[0019] Determine the user's second voice and / or second image;

[0020] Perform user emotion analysis on the second voice and / or the second image to obtain the user's voice emotion and / or image emotion;

[0021] Based on the voice emotion and / or image emotion, evaluation information for the previous target service mode is determined.

[0022] According to a pattern recommendation method provided by the present invention, the step of performing user emotion analysis on the second speech to obtain the user's speech emotion includes:

[0023] Perform voice emotion recognition on the second speech to obtain the user's voice emotion recognition result;

[0024] And / or, perform speech-to-text transcription on the second speech to obtain the transcribed text of the second speech;

[0025] Semantic sentiment recognition is performed on the transcribed text to obtain the semantic sentiment recognition result of the user;

[0026] Based on the voice emotion recognition results and / or the semantic emotion recognition results, the user's voice emotion is determined.

[0027] According to a pattern recommendation method provided by the present invention, the step of performing user sentiment analysis on the second image to obtain the user's image sentiment includes:

[0028] Determine the face region and / or body movement region in the second image;

[0029] Facial expression recognition is performed on the facial region to obtain the user's facial emotions;

[0030] And / or, perform motion recognition on the limb movement area to obtain the user's emotional state;

[0031] The user's image emotion is determined based on the facial expression emotion and / or the action emotion.

[0032] According to a pattern recommendation method provided by the present invention, after adjusting the current target service pattern based on historical usage, the method further includes:

[0033] The control command indicated by the current target service mode is transmitted to the terminal corresponding to the user, so as to request the terminal to recommend the current target service mode to the user after receiving the control command.

[0034] The present invention also provides a service mode recommendation device, comprising:

[0035] A voice and image acquisition unit is used to acquire the user's first voice and / or first image;

[0036] The target service mode determination unit is used to determine the current target service mode based on the acquired first voice and / or first image;

[0037] An adjustment unit is used to adjust the current target service mode based on historical usage; wherein the historical usage includes evaluation information of the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image.

[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the service mode recommendation method as described above.

[0039] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the service mode recommendation method as described above.

[0040] The service mode recommendation method, apparatus, electronic device, and storage medium provided by this invention determine the current target service mode based on a first voice and / or a first image, and adjust the current target service mode based on historical usage. The historical usage includes evaluation information of the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image. The evaluation information of the previous target service mode is used in the current service mode recommendation, thereby making the currently recommended target service mode more aligned with the user's actual needs. This improves the service mode recommendation function while optimizing the user experience. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0042] Figure 1 This is a flowchart illustrating the service model recommendation method provided by the present invention;

[0043] Figure 2 This is a general framework diagram of the service mode recommendation method provided by the present invention;

[0044] Figure 3 This is a schematic diagram of the service mode recommendation device provided by the present invention;

[0045] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0047] This invention provides a service mode recommendation method, which aims to use the evaluation information of the previous target service mode in the current service mode recommendation, so that the currently recommended target service mode can better meet the user's actual needs, thereby improving the user experience. Figure 1 This is a flowchart illustrating the service model recommendation method provided by the present invention, such as... Figure 1 As shown, the execution entity of this method is the server, and the method includes:

[0048] Step 110: Acquire the user's first voice and / or first image;

[0049] Specifically, before recommending a service model to a user, it is necessary to first determine one or more aspects of the user's needs, age, gender, etc. This information can be obtained through the user's voice and / or image, that is, it is necessary to obtain the user's first voice and / or first image. Subsequently, the service model can be recommended to the user based on the obtained first voice and / or first image.

[0050] It should be noted that the first voice and / or the first image here originate from the user's corresponding terminal, which is a device carrying voice processing and / or image acquisition functions, such as a refrigerator, washing machine, air conditioner, television, etc. Therefore, the above process is actually that the user's corresponding terminal first acquires the first voice and / or the first image through the voice module and / or image acquisition device (camera) installed on it, and then transmits the first voice and / or the first image to the server; and after that, the server can receive the first voice and / or the first image transmitted by the terminal.

[0051] Since both the first voice and / or the first image are used to represent the user's relevant information about the service mode that needs to be recommended, the first voice may be the user's recorded usage needs, and the first image may be the user's facial image, body movements, gestures, etc.

[0052] Step 120: Determine the current target service mode based on the acquired first voice and / or first image;

[0053] Specifically, after determining the user's first voice and / or first image in step 110, step 120 can be executed to determine the current target service mode based on the user's first voice and / or first image. The specific process may be to analyze the user's usage needs based on the information contained in the obtained user's first voice and / or first image, and then determine the current target service mode based on the user's usage needs.

[0054] For example, when the user's usage need is determined to be "washing shirts" through the first voice analysis, the current target service mode can be determined to be a suitable washing mode for "washing shirts". Specifically, it can be found from the pre-set adaptation relationship between various washing modes and clothes to be washed, such as single wash mode, rinsing mode, mixed mode, etc., and one or more of these washing modes can be used as the current target service mode so that this target service mode can be recommended to the user through the user's corresponding terminal in the future.

[0055] For example, when the analysis of the first image determines that the user is currently very cold or very hot, that is, when the user's usage requirement is to turn on the heating mode or the cooling mode, the current target service mode can be determined as the heating mode or the cooling mode.

[0056] Alternatively, information such as the user's age and gender can be analyzed from the first voice and / or first image to obtain the user's identity information, and then the current target service mode can be determined based on this identity information. For example, if the user is determined to be elderly through the analysis of the first voice and / or first image, a current target service mode that matches their identity information can be determined, such as opera or crosstalk programs (elderly people tend to enjoy opera and crosstalk). Alternatively, the current target service mode can be determined by combining the usage needs represented by the information contained in the user's first voice. It should be noted that in this case, considering that the user (elderly person) may have limited language expression ability, dialect services can be provided so that the user can accurately express their usage needs in their familiar language.

[0057] Accordingly, when the user is identified as a child through the analysis of the user's first voice and / or first image, the current target service mode can be determined as a service mode that matches the user's identity information, such as a children's channel or an intellectual activity channel. Alternatively, the current target service mode can be determined by combining the usage needs represented by the information contained in the user's first voice. For example, if the usage need is "watching cartoons," the current target service mode can be determined as the channel in the children's channel that is currently playing cartoons. It should be noted that in this case, considering that the user (child)'s language expression ability is relatively lacking, fun communication methods can be provided to clarify the user's (child's) usage needs through interaction, and then provide them with intelligent services.

[0058] It should be noted that the current target service mode determined based on the user's first voice and / or first image in the embodiments of the present invention can be a single service, such as turning on the air conditioner, or a more complex scene mode, such as home mode, sleep mode, etc.

[0059] Step 130: Adjust the current target service mode based on historical usage data; wherein, historical usage data includes evaluation information of the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image.

[0060] Considering that the evaluation information of the previous target service mode within a set time period was not included in the process of determining the current target service mode, and therefore the evaluation information of users for each recommended service mode within the set time period cannot be obtained, it is impossible to adjust the recommended service mode accordingly. In other words, it is impossible to adjust the current target service mode based on user feedback, or it can be understood as being unable to recommend a service mode that meets the user's actual needs based on the feedback, resulting in a poor user experience. Based on this, in this embodiment of the invention, in order to meet the user's usage needs and optimize the user experience, the current target service mode can be adjusted after determining it by combining historical usage data, which includes the evaluation information of the previous target service mode within a set time period.

[0061] It should be noted that the evaluation information here can be the user's level of satisfaction with the previous target service mode. For example, it can be any one of "very dissatisfied," "dissatisfied," "neutral," "satisfied," or "very satisfied." It can also be evaluation information between two adjacent evaluation information, or it can be empty, meaning the user did not express an evaluation of the previous target service mode. It can also be feedback from other users on the previous target service mode. This embodiment of the invention does not specifically limit this. Furthermore, the evaluation information can be determined by the user's second voice and / or second image. That is, it can be determined based on the semantic information and tone of attitude represented by the user's second voice, and / or the facial expressions and body movements conveyed by the second image. For example, when the tone of attitude is cheerful, the semantic information is "not bad," the facial expression is relaxed, and the body movements are relaxed, the corresponding evaluation information can be determined to be "satisfied."

[0062] Specifically, the process of adjusting the current target service mode described above can be based on the evaluation information of the previous target service mode within a set time period in the historical usage data. This involves adjusting the current target service mode, i.e., filtering out service modes in the current target service mode that clearly indicate poor user experience in the historical usage data, so as to make the historical usage data compatible with the current target service mode.

[0063] For example, if a user's requirement is "washing shirts", the previous target service mode is the single wash mode, and the user's evaluation of the single wash mode is unsatisfactory, the current target service mode can be determined to be a washing mode other than the single wash mode that is suitable for "washing shirts", such as the rinsing mode, the mixed mode, etc., and one or more of these washing modes can be used as the current target service mode so that this target service mode can be recommended to the user through the user's corresponding terminal in the future.

[0064] For example, if the user is determined to be elderly based on their first voice and / or first image analysis, and the previous target service mode was a crosstalk (xiangsheng) segment, and the user's evaluation of the crosstalk segment was unsatisfactory, the current target service mode can be adjusted to obtain a new current target service mode, such as a traditional opera segment. Conversely, if the user's evaluation of the crosstalk segment is satisfactory, the crosstalk segment can still be used as the current target service mode.

[0065] Accordingly, when the user is determined to be a child through the analysis of the user's first voice and / or first image, the previous target service mode is the Animal World channel, and the user's evaluation information for the Animal World channel is unsatisfactory, the previous target service mode (Animal World channel) can be removed from the current target service mode to obtain a new current target service mode, such as a children's channel or an intellectual activity channel.

[0066] It should be noted that, in addition to the evaluation information mentioned above, the user's historical usage data may also include the user's usage habits. These usage habits can indicate the service mode that the user usually uses and the frequency of use of the corresponding service mode. Therefore, the user's usage habits can also be applied to the process of adjusting the target service mode mentioned above, so as to provide a service mode that is closer to the user's usage habits, thereby achieving the goal of optimizing the user experience.

[0067] For example, if a user's requirement is "washing shirts", the previous target service mode is single wash mode, the user's evaluation information for single wash mode is dissatisfaction, and the user's usage habits indicate that the user usually uses the "hybrid mode" service mode, it can be further determined whether the "hybrid mode" is suitable for washing shirts. If the "hybrid mode" is suitable for "washing shirts", it can be directly adopted as the current target service mode.

[0068] The service mode recommendation method provided by this invention determines the current target service mode based on a first voice and / or a first image, and adjusts the current target service mode based on historical usage. The historical usage includes evaluation information of the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image. The evaluation information from the previous target service mode is used in the current service mode recommendation, thereby making the currently recommended target service mode more aligned with the user's actual needs. This improves the service mode recommendation function while optimizing the user experience.

[0069] Based on the above embodiments, step 130, which adjusts the current target service mode based on historical usage, further includes:

[0070] Control the user's corresponding terminal to recommend the current target service mode to the user.

[0071] Specifically, after adjusting the current target service mode based on historical usage, the system can control the user's corresponding terminal to recommend the current target service mode to the user. The specific process includes the following steps:

[0072] First, determine the control commands indicated by the current target service mode;

[0073] Subsequently, the control command is returned to the terminal transmitting the first voice and / or the first image, requesting that the terminal recommend the current target service mode to the user after receiving the control command returned by the server.

[0074] It should be noted that after receiving the control command, the terminal can recommend the current target service mode to the user through a pre-set recommendation method, such as broadcasting recommendations or displaying recommendations. That is, the terminal can broadcast relevant information about the current target service mode through the voice module installed on the terminal, and / or display relevant information about the current target service mode through the display screen.

[0075] Based on the above embodiments, step 120 includes:

[0076] Based on the acquired first voice and / or first image, determine the user's identity information and / or target service mode usage requirements;

[0077] Based on identity information and / or the usage requirements of the target service mode, determine the current target service mode.

[0078] Specifically, step 120, the process of determining the current target service mode based on the acquired user's first voice and / or first image, may include the following situations:

[0079] Firstly, the user's identity information can be determined based on the user's first image. That is, the user's identity can be identified by using the user's facial features in the first image as a basis.

[0080] The user's identity information can also be determined based on their first voice recording. This process can be done by directly identifying the user's identity information through the voiceprint features in the first voice recording, or by determining the user's identity information through the sound quality of the first voice recording. For example, if the voice is relatively lively, the user can be preliminarily identified as a young person; if the voice is relatively immature, the user can be preliminarily identified as a child; if the voice is relatively deep, the user can be preliminarily identified as a young person. It should be noted that the accuracy of the identity information determined by this method is not high. Therefore, the user's first image can be used for further identity verification.

[0081] Furthermore, the user's identity information can be determined by combining the user's first voice and first image. That is, the identity information determined by the first voice and the identity information determined by the first image are merged and summarized, and the final identity information is determined based on the fusion and summary results. Combining data from different levels to determine the user's identity information can result in a higher accuracy rate of the final identity information, thereby helping to improve the accuracy of subsequent service model recommendation functions and optimize the user experience.

[0082] Subsequently, the current target service mode can be determined based on the user's identity information obtained through any of the three methods mentioned above. That is, the service mode that matches the user's identity information can be determined and used as the current target service mode. For example, if the service modes that match the user's identity information include children's channels, animal world channels, and intellectual activity channels, one or more of them can be used as the current target service mode.

[0083] Secondly, user needs can be determined based on the user's first voice, that is, the user needs of the target service mode. In other words, the user needs are analyzed based on the user's first voice, so as to obtain the user needs of the target service mode.

[0084] Alternatively, the user's first image can be used to determine the target service mode usage requirements. Specifically, the user's needs can be analyzed based on the body movements and / or gestures in the first image to determine the target service mode usage requirements.

[0085] Furthermore, the target service mode usage requirements can be determined by combining the user's first voice and first image. That is, the target service mode usage requirements based on the first voice and the target service mode usage requirements based on the first image are integrated, and the final target service mode usage requirements are determined based on the integration results. The target service mode usage requirements determined by combining information from multiple levels can more accurately represent the user's usage intention, thereby helping to improve the accuracy of subsequent service mode recommendation functions and optimize the user experience.

[0086] Then, based on the usage requirements of the target service mode determined through any of the three methods mentioned above, the current target service mode can be determined, that is, the service mode that matches the usage requirements of the target service mode can be selected from various service modes. For example, when the usage requirement of the target service mode is "washing shirts", the current target service mode can be determined as a service mode suitable for washing shirts, such as single wash mode, rinsing mode, mixed mode, etc.

[0087] Third, the user's identity information and target service mode usage requirements can be determined through the first voice and / or the first image. The process of determining the user's identity information and target service mode usage requirements has been explained in detail above and will not be repeated here. Subsequently, the current target service mode can be determined based on the user's identity information and target service mode usage requirements. That is, by combining the user's identity information and target service mode usage requirements, a service mode that matches both can be determined and used as the current target service mode.

[0088] Based on the above embodiments, in step 120, determining the user's identity information based on the acquired first voice and / or first image includes:

[0089] Voiceprint extraction is performed on the first speech to obtain the user's voiceprint features;

[0090] Based on voiceprint features, determine the user's voice identity information;

[0091] And / or, perform user identification on the first image to obtain the user's image identity information;

[0092] The user's identity information is determined based on voice identity information and / or image identity information.

[0093] Considering the unique nature of voiceprints—that each person's voiceprint is unique and different from other people's—this embodiment of the invention can apply voiceprint features to determine a user's identity information. The specific process includes the following steps:

[0094] First, voiceprint extraction is performed on the first speech to determine the user's voiceprint features, that is, information about the user's voiceprint features is extracted from the first speech to obtain the user's voiceprint features;

[0095] Subsequently, the user's identity information can be directly determined based on the user's voiceprint characteristics. Specifically, the user's voiceprint characteristics can be matched with the voiceprint characteristics of users who have been pre-stored / recorded / registered, and the identity information of the user corresponding to the successfully matched voiceprint characteristics can be used as the user's identity information.

[0096] It should be noted that since the user's identity information obtained at this time is determined by the user's first voice, it is called the user's voice identity information; and considering that the accuracy of the user's voice identity information determined based on the user's voiceprint features is relatively high, therefore, in this embodiment of the invention, the user's voice identity information can be directly used as the user's identity information.

[0097] Alternatively, user identity information can be determined through facial features with high recognition rates. Specifically, the user's identity can be determined by performing identity recognition on the user's first image. That is, the user's facial features in the first image are used as a basis for identity recognition, and the user's identity information is finally obtained. Since this identity information is determined based on the user's first image, it is called image identity information. Because the recognition rate of facial features is very high, the image identity information obtained can be directly used as the user's identity information.

[0098] In addition, considering that there are a very small number of people with similar facial features, in order to ensure the accuracy of the obtained user identity information and the precision of subsequent service mode recommendations, this embodiment of the invention can also combine the user's voice identity information with the user's image identity information, and use both to jointly determine the user's identity information, so that the accuracy of the user's identity information can be guaranteed to the greatest extent.

[0099] Based on the above embodiments, the evaluation information is determined based on the following steps:

[0100] Determine the user's second voice and / or second image;

[0101] Perform user emotion analysis on the second speech and / or the second image to obtain the user's speech emotion and / or image emotion;

[0102] Based on voice emotion and / or image emotion, determine the evaluation information for the previous target service mode.

[0103] Specifically, since the evaluation information of the previous target service mode is used in the current service mode recommendation in this embodiment of the invention, after determining each target service mode, it is also necessary to obtain the user's evaluation information for that target service mode in order to improve the server's service mode recommendation function and thereby optimize the user experience. The process of determining this evaluation information specifically includes the following steps:

[0104] After the user's terminal switches to the previous target service mode and starts executing the service mode, the second voice and / or second image fed back by the user based on the previous target service mode can be collected first through the voice module and / or image acquisition device installed on the terminal.

[0105] Subsequently, emotion analysis can be performed on the user's second speech to obtain the user's vocal emotion. Specifically, the user's tone and attitude can be analyzed with reference to the second speech, as well as the user's emotion represented by the semantics of the second speech. The user's vocal emotion can be determined by combining these two or by either one. For example, when the user's tone and attitude are obviously unpleasant (disappointment, impatience, etc.), and / or the semantics of the second speech are "not so good", "not what I expected", "just so-so", etc., it can be determined that the vocal emotion reflected by the user's second speech is unpleasant or very unpleasant.

[0106] Emotional analysis can also be performed on users based on their second images to obtain their image emotions. Specifically, the user's facial expressions and / or body movements in the second image can be used as a benchmark to analyze the user's emotions and determine the user's image emotions. For example, when the user has a smile, a relaxed facial expression, and / or relaxed, natural and not stiff body movements in the second image, it can be determined that the image emotions reflected by the user's second image are good or excellent.

[0107] It can also perform user emotion analysis on the user's second voice and second image separately to obtain the user's voice emotion and image emotion;

[0108] Subsequently, the user's evaluation of the previous target service mode can be determined based on the user's voice emotion. For example, if the user's voice emotion is poor or very poor, the user's evaluation of the previous target service mode is dissatisfied or very dissatisfied. Alternatively, the user's evaluation of the previous target service mode can be determined based on the user's image emotion. For example, if the user's image emotion is good or excellent, the user's evaluation of the previous target service mode is satisfied or very satisfied. Or, the user's evaluation of the previous target service mode can be determined by combining the user's voice emotion and image emotion. The evaluation information determined by aggregating data from two different levels is more accurate and more conducive to the optimization of subsequent service mode recommendation functions based on this evaluation information, as well as the improvement of user experience.

[0109] It should be noted that when determining the evaluation information for the previous target service model, it is not limited to the current user, but can also include other users.

[0110] Based on the above embodiments, user emotion analysis is performed on the second voice to obtain the user's voice emotion, including:

[0111] Perform voice emotion recognition on the second speech to obtain the user's voice emotion recognition result;

[0112] And / or, perform speech-to-text transcription on the second speech to obtain the transcribed text of the second speech;

[0113] Semantic sentiment recognition is performed on the transcribed text to obtain the user's semantic sentiment recognition results;

[0114] Determine the user's voice emotion based on the results of voice emotion recognition and / or semantic emotion recognition.

[0115] Specifically, the process of performing user emotion analysis on the second speech to obtain the user's voice emotion includes the following steps:

[0116] First, the second speech is transcribed into text, thus obtaining the transcribed text of the second speech. This transcription can be achieved using conventional speech transcription techniques.

[0117] Then, semantic emotion recognition can be performed on the transcribed text of the second speech to identify the semantic information representing the user's emotions in the transcribed text, such as "not bad," "satisfied," "very satisfied," "not so good," etc., and finally obtain the user's emotion recognition result. It should be noted that since the emotion recognition result obtained at this time is determined based on the semantic information of the transcribed text of the second speech, it is called the semantic emotion recognition result.

[0118] Alternatively, voice emotion recognition can be performed directly on the user's second speech to obtain the voice emotion recognition result, that is, to identify information such as tone and attitude in the second speech that can indicate the user's emotions, so as to determine the user's voice emotion recognition result;

[0119] The above two methods can also be combined to perform both voice emotion recognition and semantic emotion recognition, thereby obtaining the user's voice emotion recognition results and semantic emotion recognition results respectively.

[0120] Subsequently, the user's voice emotion can be determined based on the user's voice emotion recognition results and / or semantic emotion recognition results. That is, the user's voice emotion recognition results or semantic emotion recognition results can be directly used as the user's voice emotion; or the user's voice emotion recognition results and semantic emotion recognition results can be combined to jointly determine the user's voice emotion. That is, the user's voice emotion recognition results and semantic emotion recognition results are fused and summarized, and the user's voice emotion is determined based on the fused and summarized results.

[0121] Based on the above embodiments, user emotion analysis is performed on the second image to obtain the user's image emotion, including:

[0122] Identify the face region and / or body movement region in the second image;

[0123] Facial expression recognition is performed on the face area to obtain the user's emotional expression;

[0124] And / or, perform motion recognition on the body movement area to obtain the user's emotional state;

[0125] Determine the user's image emotion based on facial expression emotion and / or action emotion.

[0126] Specifically, the process of performing user sentiment analysis on the second image to obtain the user's image sentiment includes the following steps:

[0127] Since facial expressions and body movements can both represent a user's emotions, after obtaining the user's second image, the user's face area and / or body movement area can be determined from the second image.

[0128] Then, facial expression recognition can be performed on the face area to identify the emotion corresponding to the user's facial expression, thereby obtaining the user's emotional expression; for example, when the user's facial expression corresponding to the face area is a hearty laugh, it can be determined that the user's emotional expression is excellent; correspondingly, when the user's facial expression is a frown and / or a downturned mouth, it can be determined that the user's emotional expression is poor.

[0129] It can also perform motion recognition on the user's body movement area to identify the user's body movements and the emotions they represent, and finally obtain the user's emotional state based on the motion. For example, when the body movement area corresponds to clapping, it can be determined that the user's emotional state is excellent; correspondingly, when the body movement is stomping, it can be determined that the user's emotional state is poor.

[0130] It can also perform facial expression recognition on the face area to obtain the user's emotional expression, and perform motion recognition on the body movement area to obtain the user's emotional movement.

[0131] Subsequently, the user's image emotion can be determined based on the user's facial expression and / or action emotion. That is, the user's facial expression emotion can be directly used as the user's image emotion, or the user's action emotion can be directly used as the user's image emotion, or the user's image emotion can be determined based on both. This embodiment of the invention does not make specific limitations in this regard.

[0132] Based on the above embodiments, step 130, which adjusts the current target service mode based on historical usage, further includes:

[0133] The control command indicated by the current target service mode is transmitted to the user's corresponding terminal, so that the terminal can recommend the current target service mode to the user after receiving the control command.

[0134] Specifically, after adjusting the current target service model based on historical usage, a series of steps are required to recommend the current target service model to users, including:

[0135] First, the server generates control instructions based on the current target service mode and returns these control instructions to the terminal that acquired the first voice and / or the first image.

[0136] Subsequently, the terminal receives the control command and recommends the current target service mode to the user based on the control command. Specifically, after receiving the control command, the terminal switches modes according to the user's actual situation, prepares to enter the current target service mode, and recommends the current target service mode to the user through a pre-set recommendation method, such as broadcast recommendation or display recommendation. That is, the relevant information of the current target service mode can be broadcast through the voice module installed on the terminal, and / or the relevant information of the current target service mode can be displayed on the display screen.

[0137] Furthermore, in order for the terminal to officially switch to the current target service mode, it is also necessary to generate serial port data according to the control instructions of the current target service mode and send it to the terminal's control board. After receiving the serial port data, the control board can control the terminal to switch to the current target service mode through the set parameters, so that the terminal can start the current target service mode after receiving the start instruction returned by the user based on the current target service mode.

[0138] After that, the normal execution process of the current target service mode begins, and evaluation information of the current target service mode within a set time period can be collected through the voice module and / or image acquisition device. This evaluation information is determined through the second voice and / or the second image so that the next service mode recommendation work can be carried out, and the above process is repeated.

[0139] Figure 2 This is a general framework diagram of the service mode recommendation method provided by the present invention, as shown below. Figure 2 As shown, the user's terminal first acquires the first voice and / or the first image through the voice module and / or image acquisition device installed on it; then, it transmits the first voice and / or the first image to the server.

[0140] After receiving the first voice and / or the first image, the server can first determine the target service mode usage requirements based on the first voice and / or the first image; it can also extract the voiceprint of the first voice to obtain the user's voiceprint features and determine the user's voice identity information based on the user's voiceprint features; and / or perform user identity recognition on the first image to obtain the user's image identity information; and determine the user's identity information based on the user's voice identity information and / or image identity information.

[0141] Subsequently, the current target service mode is determined based on the user's identity information and / or the user's needs for the target service mode.

[0142] Subsequently, the current target service mode is adjusted based on historical usage data; the historical usage data includes evaluation information of the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image.

[0143] After that, the server can control the user's terminal to recommend the current target service mode to the user. Specifically, the server can generate a control command based on the current target service mode and return the control command to the user's terminal.

[0144] Upon receiving the control command, the terminal can recommend the current target service mode to the user. Specifically, after receiving the control command, it can switch modes according to the user's actual situation, prepare to enter the current target service mode, and recommend the current target service mode to the user through pre-set recommendation methods, such as broadcast recommendation or display recommendation.

[0145] Furthermore, after recommending the current target service mode to the user, serial port data needs to be generated according to the control instructions of the current target service mode and sent to the terminal's control board. After receiving the serial port data, the control board can control the terminal to switch to the current target service mode through the set parameters, so that the terminal can start the current target service mode after receiving the start instruction returned by the user based on the current target service mode.

[0146] After that, the normal execution process of the current target service mode begins, and evaluation information of the current target service mode within a set time period can be collected through the voice module and / or image acquisition device. This evaluation information is determined through the second voice and / or the second image so that the next service mode recommendation work can be carried out, and the above process is repeated.

[0147] The service mode recommendation method provided by this invention determines the current target service mode based on a first voice and / or a first image, and adjusts the current target service mode based on historical usage. The historical usage includes evaluation information of the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image. The evaluation information from the previous target service mode is used in the current service mode recommendation, thereby making the currently recommended target service mode more aligned with the user's actual needs. This improves the service mode recommendation function while optimizing the user experience.

[0148] The service mode recommendation device provided by the present invention is described below. The service mode recommendation device described below can be referred to in correspondence with the service mode recommendation method described above.

[0149] Figure 3 This is a schematic diagram of the service mode recommendation device provided by the present invention, as shown below. Figure 3 As shown, the device includes:

[0150] The voice and image acquisition unit 310 is used to acquire the user's first voice and / or first image;

[0151] The target service mode determination unit 320 is used to determine the current target service mode based on the acquired first voice and / or first image;

[0152] The adjustment unit 330 is used to adjust the current target service mode based on historical usage; wherein the historical usage includes evaluation information of the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image.

[0153] The service mode recommendation device provided by this invention determines the current target service mode based on a first voice and / or a first image, and adjusts the current target service mode based on historical usage. The historical usage includes evaluation information of the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image. The evaluation information from the previous target service mode is used in the current service mode recommendation, thereby making the currently recommended target service mode more aligned with the user's actual needs. This improves the service mode recommendation function while optimizing the user experience.

[0154] Based on the above embodiments, the target service mode determination unit 320 is used for:

[0155] Based on the first voice and / or the first image, determine the user's identity information and / or target service mode usage requirements;

[0156] Based on the identity information and / or the usage requirements of the target service mode, determine the current target service mode.

[0157] Based on the above embodiments, the target service mode determination unit 320 is used for:

[0158] Voiceprint extraction is performed on the first speech to obtain the user's voiceprint features;

[0159] Based on the voiceprint features, the user's voice identity information is determined;

[0160] And / or, perform user identification on the first image to obtain the user's image identity information;

[0161] The user's identity information is determined based on the voice identity information and / or the image identity information.

[0162] Based on the above embodiments, the device further includes an evaluation information determination unit, used for:

[0163] Determine the user's second voice and / or second image;

[0164] Perform user emotion analysis on the second voice and / or the second image to obtain the user's voice emotion and / or image emotion;

[0165] Based on the voice emotion and / or image emotion, evaluation information for the previous target service mode is determined.

[0166] Based on the above embodiments, the evaluation information determination unit is used for:

[0167] Perform voice emotion recognition on the second speech to obtain the user's voice emotion recognition result;

[0168] And / or, perform speech-to-text transcription on the second speech to obtain the transcribed text of the second speech;

[0169] Semantic sentiment recognition is performed on the transcribed text to obtain the semantic sentiment recognition result of the user;

[0170] Based on the voice emotion recognition results and / or the semantic emotion recognition results, the user's voice emotion is determined.

[0171] Based on the above embodiments, the evaluation information determination unit is used for:

[0172] Determine the face region and / or body movement region in the second image;

[0173] Facial expression recognition is performed on the facial region to obtain the user's facial emotions;

[0174] And / or, perform motion recognition on the limb movement area to obtain the user's emotional state;

[0175] The user's image emotion is determined based on the facial expression emotion and / or the action emotion.

[0176] Based on the above embodiments, the device further includes a control unit, configured to:

[0177] The control command indicated by the current target service mode is transmitted to the terminal corresponding to the user, so as to request the terminal to recommend the current target service mode to the user after receiving the control command.

[0178] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can invoke logical instructions in the memory 430 to execute a service mode recommendation method, which includes: acquiring a user's first voice and / or first image; determining a current target service mode based on the acquired first voice and / or first image; and adjusting the current target service mode based on historical usage data; wherein the historical usage data includes evaluation information for the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image.

[0179] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0180] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, wherein when the program instructions are executed by a computer, the computer is able to execute the service mode recommendation method provided by the above methods, the method comprising: acquiring a user's first voice and / or first image; determining a current target service mode based on the acquired first voice and / or first image; adjusting the current target service mode based on historical usage; wherein the historical usage includes evaluation information for the previous target service mode within a set time period, the evaluation information being determined based on the user's second voice and / or second image.

[0181] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program implements a service mode recommendation method provided by the methods described above. The method includes: acquiring a user's first voice and / or first image; determining a current target service mode based on the acquired first voice and / or first image; and adjusting the current target service mode based on historical usage. The historical usage includes evaluation information for the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image.

[0182] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0183] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0184] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A service model recommendation method, characterized in that, include: Acquire the user's first voice and first image; Based on the acquired first voice and first image, the user's identity information is determined; Based on the first voice recording, determine the usage requirements of the voice service mode; Based on the body movements and / or gestures in the first image, the user's needs are analyzed to obtain the image service mode usage requirements; By integrating the usage requirements of the voice service mode and the usage requirements of the image service mode, the usage requirements of the target service mode are obtained. Based on the identity information and the usage requirements of the target service mode, the current target service mode is determined; Adjust the current target service mode based on historical usage; The historical usage includes evaluation information for the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image.

2. The service model recommendation method according to claim 1, characterized in that, The step of determining the user's identity information based on the acquired first voice and first image includes: Voiceprint extraction is performed on the first speech to obtain the user's voiceprint features; Based on the voiceprint features, the user's voice identity information is determined; Perform user identification on the first image to obtain the user's image identity information; The user's identity information is determined based on the voice identity information and the image identity information.

3. The service model recommendation method according to claim 1 or 2, characterized in that, The evaluation information is determined based on the following steps: Determine the user's second voice and / or second image; Perform user emotion analysis on the second voice and / or the second image to obtain the user's voice emotion and / or image emotion; Based on the voice emotion and / or image emotion, evaluation information for the previous target service mode is determined.

4. The service model recommendation method according to claim 3, characterized in that, The step of performing user emotion analysis on the second voice to obtain the user's voice emotion includes: Perform voice emotion recognition on the second speech to obtain the user's voice emotion recognition result; And / or, perform speech-to-text transcription on the second speech to obtain the transcribed text of the second speech; Semantic sentiment recognition is performed on the transcribed text to obtain the semantic sentiment recognition result of the user; Based on the voice emotion recognition results and / or the semantic emotion recognition results, the user's voice emotion is determined.

5. The service model recommendation method according to claim 3, characterized in that, The step of performing user sentiment analysis on the second image to obtain the user's image sentiment includes: Determine the face region and / or body movement region in the second image; Facial expression recognition is performed on the facial region to obtain the user's facial emotions; And / or, perform motion recognition on the limb movement area to obtain the user's emotional state; The user's image emotion is determined based on the facial expression emotion and / or the action emotion.

6. The service model recommendation method according to claim 1 or 2, characterized in that, The process of adjusting the current target service mode based on historical usage also includes: The control command indicated by the current target service mode is transmitted to the terminal corresponding to the user, so as to request the terminal to recommend the current target service mode to the user after receiving the control command.

7. A service mode recommendation device, characterized in that, include: A voice and image acquisition unit is used to acquire the user's first voice and first image; The target service mode determination unit is used to determine the user's identity information based on the acquired first voice and first image; Based on the first voice recording, determine the usage requirements of the voice service mode; Based on the body movements and / or gestures in the first image, the user's needs are analyzed to obtain the image service mode usage needs; the voice service mode usage needs and the image service mode usage needs are fused to obtain the target service mode usage needs; based on the identity information and the target service mode usage needs, the current target service mode is determined. An adjustment unit is used to adjust the current target service mode based on historical usage; wherein the historical usage includes evaluation information of the previous target service mode within a set time period, and the evaluation information is determined based on the user's second voice and / or second image.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the service mode recommendation method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the service mode recommendation method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Music recommendation method, apparatus, storage medium and terminal device

    CN109299318A

  • Media content recommendation method and device

    CN113574525A