Human-computer interaction method, device, equipment and medium

By combining voice and image data, using emotion analysis and expression recognition technology to identify users' emotional state and intentions, it solves the problem that machines find it difficult to accurately understand user needs, and improves the accuracy and user experience of reply.

CN120071971APending Publication Date: 2025-05-30CHONGQING JINKANG NEW ENERGY VEHICLE CO LTD

Patent Information

Application Number
CN202510327092.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the user interacts with the machine with voice, it is difficult for the machine to accurately understand the user's needs, resulting in inaccurate responses and reduce user experience.

Method used

By collecting the user's voice and images, combining speech emotion analysis and facial expression recognition technology, the user's emotional state is identified and matched with the speech intention calculation, the user's intention is determined to improve the accuracy of the reply.

Benefits of technology

It improves the accuracy of the machine to understand user needs, enhances the accuracy of reply and conversation quality, thereby improving the user's interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071971A_ABST
    Figure CN120071971A_ABST
Patent Text Reader

Abstract

The invention discloses a man-machine interaction method and device, equipment and a medium, and relates to the technical field of natural language processing, and the method comprises the steps: responding to a dialogue instruction triggered by a user, and collecting the voice of the user and an image of the user; the method comprises the following steps: analyzing voice of a user through a voice emotion analysis technology to obtain a first result, recognizing an image of the user through a facial expression recognition technology to obtain a second result, and obtaining an emotion state of the user according to the first result and the second result; judging whether the intention of the user is matched with the intention in the candidate intention set or not according to the voice of the user and the emotional state of the user; if the intention of the user is matched with the first intention in the candidate intention set, reply content corresponding to the first intention is replied to the user, and the method can improve the accuracy of a machine to understand user requirements, so that the accuracy of reply of the machine to the user is improved, and the interaction experience of the user is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and particularly to a human-computer interaction method, device, equipment and medium. Background Art

[0002] With the development of vehicle technology, vehicles are equipped with intelligent cockpit systems that support users to interact with vehicles by voice. During the process of voice interaction between users and machines (such as vehicles), there is a problem that machines cannot accurately understand users' needs, which in turn leads to inaccurate responses from machines to users and reduces the interaction experience of users. Summary of the Invention

[0003] This application provides a human-computer interaction method, device, equipment and medium, which can improve the accuracy of machines in understanding users' needs, thereby improving the accuracy of machine responses to users and enhancing the interaction experience of users.

[0004] To achieve the above object, this application adopts the following technical solutions: In a first aspect, this application provides a human-computer interaction method, including: In response to a dialogue instruction triggered by a user, collect the user's voice and the user's image; Analyze the user's voice through voice emotion analysis technology to obtain a first result, identify the user's image through facial expression recognition technology to obtain a second result, and obtain the user's emotional state according to the first result and the second result; Judge whether the user's intention matches the intention in the candidate intention set according to the user's voice and the user's emotional state; If the user's intention matches the first intention in the candidate intention set, reply to the user with the reply content corresponding to the first intention.

[0005] Optionally, judging whether the user's intention matches the intention in the candidate intention set includes: Calculate the matching degree between the user's intention and each intention in the candidate intention set; If the first intention with the highest matching degree with the user's intention is greater than the matching degree threshold, it is determined that the user's intention matches the first intention in the candidate intention set.

[0006] Optionally, the method further includes: If the first intention with the highest matching degree with the user's intention is less than or equal to the matching degree threshold, it is determined that the user's intention does not match any intention in the candidate intention set.

[0007] Optionally, the method further includes: If the user's intention does not match any of the intentions in the candidate intention set, the question content corresponding to the first intention is replied to the user.

[0008] Optionally, replying to the user with the reply content corresponding to the first intention includes: Determine the first text to be replied to the user according to the first intention; Determine the first text speech type and the first voice emotion type to be replied to the user according to the user's emotional state; Update the first text by using the first text speech type to obtain a second text; Process the second text by using the first voice emotion type to obtain a first target reply content; Reply the first target reply content to the user.

[0009] Optionally, replying to the user with the question content corresponding to the first intention includes: Determine the third text to be asked to the user according to the first intention; Determine the second text speech type and the second voice emotion type to be asked to the user according to the user's emotional state; Update the third text by using the second text speech type to obtain a fourth text; Process the fourth text by using the second voice emotion type to obtain a second target reply content; Reply the second target reply content to the user.

[0010] Optionally, the method further includes: Obtain the feedback information of the user for the conversation; Record the user's conversation preference according to the feedback information, where the conversation preference is used to record the corresponding relationship between the user's reference emotional state, reference text speech type, and reference voice emotion type.

[0011] In a second aspect, the present application provides a human-computer interaction device, and the device includes: An acquisition module, configured to acquire the user's voice and the user's image in response to a conversation instruction triggered by the user; An emotion recognition module, configured to analyze the user's voice through voice emotion analysis technology to obtain a first result, recognize the user's image through facial expression recognition technology to obtain a second result, and obtain the user's emotional state according to the first result and the second result; A judgment module, configured to judge whether the intention of the user matches the intention in the candidate intention set according to the voice of the user and the emotional state of the user; A dialogue module, configured to, if the intention of the user matches the first intention in the candidate intention set, reply to the user with the reply content corresponding to the first intention.

[0012] In a third aspect, the present application provides a computing device, including a memory and a processor; Wherein, one or more computer programs are stored in the memory, and the one or more computer programs include instructions; when the instructions are executed by the processor, the computing device is caused to execute the method according to any one of the first aspects.

[0013] In a fourth aspect, the present application provides a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute the method according to any one of the first aspects.

[0014] It can be seen from the above technical solutions that the present application has at least the following beneficial effects: The present application provides a human-computer interaction method, which includes: in response to a dialogue instruction triggered by a user, collecting the voice of the user and the image of the user, then using language emotion analysis technology to analyze the voice of the user to obtain a first result, using facial expression recognition technology to recognize the image of the user to obtain a second result, determining the emotional state of the user based on the first result and the second result, and then judging whether the intention of the user matches the intention in the candidate intention set based on the voice and emotional state of the user. If there is a first intention in the candidate intention set that matches the intention of the user, the user is replied with the reply content corresponding to the first intention. In this method, the dialogue with the user is not only based on the voice of the user (which can also be the text converted from the voice), but also the emotion of the user is considered, and the intention of the user is determined jointly based on the emotion and the voice. In a multi-modal manner, the determined intention of the user is more accurate, thereby improving the accuracy of the machine's understanding of the user's needs. And, when it is determined that the intention of the user is clear (there is a matching first intention), the machine replies to the user, further improving the quality of the dialogue, thereby improving the user's dialogue experience.

[0015] It should be understood that the description of technical features, technical solutions, beneficial effects or similar language in this application does not imply that all features and advantages can be realized in any single embodiment. On the contrary, it is understood that the description of features or beneficial effects means that specific technical features, technical solutions or beneficial effects are included in at least one embodiment. Therefore, the description of technical features, technical solutions or beneficial effects in this specification does not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions and beneficial effects described in the present embodiment can also be combined in any appropriate manner. Those skilled in the art will understand that the embodiment can be realized without one or more specific technical features, technical solutions or beneficial effects of a specific embodiment. In other embodiments, additional technical features and beneficial effects can also be identified in a specific embodiment that does not embody all embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A flowchart of a human-computer interaction method provided in an embodiment of the present application; Figure 2 A schematic diagram of a human-computer interaction device provided in an embodiment of the present application; Figure 3 A schematic diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0017] The terms "first", "second", "third", etc. in the specification of this application and the accompanying drawings are used to distinguish different objects rather than to limit a specific order.

[0018] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a specific way.

[0019] In order to make the description of the following embodiments clear and concise, a brief introduction to the related technology is first given: Human-computer interaction refers to the interactive communication between humans and computers through natural language. There are many application scenarios for human-computer interaction, such as smart cockpits, smart customer service, voice assistants, intelligent question-and-answer systems, etc.

[0020] Taking the smart cockpit scenario as an example, in this application scenario, the vehicle communicates with the user only based on the text converted from the user's voice. The data it relies on is relatively single, and there is an inability to accurately understand the user's needs, which leads to inaccurate responses from the vehicle to the user, poor user satisfaction, and reduced user conversation experience.

[0021] In view of this, an embodiment of the present application provides a human-computer interaction method, which can be executed by a processing device, which can be a terminal or a server. The terminal includes but is not limited to a smart phone, a tablet computer, a laptop computer, a personal digital assistant or a smart wearable device. The server can be a cloud server, for example, a central server in a central cloud computing cluster, or an edge server in an edge cloud computing cluster. Of course, the server can also be a server in a local data center. The local data center refers to a data center directly controlled by the user. In the human-computer interaction method provided in the embodiment of the present application, the processing device does not rely solely on the user's voice (or the text converted from the voice) to communicate with the user, but also considers the user's emotions, and determines the user's intention based on the emotions and voice. Through a multimodal approach, the determined user's intention is more accurate, thereby improving the accuracy of the processing device in understanding the user's needs. In addition, when it is determined that the user's first intention is clear, the processing device will reply to the user, further improving the quality of the conversation, thereby improving the user's interactive experience.

[0022] In order to make the technical solution of the present application clearer and easier to understand, the technical solution of the present application is introduced below in conjunction with the accompanying drawings. Figure 1 As shown, the figure is a flow chart of a human-computer interaction method provided in an embodiment of the present application, the method comprising: S101: A processing device collects a user's voice and an image of the user in response to a dialogue instruction triggered by the user.

[0023] The conversation instruction may be an instruction to wake up the voice conversation. In some examples, the user may trigger the conversation instruction by voice, for example, the user may say a preset voice, and the preset voice may be "Hello, Xiao A", where "Xiao A" is only an example. In other examples, the user may also trigger the conversation instruction by triggering a preset operation on a preset button, where the preset button may be a power button or a voice assistant button, and the preset operation may be pressing, clicking, etc. For example, the user may trigger the conversation instruction by pressing the voice assistant button.

[0024] After the user triggers the dialogue command, the processing device may respond to the dialogue command and collect the user's voice and the user's image. The user's voice may be the content of the user's speech, and the user's image may be an image including the user's facial expression. In other examples, the collected user's voice may include not only the user's voice information but also the noise in the environment, and the collected user's image may include not only the user's image information but also the image of the user's surrounding environment.

[0025] In some embodiments, the processing device may collect the user's speech through at least one sound pickup device (such as a microphone), and collect the user's image through at least one image acquisition device (such as a camera).

[0026] In some embodiments, after the processing device collects the user's image and the user's speech, it may align the data of different modalities (speech and image), for example, align them according to the time axis, so as to ensure the synchronization of multi-modal data.

[0027] S102. The processing device analyzes the user's speech through speech emotion analysis technology to obtain a first result, and recognizes the user's image through facial expression recognition technology to obtain a second result. Based on the first result and the second result, the emotional state of the user is obtained.

[0028] Speech emotion analysis technology is a technology that processes and analyzes speech signals to identify the emotional state of the speaker. Since speech emotion analysis technology is a relatively mature technical means, it will not be elaborated here. Among them, the first result may be the first emotional state obtained by analyzing the user's speech based on speech emotion analysis technology.

[0029] In some examples, the processing device may use speech-to-text technology to convert speech to text, then use large language models such as BERT and GPT for semantic understanding, and finally use emotion analysis APIs (application programming interfaces), emotion classification models based on CNN (Convolutional Neural Network) and RNN (Recurrent Neural Network) for emotion recognition.

[0030] Facial expression recognition technology is a technology that uses computer vision technology to analyze and recognize human facial expressions. This technology can automatically extract the features of facial expressions and classify them into different emotional states, such as expressions of happiness, sadness, anger, surprise, disgust, and fear. The second result may be the second emotional state obtained by recognizing the user's image based on facial expression recognition technology.

[0031] In some examples, the processing device may use OpenCV and deep learning frameworks (such as TensorFlow or PyTorch) to perform emotion recognition on facial expressions.

[0032] After obtaining the first result and the second result, the processing device may obtain the user's emotional state based on the first result and the second result. In some examples, the processing device may determine whether the first result and the second result are consistent. If the first result and the second result are consistent, the emotional state represented by the first result or the second result is used as the user's emotional state. If the first result and the second result are inconsistent, the processing device may obtain a voice emotion weight and an image emotion weight. If the voice emotion weight is greater than the image emotion weight, the emotional state represented by the first result is determined as the user's emotional state. If the image emotion weight is greater than the voice emotion weight, the emotional state represented by the second result is determined as the user's emotional state.

[0033] Among them, the voice emotion weight and the image emotion weight may be preset by the user, and the voice emotion weight and the image emotion weight are not equal.

[0034] In some other embodiments, the processing device may also use sample data to train an emotion recognition model. Among them, the sample data includes sample feature data and sample label data. Among them, the sample feature data may include sample voice and sample image, and the sample label data may include sample emotion data. Among them, the sample image may be a sample image sequence or a sample video.

[0035] After completing the training of the emotion recognition model using the sample data, the processing device may input the user's voice and image to the emotion recognition model. The trained emotion recognition model may output the user's emotional state. For example, the trained emotion recognition model may perform feature extraction on the above-mentioned multimodal data, perform feature extraction on the voice and the image respectively to obtain voice features and image features, and then fuse these two data to obtain a fused feature to represent the user state. Based on this fused feature, the output result of the emotion recognition model, that is, the emotional state, is obtained.

[0036] Among them, in the process of determining the user's emotional state, the processing device may also consider information such as the user's historical conversation record. For example, based on the historical conversation record, it can be determined that the user is angry because the question cannot be accurately answered, thereby increasing the weight of determining the user's emotional state as the angry state.

[0037] S103. The processing device determines whether the user's intention matches the intention in the candidate intention set according to the user's voice and the user's emotional state.

[0038] If the user's intention matches the intention in the candidate intention set, S104 is executed. If the user's intention does not match the intention in the candidate intention set, S105 is executed.

[0039] After obtaining the user's emotional state, the processing device can determine whether the user's intention matches the intention in the candidate intention set based on the user's speech and emotional state.

[0040] In some embodiments, the processing device can use sample cases and input them into the large model in the form of prompt engineering, and then verify them using standard data until the verification passes. Among them, the sample cases include multimodal data and corresponding true intentions, and the multimodal data includes speech data and emotional state data. The large model refers to a machine learning model with a Transformer or its variant as the architecture, having a large number of parameters (usually billions or even trillions of parameters) and a complex computing structure.

[0041] After the processing device fine-tunes the large model using the sample cases, the processing device can input the user's speech and emotional state into the fine-tuned large model, and the large model can output the user's intention.

[0042] Next, the processing device can calculate the matching degree between the user's intention and each intention in the candidate intention set. Exemplarily, the processing device can represent the text corresponding to the intention as a vector in the vector space. For example, the text is converted into a vector by using the bag-of-words model or TF-IDF, etc., and then the cosine similarity between the two vectors is calculated, and this cosine similarity is used as the matching degree.

[0043] In some other embodiments, in order to improve the accuracy of calculating the matching degree between intentions, the processing device can also obtain the grammatical roles (such as subject, predicate, object, etc.) of the words in the text corresponding to the intention and the dependency relationship with other words, so as to more accurately judge the meaning of the words in the text. For example, "apple" is the object in "I ate an apple", referring to a kind of fruit; while in "I want an iPhone", it is an attributive, referring to a brand. By analyzing the dependency relationship between "apple" and words such as "ate" and "phone" and the grammatical role of "apple", the different meanings of "apple" can be clarified, and then the intention matching degree can be judged more accurately.

[0044] Exemplarily, the grammatical role and the dependency relationship can be used as correction coefficients to correct the cosine similarity, and the corrected value is used as the matching degree. For example, when there are words with the same text but different grammatical roles, the correction coefficient can be less than 1, so as to reduce the cosine similarity; when there are words with the same text and the same grammatical role, the correction coefficient can be greater than or equal to 1, so as to increase the cosine similarity or keep the cosine similarity unchanged; when there are words with the same text but different dependency relationships, the correction coefficient can be less than 1, so as to reduce the cosine similarity; when there are words with the same text and the same dependency relationship, the correction coefficient can be greater than or equal to 1, so as to increase the cosine similarity or keep the cosine similarity unchanged.

[0045] Then, the processing device can find the first intent with the highest matching degree with the user's intent from the candidate intent set, and determine whether the matching degree between the first intent and the user's intent is greater than the matching degree threshold. If the matching degree between the first intent and the user's intent is greater than the matching degree threshold, it means that the user's intent matches the first intent in the candidate intent set; if the matching degree between the first intent and the user's intent is less than or equal to the matching degree threshold, it means that the user's intent does not match any intent in the candidate intent set. Among them, the matching degree threshold can be 70% or 80%. This application does not specifically limit the matching degree threshold.

[0046] S104. The processing device replies to the user with the reply content corresponding to the first intent.

[0047] If the user's intent matches the first intent in the candidate intent set, the processing device replies to the user with the reply content corresponding to the first intent.

[0048] In some embodiments, the processing device can pre-obtain an intent reply library, which includes the mapping relationship of reference intent, reference question content, and reference reply content. The intent reply library is shown in Table 1 below.

[0049] Table 1:

[0050] Taking the intent Figure 1 in Table 1 as the first intent as an example, when the processing device determines that the user's intent matches the intent Figure 1 in Table 1 above, it can determine the corresponding question content 1 and reply content 1 based on Table 1 above. For the convenience of understanding, the following takes the processing device determining that the user's intent matches the intent Figure 1 as an example for introduction. Since the processing device determines that the user's intent matches the intent Figure 1 , that is, the processing device can clarify the user's intent, and then can directly reply based on the user's intent. The processing device can directly reply to the user with the reply content 1 Figure 1 corresponding to the intent. Figure 1

[0051] In some other embodiments, after obtaining the reply content 1, the processing device can also adjust the reply content 1 based on the user's emotional state and then reply to the user, so as to improve the user's conversation experience.

[0052] ​Specifically, the processing device may first determine the first text to reply to the user based on the first intention, and the first text may be the above-mentioned reply content 1, and then determine the first text speech type and the first voice emotion type to reply to the user based on the user's emotional state. Among them, the processing device may pre-establish a mapping relationship between a reference emotional state, a reference text speech type, and a reference voice emotion type, that is, different emotional states correspond to different speech types and voice emotion types. Among them, the mapping relationship between the reference emotional state, the reference text speech type, and the reference voice emotion type is shown in Table 2 below.

[0053] Table 2:

[0054] Taking the user's emotional state as emotional state 1 in the above Table 2 as an example, the processing device can determine the speech type 1 and the voice emotion type 1 corresponding to the emotional state 1 based on the above Table 2. For example, when the user's emotional state is anger, the speech type corresponding to the angry emotional state is the soothing type, and the voice emotion type corresponding to the angry emotional state is the quiet type.

[0055] Next, the processing device can use the first text speech type to update the first text to obtain the second text, use the first voice emotion type to process the second text to obtain the first target reply content, and reply the first target reply content to the user.

[0056] For example, the first text may be "The current function is being optimized and improved, please be patient." When the first text is of a soothing type, the processing device may update the first text to a second text of "Thank you very much for your support and patience! The current function is being optimized and upgraded in full swing. We are well aware that this may cause you some inconvenience, but please rest assured that our team is doing our best to bring you a smoother, more convenient and efficient user experience, and will soon see you in a brand new look." When the first voice emotion type is a quiet type, the processing device processes the second text based on the quiet type to obtain the first target reply content in the form of replying to the user with the second text in a quiet tone.

[0057] S105: The processing device replies to the user with the question content corresponding to the first intention.

[0058] If the user's intention does not match any of the candidate intentions, the processing device replies to the user with the question content corresponding to the first intention.

[0059] As introduced in the above embodiments, the first intention has the highest matching degree with the user's intention, that is, the first intention is the closest to the user's intention. In this case, the processing device cannot clearly determine the user's intention. Therefore, instead of directly replying to the user, it asks the user again based on the first intention that is closest to the user's intention, so as to clarify the user's intention.

[0060] The following is an example. In some scenarios, the user says "Make it quieter". The processing device cannot clearly identify the user's true intention by recognizing "Make it quieter" and the user's emotional state. Therefore, the processing device can continue to ask the user. The first intention closest to the user's intention is to mute. Therefore, the processing device can ask the user "Do you need me to mute it for you?" (i.e., the question content corresponding to the first intention). In this way, when the user replies affirmatively (e.g., yes), the processing device can clarify that the user's intention is to mute, and then the processing device can execute instructions related to the mute intention. For example, it can control the vehicle head unit to mute. Further, if the user replies negatively (e.g., no), the intention that is the second closest to the user's intention is to close the window. The processing device can ask the user "Do you need me to close all the windows for you?" In this way, when the user replies affirmatively (e.g., yes), the processing device can clarify that the user's intention is to close the window, and then the processing device can execute instructions related to the intention of closing the window, and so on.

[0061] Of course, in the above scenario, the processing device can collect the noise inside the vehicle. When it is determined that the noise inside the vehicle exceeds the noise threshold, it can first ask the user "Do you need me to close all the windows for you?" because in most cases, the noise inside the vehicle is caused by wind noise after the windows are opened. Asking whether to close the windows first can help the user solve the problem more quickly, thereby improving the user's dialogue experience and interaction experience.

[0062] In some embodiments, after the processing device determines the third text to ask the user according to the first intention, it can also update the third text based on the user's emotional state to obtain the fourth text, and process the fourth text using the second voice emotion type to obtain the second target reply content, and reply the second target reply content to the user.

[0063] For example, if the user's emotional state is happy, the second text speech type corresponding to the emotional state is a friendly type, and the second voice emotion type corresponding to the emotional state is a relaxed type. The third text may be "Do you need help to close all the windows?" In the case where the second text speech type is a friendly type, the processing device may update the third text to a fourth text "It seems a bit windy outside, do you want me to help you close all the windows, so that you can have a more comfortable and relaxing time in the car?" In the case where the second voice emotion type is a relaxed type, the processing device processes the fourth text based on the relaxed type to obtain the second target reply content in the form of replying to the user with the second text in a relaxed tone.

[0064] In some embodiments, the processing device can obtain the user's feedback information on the dialogue, which can be explicit or invisible, wherein the explicit feedback information can be the user clicking the "satisfied" button or inputting satisfied feedback in text form, and the invisible feedback can be the user's pause, semantic change, reaction speed, etc. In addition, the feedback information can be the user's evaluation of the entire dialogue, or it can be an evaluation or suggestion for each reply of the processing device. In this way, the processing device can record the dialogue preferences of each user for different users, wherein the dialogue preferences are used to record the corresponding relationship between the user's reference emotional state, reference text speech type, and reference voice emotion type. In this way, the processing device can use different text speech types and voice emotion types for correction according to the emotional state of different users, so as to better adapt to each user and improve the dialogue experience and interaction satisfaction of each user.

[0065] Based on the above description, the embodiment of the present application provides a human-computer interaction method, in which the method does not rely solely on the user's voice (or the text converted from the voice) to communicate with the user, but also takes into account the user's emotions, and determines the user's intention based on the emotions and voice. Through a multimodal approach, the determined user's intention is more accurate, thereby improving the accuracy of the machine's understanding of user needs. In addition, the machine will only reply to the user when it is determined that the user's intention is clear (there is a matching first intention), which further improves the quality of the conversation and thus improves the user's conversation experience.

[0066] Furthermore, by introducing multi-modal data fusion technology (including various information such as voice, image, emotion, etc.), it is possible to comprehensively understand the user's needs and emotional state. Compared with traditional single-modal data, it can dynamically adjust the conversation strategy according to the user's real-time input and emotional changes, so as to provide more personalized and intelligent services, significantly improving the user's conversation experience; based on the fusion of multi-modal data, it can more accurately identify the user's true intention, and can make reasonable inferences whether in the case of clear intention or ambiguous intention. Especially when facing ambiguous user intentions, by combining various information such as the user's emotion, tone, and historical records, the content replied to the user can be dynamically adjusted, thereby further clarifying the user's intention and making the user interaction experience better; through emotion state recognition, not only can the emotional information in the user's voice be recognized, but also various emotional signals such as the user's facial expression and voice tone can be analyzed. In this way, the user's emotional changes (such as anger, anxiety, pleasure, etc.) can be accurately identified, and the emotional type of the replied voice (the tone of the reply) and the text conversation strategy type (the replied text) can be adjusted according to different emotional states, improving the user's emotional experience and avoiding emotional conflicts and discomfort; the conversation strategy generation method in this technical solution can be to generate conversation strategies that meet the user's needs based on a deep learning model (such as GPT), and continuously adjust the generation strategy through a real-time feedback and optimization module. Through continuous learning and optimization, the naturalness and fluency of the conversation strategy can be improved, enabling the user to feel more personalized care and attention during the conversation; with the introduction of multi-modal data fusion and deep learning models, it can better adapt to the changes of different users, different scenarios, and different situations. For example, by analyzing context information such as the user's emotional changes and topic transitions, the conversation strategy can be adjusted in real time to cope with complex or sudden conversation situations. In addition, based on the dynamic optimization mechanism of user feedback, the conversation strategy can be gradually optimized during long-term use to ensure that the conversation strategy always meets the user's needs and expectations; through multi-modal refined services and personalized response strategies, it can better meet the diverse needs of customers, improve the overall satisfaction of customers, and make the user feel that the machine can "understand" their needs and emotions, thereby increasing the trust and dependence on the machine, and enhancing the user's loyalty and usage stickiness; because it can automatically identify the user's needs and generate personalized conversation strategies, the need for manual customer service intervention is greatly reduced. The system can handle a large number of standardized problems and effectively handle complex emotional interactions, reducing the investment in human resources and improving the overall operation efficiency and cost-effectiveness.

[0067] As described above in conjunction with Figure 1 the human-computer interaction method provided by the embodiments of the present application has been introduced in detail. Next, the devices and equipment provided by the embodiments of the present application will be introduced with reference to the accompanying drawings.

[0068] As Figure 2As shown, this figure is a schematic diagram of a human-computer interaction device provided in an embodiment of the present application, and the device includes: The acquisition module 201 is used to acquire the user's voice and the user's image in response to a dialogue instruction triggered by the user; The emotion recognition module 202 is used to analyze the user's voice by using a voice emotion analysis technology to obtain a first result, recognize the user's image by using a facial expression recognition technology to obtain a second result, and obtain the user's emotional state according to the first result and the second result; A judgment module 203, configured to judge whether the user's intention matches the intention in the candidate intention set according to the user's voice and the user's emotional state; The dialogue module 204 is used to reply to the user with the reply content corresponding to the first intention in the candidate intention set if the user's intention matches the first intention in the candidate intention set.

[0069] Optionally, the judgment module 203 is specifically used to calculate the matching degree between the user's intention and each intention in the candidate intention set; if the first intention with the highest matching degree with the user's intention is greater than a matching degree threshold, it is determined that the user's intention matches the first intention in the candidate intention set.

[0070] Optionally, the judgment module 203 is further used to determine that the user's intention does not match any intention in the candidate intention set if the first intention with the highest matching degree with the user's intention is less than or equal to a matching degree threshold.

[0071] Optionally, the dialogue module 204 is further configured to reply to the user with the question content corresponding to the first intention if the user's intention does not match any intention in the candidate intention set.

[0072] Optionally, the dialogue module 204 is specifically used to determine a first text to reply to the user based on the first intention; determine a first text speech type and a first voice emotion type to reply to the user based on the emotional state of the user; update the first text using the first text speech type to obtain a second text; process the second text using the first voice emotion type to obtain a first target reply content; and reply to the user with the first target reply content.

[0073] Optionally, the dialogue module 204 is specifically configured to determine a third text for asking the user according to the first intention; determine a second text formula type and a second voice emotion type for asking the user according to the emotional state of the user; update the third text by using the second text formula type to obtain a fourth text; process the fourth text by using the second voice emotion type to obtain a second target reply content; and reply the second target reply content to the user.

[0074] Optionally, the device further includes an acquisition module and a recording module; The acquisition module is configured to acquire feedback information of the user for the dialogue; The recording module is configured to record the dialogue preference of the user according to the feedback information, where the dialogue preference is used to record the corresponding relationship between the reference emotional state, the reference text formula type, and the reference voice emotion type of the user.

[0075] The human-computer interaction device according to the embodiment of the present application may correspondingly execute the method described in the embodiment of the present application, and the above other operations and / or functions of each module / unit of the human-computer interaction device are respectively to implement Figure 1 the corresponding processes of the respective methods in the illustrated embodiments, and for the sake of brevity, will not be described in detail herein.

[0076] The embodiment of the present application further provides a computing device. As Figure 3 shown, this figure is a schematic diagram of a computing device provided by the embodiment of the present application. The computing device 300 includes a bus 301, a processor 302, a communication interface 303, and a memory 304. The processor 302, the memory 304, and the communication interface 303 communicate with each other through the bus 301.

[0077] The bus 301 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 3 only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.

[0078] The processor 302 can be any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0079] The communication interface 303 is used for external communication.

[0080] The memory 304 can include volatile memory, such as random access memory (RAM). The memory 304 can also include non-volatile memory, such as read-only memory (ROM), flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0081] Executable code is stored in the memory 304, and the processor 302 executes the executable code to perform the foregoing human-computer interaction method.

[0082] Specifically, in the case of implementing Figure 2 the illustrated embodiment, and Figure 2 when each module or unit of the human-computer interaction device described in the embodiment is implemented by software, the software or program code required to execute the functions of each module / unit in Figure 2 can be partially or entirely stored in the memory 304. The processor 302 executes the program code corresponding to each unit stored in the memory 304 to perform the foregoing human-computer interaction method.

[0083] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state drive). The computer-readable storage medium includes instructions that direct the computing device to execute the foregoing human-computer interaction method.

[0084] An embodiment of the present application also provides a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the processes or functions described in the embodiments of the present application are generated in whole or in part.

[0085] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, or data center to another website, computer, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.).

[0086] When the computer program product is executed by a computer, the computer executes any of the foregoing human-computer interaction methods. The computer program product may be a software installation package. In the case where any of the foregoing human-computer interaction methods needs to be used, the computer program product may be downloaded and executed on the computer.

[0087] The descriptions of the processes or structures corresponding to the above respective drawings each have their own emphases. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.

[0088] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application.

Claims

1. A human-computer interaction method, characterized in that: The method comprises: In response to a dialogue instruction triggered by a user, collecting the user's voice and the user's image; Analyze the user's voice by voice emotion analysis technology to obtain a first result, recognize the user's image by facial expression recognition technology to obtain a second result, and obtain the user's emotional state based on the first result and the second result; Determining whether the user's intention matches an intention in a candidate intention set according to the user's voice and the user's emotional state; If the user's intention matches the first intention in the candidate intention set, a reply content corresponding to the first intention is replied to the user.

2. The method according to claim 1, characterized in that Determining whether the user's intent matches an intent in the candidate intent set includes: Calculate the matching degree between the user's intent and each intent in the candidate intent set; If the first intention with the highest matching degree with the user's intention is greater than the matching degree threshold, it is determined that the user's intention matches the first intention in the candidate intention set.

3. The method according to claim 2, characterized in that The method further comprises: If the first intent with the highest matching degree with the user's intent is less than or equal to the matching degree threshold, it is determined that the user's intent does not match any intent in the candidate intent set.

4. The method according to claim 3, characterized in that The method further comprises: If the user's intention does not match any intention in the candidate intention set, the question content corresponding to the first intention is replied to the user.

5. The method according to claim 1, characterized in that: The replying to the user with the reply content corresponding to the first intention includes: Determining a first text to reply to the user according to the first intention; Determining a first text speech type and a first voice emotion type to reply to the user according to the emotional state of the user; Using the first text speech type to update the first text to obtain a second text; Using the first voice emotion type, processing the second text to obtain a first target reply content; Reply the first target reply content to the user.

6. The method according to claim 4, characterized in that The replying to the user the question content corresponding to the first intention includes: Determining a third text for asking a question to the user according to the first intention; Determining a second text speech type and a second voice emotion type for asking questions to the user according to the user's emotional state; Using the second text speech type to update the third text to obtain a fourth text; Using the second voice emotion type, processing the fourth text to obtain a second target reply content; Reply the second target reply content to the user.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Obtaining feedback information from the user regarding the conversation; The user's dialogue preference is recorded according to the feedback information, and the dialogue preference is used to record the corresponding relationship between the user's reference emotional state, reference text speech type, and reference voice emotion type.

8. A human-computer interaction device, characterized in that: The device comprises: A collection module, used for collecting the user's voice and the user's image in response to a dialogue instruction triggered by the user; An emotion recognition module is used to analyze the user's voice by using a voice emotion analysis technology to obtain a first result, recognize the user's image by using a facial expression recognition technology to obtain a second result, and obtain the user's emotional state according to the first result and the second result; A judgment module, used to judge whether the user's intention matches the intention in the candidate intention set according to the user's voice and the user's emotional state; The dialogue module is used to reply to the user with the reply content corresponding to the first intention in the candidate intention set if the user's intention matches the first intention in the candidate intention set.

9. A computing device, characterized in that including memory and processor; One or more computer programs are stored in the memory, and the one or more computer programs include instructions; when the instructions are executed by the processor, the computing device executes the method as claimed in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Chinese emotion evaluation unit extraction method

    CN110008477A

  • Man-computer interaction method and system and storage medium

    CN113434647A

  • Human-computer interaction method, device and system, electronic equipment and computer medium

    CN113822967A

  • Text matching method and device, storage medium and computer equipment

    CN114186558A

  • Interaction method of intelligent equipment and electronic equipment

    CN116072111A

Cited By

  • Vehicle control instruction determination method and device and vehicle

    CN120080867A

  • Vehicle control command determination method, device and vehicle

    CN120080867B