Conversation Processing Method and Device, Medium, and Electronic Device Based on Human-Computer Interaction

By identifying the word slots in the voice data and matching text data in the knowledge base, the problem of low accuracy of response results is solved, and the virtual digital people broadcast response information is realized, which improves the user experience.

CN113961680BActive Publication Date: 2025-07-18BOE INTELLIGENT IOT TECH CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111142535.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-28
Publication Date
2025-07-18
Estimated Expiration
2041-09-28

AI Technical Summary

Technical Problem

The existing human-computer interaction-based conversation processing method cannot recognize word slots in voice data, resulting in low accuracy of response results and the inability to broadcast through virtual digital people, reducing the user experience.

Method used

By recognizing the voice data to be recognized, the first word slot in the voice recognition result is extracted, and the corresponding text data is matched in the preset knowledge base, converted into voice data to be broadcast, fed back to the display terminal for broadcasting by the virtual digital person, and the text data is displayed on the display interface.

Benefits of technology

It improves the accuracy of text data, realizes broadcast of response information through virtual digital people, improves user experience, and ensures that users who cannot view text or listen to voice can also receive the corresponding response information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113961680B_ABST
    Figure CN113961680B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a conversation processing method, apparatus, medium, and electronic device based on human-computer interaction, and relates to the field of artificial intelligence technology. The method includes: when receiving the speech data to be recognized of the current user, recognizing the speech data to be recognized to obtain a speech recognition result, and segmenting the speech recognition result to obtain the first word slots included in the speech recognition result; matching the text data corresponding to the first word slots in a preset knowledge base, and converting the text data into speech data to be broadcast; feeding back the speech data to be broadcast and the text data to a display terminal, so that the display terminal calls a preset virtual digital person to broadcast the speech data to be broadcast, and displays the text data on the display interface of the display terminal. This method improves the accuracy of the text data.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] In the existing conversation processing method based on human-computer interaction, real-time emotion recognition can be performed on the user's conversation during the conversation process, and then corresponding extended responses can be selected according to the real-time emotion recognition results, and finally the response results including the extended responses are output to the customer.

[0003] However, it cannot recognize the slots in the voice data, and thus cannot generate corresponding response results according to the slots, resulting in a low accuracy of the response results; moreover, it cannot broadcast the response results through a virtual digital human, reducing the user experience.

[0004] Therefore, a new conversation processing method and device based on human-computer interaction need to be provided.

[0005] It should be noted that the information disclosed in the background art above is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0006] The purpose of the present disclosure is to provide a conversation processing method based on human-computer interaction, a conversation processing device based on human-computer interaction, a computer-readable storage medium, and an electronic device, so as to at least overcome to a certain extent the problem of low accuracy of response results caused by the limitations and defects of related technologies.

[0007] According to one aspect of the present disclosure, a conversation processing method based on human-computer interaction is provided, including:

[0008] When receiving the voice data to be recognized of the current user, recognizing the voice data to be recognized to obtain a voice recognition result, and segmenting the voice recognition result to obtain the first slots included in the voice recognition result;

[0009] Matching the text data corresponding to the first slots in a preset knowledge base, and converting the text data into voice data to be broadcast;

[0010] Feeding back the voice data to be broadcast and the text data to a display terminal, so that the display terminal calls a preset virtual digital human to broadcast the voice data to be broadcast, and displays the text data on the display interface of the display terminal.

[0011] In an exemplary embodiment of the present disclosure, before the step of receiving the voice data to be recognized of the current user, the conversation processing method based on human-computer interaction further includes:

[0012] When receiving the face image to be recognized of the current user, recognizing the face image to be recognized to obtain the face features to be recognized;

[0013] Calculate the user category of the current user according to the face features to be recognized, and when it is determined that the user category is a preset category, obtain preset text data corresponding to the preset category;

[0014] Feedback the preset text data and the preset broadcast data corresponding to the preset text data to the display terminal, so that the display terminal wakes up the preset virtual digital human to broadcast the preset broadcast data through the preset virtual digital human.

[0015] In an exemplary embodiment of the present disclosure, recognizing the face image to be recognized to obtain face features to be recognized includes:

[0016] Use a preset face detection and key point localization tool to detect the face area to be recognized in the face image to be recognized, and extract the face key points to be recognized in the face area to be recognized;

[0017] Calculate the facial attribute information of the face image to be recognized according to the face key points to be recognized, and obtain the face features to be recognized according to the facial attribute information;

[0018] Wherein, the facial attribute information includes one or more of age, gender, and facial expression.

[0019] In an exemplary embodiment of the present disclosure, calculating the user category of the current user according to the face features to be recognized includes:

[0020] Match the original face features corresponding to the face features to be recognized in a preset face database, and determine the user category of the current user according to the matching result;

[0021] Wherein, if the matching result is that there are original face features corresponding to the face features to be recognized in the preset face database, the user category of the current user is the preset category;

[0022] If the matching result is that there are no original face features corresponding to the face features to be recognized in the preset face database, the user category of the current user is a non-preset category.

[0023] In an exemplary embodiment of the present disclosure, the session processing method based on human-computer interaction further includes:

[0024] Obtain the matching time of the last match of the face features to be recognized in the preset face database, and calculate the time difference between the matching time of the last match and the current time;

[0025] When it is determined that the time difference is less than the first preset time threshold, control the display terminal to wake up the virtual digital human in a first preset manner; wherein, the first preset manner is used to indicate that the current session needs to be connected to the session corresponding to the matching time of the previous match.

[0026] In an exemplary embodiment of the present disclosure, matching the text data corresponding to the first slot in a preset knowledge base includes:

[0027] Match the slot code corresponding to the first slot in a preset knowledge base, and determine the dictionary of the first slot according to the slot code;

[0028] Determine whether the first slot is a complete slot according to the dictionary, and when it is determined that the first slot is a complete slot, match the text data corresponding to the first slot from a preset dialogue library.

[0029] In an exemplary embodiment of the present disclosure, the session processing method based on human-computer interaction further includes:

[0030] When it is determined that the first slot is an incomplete slot, determine the missing slot type of the first slot, and generate a question sentence according to the missing slot type;

[0031] Send the question sentence and the question voice data corresponding to the question sentence to the display terminal, so that the display terminal calls a preset virtual digital human to broadcast the sentence voice data, and display the question sentence on the display interface of the display terminal;

[0032] Receive the second slot obtained by the user's reply to the question sentence, and match the text data corresponding to the first slot and the second slot from a preset dialogue library.

[0033] In an exemplary embodiment of the present disclosure, identifying the to-be-identified voice data to obtain a voice recognition result includes:

[0034] Use the convolutional neural network included in a preset voice recognition model to extract the first local feature of the to-be-identified voice data;

[0035] Use the self-attention module included in the preset voice recognition model to calculate the first global feature of the to-be-identified voice data according to the first local feature respectively;

[0036] Use the fully connected layer included in the preset voice recognition model to classify the first global feature to obtain the voice recognition result of the to-be-identified voice data.

[0037] In an exemplary embodiment of the present disclosure, converting the text data into voice data to be broadcast includes:

[0038] Performing discretization processing on the text data based on a preset voice synthesis model to obtain a discrete voice synthesis result;

[0039] Calculating a voice synthesis probability distribution result according to the discrete voice synthesis result and a uniform distribution sampling result corresponding to the discrete voice synthesis result;

[0040] Performing voice synthesis on the voice synthesis probability distribution result based on a preset continuity function to obtain the voice data to be broadcast.

[0041] In an exemplary embodiment of the present disclosure, segmenting the voice recognition result to obtain a first word slot included in the voice recognition result includes:

[0042] Determining a scenario category required by the current user according to the voice recognition result, and segmenting the voice recognition result based on the scenario category to obtain a first word slot included in the voice recognition result;

[0043] Wherein, the scenario category includes one or more of a financial consultation scenario, a network point 3D tour guide scenario, and a product recommendation scenario.

[0044] In an exemplary embodiment of the present disclosure, the session processing method based on human-computer interaction further includes:

[0045] Establishing an association relationship between a user identifier of the current user, the text data, the face features to be recognized, and the user category;

[0046] Storing the text data in the knowledge base based on the association relationship, and storing the face features to be recognized in the face database based on the association relationship.

[0047] According to one aspect of the present disclosure, there is provided a session processing device based on human-computer interaction, including:

[0048] A voice recognition module, configured to recognize the voice data to be recognized when receiving the voice data to be recognized of the current user to obtain a voice recognition result, and segment the voice recognition result to obtain a first word slot included in the voice recognition result;

[0049] A voice data conversion module, configured to match text data corresponding to the first word slot in a preset knowledge base, and convert the text data into voice data to be broadcast;

[0050] The first voice data broadcast module is configured to feed back the voice data to be broadcast and the text data to a display terminal, so that the display terminal calls a preset virtual digital human to broadcast the voice data to be broadcast, and display the text data on the display interface of the display terminal.

[0051] According to one aspect of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program which, when executed by a processor, implements the method for session processing based on human-computer interaction described in any one of the above.

[0052] According to one aspect of the present disclosure, there is provided an electronic device, comprising:

[0053] a processor; and

[0054] a memory for storing executable instructions of the processor;

[0055] wherein the processor is configured to execute the method for session processing based on human-computer interaction described in any one of the above by executing the executable instructions.

[0056] For the method for session processing based on human-computer interaction provided by the embodiments of the present disclosure, on the one hand, by recognizing the voice data to be recognized to obtain a voice recognition result, segmenting the voice recognition result to obtain the first word slots included in the voice recognition result, and then matching the text data corresponding to the first word slots in a preset knowledge base, the problem in the prior art that due to the inability to recognize the word slots in the voice data, the corresponding response result cannot be generated according to the word slots, resulting in a low accuracy of the response result, is solved, and the accuracy of the text data is improved; on the other hand, by converting the text data into voice data to be broadcast, and then feeding back the voice data to be broadcast and the text data to a display terminal, so that the display terminal calls a preset virtual digital human to broadcast the voice data to be broadcast, and display the text data on the display interface of the display terminal, the broadcast of the text data corresponding to the voice data to be recognized by the virtual digital human is realized, so that users who cannot view the text or listen to the voice can receive the corresponding response information, further improving the user experience.

[0057] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0059] Figure 1 A flowchart schematically showing a method for session processing based on human-computer interaction according to an exemplary embodiment of the present disclosure.

[0060] FIG. 2(a) and FIG. 2(b) schematically show an example diagram of a rendering scene of a virtual digital human according to an exemplary embodiment of the present disclosure.

[0061] Figure 3 A block diagram schematically showing a system for session processing based on human-computer interaction according to an exemplary embodiment of the present disclosure.

[0062] Figure 4 A structural example diagram schematically showing a platform server according to an exemplary embodiment of the present disclosure.

[0063] Figure 5 A page scene example diagram schematically showing face management according to an exemplary embodiment of the present disclosure.

[0064] Figure 6 A page scene example diagram schematically showing dictionary management according to an exemplary embodiment of the present disclosure.

[0065] Figure 7 A page scene example diagram schematically showing slot management according to an exemplary embodiment of the present disclosure.

[0066] Figure 8 A page scene example diagram schematically showing skill management according to an exemplary embodiment of the present disclosure.

[0067] Figure 9 A page scene example diagram schematically showing multi-turn dialogue according to an exemplary embodiment of the present disclosure.

[0068] Figure 10 Schematically showing according to an exemplary embodiment of the present disclosure

[0069] Figure 11 A scene example diagram of a 3D holographic projection virtual digital human schematically showing according to an exemplary embodiment of the present disclosure.

[0070] Figure 12 A flowchart schematically showing another method for session processing based on human-computer interaction according to an exemplary embodiment of the present disclosure.

[0071] FIG. 13(a) and FIG. 13(b) schematically show an example diagram of a multi-screen linkage display scenario according to an exemplary embodiment of the present disclosure.

[0072] Figure 14 Schematically shown is a flowchart of another session processing method based on human-computer interaction according to an exemplary embodiment of the present disclosure.

[0073] Figure 15 Schematically shown is a block diagram of a session processing apparatus based on human-computer interaction according to an exemplary embodiment of the present disclosure.

[0074] Figure 16 Schematically shown is an electronic device for implementing the above-mentioned session processing method based on human-computer interaction according to an exemplary embodiment of the present disclosure. Detailed implementation manners

[0075] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.

[0076] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0077] With the rapid development of artificial intelligence, there are more and more AI (Artificial Intelligence) virtual digital humans. AI virtual digital humans form intelligent customer service, intelligent greeters, etc. for different industry application scenarios, and have developed into an application for the industry. In the specific application process, it can be mainly applied to technical industries such as large-scale voice Q&A, knowledge base management, natural language processing, and computer vision. The AI virtual digital human system applied to the financial industry mainly provides a professional business knowledge Q&A library for the financial industry, manages different business Q&As or daily Q&As through the knowledge base, and at the same time establishes 3D virtual human images that conform to different financial scenarios, providing scenario services such as intelligent customer service, intelligent marketing, business consultation, and interesting interaction.

[0078] In some application solutions of virtual digital humans, one is: describing real-time emotion recognition of customer conversations during the conversation process; selecting corresponding extended responses according to the real-time emotion recognition results; outputting the response results containing the extended responses to the customer; the other is: constructing a knowledge representation method and an order state machine for the context of interactive Q&A, and finally automatically generating a business logic tree by analyzing the historical conversation corpus between the intelligent customer service and the customer, realizing interactive Q&A based on a flowchart.

[0079] However, neither of the above methods involves slot recognition and the broadcast of virtual digital humans.

[0080] Based on this, in this exemplary embodiment, a session processing method based on human-computer interaction is first provided. This method can run on a server, a server cluster, a cloud server, etc.; of course, those skilled in the art can also run the method of the present disclosure on other platforms according to needs, and no special limitation is made in this exemplary embodiment. Refer to Figure 1 As shown, the session processing method based on human-computer interaction may include the following steps:

[0081] Step S110. When receiving the voice data to be recognized of the current user, recognize the voice data to be recognized to obtain a voice recognition result, and segment the voice recognition result to obtain the first slots included in the voice recognition result;

[0082] Step S120. Match the text data corresponding to the first slots in the preset knowledge base, and convert the text data into voice data to be broadcast;

[0083] Step S130. Feed back the voice data to be broadcast and the text data to the display terminal, so that the display terminal calls the preset virtual digital human to broadcast the voice data to be broadcast, and displays the text data on the display interface of the display terminal.

[0084] In the above-mentioned conversation processing method based on human-computer interaction, on the one hand, the speech recognition result is obtained by recognizing the speech data to be recognized, and the speech recognition result is segmented to obtain the first word slots included in the speech recognition result. Then, the text data corresponding to the first word slots is matched in the preset knowledge base, which solves the problem in the prior art that due to the inability to recognize the word slots in the speech data, the corresponding response result cannot be generated according to the word slots, resulting in a low accuracy rate of the response result, and improves the accuracy rate of the text data. On the other hand, by converting the text data into speech data to be broadcast, and then feeding the speech data to be broadcast and the text data back to the display terminal, so that the display terminal calls the preset virtual digital human to broadcast the speech data to be broadcast, and displays the text data on the display interface of the display terminal, realizing the broadcast of the text data corresponding to the speech data to be recognized by the virtual digital human, so that users who cannot view the text or listen to the speech can receive the corresponding response information, further improving the user experience.

[0085] Hereinafter, the conversation processing method based on human-computer interaction in the exemplary embodiments of the present disclosure will be explained and described in detail with reference to the accompanying drawings.

[0086] First, the application scenario and the invention purpose of the exemplary embodiments of the present disclosure will be explained and described. Specifically, the exemplary embodiments of the present disclosure provide a human-computer interaction conversation processing method based on a digital human knowledge base in a financial scenario. Among them, the digital human knowledge base includes a single-round question-and-answer skill management module, a multi-round question-and-answer management module, a dictionary management module, a word slot management module, and a face management module. Specifically, a knowledge base constructed based on intelligent speech, semantic technology, and face recognition algorithm with the intelligent network comprehensive management platform can be used to build a digital human system for financial domain knowledge Q&A through an ultra-realistic 8K digital human, and at the same time, the digital human can perform intelligent Q&A in different product forms.

[0087] Among them, for the virtual digital human, 3dsMax, AutoCAD and other modeling software can be used to create a digital human body model with real proportional relationships, and then the digital human body model is rendered and displayed through the software system. Specifically, the rendering display scene diagrams can refer to those shown in FIGS. 2(a) and 2(b). The present disclosure can combine the financial business Q&A knowledge base, voice interaction technology, semantic understanding technology, and 3D digital human technology to create a financial digital human that meets the requirements of scenarios such as business consultation Q&A, Q&A knowledge base management, and voice semantic analysis technology in the financial scenario, and meets the scenario requirements of business consultation, intelligent marketing, and interesting Q&A.

[0088] Secondly, the conversation processing system based on human-computer interaction involved in the exemplary embodiments of the present disclosure will be explained and described. Specifically, refer to Figure 3As shown, the conversation processing system based on human-computer interaction may include the current user 310, the display terminal (touch all-in-one computer) 320 where the virtual digital human is located, and the platform server 330; the display terminal where the virtual digital human is located is network-connected to the platform server.

[0089] Among them, the display terminal may be a display terminal based on Unity 3D. A high-definition camera, a power amplifier speaker, and a microphone array are provided on the display terminal, which are respectively used to collect the face image to be recognized of the current user, broadcast voice data, and receive the voice data to be recognized of the current user. Together with the server, they form the hardware layer in this conversation system.

[0090] At the same time, the virtual digital human can be used in various different scenarios such as intelligent greeting, business consultation, intelligent application, and interesting entertainment. It forms the business layer in the conversation system. The virtual digital human can also be used to implement voice interaction (dynamic voice broadcast, user chatting, voice interactive marketing, business knowledge Q&A) and virtual image display (3D simulation modeling, expression simulation, action display, and dressing change). It forms the software layer in the conversation system. Further, in the specific conversation processing process, the dynamic voice broadcast may include but is not limited to broadcasting various different voice data fed back by the platform server; user chatting means that when it is recognized that the voice data input by the current user but recognized is of the chatting category, corresponding single-round or multi-round conversations are output based on the user's voice input; voice interactive marketing means that when it is recognized that the voice data input by the current user but recognized is of the business consultation category, corresponding single-round or multi-round conversations are output based on the user's voice input; business knowledge Q&A means that when it is recognized that the voice data input by the current user but recognized is of the business knowledge consultation category, corresponding single-round or multi-round conversations are output based on the user's voice input.

[0091] Further, referring to Figure 4 As shown, the platform server may include a knowledge base 401, an automatic speech recognition model 402, a speech synthesis model 403, a natural language understanding (NLU, Natural Language Understanding) module 404, a speech artificial intelligence module 405, and a platform service module 406. Among them, the platform service module may include a knowledge base management module, an action management module, a face management module, and a data management module; the knowledge base may include a single-round Q&A skill management module, a multi-round Q&A management module, a dictionary management module, a slot management module, and a face management module. Specifically:

[0092] Face Management Module: The virtual digital human system (the display terminal where the virtual digital human is located) uses a camera to capture the face image to be recognized in real time. After processing through the face recognition algorithm, face attributes such as the age, gender, and expression of the current user are recognized. In addition, VIP user recognition can be performed. The face module is mainly used to preset the information of bank VIP users, and this user information can include name, nickname, gender, age, and avatar information. When the camera recognizes that the current user is a VIP customer, the digital human will conduct a customized voice broadcast for welcoming guests; for example, "Dear Ms. Wang, hello"; at the same time, exclusive product recommendation services can also be provided for this VIP user on the display interface to achieve personalized services. Among them, the specific example diagram of the face management module page can be referred to Figure 5 as shown.

[0093] Dictionary Management Module: It can be used to manage the multi-round dialogue dictionary. For example, three core words, namely Beijing, Shanghai, and Tianjin, form the city name dictionary; another example is that core words such as credit card, debit card, debit card, and VIP card can form the bank card category dictionary, etc. Among them, the specific example diagram of the personalized dictionary management page can be referred to Figure 6 as shown.

[0094] Slot Management Module: It can be used to manage the slots involved in the multi-round dialogue scenario, including slot names, slot codes, and the dictionaries corresponding to the slots, etc. Among them, the specific example diagram of the slot management page can be referred to Figure 7 as shown.

[0095] Skill Management Module: Skill management includes the knowledge base for single-round Q&A. Different skills can be set according to different types of questions, such as financial business consultation skills, 3D branch navigation skills, financial product recommendation skills, etc. When a customer asks a question through voice, such as recommending a bank card, asking about the weather, chatting, etc., the intelligent voice module will convert the voice into text and send it to the intelligent branch comprehensive management platform. After retrieving the answer through skill management, the text answer will be synthesized into voice, and finally, it will be broadcast in the form of text + voice. At the same time, for the questions included in the skill management module, similar question methods and question answers and other fields can be added, and it can also be clicked to publish after being manually input or imported as a whole data, and the knowledge base will become effective. Among them, the specific example diagram of the skill management page can be referred to Figure 8 as shown.

[0096] Multi-round Dialogue Management Module: It can manage the multi-round Q&A knowledge base by constructing a multi-round dialogue scenario. Specifically, the bound slots and the number of slots can be selected, and the standard question method and similar question methods can be supplemented. By judging whether the question contains slots, the corresponding answer can be matched and pushed. Among them, the specific example diagram of the multi-round dialogue page can be referred to Figure 9 as shown.

[0097] Next, in conjunction with FIGS. 2- Figure 9 Each step involved in the session processing method based on human-computer interaction in the exemplary embodiments of the present disclosure will be explained and described in detail.

[0098] In an exemplary embodiment of the session processing method based on human-computer interaction in the present disclosure:

[0099] In step S110, when the speech data to be recognized of the current user is received, the speech data to be recognized is recognized to obtain a speech recognition result, and the speech recognition result is segmented to obtain the first word slots included in the speech recognition result.

[0100] In this exemplary embodiment, when the display terminal receives the speech data to be recognized of the current user, the speech data can be sent to the platform server so that the platform server recognizes the speech data to be recognized to obtain a speech recognition result. Among them, recognizing the speech data to be recognized to obtain a speech recognition result may specifically include: First, the first local feature of the speech data to be recognized is extracted by using the convolutional neural network included in the preset speech recognition model; Second, the first global feature of the speech data to be recognized is calculated respectively according to the first local feature by using the self-attention module included in the preset speech recognition model; Finally, the first global feature is classified by using the fully connected layer included in the preset speech recognition model to obtain the speech recognition result of the speech data to be recognized.

[0101] It should be supplemented and explained here that the above-mentioned preset speech recognition model may be an ASR (Automatic Speech Recognition) model, and the ASR model may include 3 layers of convolutional neural network (CNN), 10 layers of self-attention module (SAB), and 2 layers of fully connected layer (FC). Of course, in the actual application process, the specific number of convolutional neural network, self-attention module and fully connected layer can also be adaptively adjusted according to actual needs, and this example does not make special limitations on this.

[0102] Further, after obtaining the speech recognition result, the speech recognition result can be segmented to obtain the first slot included in the speech recognition result. Among them, the segmentation process may specifically include: determining the scenario category required by the current user according to the speech recognition result, and segmenting the speech recognition result based on the scenario category to obtain the first slot included in the speech recognition result; wherein, the scenario category includes one or more of a financial consultation scenario, a network point 3D tour scenario, and a product recommendation scenario. For example, if the speech recognition result is: What are the credit card activities? It can be determined that the scenario category required by the current user is the bank card preferential activity scenario, and then the speech recognition result can be segmented based on this scenario category, and the first slots obtained are: credit card and activity.

[0103] In step S120, match the text data corresponding to the first slot in the preset knowledge base, and convert the text data into the speech data to be broadcast.

[0104] In the present exemplary embodiment, first, match the text data corresponding to the first slot in the preset knowledge base. Specifically, it may include: first, match the slot code corresponding to the first slot in the preset knowledge base, and determine the dictionary of the first slot according to the slot code; secondly, determine whether the first slot is a complete slot according to the dictionary, and when it is determined that the first slot is a complete slot, match the text data corresponding to the first slot from the preset dialogue library.

[0105] Further, when it is determined that the first slot is an incomplete slot, determine the missing slot type of the first slot, and generate a question sentence according to the missing slot type; send the question sentence and the question speech data corresponding to the question sentence to the display terminal, so that the display terminal calls the preset virtual digital human to broadcast the sentence speech data, and display the question sentence on the display interface of the display terminal; receive the second slot obtained by the user's reply to the question sentence, and match the text data corresponding to the first slot and the second slot from the preset dialogue library.

[0106] That is to say, after obtaining the first slot, the slot code corresponding to the first slot can be matched in the knowledge base. For example, the slot name of "credit card" is "credit_card", and the bank card type is "card_type", etc. Secondly, according to the slot name, the dictionary of the first slot is determined, and according to the slot type included in the question sentence corresponding to the first slot in the dictionary, it is determined whether the first slot is a complete slot. If so, the above text data can be directly determined. If not, the missing slot type needs to be supplemented. For example, according to the dictionary corresponding to the slot name, it can be determined that if the text data corresponding to the first slot needs to be matched, the specific city name is missing. Therefore, a question sentence "Which city's credit card activities do you want to consult?" can be generated according to the missing slot type. Then, the question sentence is converted into sentence voice data and sent to the display terminal. When the display terminal receives the sentence voice data and the question sentence, it can broadcast and display them, and then receive the second slot of the user's reply to the question sentence, which can be "Beijing" for example, and finally match the corresponding text data. For example, the matched text data can be: The credit card activities in Beijing include buy one get one free at KFC.

[0107] Secondly, after obtaining the text data, the text data can be converted into voice data to be broadcast. Specifically, it can include: First, the text data is discretized based on a preset speech synthesis model to obtain a discrete speech synthesis result. Secondly, according to the discrete speech synthesis result and the uniform distribution sampling result corresponding to the discrete speech synthesis result, the speech synthesis probability distribution result is calculated. Finally, based on a preset continuity function, the speech synthesis probability distribution result is synthesized into speech to obtain the voice data to be broadcast. Among them, the speech synthesis model (TTS model) involved in this example embodiment can be, for example, a fully convolutional speech synthesis model, such as the UFANS fully convolutional speech synthesis model. Of course, it can also be other fully convolutional speech synthesis models, and this example does not make special restrictions on this.

[0108] In step S130, the voice data to be broadcast and the text data are fed back to the display terminal, so that the display terminal calls a preset virtual digital person to broadcast the voice data to be broadcast and displays the text data on the display interface of the display terminal.

[0109] Specifically, after obtaining the voice data to be broadcast, the voice data to be broadcast and the text data are fed back to the display terminal, so that the display terminal and the virtual digital person set in the display terminal can broadcast and display the voice data to be broadcast and the text data in the form of voice + text, which is convenient for users with hearing or vision impairments to use.

[0110] It should be noted here that the product form of the virtual digital human can be based on the knowledge base built by the intelligent network comprehensive management platform, combined with the digital human front-end multimedia interaction software, and then form an overall digital human system, which is applied in multiple scenarios such as intelligent welcome, 3D navigation, product recommendation, and marketing promotion in bank branches. The specific product display can be through transparent all-in-one machine display cabinets, vertical touch all-in-one machines, 3D holographic projection, etc. In the specific display process, through the AI virtual digital human software, 3D virtual image display can be carried out, dressing can be performed, functions such as dynamic voice broadcast, chatting with customers, and voice interactive marketing can be realized; at the same time, the interaction forms of the virtual digital human include single-screen display through transparent all-in-one machine display cabinets and vertical touch all-in-one machines, 3D holographic projection display, and multi-screen linkage display, etc.

[0111] Figure 10 Schematically shows another session processing method based on human-computer interaction according to an exemplary embodiment of the present disclosure. Refer to Figure 10 As shown, the session processing method based on human-computer interaction may further include the following steps:

[0112] Step S1010, when receiving the to-be-recognized face image of the current user, recognize the to-be-recognized face image to obtain to-be-recognized face features.

[0113] In this exemplary embodiment, when receiving the to-be-recognized face image of the current user, first, use a preset face detection and key point localization tool to detect the to-be-recognized face area of the to-be-recognized face image, and extract the to-be-recognized face key points of the to-be-recognized face image in the to-be-recognized face area; secondly, according to the to-be-recognized face key points, calculate the face attribute information of the to-be-recognized face image, and obtain the to-be-recognized face features according to the face attribute information; wherein, the face attribute information includes one or more of age, gender, and facial expression. Among them, the face detection and key point localization tool can be, for example, ibug-68, and of course it can also be other localization tools, and this example does not make special limitations on this.

[0114] Step S1020, calculate the user category of the current user according to the to-be-recognized face features, and when determining that the user category is a preset category, obtain preset text data corresponding to the preset category.

[0115] In this exemplary embodiment, first, the user category of the current user is calculated according to the face features to be recognized. Specifically, it may include: matching the original face features corresponding to the face features to be recognized in a preset face database, and determining the user category of the current user according to the matching result; wherein, if the matching result is that there are original face features corresponding to the face features to be recognized in the preset face database, the user category of the current user is the preset category; if the matching result is that there are no original face features corresponding to the face features to be recognized in the preset face database, the user category of the current user is a non-preset category. It should be added here that in the specific matching process, it can be implemented by calculating the Euclidean distance, cosine value, etc. between the face features to be recognized and the original face features.

[0116] Secondly, after obtaining the user category of the current user, the corresponding preset text data can be matched. For example, when the current user is a VIP user, the preset text data may be: Hello, Ms. Wu, welcome to XX Bank, etc.; when the current user is not a VIP user, the preset text data may be: Hello, welcome to XX Bank. What can I do for you? In the specific use process, the corresponding preset text data can be configured according to actual needs, and this example does not make special limitations on this.

[0117] Step S1030, feedback the preset text data and the preset broadcast data corresponding to the preset text data to the display terminal, so that the display terminal wakes up the preset virtual digital human, and the preset virtual digital human broadcasts the preset broadcast data.

[0118] Specifically, the broadcast principle can be as follows: Deploy the virtual digital human software on a single-screen host, wake up the virtual digital human through face recognition, and the digital human conducts a welcome broadcast (broadcasts the preset broadcast data). After the broadcast is completed, the current user can conduct a voice interaction with the virtual digital human, including: five scenarios of chatting, weather, network point navigation, financial product recommendation, and financial business consultation. At the same time, for different scenarios, different 3D UIs are displayed on the software side. In addition to voice interaction, the digital human also supports touch click operations. For example, for financial scenario recommendations, specific financial products can be displayed through touch clicks; of course, the virtual digital human can also be displayed in the form of a 3D holographic projection visit. The 3D holographic projection display uses a regular square pyramid for holographic projection to create an immersive interactive experience. The 3D projection digital human can also conduct voice interaction, and different 3D UI scenarios are switched and displayed; at the same time, it can be paired with Kinect for somatosensory interaction. The digital human supports the imitation of the same actions to increase the interesting interactivity; among them, the 3D holographic projection digital human can be specifically as Figure 11 shown.

[0119] Figure 10 In the schematically illustrated conversation processing method based on human-computer interaction, on the one hand, an exclusive welcome message can be provided for VIP users, thereby enhancing the user experience; on the other hand, when a face image is detected, the virtual digital human can be awakened, thereby avoiding the problem of low user satisfaction caused by the failure to receive users in a timely manner.

[0120] Figure 12 Another conversation processing method based on human-computer interaction according to an exemplary embodiment of the present disclosure is schematically illustrated. Refer to Figure 12 As shown, the conversation processing method based on human-computer interaction may further include the following steps:

[0121] Step S1210, obtaining the matching time of the last match of the to-be-recognized face feature in the preset face database, and calculating the time difference between the matching time of the last match and the current time;

[0122] Step S1220, when determining that the time difference is less than a first preset time threshold, controlling the display terminal to wake up the virtual digital human in a first preset manner; wherein, the first preset manner is used to indicate that the current conversation needs to be connected to the conversation corresponding to the matching time of the last match.

[0123] Hereinafter, steps S1210 and S1220 will be explained and described. Specifically, by calculating the time difference, it can be obtained whether the current user has been transferred to the current display after being recognized successfully on other display terminals; for example, if the time difference is less than the first preset time threshold, it means that the current user has a conversation behavior with the virtual digital human in other display terminals. Then, the virtual digital human can control the appearance manner of the virtual digital human based on the positional relationship between the display terminal where the last conversation with the current user was conducted and the current display terminal; for example, from left to right, from right to left, from top to bottom, or from bottom to top, etc. This example does not make special restrictions on this. That is, Figure 12 The illustrated conversation processing method based on human-computer interaction realizes the multi-screen linkage display of the virtual digital human. Further, through the multi-screen linkage display, an immersive experience of the digital human in bank branches can be realized. A digital human software is deployed on each screen, and a face recognition camera is configured. For example, when the current user is recognized successfully on the No. 1 display terminal, the virtual digital human is awakened for voice interaction; when the current user leaves the No. 1 display terminal and the virtual digital human in the No. 1 display terminal exits and hides, and walks to the No. 2 screen, and the face is recognized successfully by the No. 2 camera, at this time, the virtual digital human enters from the direction of the display interface of the No. 1 display terminal with UI special effects, so that the current user can continue to have voice interaction with the virtual digital human on the No. 2 display terminal, and so on; among them, the specific application scenario diagrams can refer to those shown in FIGS. 13(a) and 13(b).

[0124] In a specific implementation process, when the same current user moves from the first display terminal to the second display terminal, the camera mounted on the second display terminal will compare with the original facial features. If it matches the user information of the most recent time, including the text information of the voice interaction with the digital human, it supports continuous multi-round conversations. When the current user 1 leaves the first display terminal and another user 2 enters the second display terminal, the camera mounted on the second display terminal recognizes it as a new user, and the digital human wake-up voice question and answer is normally carried out. That is, when it is determined that the time difference is greater than or equal to the first preset time threshold, step S1030 is directly executed without waking up the virtual digital human through the first preset method.

[0125] Figure 14 Schematically shows another session processing method based on human-computer interaction according to an exemplary embodiment of the present disclosure. Refer to Figure 14 As shown, the session processing method based on human-computer interaction may further include the following steps:

[0126] Step S1410, establishing an association relationship between the user identifier of the current user, the text data, the facial features to be recognized, and the user category;

[0127] Step S1420, storing the text data into the knowledge base based on the association relationship, and storing the facial features to be recognized into the face database based on the association relationship.

[0128] Hereinafter, steps S1410 and S1420 will be explained and described. Specifically, first, a user identifier corresponding to the current user is generated (if it is a new user, a corresponding user identifier is generated, and the user identifier can be generated according to information such as facial features, name, and gender). Then, an association relationship between the user identifier, the text data, the facial features to be recognized, and the user category is established and stored, which can facilitate the next matching.

[0129] So far, it can be known that the session processing method based on human-computer interaction provided by the exemplary embodiment of the present disclosure can provide a three-dimensional digital human image in a three-dimensional perspective, which is different from the presentation method in the form of a 2D digital human video. The 3D real-time rendering method can support various product forms of the digital human: transparent display cabinets, vertical touch all-in-ones, and holographic projections. Moreover, through face recognition technology, the customer's gender, age, and VIP customer recognition can be achieved to achieve precise marketing for each individual. Further, through the intelligent network comprehensive management platform, a large amount of knowledge bases are constructed, enabling the digital human to have professional knowledge in various industries and facilitating the expansion from the financial field to other industry fields such as transportation, politics, education, and parks.

[0130] The exemplary embodiment of the present disclosure also provides a session processing device based on human-computer interaction. Refer toFigure 15 As shown in the figure, the conversation processing device based on human-computer interaction may include a speech recognition module 1510, a speech data conversion module 1520, and a first speech data broadcast module 1530. Among them:

[0131] The speech recognition module 1510 may be configured to, when receiving the speech data to be recognized of the current user, recognize the speech data to be recognized to obtain a speech recognition result, and segment the speech recognition result to obtain a first word slot included in the speech recognition result;

[0132] The speech data conversion module 1520 may be configured to match text data corresponding to the first word slot in a preset knowledge base, and convert the text data into speech data to be broadcast;

[0133] The first speech data broadcast module 1530 may be configured to feedback the speech data to be broadcast and the text data to a display terminal, so that the display terminal invokes a preset virtual digital person to broadcast the speech data to be broadcast, and display the text data on the display interface of the display terminal.

[0134] In an exemplary embodiment of the present disclosure, the conversation processing device based on human-computer interaction may further include:

[0135] A face recognition module, which may be configured to, when receiving the face image to be recognized of the current user, recognize the face image to be recognized to obtain a face feature to be recognized;

[0136] A user category calculation module, which may be configured to calculate the user category of the current user according to the face feature to be recognized, and when determining that the user category is a preset category, obtain preset text data corresponding to the preset category;

[0137] The second speech data broadcast module feeds back the preset text data and preset broadcast data corresponding to the preset text data to the display terminal, so that the display terminal wakes up the preset virtual digital person to broadcast the preset broadcast data through the preset virtual digital person.

[0138] In an exemplary embodiment of the present disclosure, recognizing the face image to be recognized to obtain a face feature to be recognized includes:

[0139] Using a preset face detection and key point localization tool, detecting a face region to be recognized of the face image to be recognized, and extracting key points of the face to be recognized of the face image to be recognized in the face region to be recognized;

[0140] Calculate the facial attribute information of the to-be-recognized face image according to the to-be-recognized facial key points, and obtain the to-be-recognized face features according to the facial attribute information;

[0141] Among them, the facial attribute information includes one or more of age, gender, and facial expression.

[0142] In an exemplary embodiment of the present disclosure, calculating the user category of the current user according to the to-be-recognized face features includes:

[0143] Match the original face features corresponding to the to-be-recognized face features in a preset face database, and determine the user category of the current user according to the matching result;

[0144] Among them, if the matching result is that there are original face features corresponding to the to-be-recognized face features in the preset face database, the user category of the current user is the preset category;

[0145] If the matching result is that there are no original face features corresponding to the to-be-recognized face features in the preset face database, the user category of the current user is a non-preset category.

[0146] In an exemplary embodiment of the present disclosure, the session processing device based on human-computer interaction may further include:

[0147] A time difference calculation module, which can be used to obtain the matching time of the last match of the to-be-recognized face features in the preset face database, and calculate the time difference between the matching time of the last match and the current time;

[0148] A first control module, which can be used to control the display terminal to wake up the virtual digital person in a first preset manner when it is determined that the time difference is less than a first preset time threshold; wherein, the first preset manner is used to indicate that the current session needs to be continued with the session corresponding to the matching time of the last match.

[0149] In an exemplary embodiment of the present disclosure, matching the text data corresponding to the first slot in a preset knowledge base includes:

[0150] Match the slot code corresponding to the first slot in a preset knowledge base, and determine the dictionary of the first slot according to the slot code;

[0151] Determine whether the first slot is a complete slot according to the dictionary, and when it is determined that the first slot is a complete slot, match the text data corresponding to the first slot in a preset dialogue library.

[0152] In an exemplary embodiment of the present disclosure, the session processing device based on human-computer interaction may further include:

[0153] A question sentence generation module, which can be used to determine the type of the missing word slot when determining that the first word slot is an incomplete word slot, and generate a question sentence according to the type of the missing word slot;

[0154] A third voice data broadcast module, which can be used to send the question sentence and the question voice data corresponding to the question sentence to the display terminal, so that the display terminal calls a preset virtual digital person to broadcast the sentence voice data, and display the question sentence on the display interface of the display terminal;

[0155] A text data matching module, which can be used to receive the second word slot obtained by the user's reply to the question sentence, and match the text data corresponding to the first word slot and the second word slot from a preset dialogue library.

[0156] In an exemplary embodiment of the present disclosure, the speech recognition result obtained by recognizing the to-be-recognized speech data includes:

[0157] Extracting a first local feature of the to-be-recognized speech data by using a convolutional neural network included in a preset speech recognition model;

[0158] Calculating a first global feature of the to-be-recognized speech data respectively according to the first local feature by using a self-attention module included in the preset speech recognition model;

[0159] Classifying the first global feature by using a fully connected layer included in the preset speech recognition model to obtain the speech recognition result of the to-be-recognized speech data.

[0160] In an exemplary embodiment of the present disclosure, converting the text data into to-be-broadcast speech data includes:

[0161] Performing discretization processing on the text data based on a preset speech synthesis model to obtain a discrete speech synthesis result;

[0162] Calculating a speech synthesis probability distribution result according to the discrete speech synthesis result and a uniform distribution sampling result corresponding to the discrete speech synthesis result;

[0163] Performing speech synthesis on the speech synthesis probability distribution result based on a preset continuity function to obtain the to-be-broadcast speech data.

[0164] In an exemplary embodiment of the present disclosure, segmenting the speech recognition result to obtain the first word slot included in the speech recognition result includes:

[0165] Determine the scenario category required by the current user according to the speech recognition result, and segment the speech recognition result based on the scenario category to obtain the first slot included in the speech recognition result;

[0166] Among them, the scenario category includes one or more of a financial consultation scenario, a network 3D tour scenario, and a product recommendation scenario.

[0167] In an exemplary embodiment of the present disclosure, the session processing device based on human-computer interaction may further include:

[0168] An association relationship establishment module, which can be used to establish an association relationship between the user identifier of the current user, the text data, the face feature to be recognized, and the user category;

[0169] A data storage module, which can be used to store the text data into the knowledge base based on the association relationship, and store the face feature to be recognized into the face database based on the association relationship.

[0170] The specific details of each module in the above session processing device based on human-computer interaction have been described in detail in the corresponding session processing method based on human-computer interaction, so they will not be repeated here.

[0171] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into being embodied by multiple modules or units.

[0172] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that these steps must be executed in this specific order, or that all the steps shown must be executed to achieve the desired result. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution, etc.

[0173] In an exemplary embodiment of the present disclosure, an electronic device capable of implementing the above method is also provided.

[0174] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.

[0175] Reference will now be made to Figure 16 to describe the electronic device 1600 according to such an embodiment of the present disclosure. Figure 16 The shown electronic device 1600 is merely an example and should not impose any limitation on the functions and the scope of use of the embodiments of the present disclosure.

[0176] As Figure 16 shown, the electronic device 1600 is presented in the form of a general-purpose computing device. The components of the electronic device 1600 may include, but are not limited to: at least one of the above-mentioned processing units 1610, at least one of the above-mentioned storage units 1620, a bus 1630 connecting different system components (including the storage unit 1620 and the processing unit 1610), and a display unit 1640.

[0177] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 1610, so that the processing unit 1610 executes the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification. For example, the processing unit 1610 can execute steps S110 as shown in Figure 1 : when receiving the voice data to be recognized of the current user, recognizing the voice data to be recognized to obtain a voice recognition result, and segmenting the voice recognition result to obtain the first word slot included in the voice recognition result; step S120: matching the text data corresponding to the first word slot in a preset knowledge base, and converting the text data into voice data to be broadcast; step S130: feeding back the voice data to be broadcast and the text data to a display terminal, so that the display terminal calls a preset virtual digital human to broadcast the voice data to be broadcast, and displays the text data on the display interface of the display terminal.

[0178] The storage unit 1620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 16201 and / or a cache storage unit 16202, and may further include a read-only storage unit (ROM) 16203.

[0179] The storage unit 1620 may also include a program / utility 16204 having a set (at least one) of program modules 16205. Such program modules 16205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0180] The bus 1630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0181] The electronic device 1600 may also communicate with one or more external devices 1700 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 1600, and / or communicate with any device that enables the electronic device 1600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 1650. Moreover, the electronic device 1600 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1660. As shown in the figure, the network adapter 1660 communicates with other modules of the electronic device 1600 through the bus 1630. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0182] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by the way of software combined with necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.

[0183] In an exemplary embodiment of the present disclosure, there is also provided a computer-readable storage medium, on which a program product capable of implementing the above-described method of this specification is stored. In some possible implementation manners, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0184] A program product for implementing the above method according to an embodiment of the present disclosure is described. It may be in the form of a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the present disclosure is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0185] The program product may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0186] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal may take various forms, including but not limited to an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0187] The program code contained on the readable medium may be transmitted by any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the above.

[0188] Program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., through the Internet using an Internet service provider).

[0189] In addition, the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed, for example, synchronously or asynchronously in multiple modules.

[0190] Other embodiments of the present disclosure will be readily apparent to those skilled in the art after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not invented by the present disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the claims.

Claims

1. A conversation processing method based on human-computer interaction, characterized in that Including: When receiving the speech data to be recognized of the current user, recognizing the speech data to be recognized to obtain a speech recognition result, determining the scenario category required by the current user according to the speech recognition result, and segmenting the speech recognition result based on the scenario category to obtain the first word slot included in the speech recognition result; wherein, the scenario category includes one or more of a financial consultation scenario, a network point 3D navigation scenario, and a product recommendation scenario; Matching the text data corresponding to the first word slot in a preset knowledge base, and converting the text data into speech data to be broadcast; wherein, the text data is obtained by the following method: matching the word slot code corresponding to the first word slot in a preset knowledge base, and determining the dictionary of the first word slot according to the word slot code; determining whether the first word slot is a complete word slot according to the dictionary, and when determining that the first word slot is a complete word slot, matching the text data corresponding to the first word slot in a preset dialogue library; when determining that the first word slot is an incomplete word slot, determining the missing word slot type of the first word slot, and generating a question sentence according to the missing word slot type; sending the question sentence and the corresponding question speech data to a display terminal, so that the display terminal calls a preset virtual digital person to broadcast the sentence speech data, and displays the question sentence on the display interface of the display terminal; receiving the second word slot obtained by the user's reply to the question sentence, and matching the text data corresponding to the first word slot and the second word slot in a preset dialogue library; Feeding back the speech data to be broadcast and the text data to the display terminal, so that the display terminal calls a preset virtual digital person to broadcast the speech data to be broadcast, and displays the text data on the display interface of the display terminal.

2. The method for session processing based on human-computer interaction according to claim 1, wherein Before the step of receiving the speech data to be recognized of the current user, the session processing method based on human-computer interaction further includes: When receiving the face image to be recognized of the current user, recognizing the face image to be recognized to obtain face features to be recognized; Calculating the user category of the current user according to the face features to be recognized, and when determining that the user category is a preset category, obtaining the preset text data corresponding to the preset category; Feeding back the preset text data and the corresponding preset broadcast data to the display terminal, so that the display terminal wakes up the preset virtual digital person to broadcast the preset broadcast data through the preset virtual digital person.

3. The session processing method based on human-computer interaction according to claim 2, wherein Recognizing the face image to be recognized to obtain face features to be recognized, including: Using a preset face detection and key point localization tool to detect the face area to be recognized of the face image to be recognized, and extracting the face key points to be recognized of the face image to be recognized in the face area to be recognized; Calculating the face attribute information of the face image to be recognized according to the face key points to be recognized, and obtaining the face features to be recognized according to the face attribute information; The facial attribute information includes one or more of age, gender and facial expression.

4. The session processing method based on human-computer interaction according to claim 2, wherein Calculating the user category of the current user according to the facial features to be recognized includes: Matching original facial features corresponding to the facial features to be identified in a preset facial database, and determining the user category of the current user according to the matching results; Wherein, if the matching result is that the preset face database contains original face features corresponding to the face features to be identified, then the user category of the current user is the preset category; If the matching result is that the original facial features corresponding to the facial features to be identified do not exist in the preset facial database, then the user category of the current user is a non-preset category.

5. The session processing method based on human-computer interaction according to claim 4, characterized in that, The human-computer interaction-based conversation processing method further includes: Obtaining the last matching time of the facial feature to be identified in the preset face database, and calculating the time difference between the last matching time and the current time; When it is determined that the time difference is less than a first preset time threshold, the display terminal is controlled to wake up the virtual digital human in a first preset manner, wherein the first preset manner is used to indicate that the current session needs to be connected to the session corresponding to the matching time of the last match.

6. The method for session processing based on human-computer interaction according to claim 1, characterized in that Recognizing the speech data to be recognized to obtain a speech recognition result, including: Extracting a first local feature of the speech data to be recognized by using a convolutional neural network included in a preset speech recognition model; Calculating the first global features of the speech data to be recognized according to the first local features respectively by using the self-attention module included in the preset speech recognition model; The first global feature is classified using the fully connected layer included in the preset speech recognition model to obtain a speech recognition result of the speech data to be recognized.

7. The session processing method based on human-computer interaction according to claim 1, wherein Converting the text data into voice data to be broadcasted includes: Discretize the text data based on a preset speech synthesis model to obtain a discrete speech synthesis result; Calculating a speech synthesis probability distribution result according to the discrete speech synthesis result and a uniformly distributed sampling result corresponding to the discrete speech synthesis result; The speech synthesis probability distribution result is subjected to speech synthesis based on a preset continuity function to obtain the speech data to be broadcasted.

8. The session processing method based on human-computer interaction according to claim 2, wherein The human-computer interaction-based conversation processing method further includes: Establishing an association relationship between the user identification of the current user and the text data, the facial features to be identified, and the user category; The text data is stored in the knowledge base based on the association relationship, and the facial features to be identified are stored in a facial database based on the association relationship.

9. A session processing device based on human-computer interaction, characterized in that, include: A speech recognition module, which is configured to, when receiving the speech data to be recognized of the current user, recognize the speech data to be recognized to obtain a speech recognition result, determine the scenario category required by the current user according to the speech recognition result, and segment the speech recognition result based on the scenario category to obtain the first slot included in the speech recognition result; wherein, the scenario category includes one or more of a financial consultation scenario, a network point 3D navigation scenario, and a product recommendation scenario; A speech data conversion module, which is configured to match the text data corresponding to the first slot in a preset knowledge base and convert the text data into speech data to be broadcast; wherein, the text data is obtained in the following manner: match the slot code corresponding to the first slot in the preset knowledge base, and determine the dictionary of the first slot according to the slot code; determine whether the first slot is a complete slot according to the dictionary, and when it is determined that the first slot is a complete slot, match the text data corresponding to the first slot in the preset dialogue library; when it is determined that the first slot is an incomplete slot, determine the missing slot type of the first slot, and generate a question sentence according to the missing slot type; send the question sentence and the question speech data corresponding to the question sentence to a display terminal, so that the display terminal calls a preset virtual digital human to broadcast the sentence speech data and display the question sentence on the display interface of the display terminal; receive the second slot obtained by the user's reply to the question sentence, and match the text data corresponding to the first slot and the second slot in the preset dialogue library; A first speech data broadcast module, which is configured to feedback the speech data to be broadcast and the text data to a display terminal, so that the display terminal calls a preset virtual digital human to broadcast the speech data to be broadcast and display the text data on the display interface of the display terminal.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the human-computer interaction-based conversation processing method according to any one of claims 1-8.

11. An electronic device, characterized in that, Comprising: A processor; And A memory, which is configured to store the executable instructions of the processor; Wherein, the processor is configured to execute the human-computer interaction-based conversation processing method according to any one of claims 1-8 by executing the executable instructions.

Citation Information

Patent Citations

  • A session robotic system

    CN101187990A

  • Expression interaction method and device, computer device and computer readable storage medium

    CN110363079A

  • Improved end-to-end speech recognition method

    CN111048082A

  • Multi-round dialogue voice interaction method and system, storage medium and electronic equipment

    CN111680144A