Voice interaction method and apparatus, client, and computer-readable storage medium
By highlighting the target language display elements in the client display window and displaying the text of the voice data and the response text data in the target language, the problem that existing voice response systems cannot display the user's language requirements is solved, thus achieving efficient voice interaction and improving the user experience.
Patent Information
- Application Number
- PCT/CN2024/090968
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-04-30
- Publication Date
- 2025-11-06
AI Technical Summary
Existing voice response systems cannot display users' language needs, resulting in low voice interaction quality and efficiency, which affects user experience.
By highlighting the target language display elements in the client display window and displaying the text data and response text data of the voice data in the target language, and combining the voice format for responses, the server analyzes the language type and text data of the voice data.
It improves the effectiveness and efficiency of voice interaction, enhances the user's interactive experience, and enables users to effectively identify and understand the language type and response information of voice data.
Smart Images

Figure CN2024090968_06112025_PF_FP_ABST
Abstract
Description
Voice interaction method, device, client and computer readable storage medium TECHNICAL FIELD
[0001] The present disclosure belongs to the technical field of voice, and particularly relates to a voice interaction method, device, client and computer readable storage medium, and electronic equipment. BACKGROUND
[0002] With the development of science and technology, some service robots for serving users appear. The service robots can be applied in multiple fields, such as hotels and airports, and provide services including inquiry and other functions for users through a voice response system.
[0003] The content of the background art is only for understanding the scheme and cannot be considered as recognition of the prior art.
[0004] SUMMARY
[0005] The voice interaction method, device, client and computer readable storage medium, and electronic equipment are proposed for the problems in the prior art, can feed back and display response information in the language required by the user, the user can effectively receive and understand the response information, the voice interaction effect and efficiency are improved, and the user experience is improved.
[0006] The present disclosure provides the following scheme.
[0007] In a first aspect, the present disclosure provides a voice interaction method applied to a client, a display window of the client comprising display elements in one or more languages, the method comprising:
[0008] receiving first voice data;
[0009] determining that a language category of the first voice data is a target language;
[0010] when the display elements in the one or more languages comprise a target display element in the target language, highlighting the target display element;
[0011] displaying text data of the first voice data in the target language in a preset text display area of the display window.
[0012] In some possible embodiments, the method further comprises:
[0013] obtaining reply text data of the text data, the language category of the reply text data being the target language;
[0014] highlighting the target display element;
[0015] display the reply text data in the target language in a preset text display area of the display window.
[0016] In some possible embodiments, the method further includes:
[0017] generating second speech data corresponding to the reply text data in the target language;
[0018] playing the second speech data in the target language.
[0019] In some possible embodiments, the obtaining the reply text data of the text data includes:
[0020] sending the text data and a language type identifier of the target language to a first server, wherein the first server is configured to obtain the reply text data of the text data, and the language type of the reply text data is the target language;
[0021] receiving the reply text data.
[0022] In some possible embodiments, the first server is specifically configured to store or connect a language model, the language model includes one or more sub-language models, the first server is configured to send the text data to a target language sub-model of the language model according to the language type identifier of the target language, and to receive the reply text data corresponding to the text data, the target language sub-model corresponds to the target language, the reply text data is generated by the target language sub-model, and the language type of the reply text data is the target language.
[0023] In some possible embodiments, the display window further includes a virtual human display area configured to display a virtual human image, and the method further includes:
[0024] determining a current state of the client, and controlling the virtual human image to display an action corresponding to the current state.
[0025] In some possible embodiments, the method further includes:
[0026] when the display elements in the one or more languages do not include the target display element, displaying a language prompt element in a preset area, the language prompt element being configured to prompt that the language type of the first speech data is the target language.
[0027] In some possible embodiments, the determining that the language type of the first speech data is the target language includes:
[0028] sending the first voice data to a second server, the second server being configured to analyze a language category of the first voice data and text data of the first voice data;
[0029] receiving an analysis result returned by the second server, the analysis result including a language identification of the target language and the text data of the first voice data.
[0030] In some possible embodiments, the display element includes a language icon and / or a language identification, the language icon and / or the language identification being configured to represent the language category.
[0031] In some possible embodiments, the language identification is displayed in the language corresponding to the display element.
[0032] In a second aspect, the present disclosure provides a voice interaction device, applied to a client, the client including a voice receiving unit, a terminal processing unit and a display unit, a display window of the display unit including display elements in one or more languages, the device including:
[0033] the voice receiving unit being configured to receive first voice data;
[0034] the terminal processing unit being configured to determine that a language category of the first voice data is a target language;
[0035] the display unit being configured to highlight a target display element in the display elements in the one or more languages when the target display element is in the target language, and display text data of the first voice data in the target language in a preset text display area of the display window.
[0036] In some possible embodiments, the device further includes:
[0037] the obtaining unit being configured to obtain reply text data of the text data, the language category of the reply text data being the target language;
[0038] the display unit being configured to highlight the target display element, and
[0039] display the reply text data in the target language in the preset text display area of the display window.
[0040] In some possible embodiments, the device further includes:
[0041] the voice generating unit being configured to generate second voice data corresponding to the reply text data in the target language;
[0042] the voice playing unit being configured to play the second voice data in the target language.
[0043] In some possible embodiments, the obtaining unit comprises:
[0044] a first sending sub-unit configured to send the text data and the language type identifier of the target language to a first server, wherein the first server is configured to obtain reply text data of the text data, and the language type of the reply text data is the target language;
[0045] a text receiving sub-unit configured to receive the reply text data.
[0046] In some possible embodiments, the first server is specifically configured to store or connect a language model, the language model comprises one or more sub-language models, the first server is configured to send the text data to a target language sub-model of the language model according to the language type identifier of the target language, and configured to receive reply text data corresponding to the text data, the target language sub-model corresponds to the target language, and the reply text data is generated by the target language sub-model and has the language type of the target language.
[0047] In some possible embodiments, the display window further comprises a virtual human display area configured to display a virtual human image, and the apparatus further comprises:
[0048] a terminal processing unit configured to determine a current state of the client, and control the virtual human image to perform an action corresponding to the current state on the display unit.
[0049] In some possible embodiments, the display unit is further configured to:
[0050] when the display element of the one or more languages does not comprise the target display element, display a language prompt element in a preset area, the language prompt element is configured to prompt that the language type of the first voice data is the target language.
[0051] In some possible embodiments, the apparatus further comprises a second sending unit and a receiving unit, and the terminal processing unit is configured to:
[0052] connect the second sending unit, the second sending unit is configured to send the first voice data to a second server, and the second server is configured to analyze the language type of the first voice data and text data of the first voice data;
[0053] connect the receiving unit, the receiving unit is configured to receive an analysis result returned by the second server, the analysis result comprises the language type identifier of the target language and the text data of the first voice data.
[0054] In some possible embodiments, the display element comprises a language icon and / or a language identifier, the language icon and / or the language identifier being used to represent a language category.
[0055] In some possible embodiments, the language identifier is displayed in a language corresponding to the display element.
[0056] In a third aspect, the present disclosure provides a client, comprising a display unit, a processor, a memory and a bus, the memory storing machine readable instructions executable by the processor, when the electronic device is running, the processor communicates with the memory and the display unit through the bus, and the machine readable instructions are executed by the processor to perform the voice interaction control method according to any one of the first aspect.
[0057] In a fourth aspect, the present disclosure provides a computer readable storage medium, the computer readable storage medium storing a computer program, when the computer program is run by a processor, the computer program performs the voice interaction control method according to any one of the first aspect.
[0058] In a fifth aspect, the present disclosure provides a computer program, when the computer program product is run on an electronic device, the computer program product makes the electronic device perform the digital human interaction control method according to any one of the first aspect.
[0059] In a sixth aspect, the present disclosure provides an electronic device, comprising:
[0060] one or more processors;
[0061] a screen for displaying an interface, a display window of the interface comprising display elements in one or more languages;
[0062] a receiver for receiving voice data;
[0063] and a memory, the memory storing codes;
[0064] when the codes are executed by the one or more processors, the electronic device performs the digital human interaction control method according to any one of the first aspect.
[0065] In some possible embodiments, the electronic device further comprises a speaker for playing voice data.
[0066] The voice interaction method provided by the embodiments of the present disclosure can effectively receive and understand input information and response information, and feed back and display the information input by the user in the language required by the user on the interface. On the one hand, the language type of the voice data input by the user and the corresponding text data are displayed on the user interface, so that the user can know whether the voice input by the user can be effectively recognized and whether the language type of the voice input by the user is accurate. On the other hand, the response information of the user input information can be displayed on the interface in the same language type, so that the user can effectively receive and understand the response information. The embodiments of the present disclosure can improve the voice interaction effect and efficiency, and improve the interaction experience of the user.
[0067] Other advantages of the present disclosure will be described in more detail in conjunction with the following description and drawings.
[0068] It should be understood that the above description is only a summary of the technical solutions of the present disclosure, so that the technical means of the present disclosure can be more clearly understood, and the content of the description can be implemented. In order to make the above and other purposes, characteristics and advantages of the present disclosure more obvious and easy to understand, the specific embodiments of the present disclosure are described below. BRIEF DESCRIPTION OF DRAWINGS
[0069] The advantages and benefits described herein, as well as other advantages and benefits, will be apparent to those of ordinary skill in the art by reading the following detailed description of exemplary embodiments. The drawings are for the purpose of illustrating exemplary embodiments only and are not to be considered as limiting of the present disclosure. Moreover, the same reference numerals are used throughout the several views of the drawings to refer to same or like parts. In the drawings:
[0070] FIG. 1 is a flowchart of a voice interaction method provided by an embodiment of the present disclosure;
[0071] FIG. 2 is a schematic diagram of a graphical user interface provided by an embodiment of the present disclosure;
[0072] FIG. 3 is another schematic diagram of a graphical user interface provided by an embodiment of the present disclosure;
[0073] FIG. 4 is another schematic diagram of a graphical user interface provided by an embodiment of the present disclosure;
[0074] FIG. 5 is another schematic diagram of a graphical user interface provided by an embodiment of the present disclosure;
[0075] FIG. 6 is another schematic diagram of a graphical user interface provided by an embodiment of the present disclosure;
[0076] FIG. 7 is a flowchart of another voice interaction method provided by an embodiment of the present disclosure;
[0077] FIG. 8 is a flowchart of another voice interaction method provided by an embodiment of the present disclosure;
[0078] FIG. 9 is a schematic diagram of another graphical user interface according to an embodiment of the present disclosure;
[0079] FIG. 10 is a flowchart of another voice interaction method according to an embodiment of the present disclosure;
[0080] FIG. 11 is a schematic diagram of another graphical user interface according to an embodiment of the present disclosure;
[0081] FIG. 12 is a schematic diagram of a voice interaction device according to an embodiment of the present disclosure;
[0082] FIG. 13 is a schematic diagram of another voice interaction device according to an embodiment of the present disclosure;
[0083] FIG. 14 is a schematic diagram of an electronic device according to an embodiment of the present disclosure;
[0084] FIG. 15 is a schematic diagram of another electronic device according to an embodiment of the present disclosure;
[0085] In the drawings, identical or corresponding reference signs indicate identical or corresponding parts. DETAILED DESCRIPTION
[0086] Exemplary embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.
[0087] In the description of the embodiments of the present disclosure, it should be understood that terms such as "include" or "have" are intended to indicate that there can be additional features, numbers, steps, actions, components, parts, or combinations thereof in the specification, and do not exclude the presence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.
[0088] Unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" herein is only a description of the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone.
[0089] The terms "first", "second", etc. are used only for the purpose of distinguishing similar or identical technical features, and cannot be understood as indicating or implying relative importance or quantity thereof. Thus, features defined with "first", "second" etc. can explicitly or implicitly include one or more of these features. In the description of the embodiments of the present disclosure, the meaning of the term "plurality" is two or more, unless otherwise specified.
[0090] It should also be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0091] With the development of science and technology, some service robots for serving users have appeared. Service robots can be applied in many fields, such as hotels and airports, etc., and provide services including inquiry for users through voice response systems.
[0092] The inventor found that there are some problems: there is a need for multi-language response in the above-mentioned fields at present, but the existing voice response system cannot display the language needs of users, and cannot feed back and display response information in the language required by users, so that users cannot effectively receive and understand the response information, resulting in low voice interaction effect and poor efficiency, which affects the user experience. In order to avoid the above-mentioned problems as much as possible, the present disclosure provides a series of solutions.
[0093] As shown in FIG. 1, a flowchart of a voice interaction method provided by some embodiments of the present disclosure is shown, which is applied to a client, and the client includes a display window, and the display window includes display elements of one or more languages. In some examples, the one or more languages include but are not limited to Chinese, English, Arabic, Thai, etc., which are not particularly limited by the present disclosure.
[0094] In some embodiments, the display window is used to display a graphical user interface, and in the graphical user interface, the display elements of one or more languages are included. In some examples, the display elements of different languages are different.
[0095] In some examples, the display elements include language icons and / or language identifiers, and the language icons and / or language identifiers are used to represent language categories. Here, "and / or" represents three ways: ① language icons and language identifiers exist at the same time, ② language icons and language identifiers exist alternatively. For example, as shown in FIG. 2, the display elements include language icons and language identifiers, and the language icons and language identifiers represent language categories at the same time. For another example, as shown in FIG. 3, the display elements include language icons, and the language icons represent language categories. For another example, as shown in FIG. 4, the display elements include language identifiers, and the language identifiers represent language categories.
[0096] In some examples, the language identifier displays the language corresponding to the display element. As shown in FIG. 5, the display element includes a language icon and a language identifier. The language icon is a graphical icon representing a language category, and the language identifier is text displayed in the corresponding language. It should be noted that the language icon can also be an icon of other graphics, such as a national flag as a language icon, and the embodiments of the present disclosure do not limit the form of the language icon.
[0097] As shown in FIG. 1, the language interaction method provided by some embodiments of the present disclosure at least includes 101-104.
[0098] 101. Receive first voice data.
[0099] The client receives the first voice data issued by the user. In some examples, the language category of the first voice data is the language category commonly used by the user, which can usually be the first language, i.e., the native language. In other examples, the language category of the first voice data issued by the user is an internationally common language, such as English. The embodiments of the present disclosure do not limit the language category used by the user.
[0100] 102. Determine that the language category of the first voice data is a target language.
[0101] After the client receives the first voice data, the language category of the first voice data is analyzed and determined. In the embodiments of the present disclosure, the determined language category of the first voice data is referred to as the target language.
[0102] 103. When the display element of one or more languages includes a target display element of the target language, highlight the target display element.
[0103] 104. Display the text data of the first voice data in the target language in a preset text display area of the display window.
[0104] In some embodiments, 103 and 104 do not distinguish the order, and both can be executed at the same time, or can not be executed at the same time (for example, 103 is executed before 104, and for example, 103 is executed after 104). In an example of the present disclosure, both are executed at the same time, as shown in FIG. 6. When the target language is English, the language icon and the language identifier of English are boxed by a round rectangle to highlight, and the English text of the first voice data is displayed in the preset text display area.
[0105] The voice interaction method provided by the embodiments of the present disclosure can feed back and display the information input by the user in the language required by the user on the interface, so that the language category of the voice data input by the user and the corresponding text data are displayed on the user interface, so that the user can know whether the voice input by the user can be effectively recognized and whether the language category of the voice input by the user is accurate. The embodiments of the present disclosure can improve the voice interaction effect and efficiency, and improve the interaction experience of the user.
[0106] Referring to FIG. 7, some embodiments of the present disclosure further provide a language interaction method, which includes 701-707. The embodiments provide a solution for a client to respond to voice data after receiving the voice data from a user.
[0107] 701. Receive first voice data.
[0108] The client receives first voice data from a user. In some examples, the language category of the first voice data is a language category commonly used by the user, which can be a first language, i.e., a native language. In other examples, the language category of the first voice data from the user is an internationally used language, such as English. Embodiments of the present disclosure do not limit the language category used by the user.
[0109] 702. Determine that the language category of the first voice data is a target language.
[0110] After receiving the first voice data, the client analyzes and determines the language category of the first voice data, which is referred to as a target language in embodiments of the present disclosure.
[0111] In some examples, after receiving the first voice data, the client parses the first voice data to obtain text data of the voice data.
[0112] In other examples, to reduce the resource load rate of the client and improve the running efficiency of the client, the client further determines the language category of the first voice data by using a server. As shown in FIG. 8, 702 includes 7021-7022.
[0113] 7021. Send the first voice data to a second server, which is configured to analyze the language category of the first voice data and text data of the first voice data.
[0114] 7022. Receive an analysis result returned by the second server, the analysis result including a language category identifier of the target language and the text data of the first voice data.
[0115] 703. When one or more language display elements include a target display element of the target language, highlight the target display element.
[0116] 704. Display the text data of the first voice data in the target language in a preset text display area of a display window.
[0117] 701-704 are that after the client receives the first voice data, the client feeds back and displays the information input by the user in the language required by the user on the interface. Then the client answers the first voice data in the same language as the first voice data. In some examples, the manner of answering the first voice data at least includes 705-707.
[0118] 705, obtaining reply text data of the text data, and the language type of the reply text data is the target language.
[0119] 706, highlighting the target display element.
[0120] 707, displaying the reply text data in the target language in the preset text display area of the display window.
[0121] In some embodiments, 706 and 707 do not distinguish the order, and both can be executed at the same time or not at the same time (for example, 706 is executed before 707, and for example, 706 is executed after 707). In an example of the present disclosure, both are executed at the same time. As shown in FIG. 9, when the target language is English, the language icon and the language type of English are boxed by a circular rectangle to highlight, and the reply text data of the first voice data is displayed in the preset text display area, and the reply text data is displayed in English, that is, English text.
[0122] The voice interaction method provided by the embodiments of the present disclosure can effectively receive, understand, and answer the input information and display the information input by the user in the language required by the user on the interface. On the one hand, the language type of the voice data input by the user and the corresponding text data are displayed on the user interface, so that the user can know whether the voice input by the user can be effectively recognized and whether the language type of the voice input by the user is accurate. On the other hand, the answer information of the user input information can be displayed on the interface in the same language type, so that the user can effectively receive and understand the answer information. The embodiments of the present disclosure can improve the voice interaction effect and efficiency, and improve the interaction experience of the user.
[0123] In some possible embodiments, the voice data input by the user is also answered in the same language type as the first voice data in the form of voice. As shown in FIG. 10, the method provided by the embodiments of the present disclosure further includes 708 and 709.
[0124] 708, generating second voice data corresponding to the reply text data in the target language.
[0125] The generated second voice data is the answer voice of the first voice data, and the language type of the second voice data is the target language.
[0126] 709, playing the second voice data in the target language.
[0127] In some embodiments, 708, 709 and 706, 707 are not distinguished in sequence, 708, 709 and 706, 707 can be executed at the same time, or can not be executed at the same time (for example, 708, 709 is executed before 706, 707, and for example, 708, 709 is executed after 706, 707). The present disclosure is not particularly limited.
[0128] The embodiments of the present disclosure can use the voice form to answer the voice data input by the user in the same language as the first voice data, and listen to the reply text data in the same language as the first voice data while making the user watch the reply text data in the language of the first voice data, so as to realize the listening and watching of the reply data, improve the possibility of the user effectively receiving and understanding the reply information, and improve the voice interaction effect and efficiency.
[0129] In some possible embodiments described above, 705 "obtains the reply text data of the text data" includes 7051-7052.
[0130] 7051, the text data and the language identification of the target language are sent to the first server; wherein the first server is used to obtain the reply text data of the text data, and the language type of the reply text data is the target language. The language identification is used to represent the language type of the target language, which can be text, string, or identification composed of letters and numbers, and the present disclosure is not particularly limited.
[0131] In some examples, the client sends the text data and the language identification of the target language to the first server, so that the first server analyzes the text data in the target language according to the language identification, and generates the reply text data in the target language.
[0132] In other examples, the client sends the text data and the language identification of the target language to the first server, and the first server analyzes the text data in the target language according to the language identification by using a language model, and generates the reply text data in the target language. In one example, the first server is specifically used to store or connect the language model, the language model includes one or more sub language models, the first server is used to send the text data to the target language sub model of the language model according to the language identification of the target language, and is used to receive the reply text data corresponding to the text data, the target language sub model corresponds to the target language, the reply text data is generated by the target language sub model, and the language type is the target language.
[0133] For example:
[0134] The language identification of the target language is English, and the language model connected by the first server includes three sub-language models, which are an English sub-language model, an Arabic sub-language model, and a Thai sub-language model. The English sub-language model is used to analyze and generate English text, the Arabic sub-language model is used to analyze and generate Arabic text, and the Thai sub-language model is used to analyze and generate Thai text. When the first server receives English text data and the language identification "English", the English text data is sent to the English sub-language model of the language model, the English sub-language model analyzes the text data, generates reply text data in English, and returns the reply text data in English to the first server. The first server sends the reply text data to the client.
[0135] 7052、Receiving the reply text data.
[0136] In some embodiments of the present disclosure, the display window further includes a virtual person display area for displaying a virtual person image. The voice interaction method provided by the embodiments of the present disclosure further displays the current state of the client in the virtual person image in the display window. Specifically, the voice interaction method further includes: determining the current state of the client, and controlling the virtual person image to display an action corresponding to the current state.
[0137] In some examples, the state of the client includes at least a receiving state, a thinking state, and a responding state.
[0138] The receiving state, also known as the listening state, is a state in which the client receives user input. For example, the state of the client before or at the time of executing the aforementioned 701.
[0139] The thinking state is a state in which the client understands user input and / or obtains reply data. For example, the state of the client when executing the aforementioned 702 or 705.
[0140] The responding state, also known as the reply state, is a state in which the client replies or responds to the user. For example, the state of the client when executing the aforementioned 706 or 707.
[0141] The different states correspond to different actions. When the client determines its current state, the virtual person image is controlled to display the corresponding action in the display window. For example, the receiving state is a listening action, the thinking state is a chin-thought action, and the responding state is a reply action. The actions of each state can be set according to requirements, and the embodiments of the present disclosure are not particularly limited.
[0142] In one example, when the client receives the first voice data input by the user, the receiving state is triggered, and the virtual human image can be controlled to show the corresponding action of the receiving state. When the client obtains the reply text data of the text data, the thinking state is triggered, and the virtual human image can be controlled to show the corresponding action of the thinking state. When the client performs 706 or 707, the answering state is triggered, and the virtual human image can be controlled to show the corresponding action of the answering state.
[0143] In combination with the foregoing embodiments, to better interact with the user in voice, when the target language is determined, the virtual human image corresponding to the target language can be displayed on the display window. Different virtual human images correspond to different language categories. In one example, language category 1 corresponds to virtual human image 1, and language category 2 corresponds to virtual human image 2. The present disclosure sets virtual human images corresponding to various language categories as needed, and the present disclosure is not particularly limited.
[0144] In combination with the foregoing embodiments, in some other embodiments, if the display elements of one or more languages do not include the target display element, a language prompt element is displayed in a preset area, and the language prompt element is used to prompt that the language category of the first voice data is the target language. As shown in the interface of FIG. 11, there is no Korean in the display elements of the display window, and a prompt is displayed in the preset area of the display window to prompt that the language category of the first voice data is Korean. Through this solution, when there is no display element of the language category in the display window, the language category input by the user can be informed to the user through the language prompt element, and the interaction effect and user experience in this case are improved.
[0145] As shown in FIG. 12, the present disclosure also provides a voice interaction device, which is applied to a client. The client 1200 includes a voice receiving unit 1201, a terminal processing unit 1202, and a display unit 1203. The display window of the display unit includes display elements of one or more languages. Each unit is explained as follows.
[0146] The voice receiving unit 1201 is configured to receive first voice data.
[0147] The terminal processing unit 1202 is configured to determine that the language category of the first voice data is a target language.
[0148] The display unit 1203 is configured to, when the display elements of one or more languages include a target display element of the target language, highlight the target display element; and display text data of the first voice data in the target language in a preset text display area of the display window.
[0149] As shown in FIG. 13, in some possible embodiments, the device further includes:
[0150] The acquisition unit 1204 is configured to acquire reply text data of the text data, and a language type of the reply text data is a target language.
[0151] The display unit 1203 is configured to highlight the target display element.
[0152] The reply text data is displayed in the target language in the preset text display area of the display window.
[0153] In some possible embodiments, the apparatus further includes:
[0154] The speech generation unit 1205 is configured to generate second speech data corresponding to the reply text data in the target language.
[0155] The speech playing unit 1206 is configured to play the second speech data in the target language.
[0156] In some possible embodiments, the acquisition unit 1204 includes:
[0157] The first sending sub-unit is configured to send the text data and the language type identification of the target language to a first server, where the first server is used to acquire reply text data of the text data, and a language type of the reply text data is the target language.
[0158] The text receiving sub-unit is configured to receive the reply text data.
[0159] In some possible embodiments, the first server is specifically configured to store or connect a language model, the language model includes one or more sub-language models, the first server is used to send the text data to a target language sub-model of the language model according to the language type identification of the target language, and is used to receive reply text data corresponding to the text data, the target language sub-model corresponds to the target language, the reply text data is generated by the target language sub-model, and the language type is the target language.
[0160] In some possible embodiments, the display window further includes a virtual human display area, the virtual human display area is used to display a virtual human image,
[0161] The terminal processing unit 1202 is configured to determine a current state of the client, and control the virtual human image to display an action corresponding to the current state on the display unit.
[0162] In some possible embodiments, the display unit 1203 is further configured to:
[0163] When the display element of one or more languages does not include the target display element, a language prompt element is displayed in the preset area, and the language prompt element is used to prompt that the language type of the first speech data is the target language.
[0164] In some possible embodiments, the terminal processing unit 1202 is further configured to:
[0165] The second sending unit 1207 is connected to the second sending unit and is configured to send the first voice data to a second server, the second server being configured to analyze a language type of the first voice data and text data of the first voice data.
[0166] The receiving unit 1208 is connected to the receiving unit and is configured to receive an analysis result returned by the second server, the analysis result including a language type identification of a target language and the text data of the first voice data.
[0167] In some possible embodiments, the display element includes a language icon and / or a language type identification, the language icon and / or the language type identification being used to represent the language type.
[0168] In some possible embodiments, the language type identification is displayed in the language corresponding to the display element.
[0169] The embodiment of the present disclosure provides a client, including a display unit, a processor, a memory and a bus, the memory stores machine readable instructions executable by the processor, when the electronic device is running, the processor communicates with the memory and the display unit through the bus, and the machine readable instructions are executed by the processor to perform the voice interaction control method in any one of the preceding embodiments.
[0170] The embodiment of the present disclosure provides a computer readable storage medium, the computer readable storage medium stores a computer program, and the computer program is executed by the processor to perform the voice interaction control method in any one of the preceding embodiments.
[0171] The embodiment of the present disclosure provides a computer program, when the computer program product is executed on the electronic device, so that the electronic device performs the voice interaction control method in any one of the preceding embodiments.
[0172] As shown in FIG. 14, the embodiment of the present disclosure further provides an electronic device, including:
[0173] one or more processors;
[0174] a screen for displaying an interface, a display window of the interface including display elements in one or more languages;
[0175] a receiver for receiving voice data;
[0176] and a memory, the memory storing codes;
[0177] When the code is executed by the one or more processors, the electronic device is caused to perform the digital human interaction control method of any one of the first aspect.
[0178] In some possible embodiments, as shown in FIG. 15, the electronic device further includes a loudspeaker for playing voice data. When the code is executed by the one or more processors, the electronic device is caused to perform the digital human interaction control method of any one of the first aspect.
[0179] In the description of the present specification, the description made with reference to the terms “some possible embodiments”, “some embodiments”, “an example”, “a specific example”, or “some examples” and the like means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure, and the above terms do not necessarily indicate the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present specification and the features of the different embodiments or examples without contradiction.
[0180] Regarding the method flowchart of the embodiments of the present disclosure, certain operations are described as different steps executed in a certain order. Such flowcharts are illustrative rather than limiting. Certain steps described herein can be grouped together and executed in a single operation, or certain steps can be split into multiple sub-steps, and certain steps can be executed in an order different from that shown herein. Each step shown in the flowchart can be implemented in any manner by any circuit structure and / or tangible mechanism (for example, by software running on a computer device, hardware (for example, processor or chip implemented logic functions), etc., and / or any combination thereof) in any way.
[0181] Those skilled in the art can understand that in the method described in the above specific embodiments, the writing order of each step does not mean a strict execution order, and the specific execution order of each step should be determined by its function and possible inherent logic.
[0182] It should be noted that the apparatus in the embodiments of the present disclosure can implement each process of the embodiments of the foregoing method and achieve the same effects and functions, which will not be described here.
[0183] Computer-readable media includes permanent and non-permanent, movable and non-movable media that can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer-readable storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory, read-only memory, electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device. In addition, although the operations of the methods of the present disclosure are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in that particular order, or that all of the shown operations must be performed to achieve the desired results. In addition, some steps can be omitted, a plurality of steps can be combined into one step, and / or one step can be divided into a plurality of sub-steps.
[0184] Although the spirit and principles of the present disclosure have been described above with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the disclosed specific embodiments, and the division of aspects does not mean that the features in these aspects cannot be combined. The present disclosure is intended to cover various modifications and equivalent arrangements included in the spirit and scope of the appended claims.
Claims
1. A voice interaction method, characterized in that, Applied to a client, a display window of the client includes display elements of one or more languages, and the method comprises: receiving first voice data; determining that the language category of the first voice data is a target language; when the display elements of one or more languages include a target display element of the target language, highlighting the target display element; displaying text data of the first voice data in the target language in a preset text display area of the display window.
2. The method of claim 1, wherein, The method further comprises: obtaining reply text data of the text data, the language category of the reply text data being the target language; highlighting the target display element; displaying the reply text data in the target language in a preset text display area of the display window.
3. The method of claim 2, wherein, The method further comprises: generating second voice data corresponding to the reply text data in the target language; playing the second voice data in the target language.
4. The method of claim 2, wherein, The obtaining of the reply text data of the text data comprises: sending the text data and the language category identification of the target language to a first server; wherein the first server is configured to obtain the reply text data of the text data, and the language category of the reply text data is the target language; receiving the reply text data.
5. The method of claim 4, wherein, The first server is specifically configured to store or connect a language model, the language model including one or more sub-language models, and the first server is configured to send the text data to a target language sub-model of the language model according to the language category identification of the target language, and to receive the reply text data corresponding to the text data, the target language sub-model corresponding to the target language, the reply text data being generated by the target language sub-model and having the language category of the target language.
6. The method of claim 1, wherein, The display window further includes a virtual person display area for displaying a virtual person image, and the method further comprises: determining the current state of the client, and controlling the virtual person image to display actions corresponding to the current state.
7. The method of claim 1, wherein, The method further comprises: when the display elements of one or more languages do not include the target display element, displaying a language prompt element in a preset area, the language prompt element being configured to prompt that the language category of the first voice data is the target language.
8. The method of claim 1, wherein, The determination that the language category of the first voice data is the target language comprises: sending the first voice data to a second server, the second server being configured to analyze the language category of the first voice data and the text data of the first voice data; receiving an analysis result returned by the second server, the analysis result including the language category identification of the target language and the text data of the first voice data.
9. The method according to any of claims 1 to 8, characterized in that, The display element includes a language icon and / or a language category identification, the language icon and / or the language category identification being configured to represent the language category.
10. The method of claim 9, wherein, The language category identification is displayed in the language corresponding to the display element.
11. A voice interaction device, characterized by Applied to a client, the client includes a voice receiving unit, a terminal processing unit and a display unit, a display window of the display unit includes display elements of one or more languages, and the device comprises: The voice receiving unit is configured to receive first voice data; The terminal processing unit is configured to determine that a language category of the first voice data is a target language; The display unit is configured to highlight a target display element of the target language when the display element of the one or more languages includes the target display element; and display text data of the first voice data in the target language in a preset text display area of the display window.
12. The apparatus of claim 11, wherein, The device further comprises: The acquisition unit is configured to acquire reply text data of the text data, and a language category of the reply text data is the target language; The display unit is configured to highlight the target display element; and Display the reply text data in the target language in a preset text display area of the display window.
13. The apparatus of claim 12, wherein, The device further comprises: The voice generation unit is configured to generate second voice data corresponding to the reply text data in the target language; The voice playing unit is configured to play the second voice data in the target language.
14. The apparatus of claim 12, wherein, The acquisition unit comprises: The first sending sub-unit is configured to send the text data and a language category identifier of the target language to a first server; wherein the first server is used to acquire reply text data of the text data, and a language category of the reply text data is the target language; The text receiving sub-unit is configured to receive the reply text data.
15. The apparatus of claim 14, wherein, The first server is specifically used to store or connect a language model, the language model comprises one or more sub-language models, the first server is used to send the text data to a target language sub-model of the language model according to the language category identifier of the target language, and is used to receive reply text data corresponding to the text data, the target language sub-model corresponds to the target language, the reply text data is generated by the target language sub-model, and the language category is the target language.
16. The apparatus of claim 11, wherein, The display window further comprises a virtual person display area for displaying a virtual person image, and the device further comprises: The terminal processing unit is configured to determine a current state of the client, and control the virtual person image to perform an action corresponding to the current state on the display unit.
17. The apparatus of claim 11, wherein, The display unit is further configured to: When the display element of the one or more languages does not include the target display element, display a language prompt element in a preset area, and the language prompt element is used to prompt that the language category of the first voice data is the target language.
18. The apparatus of claim 11, wherein, Further comprising a second sending unit and a receiving unit, and the terminal processing unit is configured to: The second sending unit is connected to the terminal processing unit, and is configured to send the first voice data to a second server, and the second server is used to analyze a language category of the first voice data and text data of the first voice data; The receiving unit is connected to the terminal processing unit, and is configured to receive an analysis result returned by the second server, and the analysis result includes a language category identifier of the target language and text data of the first voice data.
19. The apparatus of any of claims 11-18, wherein, The display element includes a language icon and / or a language identifier, the language icon and / or the language identifier being used to represent a language category.
20. The apparatus of claim 19, wherein, The language identifier is displayed in a language corresponding to the display element.
21. A client, comprising: The electronic device includes a display unit, a processor, a memory, and a bus, the memory storing machine readable instructions executable by the processor, the processor in communication with the memory and the display unit via the bus when the electronic device is running, the machine readable instructions being executed by the processor to perform the digital human interaction control method of any one of claims 1-10.
22. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, the computer program being executed by the processor to perform the digital human interaction control method of any one of claims 1-10.
23. A computer program, characterized in that, The computer program product, when executed on the electronic device, causes the electronic device to perform the digital human interaction control method of any one of claims 1-10.
24. An electronic device, comprising: The computer program product, when executed on the electronic device, causes the electronic device to perform the digital human interaction control method of any one of claims 1-10. The computer program product, when executed on the electronic device, causes the electronic device to perform the digital human interaction control method of any one of claims 1-10. The computer program product, when executed on the electronic device, causes the electronic device to perform the digital human interaction control method of any one of claims 1-10. The computer program product, when executed on the electronic device, causes the electronic device to perform the digital human interaction control method of any one of claims 1-10. The computer program product, when executed on the electronic device, causes the electronic device to perform the digital human interaction control method of any one of claims 1-10. The computer program product, when executed on the electronic device, causes the electronic device to perform the digital human interaction control method of any one of claims 1-10.
25. The electronic device of claim 24, wherein, The computer program product, when executed on the electronic device, causes the electronic device to perform the digital human interaction control method of any one of claims 1-10. The computer program product, when executed on the electronic device, causes the electronic device to perform the digital human interaction control method of any one of claims 1-10. The computer program product, when executed on the electronic device, causes the electronic device to perform the digital human interaction control method of any one of claims 1-10.
Citation Information
Patent Citations
Method and device for controlling interaction of intelligent equipment
CN109949795A
Multilingual intelligent voice conversation method and system
CN111128126A
Speech recognition method, device, terminal and storage medium
CN111261144A
Voice interaction method and device, electronic equipment and medium
CN114582339A
Man-machine customer service interaction method and device, terminal equipment and storage medium
CN115086257A