Method, apparatus, device, storage medium and product for speech processing

By switching input modes in the voice input area, the system can distinguish and process voice input that includes both dialogue and non-dialogue content, thus solving the problem of limited voice input methods and improving the convenience and practicality of voice input.

CN119181362BActive Publication Date: 2025-11-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411289339.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2025-11-25
Estimated Expiration
2044-09-13

AI Technical Summary

Technical Problem

Existing voice input methods are limited and cannot effectively distinguish between conversational and non-conversational content in voice messages, affecting convenience and practicality.

Method used

By responding to different operations in the voice input area, the voice input mode can be flexibly switched to receive and distinguish between voice input containing dialogue content and non-dialogue content, and then presented in the recognition results.

Benefits of technology

It improves the convenience and practicality of voice input, allowing users to input various content via voice and providing a better user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119181362B_ABST
    Figure CN119181362B_ABST
Patent Text Reader

Abstract

According to an embodiment of the present disclosure, a method, device, equipment, storage medium and product for voice processing are provided. The method comprises: in response to a first type of operation on a voice input area, starting a first input mode to receive voice input; in the first input mode, in response to a second type of first operation on the voice input area, switching to a second input mode to receive voice input; and presenting a voice recognition result corresponding to the received voice input. Thus, different input modes of voice can be flexibly switched. This enables the user to input various contents for different purposes through voice, thereby improving the convenience and practicability of voice input and providing a better user experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Example embodiments of the present disclosure generally relate to the field of computers, and in particular, to a method, an apparatus, a device, a storage medium and a product for speech processing. BACKGROUND

[0002] Speech input is a common input method in various applications such as entertainment, socialization, etc. The speech input by a user can be converted into text or the like by a speech recognition model. However, the current speech input method is very single. How to improve the richness and practicality of speech input is a technical problem to be explored at present. SUMMARY

[0003] In a first aspect of the present disclosure, a speech processing method is provided. The method comprises: in response to a first type of operation on a speech input area, starting a first input mode to receive speech input; in the first input mode, in response to a second type of first operation on the speech input area, switching to a second input mode to receive speech input; and presenting a speech recognition result corresponding to the received speech input, in which a first recognition result corresponding to the speech received in the first input mode is distinguishable from a second recognition result corresponding to the speech received in the second mode.

[0004] In a second aspect of the present disclosure, an apparatus for speech processing is provided. The apparatus comprises: a starting module configured to, in response to a first type of operation on a speech input area, start a first input mode to receive speech input; a first switching module configured to, in the first input mode, in response to a second type of first operation on the speech input area, switch to a second input mode to receive speech input; and a result presentation module configured to present a speech recognition result corresponding to the received speech input, in which a first recognition result corresponding to the speech received in the first input mode is distinguishable from a second recognition result corresponding to the speech received in the second mode.

[0005] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.

[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium has stored thereon a computer program, which is executable by a processor to implement the method of the first aspect.

[0007] In a fifth aspect of the disclosure, a computer program product is provided. The computer program product includes computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method of the first aspect.

[0008] It should be understood that the matters described in this section are not intended to define key or essential features of the embodiments of the disclosure, nor are they intended to limit the scope of the disclosure. Other features of the disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0009] The above and other features, aspects, and advantages of embodiments of the disclosure will become more apparent from the following detailed description in conjunction with the accompanying drawings, in which like reference numerals denote like elements, and in which:

[0010] FIG. 1 A schematic diagram showing an example environment in which embodiments of the disclosure can be implemented is shown;

[0011] FIGS. 2A-2D A schematic diagram showing a dialogue interface with a virtual object according to some embodiments of the disclosure is shown;

[0012] FIG. 3 A schematic diagram showing a process of a presentation style transition of a voice input area according to some embodiments of the disclosure is shown;

[0013] FIG. 4 A schematic diagram showing a process of generating a voice recognition result according to some embodiments of the disclosure is shown;

[0014] FIG. 5 A schematic diagram showing an interaction process between an ASR manager, an ASR root session, a voice server, and an ASR sub-session according to some embodiments of the disclosure is shown;

[0015] FIG. 6 A flowchart showing a process for voice processing according to some embodiments of the disclosure is shown;

[0016] FIG. 7 A block diagram showing an apparatus for voice processing according to some embodiments of the disclosure is shown; and

[0017] FIG. 8 A block diagram showing an apparatus capable of implementing embodiments of the disclosure is shown. DETAILED DESCRIPTION

[0018] It can be understood that, before using the technical solutions disclosed in the embodiments of the disclosure, the type of personal information involved in the disclosure, the scope of use, the use scenario, etc. should be informed to the user and the authorization of the user should be obtained in a proper manner according to relevant laws and regulations.

[0019] For example, in response to receiving an active request of a user, a prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed by the user will need to acquire and use personal information of the user. Thus, the user can autonomously select whether to provide the personal information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solution of the present disclosure according to the prompt information.

[0020] As an optional but non-limiting implementation manner, in response to receiving an active request of a user, the manner of sending a prompt information to the user may, for example, be a pop-up window manner, and the prompt information may, for example, be presented in the pop-up window in a text manner. In addition, the pop-up window may, for example, also carry a selection control for the user to select “agree” or “disagree” to provide the personal information to the electronic device.

[0021] It can be understood that the above notification and acquisition of user authorization process is only illustrative, and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0022] It can be understood that the data (including but not limited to the data itself, acquisition or use of the data) involved in the technical solution should comply with the requirements of the relevant laws and regulations and the relevant provisions.

[0023] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments described herein, rather, these embodiments are provided to make the present disclosure more thorough and complete. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes, and are not intended to limit the scope of protection of the present disclosure.

[0024] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. Furthermore, embodiments described in any section / subsection can be combined with any other embodiment described in the same section / subsection and / or a different section / subsection in any manner.

[0025] In this document, unless explicitly stated otherwise, performing a step “in response to A” does not mean that the step is performed immediately after A, but can include one or more intermediate steps.

[0026] In the description of embodiments of the disclosure, the term "includes" and its conjugates are to be interpreted as open-ended terms that mean "includes, but is not limited to." The term "based on" is to be interpreted as "based, at least in part, on." The term "one embodiment" or "the embodiment" are to be interpreted as "at least one embodiment." The term "some embodiments" is to be interpreted as "at least some embodiments." Other explicit or implicit definitions can also be included below. The terms "first", "second", etc. can refer to different or the same objects. Other explicit and implicit definitions can also be included below.

[0027] As used herein, the term "model" can learn the association between the corresponding input and output from the training data, so that the corresponding output can be generated for a given input after the training is completed. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes input and provides a corresponding output by using multiple layers of processing units. In this document, "model" can also be referred to as "machine learning model", "machine learning network" or "network", which are used interchangeably in this document. A model can also include different types of processing units or networks.

[0028] As briefly mentioned above, voice input is a common input method in various applications such as entertainment, social interaction, etc. The voice input by the user can be converted into text or other forms through a voice recognition model. However, the current voice input method is very single, which can only process the voice input by the user as a whole. When the voice input by the user contains multiple parts for different purposes, it cannot be distinguished. Taking the interaction between the user and the virtual object as an example, the user needs to input the dialogue content through the voice, and sometimes needs to input non-dialogue content such as actions, expressions, emotions, etc. so that the virtual object can better understand the dialogue content. However, the current voice input method cannot distinguish the dialogue content and the non-dialogue content in the voice. This has a certain impact on the convenience, practicality and user experience of voice input.

[0029] Embodiments of the present disclosure propose a scheme for voice processing. According to various embodiments of the present disclosure, in response to a first type of operation on a voice input area, a first input mode is started to receive voice input. Then, in the first input mode, in response to a second type of first operation on the voice input area, a second input mode is switched to receive voice input. Finally, a voice recognition result corresponding to the received voice input is presented, and in the voice recognition result, a first recognition result corresponding to the voice received in the first input mode is distinguishable from a second recognition result corresponding to the voice received in the second mode.

[0030] In this way, different input modes of voice can be flexibly switched. Voice recognition is performed on the voice received in different input modes, and the presentation of the recognition results in different input modes is distinguishable. This enables the user to input various content for different uses by voice, thereby improving the convenience and practicality of voice input, and providing a better user experience.

[0031] Example Environment

[0032] FIG. 1 A schematic diagram illustrating an example environment 100 in which embodiments of the present disclosure can be implemented is shown. In this example environment 100, a terminal device 110 has an application 120 installed therein. A user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110.

[0033] In some embodiments, the application 120 can be any suitable application that can provide voice input related services. As one example, the application 120 can be an instant messaging (IM) application. As another example, the application 120 can be an entertainment type application, e.g., a virtual object interaction application, etc. It should be appreciated that the above are merely examples and are not intended to be limiting in any way, and the application 120 can be any suitable type of voice input enabled application. In some embodiments, the application 120 can be a native application installed on the terminal device 110, or a web application provided by a server 130. FIG. 1 In the environment 100, if the application 120 is in an active state, the terminal device 110 can present an interface 150 of the application 120. The interface 150 can include various pages that the application 120 can provide, such as an interface for a conversation with another user, a conversation interface with a virtual object, a selection interface of virtual objects, etc.

[0034] In some embodiments, the terminal device 110 communicates with the server 130 to implement the provision of services of the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media player, a multimedia player, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an electronic book device, a game device, or any combination thereof, including accessories and peripherals of these devices, or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface to the user (such as a "wearable" circuit, etc.). The server 130 can be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, etc.

[0035] It should be appreciated that the structure and functionality of the various elements in the environment 100 are described for illustrative purposes only and are not intended to imply any limitation on the scope of the present disclosure.

[0036] Some example embodiments of the present disclosure will be described hereinafter with reference to the accompanying drawings. It should be appreciated that the pages shown in the accompanying drawings are merely examples and various page designs can actually exist. The various graphical elements in the pages can have different arrangements and different visual representations, one or more elements among them can be omitted or replaced, and one or more other elements can also exist. The embodiments of the present disclosure are not limited in this respect. In addition, in the following, the example embodiments will be mainly described with respect to the terminal device 110. It should be appreciated that the actions described with respect to the terminal device 110 can be performed by the application 120 on the terminal device 110, or can be performed by the application 120 in cooperation with its server (e.g., the server 130).

[0037] Reference is first made to FIGS. 2A-2D Example embodiments of the present disclosure for speech processing are described. FIGS. 2A-2D Schematic diagrams of conversation interfaces 200A to 200D with virtual objects are shown in accordance with some embodiments of the present disclosure.

[0038] It should be noted that the methods for speech processing of the present disclosure can be applied to various applications such as entertainment, IM, social, etc. The object to which the speech input is directed can be any suitable type of object such as other users or virtual objects. In this context, the virtual object may, for example, be a digital robot implemented by means of a machine learning model (e.g., a language model) such as a digital assistant or a virtual chat object. For ease of understanding, the following will mainly be described by way of example with respect to the interaction application with virtual objects, but it should be understood that this is merely exemplary and is not intended to imply any limitation. The embodiments for speech processing described with respect to the interaction application with virtual objects can be applied to other types of applications.

[0039] In some embodiments, the terminal device 110 can present a conversation interface for the user 140 to have a conversation with a virtual object. FIG. 2A An example of a conversation interface 200A with a virtual object is shown by way of example of having a conversation with a virtual object A. The user 140 can have a conversation with the virtual object A through the interface 200A in text, voice, etc.

[0040] In some embodiments, the interface 200A includes a voice input region 201. The user 140 can provide voice input through an operation on the voice input region 201 to have a conversation with the virtual object A. In some embodiments, if a first type of operation on the voice input region 201 is detected, the terminal device 110 can start a first input mode to receive voice input. For example, the voice input region 201 can include a voice input control, and the first type of operation can be a press operation. The user 140 can start the first input mode by pressing the voice input control. In this input mode, the user can release the press on the voice input control to send voice, or slide up to cancel the voice input. Taking the user having a conversation with the virtual object A as an example, the first input mode can be a conversation content input mode. In this mode, the user can input conversation content such as “hello”.

[0041] In some embodiments, in the first input mode, the voice input region 201 is presented in a first style. Referring to FIG. 2B An example of a conversation interface 200B with a virtual object is shown. In the interface 200B, the voice input region 201 is presented in the first style. In this input mode, the voice input region 201 in the first style includes a black input box and two gray circles.

[0042] In some embodiments, in the first input mode, if a first operation of a second type on the voice input region 201 is detected, the terminal device 110 can switch to a second input mode to receive voice input. For example, the first operation of the second type can be a sliding operation. The user 140 can switch the first input mode to the second input mode by performing a sliding operation to the left side in the voice input region 201. Taking the user having a conversation with the virtual object A as an example, the second input mode can be an input mode of non-conversation content such as action, expression, emotion, etc. In this mode, the user can input non-conversation content such as “smiling face” and “waving hand”.

[0043] In some embodiments, in the second input mode, the voice input region 201 is presented in a second style. Referring to FIG. 2C An example of a conversation interface 200C with a virtual object is shown. In the interface 200C, the voice input region 201 is presented in the second style. In this input mode, the voice input region 201 in the second style includes a white input box and two gray circles. Generally, the second style can be different from the first style in color, shape, etc., to facilitate the user 140 to distinguish the current input mode.

[0044] In some embodiments, the terminal device 110 can present prompt information above the voice input region 201. For example, FIG. 2B and FIG. 2CAs shown, a prompt information 202 can be presented above the voice input area 201. For example, in the dialog content input mode, the prompt information 202 can be set as "Please input dialog", and in the non-dialog content input mode, the prompt information 202 can be set as "Please input action".

[0045] In some embodiments, in the second input mode, if a second operation of a second type on the voice input area 201 is detected by the user 140, the terminal device 110 can switch to the first input mode to receive voice input. The second operation is of the same type as the first operation but in the opposite direction. For example, when the first operation is a sliding operation to the left side, the second operation can be a sliding operation to the right side. In some embodiments, the user 140 can also stop the operation on the voice input area in the second input mode to end the voice input and send the voice without returning to the first input mode.

[0046] In some embodiments, the terminal device 110 can present a voice recognition result corresponding to the received voice input. The voice recognition result can be presented after the voice input is ended, or can be presented step by step while the voice input is ongoing. In the voice recognition result, a first recognition result corresponding to the voice received in the first input mode is distinguishable from a second recognition result corresponding to the voice received in the second mode. The first recognition result and the second recognition result can include any suitable type of content, such as text, symbols, etc. The presentation effect of the recognition results in different input modes can be distinguished in any suitable manner.

[0047] In some embodiments, the recognition results in different input modes can be separated by a separator. In such embodiments, the recognition results in different input modes can have the same type of content (e.g., both are text), or can have different types of content (e.g., the recognition result in one input mode is text, and the recognition result in another input mode is symbols, etc.). For example, the second recognition result can have a first separator and a second separator at the beginning and the end, respectively. For example, the first separator can be a left bracket "(", and the second separator can be a right bracket ")". For example, the first recognition result can be presented between the left bracket and the right bracket. FIG. 2D An example of the dialog interface 200D with a virtual object is shown. In the interface 200D, a voice recognition result "Hello (smiling face)" is presented in the voice input area 201. "Hello" is a first recognition result corresponding to the voice received in the first input mode (dialog content input mode), and "smiling face" is a second recognition result corresponding to the voice received in the second input mode (non-dialog content input mode). The non-dialog content in the voice recognition result is distinguished from the dialog content by left and right brackets. The punctuation before and after the brackets is not presented.

[0048] Alternatively or additionally, in some embodiments, the recognition results in different input modes can be different types of content. Illustratively, the first recognition result can include text corresponding to the speech received in the first input mode, and the second recognition result can include a symbol (e.g., an emoticon, a gesture symbol, etc.) matching the speech received in the second input mode. For example, the text corresponding to the speech received in the second input mode can be recognized first, and the symbol can be presented based on semantic understanding of the recognized text. Continuing the example above, the "smiling face" in the second recognition result can be presented as an emoticon of "smile" based on semantic understanding. The presented symbol (e.g., an emoticon or a gesture symbol) can be selected from a plurality of symbols obtained in advance, or can be generated based on semantic understanding. In this document, embodiments are mainly described with text as the recognition result, but it should be understood that this is only illustrative and is not intended to be any limitation.

[0049] In some embodiments, the terminal device 110 can present a first preset motion effect during a process in which the speech input area 201 is transformed from the first style shown in FIG. 2B to the second style shown in FIG. 2C In some embodiments, the terminal device 110 can present a second preset motion effect during a process in which the speech input area 201 is transformed from the second style shown in FIG. 2C to the first style shown in FIG. 2B In some embodiments, the operation of triggering the switching between the first input mode and the second input mode can be a sliding operation. The presentation of the first preset motion effect or the second preset motion effect can be in response to the sliding operation satisfying a preset condition. The preset condition can include that a sliding distance of the sliding operation is greater than or equal to a preset distance, for example, the sliding distance is greater than or equal to 50pt. Alternatively or additionally, the preset condition can include that a sliding speed of the sliding operation is greater than or equal to a preset speed, for example, the sliding speed is greater than or equal to 100pt / s.

[0050] Reference is made to FIG. 3 showing one example of a process 300 of speech input area style transformation. As shown in FIG. 3 FIG. 3, the process 300 includes the following steps. FIG. 3As shown, in the dialogue content input mode, the voice input area is presented in a first style as shown in 301. The circles on the left and right sides of the input box can be used to guide the user to switch the input mode. When the user 140 performs a side swipe (306) operation (i.e. the first operation described above) on the voice input area, and the sliding distance of the side swipe operation reaches 50pt or the sliding speed reaches 100pt / s, the terminal device 110 is triggered (307) to switch to the non-dialogue content input mode. In the switching process, the left circle of the voice input area is gradually enlarged by 50-130pt (i.e. the first preset motion effect described above) in the manner shown in 302, and the duration of the first preset motion effect can be 100ms. When switching to the non-dialogue input mode, the voice input area is presented in a second style as shown in 303. When the user 140 performs a back swipe (308) operation (i.e. the second operation described above) on the voice input area, and the sliding distance of the back swipe operation reaches 50pt or the sliding speed reaches 100pt / s, the terminal device 110 is triggered (309) to switch to the dialogue content input mode. In the switching process, the left circle of the voice input area is gradually reduced to 30pt (i.e. the second preset motion effect described above) in the manner shown in 304, and the duration of the second motion effect can be 100ms. When switching back to the non-dialogue input mode, the voice input area is presented in the first style as shown in 305. In the presentation process of the first motion effect and the second motion effect, the size, color and other parameters of the left circle and the right circle can be designed as needed. It should be noted that the above operation of triggering mode switching and the motion effect in the switching process are only exemplary and are not intended to limit in any way.

[0051] In some embodiments, if it is detected that the current input mode is switched from the first input mode to the second input mode, the terminal device 110 can provide the first voice received in the first input mode to the speech recognition model in the server 130 to obtain a first recognition result corresponding to the first voice. In some embodiments, if it is detected that the current input mode is switched from the second input mode to the first input mode, the terminal device 110 can provide the second voice received in the second input mode to the speech recognition model in the server to obtain a second recognition result corresponding to the second voice. In the case where the first recognition result and the second recognition result are both texts, in the process of switching the first input mode to the second input mode and switching the second input mode to the first input mode, the terminal device 110 can respond to the sliding operation by sending a corresponding instruction to the server 130 through the software development function block to mark that a separator will be generated at this point.

[0052] In some embodiments, the server 130 can combine the first recognition result and the second recognition result as at least part of the speech recognition result. In the case where the terminal device 110 receives multiple segments of the first speech or multiple segments of the second speech, the combination order of the recognition results corresponding to each segment of speech is consistent with the input order of the corresponding speech.

[0053] Reference is made to FIG. 4 An example of the generation process 400 of the speech recognition result is shown. As FIG. 4 As shown, in the process of one speech input, based on the switching of the input mode, the speech is divided into 5 segments of conversation for processing respectively, i.e. conversation-1 in block 401 to conversation-5 in block 405. Conversation-1, conversation-3 and conversation-5 are the conversations obtained in the first input mode, and conversation-2 and conversation-4 are the conversations obtained in the second input mode. Multiple conversations can exist simultaneously, for example, at the time point shown by the dashed line 407, conversation-2 to conversation-4 exist simultaneously. Each conversation is received by the recorder as a segment of speech, and then a corresponding segment of text is obtained by using the speech recognition model. The server 130 inserts a left bracket and a right bracket at the beginning and the end of the "2-text" and the "4-text" obtained in the second input mode respectively, and combines the segments of text to obtain the speech recognition result ASR(Automatic Speech Recognition, speech recognition)-text, where ASR-text = 1-text + (2-text) + 3-text + (4-text) + 5-text. It should be understood that the text in the brackets can be presented as a corresponding symbol.

[0054] In order to more clearly understand the speech processing scheme of the present disclosure, reference is made to FIG. 5 An example of the interaction process 500 between the ASR manager, the ASR root conversation, the speech server and the ASR sub-conversation is described.

[0055] In FIG. 5 In the example, the ASR manager 520 can be one method or operation in the terminal device 110, and the speech server 530 can be one method or operation in the server 130.

[0056] Specifically, FIG. 5 An interaction process 500 between the ASR manager 520, the ASR root conversation 525, the speech server 530 and the ASR sub-conversation 535 in the process of one speech input of the user 140 is shown. As FIG. 5As shown, the user presses (501) the input box through the ASR manager 520 to start the first input mode, such as the voice input mode. The ASR root session 525 creates (502) the voice recorder receiver and initializes (503) the software development function block. The ASR root session 525 establishes (504) a connection to the voice server 530 and streams (505) the audio. The voice server 530 returns (506) the recognition result to the ASR root session 525 after recognizing the text based on the audio.

[0057] Continuing FIG. 5 The process, when the user switches the input mode through the side swipe and the like, such as the input (507) action, the ASR root session 525 creates (508) the ASR sub-session 535. It can be understood that the ASR root session 525 creates an ASR sub-session 535 each time the user switches the input mode. The ASR sub-session 535 also creates (509) the voice recorder receiver and initializes (510) the software development function block. The ASR sub-session 535 establishes (511) a connection to the voice server 530 and streams (512) the audio. The voice server 530 returns (513) the recognition result to the ASR sub-session 535 after recognizing the text based on the audio. The sub-ASR results of one or more ASR sub-sessions 535 are aggregated (514) to obtain the voice recognition result. The ASR root session 525 returns (515) the recognition result to the ASR manager 520, and the voice recognition result is presented in the terminal device 110. Finally, the ASR manager 520 sends (516) the information of releasing the resources to the ASR root session 525, and the ASR root session 525 sends (517) the information of releasing the resources to each ASR sub-session 535, completing the process of one voice input.

[0058] In some embodiments, the voice input received by the terminal device 110 includes one or more pieces of first voice received in the first input mode and one or more pieces of second voice received in the second input mode, and the voice recognition result includes one or more pieces of first recognition result corresponding to the one or more pieces of first voice respectively and one or more pieces of second recognition result corresponding to the one or more pieces of second voice respectively. It should be noted that after starting the first input mode, the user may not input the first voice, and at this time the first piece of voice received by the terminal device 110 is the second voice. For example, in the process of one voice input, the action content such as "(wave hand)" can be received at the beginning.

[0059] In some embodiments, the voice input is issued by the user to the virtual object in the interaction between the user and the virtual object, and the one or more pieces of first recognition result includes the dialogue content of the user to the virtual object, and the one or more pieces of second recognition result includes the explanation information of the dialogue content.

[0060] In some embodiments, the server 130 can generate a reply to the voice input by using a language model based on the voice recognition results. The second recognition result in the voice recognition results is used to assist the language model to understand the first recognition result adjacent to (i.e. before or after) the second recognition result. Taking the application of interaction with the virtual object as an example, the first recognition result can be the conversation content "I am very happy today" input by the user, and the second recognition result can be the expression "(depressed)" input by the user. When generating the reply by using the voice recognition model, the expression "(depressed)" can assist the model to understand the real emotion implied by the user in the conversation "I am very happy today".

[0061] In some embodiments, the server 130 can generate a first part of the model prompt information based on one or more first recognition results, and the first part is used to provide the conversation content. The server 130 can also generate a second part of the model prompt information based on one or more second recognition results, and the second part is used to provide the explanation information to the conversation content. The server 130 can provide the model prompt information to the language model to obtain the output of the language model, and determine the reply based on the output of the language model. The language model can be trained based on the sample conversation content of the user's conversation with the virtual object.

[0062] In summary, by the present disclosure, different input modes of voice can be flexibly switched, and distinguishable voice recognition results can be presented according to the voice received in different input modes, so that the user can input various contents for different purposes by voice, and the convenience and practicability of voice input are improved, and a better user experience is provided.

[0063] Example Process

[0064] FIG. 6 A flowchart of a process 600 for voice processing according to some embodiments of the present disclosure is shown. The process 600 can be implemented at the terminal device 110 or coordinated by the server 130. The process 600 is described below with reference to 1.

[0065] At block 610, in response to the first type of operation on the voice input area, the terminal device 110 starts the first input mode to receive the voice input.

[0066] At block 620, in the first input mode, in response to the second type of first operation on the voice input area, the terminal device 110 switches to the second input mode to receive the voice input.

[0067] At block 630, the terminal device 110 presents speech recognition results corresponding to the received speech input, in which a first recognition result corresponding to speech received in the first input mode is distinguishable from a second recognition result corresponding to speech received in the second mode.

[0068] In some embodiments, the second recognition result has a first delimiter and a second delimiter at the beginning and the end, respectively.

[0069] In some embodiments, the first recognition result includes text corresponding to speech received in the first input mode, and the second recognition result includes symbols matching speech received in the second input mode.

[0070] In some embodiments, the process 600 further includes, in the second input mode, switching to the first input mode to receive speech input in response to a second operation of a second type on the speech input area, wherein the second operation is opposite to the first operation.

[0071] In some embodiments, in the first input mode, the speech input area is presented in a first style, and in the second input mode, the speech input area is presented in a second style.

[0072] In some embodiments, the process 600 further includes at least one of: presenting a first preset motion effect in a process in which the speech input area transitions from the first style to the second style, or presenting a second preset motion effect in a process in which the speech input area transitions from the second style to the first style.

[0073] In some embodiments, the operation that triggers switching between the first input mode and the second input mode is a sliding operation, and the presentation of the first preset motion effect or the second preset motion effect is in response to the sliding operation satisfying at least one of: a sliding distance of the sliding operation being greater than or equal to a preset distance, or a sliding speed of the sliding operation being greater than or equal to a preset speed.

[0074] In some embodiments, the speech recognition results are generated by: in response to switching from the first input mode to the second input mode, providing first speech received in the first input mode to a speech recognition model to obtain a first recognition result corresponding to the first speech; in response to switching from the second input mode to the first input mode, providing second speech received in the second input mode to the speech recognition model to obtain a second recognition result corresponding to the second speech; and merging the first recognition result and the second recognition result as at least part of the speech recognition results.

[0075] In some embodiments, the received speech input includes one or more segments of first speech received in the first input mode and one or more segments of second speech received in the second input mode, the speech recognition result includes one or more segments of first recognition result corresponding to the one or more segments of first speech respectively and one or more segments of second recognition result corresponding to the one or more segments of second speech respectively, and the process 600 further includes: generating a reply to the speech input based on the speech recognition result by using the language model, wherein the second recognition result in the speech recognition result is used to assist the language model to understand the first recognition result adjacent to the second recognition result.

[0076] In some embodiments, the speech input is uttered by the user to the virtual object in the user's interaction with the virtual object, and the one or more segments of first recognition result includes a conversation content of the user to the virtual object, and the one or more segments of second recognition result includes explanation information to the conversation content.

[0077] In some embodiments, generating the reply to the speech input by using the language model includes: generating a first part of model prompt information based on the one or more segments of first recognition result, the first part being used to provide the conversation content; generating a second part of model prompt information based on the one or more segments of second recognition result, the second part being used to provide the explanation information to the conversation content; providing the model prompt information to the language model to obtain an output of the language model; and determining the reply based on the output of the language model.

[0078] Example Devices and Apparatus

[0079] FIG. 7 A schematic structural block diagram of an apparatus 700 for speech processing according to some embodiments of the present disclosure is shown. The apparatus 700 can be implemented as or included in the terminal device 110. Various modules / components in the apparatus 700 can be implemented by hardware, software, firmware, or any combination thereof.

[0080] As shown in FIG. 7 The apparatus 700 includes a starting module 710 configured to start a first input mode to receive speech input in response to a first type of operation on a speech input area. The apparatus 700 further includes a first switching module configured to switch to a second input mode to receive speech input in response to a second type of first operation on the speech input area in the first input mode. The apparatus 700 further includes a result presenting module configured to present a speech recognition result corresponding to the received speech input, in which first recognition result corresponding to speech received in the first input mode is distinguishable from second recognition result corresponding to speech received in the second mode.

[0081] In some embodiments, the second recognition result has a first delimiter and a second delimiter at the beginning and the end, respectively.

[0082] In some embodiments, the first recognition result includes text corresponding to the speech received in the first input mode, and the second recognition result includes symbols matching the speech received in the second input mode.

[0083] In some embodiments, the apparatus 700 further includes a second switching module configured to, in the second input mode, switch to the first input mode to receive the speech input in response to a second operation of a second type on the speech input area, wherein the second operation is opposite to the first operation direction.

[0084] In some embodiments, in the first input mode, the speech input area is presented in a first style, and in the second input mode, the speech input area is presented in a second style.

[0085] In some embodiments, the apparatus 700 further includes a first animation module configured to present a first preset animation in a process in which the speech input area transitions from the first style to the second style. The apparatus 700 further includes a second animation module configured to present a second preset animation in a process in which the speech input area transitions from the second style to the first style.

[0086] In some embodiments, the operation that triggers the switching between the first input mode and the second input mode is a sliding operation, and the presentation of the first preset animation or the second preset animation is in response to the sliding operation satisfying at least one of: a sliding distance of the sliding operation being greater than or equal to a preset distance, or a sliding speed of the sliding operation being greater than or equal to a preset speed.

[0087] In some embodiments, the apparatus 700 further includes a speech recognition module configured to, in response to switching from the first input mode to the second input mode, provide a first speech received in the first input mode to a speech recognition model to obtain a first recognition result corresponding to the first speech; in response to switching from the second input mode to the first input mode, provide a second speech received in the second input mode to the speech recognition model to obtain a second recognition result corresponding to the second speech; and combine the first recognition result and the second recognition result as at least part of a speech recognition result.

[0088] In some embodiments, the received speech input includes one or more segments of first speech received in a first input mode and one or more segments of second speech received in a second input mode, the speech recognition result includes one or more segments of first recognition result corresponding to the one or more segments of first speech respectively and one or more segments of second recognition result corresponding to the one or more segments of second speech respectively, and the apparatus 700 further includes a reply module configured to generate a reply to the speech input based on the speech recognition result, wherein the second recognition result in the speech recognition result is used to assist the language model to understand the first recognition result adjacent to the second recognition result.

[0089] In some embodiments, the speech input is uttered by the user to the virtual object in the user's interaction with the virtual object, and the one or more segments of first recognition result includes a conversation content of the user to the virtual object, and the one or more segments of second recognition result includes explanation information to the conversation content.

[0090] In some embodiments, the reply module is further configured to generate a first part of the model prompt information based on the one or more segments of first recognition result, the first part being used to provide the conversation content; generate a second part of the model prompt information based on the one or more segments of second recognition result, the second part being used to provide the explanation information to the conversation content; provide the model prompt information to the language model to obtain an output of the language model; and determine the reply based on the output of the language model.

[0091] The units and / or modules included in the apparatus 700 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine executable instructions stored on a storage medium. In addition to or alternatively, some or all of the units and / or modules in the apparatus 700 can be implemented at least partially by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Products (ASSPs), System-on-a-Chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0092] FIG. 8 A block diagram of an electronic device 800 in which one or more embodiments of the disclosure can be implemented is shown. It should be understood that FIG. 8 The electronic device 800 illustrated is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein.

[0093] As FIG. 8As shown, the electronic device 800 is in the form of a general electronic device. Components of the electronic device 800 can include, but are not limited to, one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processing unit 810 can be a real or virtual processor and capable of executing various processing in accordance with programs stored in the memory 820. In a multi-processing system, multiple processing units can execute computer-executable instructions in parallel to improve the processing power of the electronic device 800.

[0094] The electronic device 800 typically includes a plurality of computer storage media. Such media can be any available media that is located either internally or externally to the electronic device 800, including, but not limited to, memory, volatile and non-volatile, removable, and non-removable. The memory 820 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically-erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 can be a removable or non-removable media, and can include machine-readable media, such as a flash drive, a magnetic disk, or any other media that can be used to store information and / or data and that can be accessed by the electronic device 800.

[0095] The electronic device 800 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 8 disk drives for reading from or writing to a removable, non- volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk (e.g., a CD-ROM, DVD, etc.). In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. The memory 820 can include a computer program product 825 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.

[0096] The communication unit 840 enables communications with other electronic devices over communication media. Additionally, the functionality of the components of the electronic device 800 can be implemented in a single computing cluster or a plurality of computer machines capable of communicating over a communication connection. As such, the electronic device 800 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networking environment.

[0097] Input device 850 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. Output device 860 can be one or more output devices, such as a display, a speaker, a printer, etc. Electronic device 800 can also communicate with one or more external devices (not shown), such as a storage device, a display device, etc., through communication unit 840, as desired, in order to communicate with a user in order to interact with electronic device 800, or to communicate with any device (e.g., a network card, a modem, etc.) that enables electronic device 800 to communicate with one or more other electronic devices. Such communication can be carried out via an input / output (I / O) interface (not shown).

[0098] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.

[0099] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0100] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0101] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0102] The computer program product of the present disclosure can be a computer program product, which is a machine-readable medium (media) having instances of the software embodied thereon, such as computer software, firmware, wireless application protocol (WAP), middleware or microcode. For example, a computer program product can be a floppy disk, a CD-ROM, a DVD, a Blu-ray Disc™, a flash drive, a memory stick, a magnetic tape, or a hard disk drive. The machine-readable medium can be a single medium, or multiple media, of the same or different type. The computer program product can be one or more computer program components embodied in medium and / or transmission signals. The computer program product can have one or more computer program components embodied in medium and / or transmission signals.

[0103] Having thus described the various implementations of the present disclosure in detail, it will be apparent to those skilled in the present technology that numerous modifications and variations of the implementations described above can be made without departing from the scope of the various implementations of the present disclosure. Accordingly, the above description is meant to be taken only by way of example, and the scope of the various implementations of the present disclosure is to be limited only by the appended claims.

Claims

1. A voice processing method, comprising: starting a first input mode to receive voice input in response to a first type of operation on a voice input area; in the first input mode, switching to a second input mode to receive voice input in response to a second type of first operation on the voice input area; presenting a voice recognition result corresponding to the received voice input, the voice recognition result comprising a first recognition result corresponding to voice received in the first input mode and a second recognition result corresponding to voice received in the second input mode, and the first recognition result and the second recognition result being presented distinguishably; and generating a reply to the voice input using a language model based on the voice recognition result, wherein the second recognition result in the voice recognition result is used to assist the language model to understand the first recognition result. 2.The method of claim 1, wherein the second recognition result has a first delimiter and a second delimiter at a beginning and an end, respectively. 3.The method of claim 1, wherein the first recognition result comprises text corresponding to voice received in the first input mode, and the second recognition result comprises symbols matching voice received in the second input mode. 4.The method of claim 1, further comprising: in the second input mode, switching to the first input mode to receive voice input in response to a second type of second operation on the voice input area, wherein the second operation is opposite to the first operation in direction. 5.The method of claim 1, wherein in the first input mode, the voice input area is presented in a first style, and in the second input mode, the voice input area is presented in a second style. 6.The method of claim 5, further comprising at least one of: presenting a first preset dynamic effect in a process in which the voice input area is transformed from the first style to the second style, or presenting a second preset dynamic effect in a process in which the voice input area is transformed from the second style to the first style. 7.The method of claim 6, wherein the operation that triggers switching between the first input mode and the second input mode is a sliding operation, and the presentation of the first preset dynamic effect or the second preset dynamic effect is in response to the sliding operation satisfying at least one of: a sliding distance of the sliding operation is greater than or equal to a preset distance, or a sliding speed of the sliding operation is greater than or equal to a preset speed. 8.The method of claim 1, wherein the voice recognition result is generated by: in response to switching from the first input mode to the second input mode, providing first voice received in the first input mode to a voice recognition model to obtain the first recognition result corresponding to the first voice. ​ ​ ​ ​ ​ in response to switching from the second input mode to the first input mode, providing a second speech received in the second input mode to the speech recognition model to obtain a second recognition result corresponding to the second speech; and merging the first recognition result and the second recognition result as at least part of the speech recognition result.

9. The method of claim 1, wherein the received speech input comprises one or more pieces of first speech received in the first input mode and one or more pieces of second speech received in the second input mode, the speech recognition result comprises one or more pieces of first recognition result corresponding to the one or more pieces of first speech respectively and one or more pieces of second recognition result corresponding to the one or more pieces of second speech respectively, and a piece of second recognition result of the one or more pieces of second recognition result is used to assist the language model to understand a piece of first recognition result adjacent to the piece of second recognition result.

10. The method of claim 9, wherein the speech input is uttered by a user in interaction with a virtual object, and the one or more pieces of first recognition result comprises a conversation content of the user to the virtual object, and the one or more pieces of second recognition result comprises interpretation information to the conversation content.

11. The method of claim 9, wherein generating a reply to the speech input using a language model comprises: generating a first part of model prompt information based on the one or more pieces of first recognition result, the first part being used to provide a conversation content; generating a second part of the model prompt information based on the one or more pieces of second recognition result, the second part being used to provide interpretation information to the conversation content; providing the model prompt information to the language model to obtain an output of the language model; and determining the reply based on the output of the language model.

12. An apparatus for speech processing, comprising: a starting module configured to start a first input mode to receive speech input in response to a first type of operation on a speech input area; a first switching module configured to switch to a second input mode to receive speech input in response to a second type of first operation on the speech input area in the first input mode; a result presenting module configured to present a speech recognition result corresponding to the received speech input, the speech recognition result comprising a first recognition result corresponding to speech received in the first input mode and a second recognition result corresponding to speech received in the second input mode, and the first recognition result and the second recognition result being presented distinguishably; and a reply module configured to generate a reply to the speech input using a language model based on the speech recognition result, wherein the second recognition result in the speech recognition result is used to assist the language model to understand the first recognition result.

13. An electronic device, comprising: at least one processing unit; and ​ ​ ​ at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1-11.

14. A computer readable storage medium having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1-11.

15. A computer program product comprising computer executable instructions, wherein the computer executable instructions implement the method according to any one of claims 1-11 when executed by a processor.

Citation Information

Patent Citations

  • Voice input method, device and terminal

    CN104331265A

  • Content input method and device

    CN109739462A