Information processing method, device, storage medium, and product
By detecting network communication capabilities, the machine learning model is invoked on the server or locally to perform speech-to-text and text-to-speech functions, solving the accuracy problem when network communication is poor and realizing efficient voice dialogue under any network conditions.
Patent Information
- Application Number
- PCT/CN2025/085163
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-26
- Filing Date
- 2025-03-26
- Publication Date
- 2026-01-02
AI Technical Summary
In situations with poor network communication capabilities, existing technologies that utilize server-side machine learning models to perform speech-to-text and text-to-speech functions suffer from poor accuracy in speech synthesis and speech recognition, failing to meet users' real-time dialogue needs.
Based on the network communication capability test results, when the network communication is good, the server-side machine learning model is used to perform speech-to-text and text-to-speech functions; when the network communication is poor, the local machine learning model is called to perform the corresponding function processing.
Under any network communication conditions, it can meet the needs of speech-to-text and text-to-speech, improving dialogue efficiency and response speed while ensuring the accuracy of functions.
Smart Images

Figure CN2025085163_02012026_PF_FP_ABST
Abstract
Description
Information processing method, device, storage medium and product
[0001] The present application claims priority to the Chinese patent application No. 202410842262.6, filed on June 26, 2024, entitled “Information processing method, device, storage medium and product”, the whole content of which is incorporated herein by reference. TECHNICAL FIELD
[0002] The example embodiments of the present disclosure generally relate to the field of computers, and in particular, to an information processing method, device, electronic device, computer-readable storage medium and computer program product. BACKGROUND
[0003] With the development of Internet technology, more and more applications or platforms provide text-to-speech (TTS) function (also known as speech synthesis function) and speech recognition (ASR) function (also known as speech-to-text function), which brings great convenience to the general public. Applications or platforms with TTS function and ASR function can provide voice dialogue function relying on TTS service and ASR service to users by using trained machine learning models. SUMMARY
[0004] In a first aspect of the present disclosure, an information processing method is provided. The method is applied to a client, and the method comprises: in response to a first request for speech-to-text and / or a second request for text-to-speech, detecting a network communication capability between the client and a server; if the network communication capability exceeds a first capability level, sending a first speech corresponding to the first request and / or a second text corresponding to the second request to the server to request the server to perform a speech-to-text function and / or a text-to-speech function; and if the network communication capability is lower than a second capability level, performing the speech-to-text function on the first speech by invoking a local first machine learning model, and / or performing the text-to-speech function on the second text by invoking a local second machine learning model.
[0005] In a second aspect of the present disclosure, an apparatus for information processing is provided. The apparatus is applied to a client, and the apparatus comprises: a communication capability detection module configured to detect a network communication capability between the client and a server in response to a first request for speech-to-text and / or a second request for text-to-speech; a first function execution module configured to send, to the server, a first speech corresponding to the first request and / or a second text corresponding to the second request to request the server to perform a speech-to-text function and / or a text-to-speech function if the network communication capability exceeds a first capability level; and a second function execution module configured to perform the speech-to-text function on the first speech by invoking a local first machine learning model and / or perform the text-to-speech function on the second text by invoking a local second machine learning model if the network communication capability is below a second capability level.
[0006] In a third aspect of the present disclosure, an electronic device is provided. The device comprises at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the electronic device to perform the method of the first aspect.
[0007] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. The medium has stored thereon a computer program which, when executed by a processor, implements the method of the first aspect.
[0008] In a fifth aspect of the present disclosure, a computer program product is provided. The product comprises a computer program which, when executed by a processor, implements the method according to the first aspect of the present disclosure.
[0009] It should be understood that nothing in this section is intended to limit the key or critical features of the embodiments of the present disclosure or limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments thereof with reference to the annexed drawings in which:
[0011] FIG. 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0012] FIG. 2 shows a schematic diagram of an example architecture of information processing according to some embodiments of the present disclosure;
[0013] FIG. 3 shows an example of a flow of information processing according to some embodiments of the present disclosure;
[0014] FIG. 4 shows a flow chart of a method for information processing according to some embodiments of the present disclosure;
[0015] FIG. 5 shows an exemplary structural block diagram of an apparatus for information processing according to some embodiments of the present disclosure; and
[0016] FIG. 6 shows a block diagram of an electronic device that can implement one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather, these embodiments are provided so that the present disclosure can be more thoroughly and completely understood. It will be appreciated that the drawings of the present disclosure are for illustrative purposes only and should not be construed as limiting the scope of the present disclosure.
[0018] In the description of embodiments of the present disclosure, the term "comprising" and its conjugations should be understood to encompass the meanings of "consisting of" and "consisting essentially of", i.e., "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions can also be included below.
[0019] In this document, unless explicitly stated otherwise, performing a step "in response to" A does not mean that the step is performed immediately after A, but can include one or more intermediate steps.
[0020] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the obtaining, use, storage or deletion of the data) should comply with the requirements of relevant laws and regulations and relevant provisions.
[0021] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of information involved in the present disclosure, the scope of use, the use scenario, etc. should be informed to the relevant user and the authorization of the relevant user should be obtained through appropriate means, wherein the relevant user can include any type of right subject, such as individuals, enterprises, groups.
[0022] For example, in response to receiving an active request of a user, a prompt information is sent to the relevant user to explicitly prompt the relevant user that the operation requested to be performed will require obtaining and using information of the relevant user, so that the relevant user can autonomously select whether to provide information to the software or hardware such as an electronic device, an application program, a server or a storage medium, etc. performing the operation of the technical solution of the present disclosure according to the prompt information.
[0023] As an optional but non-limiting implementation manner, in response to receiving an active request of a relevant user, the manner of sending a prompt information to the relevant user may, for example, be a pop-up window manner, in which the prompt information may be presented in the form of text. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide information to the electronic device.
[0024] It can be understood that the above notification and user authorization obtaining process is only illustrative and does not limit the implementation manner of the present disclosure, and other manners meeting relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0025] As used herein, the term "model" can learn an association between respective inputs and outputs from training data, such that a corresponding output can be generated for a given input after training is completed. The generation of a model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is one example of a model based on deep learning. In this document, a "model" can also be referred to as a "machine learning model", a "learning model", a "machine learning network", or a "learning network", which terms are used interchangeably herein.
[0026] A "neural network" is a machine learning network based on deep learning. A neural network is capable of processing inputs and providing corresponding outputs, which typically includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications typically include many hidden layers, thereby increasing the depth of the network. The layers of a neural network are connected in sequence, such that the output of a previous layer is provided as input to a subsequent layer, with the input layer receiving the input to the neural network and the output of the output layer as the final output of the neural network. Each layer of a neural network includes one or more nodes (also referred to as processing nodes or neurons), each of which processes input from the previous layer.
[0027] Generally, machine learning can include three stages, namely a training stage, a testing stage, and an application stage (also referred to as an inference stage). In the training stage, a given model can be trained using a large amount of training data, iteratively updating parameter values until the model is able to derive consistent inferences from the training data that satisfy an expected objective. Through training, the model can be considered to have learned an association (also referred to as a mapping) from input to output. The parameter values of the trained model are determined. In the testing stage, test inputs are applied to the trained model to determine whether the model is able to provide correct outputs, thereby determining the performance of the model. In the application stage, the model can be used to process actual inputs based on the trained parameter values to determine corresponding outputs.
[0028] FIG. 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In the example environment 100, an application 120 is installed in a client 110. A user 140 can interact with the application 120 via the client 110 and / or an attached device of the client 110. For example, the application 120 can capture speech 145 of the user 140 via a speech capture device (e.g., a microphone) of the client 110.
[0029] In embodiments of the present disclosure, the application 120 can be any suitable application with human-to-computer dialog functionality. For example, the application 120 can provide a digital assistant for human-to-computer dialog. The digital assistant can support text dialog services, speech dialog services, and content dialog in other modalities with the user 140. In some embodiments, the application 120 or the digital assistant therein can utilize a machine learning model 160 (which can include one or more machine learning models, e.g., can include a machine learning model 160-1, a machine learning model 160-2, …, a machine learning model 160-N, etc., where N is a positive integer) to support interactions with the user 140. For example, the application 120 or the digital assistant therein can utilize one or more machine learning models 160 to provide question-and-answer services to the user 140.
[0030] In the environment 100, the client 110 can present a user interface 150 of the application 120 if the application 120 is in an active state. The user interface 150 can include various pages that the application 120 is capable of providing, such as a dialog page between the user and the digital assistant (in which current and historical dialog, including text dialog content, can be presented), etc. In some embodiments, the client 110 can play speech 152 and present text 154 in the user interface 150. The speech 152 can include, for example, the speech 145 from the user 140 or speech in response to the speech 145.
[0031] The machine learning models 160 can be different types of models. In some embodiments, one or more machine learning models 160 can be built based on a language model (LM). The machine learning models used are content generative models, capable of generating corresponding outputs based on model inputs. In some embodiments, the language model based machine learning models are capable of processing model inputs in a text modality (e.g., natural language and / or machine language) and / or non-text modality (e.g., images, speech, videos, etc.), and are capable of generating desired outputs from the model inputs and a prompt word. The prompt word here is used to guide the machine learning model to generate an output that can address a user need indicated by the model inputs. In an application scenario for supporting user conversations, the input of the user 140 can be provided as at least a portion of the model inputs (other portions can include the prompt word) to the machine learning model 160. The user input is considered as a question. Based on the model output, a corresponding answer can be generated to be provided to the user 140.
[0032] In some embodiments, one or more machine learning models 160 can be speech related models, including an automatic speech recognition (ASR) model and a text-to-speech (TTS) model. The input of the ASR model is speech, and the output is text. The input of the TTS model is text, and the output is corresponding speech.
[0033] In some embodiments, the client 110 communicates with the server 130 to implement provisioning of services of the application 120. As shown in FIG. 1, the server 130 can invoke the machine learning models 160 to support human-to-computer dialog functionality between the application 120 and the user 140 based on outputs of the machine learning models 160. The client 110 can be any type of mobile terminal, fixed terminal, or portable terminal including a mobile handset, a tablet computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a game device, or any combination thereof, including accessories and peripherals of such devices, or any combination thereof. In some embodiments, the client 110 can also support any type of interface for the user (such as “wearable” circuitry, etc.). The server 130 can be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, etc. The server 130 can be implemented, for example, based on a cloud environment.
[0034] It should be understood that the structures and functions of the various elements in the environment 100 are described for illustrative purposes only, without implying any limitation on the scope of the present disclosure.
[0035] As mentioned above, an application or platform with TTS function and ASR function can utilize a trained machine learning model to provide a user with a voice dialogue function of TTS service and ASR service. The process of the application or platform invoking the machine learning model deployed at the server side is affected by network communication. In the case of poor network communication, the client side can not be able to invoke the machine learning model deployed at the server side. In addition, due to the limitation of the computing power of the client side where the application or platform is deployed, the model capability of the machine learning model deployed locally at the client side is usually weaker than that of the machine learning model deployed at the server side. When the application or platform invokes the machine learning model deployed locally at the client side to implement the TTS function / ASR function, it can cause the accuracy of the synthesized voice / text recognized by voice to be poor.
[0036] Therefore, according to an embodiment of the present disclosure, an improved scheme of information processing is provided. According to the scheme of the present disclosure, in response to a first request for voice-to-text and / or a second request for text-to-voice, the network communication capability between the client side and the server side is detected. If the network communication capability exceeds a first capability level, the first voice corresponding to the first request and / or the second text corresponding to the second request is sent to the server side to request the server side to perform the voice-to-text function and / or the text-to-voice function. If the network communication capability is lower than a second capability level, the first voice-to-text function is performed by invoking a local first machine learning model, and / or the text-to-voice function is performed by invoking a local second machine learning model.
[0037] In this way, the voice-to-text function and / or the text-to-voice function can be performed by means of the machine learning model of the cloud in the case of strong network communication capability, and the voice-to-text function and / or the text-to-voice function can be performed by means of the local machine learning model in the case of poor network communication capability. The voice-to-text requirement and / or the text-to-voice requirement of the user in any network communication can be met, and the dialogue efficiency and response speed can be improved while meeting the accuracy.
[0038] Some example embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0039] FIG. 2 shows a schematic diagram of an example architecture 200 of information processing, according to some embodiments of the present disclosure. The architecture 200 can be implemented at the client 110. For ease of discussion, the architecture 200 will be described with reference to the environment 100 of FIG. 1. It is noted that the operations performed by the aforementioned client 110 and the operations performed by the client 110 described later can be performed by a relevant application (e.g., the application 120) installed on the client 110. In some embodiments, the operations performed on the electronic device can be completed with the assistance of the server 130.
[0040] As shown in FIG. 2, the architecture 200 involves the client 110 and the server 130. The client 110 can detect a network communication capability between the client 110 and the server 130 in response to a first request for speech-to-text and / or a second request for text-to-speech. If the network communication capability between the client 110 and the server 130 is good (e.g., the network communication capability exceeds a first capability level), the client 110 sends a first speech corresponding to the first request and / or a second text corresponding to the second request to the server 130 to request the server 130 to perform an ASR function and / or a TTS function. If the network communication capability is poor (e.g., the network communication capability is lower than a second capability level), it means that the response speed of requesting the speech service from the server is slow, or even can always fail. At this time, the client 110 performs the ASR function on the first speech by invoking a first machine learning model locally, and / or performs the TTS function on the second text by invoking a second machine learning model locally. The second capability level can be, for example, a capability level lower than or equal to the first capability level. In some embodiments, the first machine learning model of the client 110 is also an ASR model, and the second machine learning model is also a TTS model.
[0041] The first speech here can be, for example, a question speech received in a question-and-answer dialogue, and the second text can be, for example, an answer text to the question speech. In the scenario of the question-and-answer dialogue, the question speech of the user is usually converted into corresponding question text to be presented in the user interface and can be used to provide input to the question-and-answer model. In addition, if the question-and-answer model supports input and output in the text modality, after the question-and-answer model determines the answer text corresponding to the question text, the TTS model can be used to continue to convert the answer text into corresponding speech to be played to the user. It can be understood that the speech in the present disclosure can be any appropriate speech of any length, any language, and any tone, and the text can be any appropriate text in the same language as the speech. The question-and-answer dialogue can be, for example, a dialogue between the user and a digital assistant, and the question-and-answer dialogue can occur in any dialogue window.
[0042] As shown in detail in FIG. 2, the client 110 can include a business module 210 and an application module 220. The business module 210 can correspond to an application that supports voice conversation. The business module 210 includes an application voice module 211 and an application message module 212. The application module 220 can include a policy module 221, an online ASR / TTS request module 222, and an offline ASR / TTS request module 223. The server 130 can include an access layer 240, a voice module 250, and a question-answering module 260. The application voice module 211 can be configured to receive voice (e.g., the voice 145) from a user (e.g., the user 140). In some embodiments, the application voice module 211 can receive voice (i.e., “user voice”, also referred to as “first voice” herein) from the user 140. At this time, the application voice module 221 determines that a first request for voice-to-text is received.
[0043] The application voice module 211 can send the user voice to the policy module 221 in the application module 220. The policy module 221 can detect a network communication capability between the client 110 and the server 130. The network communication capability can be determined based on one or more factors such as network signal strength, network link speed, network latency, etc., for measuring networking between the client 110 and the server 130. The policy module 221 can detect the network communication capability between the client 110 and the server 130 by a network probe. The policy module 221 can also determine a capability level corresponding to the network communication capability based on the detection result. The policy module 221 can determine the correspondence between the network communication capability and the capability level in advance, for example.
[0044] For example, the policy module 221 can determine that the network communication capability is in the interval [a1, b1) when it is in capability level A, in the interval [a2, b2) when it is in capability level B, in the interval [a3, b3) when it is in capability level C, and so on. It can be understood that there is no same value in different intervals. For example, there can be no same value in the interval [a1, b1) and the interval [a2, b2).
[0045] If the network communication capability exceeds the first capability level, the policy module 221 can determine that the current network communication is good, and then instruct the online ASR / TTS request module 222 to send a request to the service end 130 to request to perform the ASR function and / or the TTS function by means of the machine learning model deployed in the service end 130 (such a case can also be referred to as online ASR function and / or online TTS function). In some embodiments, the network connection between the client 110 and the service end 130 can be a long connection in compliance with the Transmission Control Protocol (TCP). The client 110 can establish a long connection with the service end 130 through a three-interaction process (also referred to as three-way handshake) specified by the TCP protocol. In response to the first request for speech-to-text and the network communication capability exceeding the first capability level, the policy module 221 can instruct the online ASR / TTS request module 222 to send the first speech corresponding to the first request (for example, the aforementioned user speech from the user) to the access layer 240 of the service end 130. It should be understood that the first capability level here can be any appropriate network capability level configured according to actual needs, to limit the network condition suitable for calling the online ASR / TTS function.
[0046] The network interface 245 in the access layer 240 can receive the user speech from the client 110, and send the received user speech to the speech module 250. The speech module 250 can perform the ASR function and / or the TTS function by means of the ASR model 271 and / or the TTS model 272 in the speech service 270. The ASR model 271 and / or the TTS model 272 are deployed in the service end 130, and can also be referred to as the ASR model 271 and / or the TTS model 272 are deployed in the cloud. Exemplarily, the speech module 250 can send the user speech to the ASR model 271, and obtain the ASR text (i.e., the first speech converted text) corresponding to the user speech from the ASR model 271.
[0047] The speech module 250 can send the obtained ASR text to the network interface 245 of the access layer 240 (this process can also be referred to as the speech module 250 returning the ASR text to the network interface 245). The access layer 240 can in turn send the obtained ASR text to the application module 220 (e.g., the online ASR / TTS request module 222 in the application module 220) of the client 110. The application module 220 (e.g., the policy module 221 in the application module 220) can send the received ASR text to the application speech module 211 of the service module 210. The application speech module 211 can send the ASR text to the application message module 212 so that the application message module 212 can present the ASR text to the user 140. For example, the application message module 212 can present the ASR text (i.e., the first text converted from the first speech) in the user interface corresponding to the question-and-answer session.
[0048] In some embodiments, the speech module 250 in the server 130 can also construct a sending message based on the obtained ASR text and send the sending message to the question-and-answer module 260. In some embodiments, to improve the accuracy of the answer generated by the question-and-answer module 260, the speech module 250 can also send the context of the first speech to the question-and-answer module 260 together with the sending message. The question-and-answer module 260 can determine the answer text (i.e., the second text) for the sending message by means of the question-and-answer model 280. The question-and-answer module 260 can send the answer text to the network interface 245 of the access layer 240 (this process can also be referred to as the question-and-answer module 260 returning the answer text to the network interface 245).
[0049] The access layer 240 can send the obtained answer text to the application module 220 (e.g., the online ASR / TTS request module 222 in the application module 220) of the client 110. The application module 220 (e.g., the policy module 221 in the application module 220) can send the received answer text to the application message module 212 of the service module 210. In some embodiments, the application message module 212 can present the answer text to the user 140. For example, the application message module 212 can present the answer text in the user interface (e.g., the conversation window) corresponding to the question-and-answer session.
[0050] In some embodiments, the application message module 212 can determine, in response to receiving the reply text, that a second request for text-to-speech is received. In some embodiments, the application message module 212 can send the reply text to the policy module 221 in the application module 220. The policy module 221 can determine, based on the network communication capability between the client 110 and the server 130, whether to request the server 130 to perform the text-to-speech function or invoke a local TTS model to perform the text-to-speech function on the reply text. In some embodiments, the policy module 221 in the application module 220, upon receiving the reply text, can determine that the TTS operation on the reply text should be implemented with the client's local model capability or invoke the online model capability at the server 130.
[0051] Similarly, if it is determined that the network communication capability exceeds the first capability level, the policy module 221 can send a second request for a second text (e.g., the reply text) to the access layer 240 of the server 130. The network interface 245 in the access layer 240 can receive the reply text from the client 110 and send the received reply text to the speech module 250. In some embodiments, to improve processing efficiency, the question-answering module 260 in the server 130 can also send the reply text directly to the speech module 250. This process can also be referred to as the question-answering module 260 returning the reply text to the speech module 250.
[0052] In some embodiments, the speech module 250 can send the reply text to the TTS model 272 and obtain, from the TTS model 272, a second speech corresponding to the reply text, which can also be referred to as a TTS speech. The speech module 250 can send the obtained TTS speech to the network interface 245 of the access layer 240 (this process can also be referred to as the speech module 250 returning the TTS speech to the network interface 245). The access layer 240 can in turn send the obtained TTS speech to the application module 220 of the client 110 (e.g., the online ASR / TTS request module 222 in the application module 220).
[0053] The application module 220 (e.g., the policy module 221 in the application module 220) can send the received TTS speech to the application speech module 211 of the service module 210. In some embodiments, the application speech module 211 can send the TTS speech to the application message module 212 so that the application message module 212 can play the TTS speech to the user 140. For example, the application message module 212 can play the TTS speech (i.e., the second text converted second speech) in the user interface (e.g., the dialog window) corresponding to the question-answering dialog.
[0054] In some embodiments, in response to determining that the network communication capability is below the second capability level, the policy module 221 can determine that the current network communication is poor to perform the ASR function and / or the TTS function by means of the machine learning model deployed at the server 130, instruct the offline ASR / TTS request module 223 to perform the ASR function and / or the TTS function by means of the machine learning model deployed locally at the client 110 (this case can also be referred to as offline ASR function and / or offline TTS function).
[0055] In some embodiments, in response to the first request for speech-to-text and the network communication capability being below the second capability level, the policy module 221 in the client 110 can instruct the offline ASR / TTS request module 223 to perform the ASR function by means of the ASR model 231 in the local model capability 230. Illustratively, the offline ASR / TTS request module 223 can send the first speech to the ASR model 231 and obtain the ASR text corresponding to the first speech from the ASR model 231.
[0056] Similarly, in some embodiments, in response to the second request for text-to-speech and the network communication capability being below the second capability level, the policy module 221 in the client 110 can instruct the offline ASR / TTS request module 223 to perform the TTS function by means of the TTS model 232 in the local model capability 230. Illustratively, the offline ASR / TTS request module 223 can send the reply text to the TTS model 232 and obtain the TTS speech corresponding to the reply text from the TTS model 232. It should be appreciated that the second capability level here can be any appropriate network capability level configured according to actual needs to define the network condition suitable for invoking the local ASR / TTS function.
[0057] Alternatively or additionally, in some embodiments, if the network communication capability is below the second capability level, the client 110 can also preferentially send a network request to the server 130 to request the online ASR / TTS function to perform conversion on the current user speech or the reply text. The network request can include the first speech for speech-to-text applied by the first request and / or the second text for text-to-speech applied by the second request. If no network response to the network request is received from the server 130 after exceeding a predetermined duration (which can be any appropriate time such as 3 seconds, 5 seconds, 10 seconds, etc.), the client 110 performs the speech-to-text function on the first speech by invoking the local ASR model and / or performs the text-to-speech function on the second text by invoking the local TTS model.
[0058] For example, the policy module 221 can send a network request including the first speech and / or the second text to the server 130 in a case where the network communication capability is determined to be below the second capability level, and instruct the offline ASR / TTS request module 223 to perform the ASR function and / or the TTS function by means of the ASR model 231 and / or the TTS model 232 in the local model capability 230 in a case where no network response to the network request is received from the server 130 after a predetermined duration of time. In this way, the machine learning model at the server 130 can still be prioritized for use in a case where the network communication capability is weak, and the local machine learning model is only switched to in a case where the request times out.
[0059] In some embodiments, in a case where the network communication capability is determined to be below the second capability level, the client 110 can first request the local ASR model 231 to generate one or more text segments in the first text corresponding to the first speech, and present the generated text segments in a timely manner. The text segments can be one word, one phrase, one sentence, etc. in the first text. The process of presenting the one or more text segments in response to the ASR model generating the one or more text segments in the first text corresponding to the first speech can also be regarded as a streaming text presentation by the client 110. While streaming the text, if the network communication capability between the client 110 and the server 130 is detected to exceed the third capability level, the client 110 can still send a network request including the first speech to the server 130 to request the server 130 to perform the ASR function on the first speech. The third capability level can be equal to, higher than, or lower than the first capability level, and the scope of embodiments of the present disclosure does not limit this.
[0060] For example, the policy module 211 in the client 110 can detect the network communication capability between the client 110 and the server 130 in real time, and generate the first text corresponding to the first speech by means of the ASR model 271 at the server 130 in a case where the network communication capability is determined to exceed the third capability level. The client 110 can replace the first text segments previously generated by means of the ASR model 231 with the first text generated by means of the ASR model 271 in response to obtaining the first text from the server 130. In this way, the client 110 can replace the text segments previously generated by means of the local ASR model 231 with the first text obtained from the server 130. In this way, in a case where the network communication capability is initially unstable or poor, the local model capability can be used to obtain the ASR result in a timely manner for presentation to the user. After the network communication capability recovers, the original text can be replaced with the more accurate ASR result obtained from the server 130 through the network request.
[0061] A similar procedure can also be applied for TTS functionality. In some embodiments, upon determining that the network communication capability is below the second capability level, the client 110 can first request the local TTS model 232 to generate one or more speech segments of the second speech corresponding to the second text, and play the generated speech segments in time. The speech segments can be, for example, speech segments corresponding to certain text segments in the second text. The first speech segment can also be, for example, a speech segment of a predetermined time length (e.g., 5 seconds, 10 seconds, or any other suitable time length) or a speech segment with complete semantics. The procedure of requesting the TTS model 232 to generate one or more speech segments of the second speech corresponding to the second text and playing the speech segments in time can also be considered as a voice streaming of the client 110. While streaming the voice, if it is detected that the network communication capability between the client 110 and the server 130 exceeds the third capability level, the client 110 can still send a network request including the second text to the server 130 to request the server 130 to perform TTS functionality on the second text.
[0062] Exemplarily, the policy module 211 in the client 110 can detect the network communication capability between the client 110 and the server 130 in real time, and upon determining that the network communication capability exceeds the third capability level, generate the second speech corresponding to the second text by means of the TTS model 272 at the server 130. If the client 110 can obtain one or more speech segments of the second speech from the server 130, the client 110 can resume playing by using the speech segments obtained from the server 130. Specifically, if the client 110 determines that the currently played speech segment (e.g., the first speech segment) corresponds to a certain speech segment (e.g., speech segment A) obtained from the server 130, and the client 110 also receives one or more speech segments (e.g., speech segment B) after the speech segment A from the server 130. At this time, the client 110 can play the speech segment after the speech segment corresponding to the first speech segment (i.e., play the speech segment B after the speech segment A) after the current speech segment (e.g., the first speech segment) is played. That is, considering that the speech segments provided by the server have higher quality, the client 110 can start from the speech segment B. The second speech is played by using the speech segments received from the server 130.
[0063] In some embodiments, if the network communication capability exceeds the first capability level, the client 110 can perform, on the basis of sending the first speech to the server 130, an ASR function on the first speech by invoking the local ASR model 231 to obtain one or more text segments corresponding to one or more speech segments in the first speech. The client 110 can replace the presented one or more text segments with the one or more text segments received from the server 130 in response to receiving the one or more text segments converted from the first speech from the server 130 while presenting the one or more text segments. That is, in the case of better network communication capability, considering that the communication between the client and the server always has a certain network delay and the possibility of network transmission failure, in order for the user to obtain feedback as soon as possible, the client still additionally utilizes the local model capability to obtain the ASR result, and when necessary (i.e., when the text segment is not received from the server in time), the client presents the partial text segment generated locally first and then replaces it with the ASR result of higher quality received from the server.
[0064] Similarly, in some embodiments, if the network communication capability exceeds the first capability level, the client 110 can also perform, on the basis of sending the second text to the server 130, a TTS function on the second text by invoking the local TTS model 232 to obtain one or more speech segments corresponding to one or more text segments in the second text. The client 110 can continue to play the second speech segment received from the server 130 after the end of the last played first speech segment in response to receiving the one or more speech segments converted from the second text from the server 130 while playing the one or more speech segments, the text segment corresponding to the second speech segment being after the text segment corresponding to the first speech segment. In some embodiments, the client 110 can also present the text segment corresponding to each of the one or more speech segments while playing the one or more speech segments. Thus, since the client 110 has obtained the second text before requesting the cloud to perform TTS, the client 110 can present the text segment and the speech segment to the user at the same time more quickly by quickly generating the speech segment by utilizing the local TTS model 232.
[0065] According to the above embodiments, in the case that the network communication capability is strong, the client 110 can still first process the first speech corresponding to the first request and / or the second text corresponding to the second request by using the local machine learning model, and first play the speech segment generated by the local machine learning model and / or present the text segment generated by the local machine learning model to the user. At the same time, the client 110 can request the server 130 to perform the ASR function and / or the TTS function by means of the cloud machine learning model. The client 110 can respond to the first speech and / or the second text generated by the cloud machine learning model obtained from the server 130, and continue playing the first speech and / or replace the second text generated by the local machine learning model with the second text generated by the cloud machine learning model.
[0066] Referring to FIG. 3, FIG. 3 shows an example of a flow 300 of information processing according to some embodiments of the present disclosure. The client 110 can determine that a first request for speech-to-text is received in response to obtaining 301 a user speech. The client 110 can detect a network communication capability between the client 110 and the server 130 in response to the first request. If the network communication capability exceeds a first capability level, the client 110 can send a first speech (i.e., the user speech) corresponding to the first request to the server 130 to request the server 130 to perform a speech-to-text function. If the network communication capability is lower than a second capability level, the client 110 can perform the speech-to-text function on the first speech by invoking a local ASR model. The client 110 can obtain 302 a first text (which can also be referred to as ASR text) corresponding to the first speech from the server 130 or from the local ASR model. The time consumed by the client 110 from receiving the first request to obtaining the first text can be referred to as ASR time consumption.
[0067] The client 110 can perform pre-processing (which can include any appropriate processing operation) on the obtained first text, and determine a message to be sent to the server 130 based on the pre-processed first text. In the process of performing pre-processing on the first text, the client 110 can determine, for example, a context of the first text. The time consumed by the client 110 in performing pre-processing can be referred to as pre-processing time consumption. The client 110 can send 303 the message to the server 130 to instruct the server 130 to process the message.
[0068] After receiving (304) the message, the service end 130 can process the message. For example, the service end 130 can process the message to determine the model input to be provided to the question and answer model. The time consumed by the service end 130 in processing the message can be referred to as message processing latency. The service end 130 can further perform pre-processing for the question and answer model call. For example, the service end 130 can perform pre-processing for the question and answer model call by obtaining the plug-in of the question and answer model, assembling the PE system of the question and answer model, and the like. The time consumed by the service end 130 in performing the pre-processing can also be referred to as pre-processing latency. After performing the pre-processing, the service end 130 can call (305) the question and answer model. The service end 130 can provide the model input determined based on the message to the question and answer model. The time consumed by the service end 130 in providing the model input to the question and answer model to obtaining the response text from the question and answer model (i.e., the time consumed by the question and answer model in outputting the response text based on the model input) is referred to as execution latency of the question and answer model. The service end 130 can obtain (306) the response text output by the question and answer model for the message from the question and answer model. The service end 130 can send the obtained response text to the client 110.
[0069] The client 110 can receive (307) the response text sent by the service end 130. The client 110 can refer to the time consumed in starting to receive the response text to receiving the response text as response receiving latency. After obtaining the response text, regardless of whether the network communication capability is higher than the first capability level or lower than the second capability level, the client 110 can first call the local TTS model to perform text-to-speech conversion on the response text. If the network communication capability is higher than the first capability level, the client 110 can also send the response text to the service end 130 to process the response text by means of the TTS model deployed in the service end 130. The process in which the client 110 processes the response text by means of the local TTS model and the TTS model in the service end 130 can be referred to as text-to-TTS (308).
[0070] It can be understood that the local TTS model can generate one or more speech segments corresponding to one or more text segments in the response text earlier than the TTS model in the service end 130, and the client 110 can first obtain the one or more speech segments generated by the local TTS model and play (309) the one or more speech segments (this process can also be referred to as playing TTS speech). The time consumed in starting to generate the speech segment by means of the TTS model after receiving the response text to receiving the first speech segment and starting to play the first speech segment can be referred to as TTS conversion latency. The response receiving latency and the TTS conversion latency can also be collectively referred to as TTS first packet latency.
[0071] It can be understood that when the client 110 starts playing the voice segment, the client 110 has not obtained the voice segment from the server 130. Thus, one or more voice segments can be quickly generated by using the local TTS model, so that the user can obtain part of the feedback content as soon as possible. After starting to receive the voice segment from the server, since the TTS function of the server is often more accurate, the voice segment with higher quality received from the server can be played from the current played time point. This helps to improve the user experience of the user receiving the voice.
[0072] In summary, according to the embodiments of the present disclosure, the voice-to-text function and / or the text-to-voice function can be performed by means of the machine learning model of the cloud in the case of strong network communication capability, and the voice-to-text function and / or the text-to-voice function can be performed by means of the local machine learning model in the case of poor network communication capability. The voice-to-text demand and / or the text-to-voice demand of the user in any network communication can be met, and the user experience of the user can be improved while the accuracy is met.
[0073] FIG. 4 shows a flowchart of a method 400 for information processing according to some embodiments of the present disclosure. The method 400 can be implemented at the client 110.
[0074] At block 410, the client 110 detects the network communication capability between the client and the server in response to the first request for voice-to-text and / or the second request for text-to-voice.
[0075] At block 420, if the network communication capability exceeds the first capability level, the client 110 sends the first voice corresponding to the first request and / or the second text corresponding to the second request to the server to request the server to perform the voice-to-text function and / or the text-to-voice function.
[0076] At block 430, if the network communication capability is lower than the second capability level, the client 110 performs the voice-to-text function on the first voice by invoking the local first machine learning model, and / or performs the text-to-voice function on the second text by invoking the local second machine learning model.
[0077] In some embodiments, the first voice is a question voice received in a question and answer dialogue, and the second text is an answer text to the question voice.
[0078] In some embodiments, the method 400 further includes presenting the first text converted from the first voice in a user interface corresponding to the question and answer dialogue, and / or playing the second voice converted from the second text in the user interface corresponding to the question and answer dialogue.
[0079] In some embodiments, the performing the speech-to-text function on the first speech by invoking the first machine learning model locally and / or the performing the text-to-speech function on the second text by invoking the second machine learning model locally comprises: if the network communication capability is below the second capability level, sending a network request to the server, the network request comprising the first speech to which the first request applies and / or the second text to which the second request applies for the text-to-speech function; and if no network response to the network request is received from the server after a predetermined duration of time has elapsed, performing the speech-to-text function on the first speech by invoking the first machine learning model locally and / or performing the text-to-speech function on the second text by invoking the second machine learning model locally.
[0080] In some embodiments, the method 400 further comprises: in response to the first machine learning model generating the first text segment in the first text corresponding to the first speech, presenting the first text segment at the client; and in response to detecting that the network communication capability between the client and the server exceeds the third capability level, sending the first speech to the server to request the server to perform the speech-to-text function on the first speech.
[0081] In some embodiments, the method 400 further comprises: in response to the second machine learning model generating the first speech segment in the second speech corresponding to the second text, playing the first speech segment at the client; and in response to detecting that the network communication capability between the client and the server exceeds the third capability level, sending the second text to the server to request the server to perform the text-to-speech function on the second text.
[0082] In some embodiments, the method 400 further comprises: if the network communication capability exceeds the first capability level, on the basis of sending the second text to the server, performing the text-to-speech function on the second text by invoking the second machine learning model to obtain one or more speech segments corresponding to one or more text segments in the second text; and while the one or more speech segments are being played at the client, if one or more speech segments converted from the second text are received from the server, after the end of the last played first speech segment, continuing to play the second speech segments received from the server, the text segments corresponding to the second speech segments being after the text segments corresponding to the first speech segments.
[0083] In some embodiments, the method 400 further comprises: while the one or more speech segments are being played, presenting the text segments corresponding to the one or more speech segments respectively at the client.
[0084] In some embodiments, the method 400 further includes: if the network communication capability exceeds the first capability level, on the basis of sending the first voice to the server, performing a text-to-speech function on the first voice by invoking the first machine learning model to obtain one or more text segments corresponding to one or more voice segments in the first voice; and while the client presents the one or more text segments, if one or more text segments converted from the first voice are received from the server, replacing the presented one or more text segments with the one or more text segments received from the server.
[0085] Embodiments of the present disclosure also provide a corresponding apparatus for implementing the above method or process. FIG. 5 shows an exemplary structural block diagram of an apparatus 500 for information processing according to some embodiments of the present disclosure. The apparatus 500 can be implemented as or included in the client 110. Various modules / components in the apparatus 500 can be implemented by hardware, software, firmware, or any combination thereof.
[0086] As shown in FIG. 5, the apparatus 500 includes a communication capability detection module 510 configured to detect a network communication capability between the client and the server in response to a first request for speech-to-text and / or a second request for text-to-speech. The apparatus 500 further includes a first function execution module 520 configured to send, if the network communication capability exceeds a first capability level, a first voice corresponding to the first request and / or a second text corresponding to the second request to the server to request the server to perform a speech-to-text function and / or a text-to-speech function. The apparatus 500 further includes a second function execution module 530 configured to perform, if the network communication capability is lower than a second capability level, a speech-to-text function on the first voice by invoking a local first machine learning model, and / or a text-to-speech function on the second text by invoking a local second machine learning model.
[0087] In some embodiments, the first voice is a question voice received in a question-and-answer dialogue, and the second text is an answer text to the question voice.
[0088] In some embodiments, the apparatus 500 further includes a presentation module configured to present the first text converted from the first voice in a user interface corresponding to the question-and-answer dialogue, and / or play the second voice converted from the second text in the user interface corresponding to the question-and-answer dialogue.
[0089] In some embodiments, the second function execution module 530 includes: a request sending module configured to send a network request to the server if the network communication capability is lower than the second capability level, the network request including a first request for applying a first speech to speech-to-text and / or a second request for applying a second text to text-to-speech; and a local model calling module configured to, if a network response to the network request is not received from the server after a predetermined duration, execute the speech-to-text function on the first speech by calling a local first machine learning model and / or execute the text-to-speech function on the second text by calling a local second machine learning model.
[0090] In some embodiments, the apparatus 500 further includes: a first text presentation module configured to present, at the client, a first text segment in the first text corresponding to the first speech in response to the first machine learning model generating the first text; and a speech sending module configured to send the first speech to the server to request the server to execute the speech-to-text function on the first speech in response to detecting that the network communication capability between the client and the server exceeds a third capability level.
[0091] In some embodiments, the apparatus 500 further includes: a first speech playing module configured to play, at the client, a first speech segment in the second speech corresponding to the second text in response to the second machine learning model generating the second text; and a text sending module configured to send the second text to the server to request the server to execute the text-to-speech function on the second text in response to detecting that the network communication capability between the client and the server exceeds a third capability level.
[0092] In some embodiments, the apparatus 500 further includes: a third function execution module configured to, if the network communication capability exceeds the first capability level, execute the text-to-speech function on the second text by calling the second machine learning model to obtain one or more speech segments corresponding to one or more text segments in the second text on the basis of sending the second text to the server; and a second speech playing module configured to, while playing the one or more speech segments at the client, if one or more speech segments converted from the second text are received from the server, continue to play the second speech segments received from the server after the end of the last played first speech segment, the text segment corresponding to the second speech segment being after the text segment corresponding to the first speech segment.
[0093] In some embodiments, the apparatus 500 further includes: a second text presentation module configured to present, at the client, the text segment corresponding to each of the one or more speech segments while playing the one or more speech segments.
[0094] In some embodiments, the apparatus 500 further includes a fourth function execution module configured to, if the network communication capability exceeds the first capability level, perform, on the basis of sending the first voice to the server, a text-to-speech function on the first voice by invoking a first machine learning model to obtain one or more text segments corresponding to one or more voice segments in the first voice; and a text replacement module configured to, while the client presents the one or more text segments, if receiving the one or more text segments converted from the first voice from the server, replace the presented one or more text segments with the one or more text segments received from the server.
[0095] The units and / or modules included in the apparatus 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, e.g., machine-executable instructions stored on a storage medium. In addition or as an alternative, part or all of the units and / or modules in the apparatus 500 can be implemented by one or more hardware logic components. As an example and not by way of limitation, example types of hardware logic components that can be used include Field-Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application-Specific Standard Products (ASSPs), System-on-a-Chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0096] It should be understood that one or more steps in the above methods can be performed by an appropriate electronic device or combination of electronic devices. Such an electronic device or combination of electronic devices may, for example, include the client 110 in FIG. 1.
[0097] FIG. 6 shows a block diagram of an electronic device 600 in which one or more embodiments of the disclosure can be implemented. It should be understood that the electronic device 600 shown in FIG. 6 is merely an example and should not be construed to limit the functionality and scope of the embodiments described herein. The electronic device 600 shown in FIG. 6 can be used to implement the client 110 of FIG. 1 or the apparatus 500 of FIG. 5.
[0098] As shown in FIG. 6, the electronic device 600 is in the form of a general electronic device. Components of the electronic device 600 can include, but are not limited to, one or more processors or processing units 610, a memory 620, a storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processing unit 610 can be a real or virtual processor and is capable of performing various processes according to programs stored in the memory 620. In a multi-processor system, multiple processing units perform computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 600.
[0099] Electronic device 600 typically includes a plurality of computer storage media. Such media can be removable and / or non-removable, and can include volatile and / or nonvolatile media. Storage 620 can be a volatile memory, such as a register, cache, or random access memory (RAM); a non-volatile memory, such as a read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), or flash memory; or some combination of volatile and / or non-volatile media. Storage device 630 can be a removable or non-removable media, and can include a machine-readable medium, such as a flash drive, magnetic disk, or any other medium that can be used to store information and / or data and that can be accessed by electronic device 600.
[0100] Electronic device 600 can further include additional removable / non-removable, volatile / nonvolatile storage media. Although not shown in FIG. 6, a disk drive or other computer-readable media drive can be provided for reading from or writing to a removable, non-removable, volatile, or non-volatile media such as a floppy disk, a magnetic tape, or an optical disk, for example. In these instances, each drive can be connected to the bus (not shown) by one or more data media interfaces. Storage 620 can include a computer-program product 625 having one or more program modules configured to carry out the various methods or actions of the various embodiments of the present disclosure.
[0101] Communication unit 640 enables communications with other electronic devices over a communication medium. Additionally, the functionality of the components of electronic device 600 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating with one another over a communication connection. As such, electronic device 600 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes.
[0102] Input device 650 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. Output device 660 can be one or more output devices, such as a display, a speaker, a printer, etc. Electronic device 600 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through communication unit 640, as needed, one or more devices that enable a user to interact with electronic device 600, or any devices (e.g., a network card, a modem, etc.) that enable electronic device 600 to communicate with one or more other electronic devices. Such communication can be enabled by an input / output (I / O) interface (not shown).
[0103] According to an example implementation of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to an example implementation of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above.
[0104] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0105] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium. The instructions stored on the computer readable storage medium can be used to program a computer, a programmable data processing apparatus, and / or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0106] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0107] The computer program product of the present disclosure can have a signal including said computer program. This signal can be electronic, electromagnetic, optical, or any other suitable type of signal. Such a signal can be provided through a communication connection, such as electrical wiring, optical fiber, wireless interface, etc. Examples of computer program products include computer program implemented on a personal computer, server, or other networked device. A non-transitory computer readable medium, such as a floppy disk, CD-ROM, DVD-ROM, Blu-ray Disc, hard disk, or memory stick, can also be used to implement the present disclosure. The computer program product of the present disclosure can also be provided as a service to download and use the computer program over a network, such as the Internet.
[0108] Having described several implementations of the present disclosure, it will be clear to those skilled in the art that many modifications, additions, and substitutions are possible without departing from the scope and spirit of the described implementations. Many modifications and variations of the present disclosure are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the appended claims, the present disclosure can be practiced otherwise than as specifically described. While the present disclosure has been described with reference to the implementation figures, it will be understood by those skilled in the art that various changes can be made and equivalents can be substituted for elements thereof without departing from the scope of the present disclosure. In addition, many modifications can be made to adapt a particular situation or material to the teachings of the disclosure without departing from its scope. Therefore, it is contemplated to cover any and all adaptations and modifications falling within the scope of the appended claims. It should also be noted that, in this document, the term "can" is used to mean various possible constructions, for example, a structure, device, or apparatus can be configured to perform one or more functions.
Claims
1. An information processing method applied to a client, the method comprising: In response to a first request for speech-to-text conversion and / or a second request for text-to-speech conversion, the network communication capability between the client and the server is detected; If the network communication capability exceeds the first capability level, send the first voice corresponding to the first request and / or the second text corresponding to the second request to the server to request the server to perform the voice-to-text function and / or text-to-speech function. as well as If the network communication capability is lower than the second capability level, the first local machine learning model is invoked to perform speech-to-text conversion on the first speech, and / or the second local machine learning model is invoked to perform text-to-speech conversion on the second text.
2. The method of claim 1, wherein the first speech is a question speech received in a question-and-answer dialogue, and the second text is a response text to the question speech.
3. The method according to claim 2, further comprising: The first text converted from the first speech is presented in the user interface corresponding to the question-and-answer dialogue, and / or The second speech, converted from the second text, is played in the user interface corresponding to the question-and-answer dialogue.
4. The method according to claim 1, wherein performing speech-to-text conversion on the first speech by invoking a local first machine learning model, and / or performing text-to-speech conversion on the second text by invoking a local second machine learning model, comprises: If the network communication capability is lower than the second capability level, a network request is sent to the server. The network request includes the first request for the first speech applied to speech-to-text and / or the second request for the second text applied to text-to-speech. as well as If no network response to the network request is received from the server after a predetermined time period, the system performs speech-to-text conversion on the first speech by invoking a local first machine learning model, and / or performs text-to-speech conversion on the second text by invoking a local second machine learning model.
5. The method according to claim 1, further comprising: In response to the first text segment generated by the first machine learning model from the first text corresponding to the first speech, the first text segment is presented at the client. as well as In response to detecting that the network communication capability between the client and the server exceeds the third capability level, the client sends the first voice message to the server to request the server to perform a speech-to-text function on the first voice message.
6. The method according to claim 1, further comprising: In response to the second machine learning model generating a first speech segment in the second speech corresponding to the second text, the first speech segment is played at the client. as well as In response to detecting that the network communication capability between the client and the server exceeds the third capability level, the client sends the second text to the server to request the server to perform text-to-speech function on the second text.
7. The method according to claim 1, further comprising: If the network communication capability exceeds the first capability level, based on sending the second text to the server, the second machine learning model is invoked to perform text-to-speech function on the second text to obtain one or more speech segments corresponding to one or more text segments in the second text; as well as While the client is playing the one or more audio segments, if it receives one or more audio segments converted from the second text from the server, after the last audio segment is played, it continues to play the second audio segment received from the server, with the text segment corresponding to the second audio segment following the text segment corresponding to the first audio segment.
8. The method according to claim 7, further comprising: While playing one or more audio clips, the corresponding text clips for each of the one or more audio clips are presented at the client.
9. The method according to claim 1, further comprising: If the network communication capability exceeds the first capability level, based on sending the first voice to the server, the first machine learning model is invoked to perform text-to-speech on the first voice to obtain one or more text segments corresponding to one or more voice segments in the first voice; and While the client is presenting the one or more text fragments, if the server receives one or more text fragments converted from the first speech, the client replaces the one or more text fragments already presented with the one or more text fragments received from the server.
10. An apparatus for information processing, applied to a client, the apparatus comprising: The communication capability detection module is configured to detect the network communication capability between the client and the server in response to a first request for speech-to-text conversion and / or a second request for text-to-speech conversion. The first function execution module is configured to send the first voice corresponding to the first request and / or the second text corresponding to the second request to the server if the network communication capability exceeds the first capability level, so as to request the server to perform the voice-to-text function and / or text-to-speech function. as well as The second function execution module is configured to, if the network communication capability is lower than the second capability level, perform speech-to-text conversion on the first speech by calling a local first machine learning model, and / or perform text-to-speech conversion on the second text by calling a local second machine learning model.
11. An electronic device, comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processor.
12. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 9.
13. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Voice recognition method and device
CN105118508A
Voice recognition system and voice recognition method through local-and-cloud combination
CN106847291A
Dialogue method and device with virtual object, client and storage medium
CN112100352A
Speech recognition processing method, system and device and storage medium
CN117095683A
Vehicle-mounted speech recognition method, system and device based on cloud collaboration and medium
CN117935793A