A voice processing method and device, electronic equipment and storage medium

By continuously sending voice segments from the client for online recognition and semantic prediction, combined with offline recognition and matching response information, the problem of slow voice processing speed is solved, resulting in faster response and a better user experience.

CN116013312BActive Publication Date: 2026-03-17APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-16
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

During human-computer voice interaction, the client needs to wait for the cloud to complete the online recognition results, which results in slow voice processing speed and susceptibility to network signal interference. In particular, offline recognition results cannot be utilized in scenarios that rely on cloud resources.

Method used

The client continuously sends the current speech segment of the speech data to be recognized to the server for online recognition and semantic prediction. After receiving the complete data, it performs offline recognition, matches the predicted semantics to determine the response information, without waiting for the online response of the complete data.

Benefits of technology

It improves voice processing speed, ensures faster response to user needs, enhances user experience, and reduces reliance on network signals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116013312B_ABST
    Figure CN116013312B_ABST
Patent Text Reader

Abstract

This disclosure provides a speech processing method, apparatus, electronic device, and storage medium, relating to the field of computer technology, specifically to speech technology, artificial intelligence, autonomous driving, cloud computing, and other technical fields. The specific implementation scheme is as follows: the client sends the current speech segment of the speech data to be recognized to the server; the server receives and recognizes the current speech segment, performs semantic prediction on the recognition result, and sends the obtained predicted semantics and the current response information generated based on the predicted semantics to the client; the client receives and caches the predicted semantics and the current response information; upon receiving complete speech data to be recognized, the target offline speech recognition result obtained by recognizing the complete speech data to be recognized is matched with all received predicted semantics; if a match is found, the current response information corresponding to the predicted semantics matching the target offline speech recognition result is determined as the target response information for the complete speech data to be recognized, thereby improving the speed of speech processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, specifically to the fields of voice technology, artificial intelligence, autonomous driving, cloud computing, etc., and particularly to a voice processing method, device, electronic device, and storage medium. Background Technology

[0002] With the development of artificial intelligence technology, the application of human-computer voice interaction is becoming more and more widespread. In the process of human-computer voice interaction, the terminal receives voice data and sends the received voice data to the cloud. The cloud recognizes the voice data and feeds back the recognition results to the terminal, which then responds to the feedback information of the recognition results. Summary of the Invention

[0003] This disclosure provides a method, apparatus, device, and storage medium for speech processing.

[0004] According to one aspect of this disclosure, a voice processing method is provided, applied to a client, comprising:

[0005] Continuously acquire speech data to be recognized, and send the current speech segment of the acquired speech data to the server;

[0006] Receive and cache the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment;

[0007] Upon receiving complete speech data to be recognized, speech recognition is performed on the complete speech data to be recognized to obtain the target offline speech recognition result, and the complete speech data to be recognized is sent to the server.

[0008] The target offline speech recognition result is matched with all received predicted semantics;

[0009] If a predicted semantic that matches the target offline speech recognition result exists, the current response information corresponding to the predicted semantic that matches the target offline speech recognition result is determined as the target response information of the complete speech data to be recognized.

[0010] According to another aspect of this disclosure, a voice processing method is provided, applied on a server side, comprising:

[0011] The receiving client sends the current voice segment while continuously acquiring the voice data to be recognized;

[0012] Perform speech recognition on the current speech segment to obtain the current online speech recognition result;

[0013] Semantic prediction is performed on the current online speech recognition results to obtain the predicted semantics;

[0014] Based on the predicted semantics, generate the current response information corresponding to the current speech segment;

[0015] The predicted semantics and the current response information are sent to the client.

[0016] According to another aspect of this disclosure, a voice processing apparatus is provided for use on a client, comprising:

[0017] The first sending module is used to continuously acquire the speech data to be recognized and send the current speech segment of the acquired speech data to the server.

[0018] The first receiving module is used to receive and cache the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment;

[0019] The first recognition module is used to perform speech recognition on the complete speech data to be recognized when it receives the complete speech data to be recognized, obtain the target offline speech recognition result, and send the complete speech data to be recognized to the server.

[0020] The matching module is used to match the target offline speech recognition result with all received predicted semantics;

[0021] The first determining module is used to determine the current response information corresponding to the predicted semantic that matches the target offline speech recognition result as the target response information of the complete speech data to be recognized when the matching module finds that there is a predicted semantic that matches the target offline speech recognition result.

[0022] According to another aspect of this disclosure, a voice processing apparatus is provided, applied on a server side, comprising:

[0023] The second receiving module is used to receive the current voice segment sent by the client while continuously acquiring voice data to be recognized;

[0024] The second recognition module is used to perform speech recognition on the current speech segment to obtain the current online speech recognition result;

[0025] The semantic prediction module is used to perform semantic prediction on the current online speech recognition result to obtain the predicted semantics;

[0026] The first generation module is used to generate current response information corresponding to the current speech segment based on the predicted semantics;

[0027] The second sending module is used to send the predicted semantics and the current response information to the client.

[0028] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0029] At least one processor; and

[0030] A memory communicatively connected to the at least one processor; wherein,

[0031] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods described in this disclosure.

[0032] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform any of the methods described in any one of this disclosure.

[0033] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described in any one of this disclosure.

[0034] The embodiments disclosed herein can improve the speed of voice processing.

[0035] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0036] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0037] Figure 1 This is a schematic diagram of the speech processing process in related technologies;

[0038] Figure 2 This is a first schematic diagram of the speech processing method according to this disclosure;

[0039] Figure 3 This is a second schematic diagram based on the speech processing method disclosed herein;

[0040] Figure 4 This is a third schematic diagram based on the speech processing method disclosed herein;

[0041] Figure 5 This is an interactive schematic diagram based on the speech processing method disclosed herein;

[0042] Figure 6 This is a schematic diagram illustrating the speech processing process according to this disclosure;

[0043] Figure 7 This is a schematic diagram of a speech processing device according to the present disclosure;

[0044] Figure 8 This is another schematic diagram of a speech processing device according to the present disclosure;

[0045] Figure 9 This is a block diagram of an electronic device used to implement the speech processing method of the embodiments of this disclosure. Detailed Implementation

[0046] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0047] During human-computer voice interaction, the client receives the user's voice data and sends it to the server. The server performs online recognition on the voice data, obtains the online recognition result, and sends the online recognition result back to the client. The client then responds to the online recognition result.

[0048] The speech processing process in related technologies, such as Figure 1 As shown, after the user finishes speaking, the client receives the complete voice data, performs offline recognition on the voice data to obtain the offline recognition result, and simultaneously sends the complete voice data to the server. The server performs online recognition on the voice data to obtain the online recognition result. Compared to the server's online recognition, the client first receives the offline recognition result. After receiving the offline recognition result, the client first checks its local semantic library to see if there is a semantic match for the offline recognition result. If there is, it directly uses the semantic match from the local semantic library as the recognition result for the voice data, and uses the response information corresponding to the semantic match from the local semantic library as the response information for the voice data, executing the corresponding instruction (i.e.,...). Figure 1 The client executes offline recognition instructions; if not, it waits for the server to return the online recognition result. The client's local semantic library can contain at least one set of semantics and the corresponding response information. After completing online recognition of the speech data, the server sends the online recognition result to the client. The client receives the online recognition result returned by the server, uses it as the recognition result for the speech data, and executes the instructions corresponding to the online recognition result (i.e.,...). Figure 1 (Execute online recognition action instructions).

[0049] In practical applications, some human-computer voice interactions rely on cloud resources, such as querying real-time information and obtaining real-time resources. In this case, the client cannot use offline recognition results during the above voice processing process and can only wait for online recognition results. The online recognition results need to be completed by the server before they can be sent to the client. The server's online recognition takes a certain amount of time, and the process of sending the online recognition results is also easily affected by network signals, which prolongs the voice processing time and thus affects the speed of voice processing.

[0050] To improve the speed of speech processing, this disclosure provides a speech processing method. A client continuously acquires speech data to be recognized and sends the current speech segment of the acquired speech data to a server. The server receives the current speech segment sent by the client while continuously acquiring speech data to be recognized, performs speech recognition on the current speech segment to obtain the current online speech recognition result, performs semantic prediction on the current online speech recognition result to obtain predicted semantics, generates current response information corresponding to the current speech segment based on the predicted semantics, and sends the predicted semantics and current response information to the client. The client receives and caches the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment. When receiving complete speech data to be recognized, the client performs speech recognition on the complete speech data to obtain a target offline speech recognition result and sends the complete speech data to the server. The client also matches the target offline speech recognition result with all received predicted semantics. If a predicted semantic matches the target offline speech recognition result, the current response information corresponding to the predicted semantic matching the target offline speech recognition result is determined as the target response information for the complete speech data to be recognized.

[0051] In this embodiment, while continuously receiving speech data to be recognized, the client sends the current speech segment from the currently received speech data to the server. The server performs speech recognition and semantic prediction on the current speech segment online and generates current response information corresponding to the current speech segment. When the client completes the reception of the complete speech data to be recognized and obtains the offline recognition result of the complete speech data to be recognized (i.e., the target offline speech recognition result), the target offline speech recognition result is matched with all received predicted semantics. If a match exists, the current response information corresponding to the predicted semantics that match the target offline speech recognition result is determined as the target response information of the complete speech data to be recognized. There is no need to wait for the server to return the online response information of the complete speech data to be recognized, thus improving the speed of speech processing.

[0052] The voice processing method provided in this disclosure can be applied to any voice interaction scenario, such as voice recognition, voice assistants, and other voice interaction products. In this disclosure, the client and server are connected via network communication. The client can be a mobile terminal, computer, smart home device, smart vehicle, speaker, or other device that supports voice recognition. The server can be a computer, cloud server, server cluster, etc.

[0053] The speech processing method provided in the embodiments of this disclosure will be described in detail below.

[0054] See Figure 2 The present disclosure provides a voice processing method applied to a client, comprising the following steps:

[0055] S210 continuously acquires the speech data to be recognized and sends the current speech segment of the acquired speech data to the server.

[0056] In one example, upon detecting a user's voice interaction command, the client continuously receives voice data to be recognized and sends the current voice segment of the received voice data to the server at a preset period or in real time. The current voice segment of the voice data to be recognized is the voice data currently received in real time. The voice data to be recognized can be electrical signals, sound waves, etc., and the preset period can be set according to actual needs, such as 1 second, 2 seconds, or 5 seconds.

[0057] In this embodiment of the disclosure, when the client detects that the user has triggered a voice recognition command, it continuously receives voice data to be recognized and sends the current voice segment of the received voice data to the server. The current voice segment of the voice data to be recognized includes all voice data received before the current moment.

[0058] In one example, the voice data to be recognized can also be continuously sent to the client device by other devices, and is not limited to the user's real-time voice data.

[0059] In one possible implementation, the client can perform real-time speech recognition on the current speech segment in the acquired speech data to obtain the current offline speech recognition result corresponding to the current speech segment. This current offline speech recognition result can be, for example, the text content corresponding to the current speech segment. For example, the client can utilize an offline recognition engine to perform real-time speech recognition on the current speech segment; alternatively, it can utilize any audio-to-text tool to perform real-time speech recognition on the current speech segment. The offline recognition engine is a pre-trained speech recognition engine capable of performing speech recognition.

[0060] S220 receives and caches the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment.

[0061] The client continuously acquires speech data to be recognized and sends the current speech segment to the server. The server performs online speech recognition on the current speech segment and semantic prediction on the result of the online speech recognition. It then returns the predicted semantics and the corresponding current response information to the client. The client receives and caches the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on all current speech segments.

[0062] In one example, speech recognition result is text information obtained by performing speech recognition on speech data, such as "today" or "today's weather"; semantic prediction is text information obtained by performing semantic prediction on speech recognition result, such as "today's weather" or "how is the weather today";

[0063] For example, the current online speech recognition result of the current speech segment is "today", the predicted semantics is "today's weather", and the corresponding current response information is today's weather data or information.

[0064] S230, upon receiving complete speech data to be recognized, performs speech recognition on the complete speech data to be recognized, obtains the target offline speech recognition result, and sends the complete speech data to be recognized to the server.

[0065] In one example, if the client does not detect any voice signal input within a preset time period, it determines that the voice input is complete; that is, all the voice data received up to the current moment is considered complete voice data. The preset time period can be set according to actual needs, such as 0.1 seconds, 1 second, or 2 seconds. The shorter the preset time period, the faster the client's response speed in determining that the voice data to be recognized has been completely received; the longer the preset time period, the more completely the client can receive the voice data to be recognized, resulting in more accurate recognition results.

[0066] Once the client confirms that it has received complete speech data to be recognized, it can use an offline recognition engine or an audio-to-text tool to perform speech recognition on the complete speech data to obtain the target offline speech recognition result. The current speech segment of the speech data to be recognized can be a part of the complete speech data to be recognized. For example, the speech recognition result obtained after recognizing the current speech segment is "Today's weather," while the speech recognition result obtained after recognizing the complete speech data is "How's the weather today?"

[0067] In one example, when the client receives complete speech data to be recognized, it can simultaneously perform speech recognition on the complete speech data and send it to the server so that the server can perform online speech analysis on the complete speech data in a timely manner.

[0068] S240, Match the target offline speech recognition result with all received predicted semantics.

[0069] In one example, before receiving the complete speech data to be recognized, the client may receive one or more sets of predicted semantics and current response information from the server. The client can perform a one-to-one or one-to-many match between the target offline speech recognition result and all the predicted semantics received from the server. Both the target offline speech recognition result and the predicted semantics are text information. The client can also calculate the text similarity between the target offline speech recognition result and each predicted semantic, and determine the predicted semantics with a text similarity greater than a preset threshold as the predicted semantics that match the target offline speech recognition result. The preset threshold can be set according to requirements, such as 0.7, 0.8, or 0.9, etc.

[0070] For example, the client performs speech recognition on the complete speech data to be recognized, and the target offline speech recognition result is: "How's the weather today?". The client receives predicted semantics from the server, including multiple predicted semantics such as: "How's the weather today?", "The weather is suitable for a picnic today?", and "It's suitable for a car wash today?". The text similarity between the target offline speech recognition result and each of the multiple predicted semantics is calculated. If the text similarity between the target offline speech recognition result and the predicted semantic "How's the weather today?" is greater than a preset threshold, then the target offline speech recognition result is determined to match the predicted semantic "How's the weather today?".

[0071] S250, if there is a predicted semantic that matches the target offline speech recognition result, the current response information corresponding to the predicted semantic that matches the target offline speech recognition result is determined as the target response information of the complete speech data to be recognized.

[0072] For example, if the target offline speech recognition result "How is the weather today?" successfully matches the predicted semantic "How is the weather today?", then the current response information corresponding to the predicted semantic "How is the weather today?" is determined as the target response information of the complete speech data to be recognized.

[0073] In this embodiment, while continuously receiving voice data to be recognized, the client sends the current voice segment from the currently received voice data to the server. The server performs online speech recognition and semantic prediction on the current voice segment and generates current response information corresponding to the current voice segment. Once the client has received the complete voice data to be recognized and obtained the offline recognition result (i.e., the target offline speech recognition result), the target offline speech recognition result is matched with all received predicted semantics. If a match exists, the current response information corresponding to the predicted semantics matching the target offline speech recognition result is determined as the target response information for the complete voice data to be recognized. This eliminates the need to wait for the server to return the online response information for the complete voice data to be recognized, thus improving the speed of voice processing and enabling faster responses to user needs. Furthermore, since voice response speed is one of the core indicators for evaluating the usability of a voice product, a faster response speed can enhance the user experience.

[0074] In one possible implementation, the above method may further include:

[0075] In the absence of a predicted semantic that matches the target offline speech recognition result, semantic parsing is performed on the target offline speech recognition result to obtain the target semantics;

[0076] Based on the target semantics and preset strategies, the target response information of the complete speech data to be recognized is determined.

[0077] In one example, when there is no match between the target offline speech recognition result and the predicted semantics, the client can perform offline semantic parsing on the target offline speech recognition result to obtain the complete semantic representation.

[0078] Identify the interaction purpose or intent corresponding to the voice data, i.e., the target semantics. For example, a semantic parser can be used to perform offline semantic parsing on the target offline speech recognition results to obtain the target semantics.

[0079] Semantics is used to represent the interaction purpose or intention corresponding to the complete speech data to be recognized.

[0080] For example, after obtaining the target semantics corresponding to the complete speech data to be recognized, the response information corresponding to the target semantics is obtained according to the target semantics and a preset strategy, and is identified as the complete speech data to be recognized.

[0081] The target response information for the identified voice data is determined. The preset strategy can be to match the target semantics with semantics in a whitelist stored locally on the client's device, or to match the target semantics with a preset semantic-response information association table. If a match is found, the response information corresponding to the matched semantics is determined as the target response information for the complete voice data to be identified. The whitelist can contain at least one set of semantics and its corresponding response information, and the semantic-response information association table...

[0082] It contains at least one set of semantic and response information associations. Matching can be determined by calculating the text similarity between the target semantic and the semantics in the whitelist or association table. For example, the target response information can be specific response content, such as operation commands like turning on the air conditioner, increasing the brightness of the lights, or playing music.

[0083] In this embodiment of the disclosure, when there is no predicted semantic matching the target offline speech recognition result...

[0084] In this case, the client prioritizes determining the target response information of the complete voice data to be recognized locally, without waiting for the server to return the online response information of the complete voice data to be recognized, which improves the speed of voice processing and enables a faster response to user needs.

[0085] In one possible implementation, the process of determining the target response information of complete speech data to be recognized based on target semantics and a preset strategy may include the following steps:

[0086] Step 1: Determine if the target semantics matches the preset whitelist.

[0087] In one example, the preset whitelist contains at least one set of preset semantics and the preset response information corresponding to the preset semantics. After obtaining the target semantics, the text similarity between the target semantics and each preset semantic can be calculated. If there is a text similarity greater than a preset threshold, it is determined that the target semantics matches the preset whitelist. If there is no text similarity greater than the preset threshold, it is determined that the target semantics does not match the preset whitelist.

[0088] Step 2: If the target semantics matches the preset whitelist, the preset response information corresponding to the matched preset semantics is determined as the target response information of the complete speech data to be recognized.

[0089] For example, if the target semantic is "play favorite songs", and the target semantic matches the preset semantic in the preset whitelist which is "play songs in favorite list", then the preset response information "instruction to play songs in favorite list" corresponding to the preset semantic is determined as the target response information of the complete voice data to be recognized.

[0090] In this embodiment of the disclosure, when the target semantics matches the preset whitelist, the preset response information corresponding to the preset semantics matched locally on the client is determined as the target response information of the complete voice data to be recognized. This eliminates the need to wait for the server to return the online response information of the complete voice data to be recognized, thereby improving the speed of voice processing.

[0091] In one possible implementation, the process of determining the target response information of the complete speech data to be recognized based on the target semantics and the preset strategy may further include: receiving and caching the online response information sent by the server after performing speech recognition and semantic parsing on the complete speech data to be recognized when the target semantics does not match the preset whitelist; and determining the online response information as the target response information of the complete speech data to be recognized.

[0092] In one example, the client can receive and cache the online response information sent by the server after performing speech recognition and semantic parsing on the complete speech data to be recognized, under any conditions. If the current response information corresponding to the predicted semantics matching the target offline speech recognition result is determined as the target response information for the complete speech data to be recognized, or if the preset response information corresponding to the preset semantics matched by the target semantics is determined as the target response information for the complete speech data to be recognized, then no processing of the online response information is required. If the target semantics does not match the preset whitelist, the online response information is determined as the target response information for the complete speech data to be recognized. The client can also receive and cache the online response information sent by the server after performing speech recognition and semantic parsing on the complete speech data to be recognized only when the target semantics does not match the preset whitelist, and in other cases, send a message to the server that does not require returning online response information.

[0093] For example, if the target semantic is "play songs from the charts", and the target semantic does not match any of the preset semantics in the preset whitelist, then wait for the server to return the online response information of the complete voice data to be recognized. After the server completes the speech recognition and semantic parsing of the complete voice data to be recognized, the client receives and caches the online response information sent by the server after performing speech recognition and semantic parsing on the complete voice data to be recognized, and determines the online response information as the target response information of the complete voice data to be recognized.

[0094] In this embodiment of the disclosure, if the target semantics obtained by the client's local parsing does not match the preset whitelist, the client waits for and receives the online response information sent by the server after performing speech recognition and semantic parsing on the complete speech data to be recognized. The online response information is then identified as the target response information for the complete speech data to be recognized, making the response result of the complete speech data to be recognized more accurate.

[0095] In one possible implementation, where the online response information does not involve timeliness issues, the client can cache the online response information received from the server and the target offline speech recognition result corresponding to the online response information in a local whitelist. This allows the client to directly determine the target response information corresponding to the speech data to be recognized from the local whitelist when the target offline speech recognition result is recognized in subsequent speech recognition processes.

[0096] In one possible implementation, the target response information of the complete voice data to be recognized may include: information to be output or instructions. The voice processing method may also include: displaying or playing the information to be output, or executing instructions.

[0097] In this embodiment of the disclosure, after obtaining the target response information corresponding to the complete speech data to be recognized, the information contained in the target response information is output, or the instructions contained in the target response information are executed, in order to respond to the complete speech data to be recognized.

[0098] In one possible implementation, the above-mentioned receiving and caching of the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment may include: after performing speech recognition on the current speech segment and obtaining the current offline speech recognition result, receiving and caching the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment.

[0099] In one example, while continuously acquiring speech data to be recognized, the client can simultaneously send the current speech segment of the acquired speech data to the server and use an offline recognition engine or audio-to-text tool to perform speech recognition on the current speech segment, obtaining the current offline speech recognition result for the current speech segment. Since local offline recognition is faster, receiving and caching the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment is an operation performed after obtaining the current offline speech recognition result locally.

[0100] In this embodiment of the disclosure, because the client's local offline recognition speed is faster, the client receives and caches the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment. This operation is performed after obtaining the current offline speech recognition result locally. That is, for the same current speech segment, the client's local offline recognition speed is faster than the server's online recognition speed.

[0101] For example, such as Figure 3 As shown in the embodiments of this disclosure, a voice processing method applied to a client may include the following steps:

[0102] S301 continuously acquires the speech data to be recognized and sends the current speech segment of the acquired speech data to the server.

[0103] S302, Receive and cache the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment;

[0104] S303: Upon receiving complete speech data to be recognized, performs speech recognition on the complete speech data to be recognized, obtains the target offline speech recognition result, and sends the complete speech data to be recognized to the server.

[0105] S304, Match the target offline speech recognition result with all received predicted semantics;

[0106] S305, if there is a predicted semantic that matches the target offline speech recognition result, the current response information corresponding to the predicted semantic that matches the target offline speech recognition result is determined as the target response information of the complete speech data to be recognized.

[0107] S306, In the absence of a predicted semantic that matches the target offline speech recognition result, perform semantic parsing on the target offline speech recognition result to obtain the target semantic;

[0108] S307, determine whether the target semantic matches the preset whitelist; wherein, the preset whitelist contains at least one set of preset semantics and the preset response information corresponding to the preset semantics;

[0109] S308, when the target semantics matches the preset whitelist, the preset response information corresponding to the matched preset semantics is determined as the target response information of the complete speech data to be recognized.

[0110] S309, when the target semantics does not match the preset whitelist, receives and caches the online response information sent by the server after performing speech recognition and semantic parsing on the complete speech data to be recognized;

[0111] S310 identifies the online response information as the target response information of the complete voice data to be recognized;

[0112] S311, Display or play the information to be output, or execute an instruction, wherein the target response information of the complete speech data to be recognized includes: the information to be output or the instruction.

[0113] See Figure 4 Another voice processing method provided in this disclosure, applied to a server, includes the following steps:

[0114] S410 receives the current voice segment sent by the client while continuously acquiring voice data to be recognized.

[0115] In one example, the client, while continuously acquiring the speech data to be recognized, will receive...

[0116] The current speech segment of the speech data to be recognized is sent to the server according to a preset period or in real time. Correspondingly, the server receives the current speech segment of the speech data to be recognized sent by the client within the preset period or in real time.

[0117] S420 performs speech recognition on the current speech segment and obtains the current online speech recognition result.

[0118] In one example, each time the server receives a current speech segment, it performs speech recognition on that segment to obtain the current online speech recognition result. For instance, the server could utilize an online speech recognition engine or any audio-to-text converter to perform speech recognition on the current speech segment.

[0119] The tool performs speech recognition on the current speech segment. The online recognition engine is a pre-trained speech recognition engine capable of performing speech recognition.

[0120] S430 performs semantic prediction on the current online speech recognition results to obtain the predicted semantics.

[0121] In one example, the server can use a semantic prediction model to perform semantic prediction on the current online speech recognition results. This semantic prediction model can be trained based on sample text and its semantics.

[0122] S440 generates current response information corresponding to the current speech segment based on predicted semantics.

[0123] For example, if the predicted semantics are "today's weather", then the data corresponding to today's weather will be packaged to generate the current response information corresponding to the current speech segment. Or if the predicted semantics are "play 0 popular songs", then the data of popular songs and the playback command will be packaged to generate the current response information corresponding to the current speech segment.

[0124] S450 sends the predicted semantics and current response information to the client.

[0125] In this embodiment of the disclosure, the server receives the client's continuously acquired speech data to be recognized.

[0126] The system receives the current speech segment, performs speech recognition on the current speech segment, and performs semantic prediction on the current online speech recognition result. Then, based on the semantic prediction, it generates a pre-defined semantic prediction result.

[0127] Semantic prediction is performed, and current response information corresponding to the current speech segment is generated. The predicted semantics and current response information are then sent to the client so that the client can receive the complete speech data to be recognized and obtain the offline recognition result of the complete speech data (i.e., the target offline speech recognition result).

[0128] In this case, the target offline speech recognition result is matched with all received predicted semantics first. If a match exists, the current response information corresponding to the predicted semantics that matches the target offline speech recognition result is determined as the target response information of the complete speech data to be recognized. There is no need to wait for the server to return the online response information of the complete speech data to be recognized, which improves the speed of speech processing and enables a faster response to user needs.

[0129] In one possible implementation, the above method may further include:

[0130] The receiving client sends complete unrecognized voice data upon receiving complete unrecognized voice data;

[0131] Perform speech recognition on the complete speech data to be recognized to obtain the target online speech recognition result;

[0132] Semantic analysis is performed on the target online speech recognition results to obtain online semantics;

[0133] Based on online semantics, generate online response information corresponding to the complete speech data to be recognized;

[0134] Send online response information to the client.

[0135] In one example, the server receives the complete speech data to be recognized from the client, performs speech recognition on the complete speech data, obtains the target online speech recognition result, and then directly performs semantic parsing on the target online speech recognition result to obtain the online semantics. Furthermore, it generates online response information corresponding to the complete speech data to be recognized based on the online semantics. For example, the server can use a semantic parser to perform semantic parsing on the target online speech recognition result to obtain the online semantics.

[0136] In this embodiment of the disclosure, the server receives complete speech data to be recognized sent by the client, performs speech recognition on the complete speech data to be recognized, performs online semantic parsing, no longer predicts semantics, and sends the online response information corresponding to the complete speech data to be recognized generated according to the online semantics to the client, so that the client can respond accurately to the complete speech data to be recognized.

[0137] In one possible implementation, the above method may further include:

[0138] Determine whether the current online speech recognition result hits a preset database;

[0139] When the current online speech recognition result hits the preset database, trigger the execution of the steps: perform semantic prediction on the current online speech recognition result to obtain a predicted semantics.

[0140] The preset database can be set according to requirements. For example, the preset database can contain multiple groups of words or phrases, etc., and the words or phrases contained can perform semantic prediction. After the server recognizes the current online speech recognition result, it can further query the preset database to determine whether the current online speech recognition result is included in the preset database. If it is included, it is determined that the current online speech recognition result hits the preset database; otherwise, it is determined that the current online speech recognition result does not hit the preset database. When the current online speech recognition result hits the preset database, perform the step of performing semantic prediction on the current online speech recognition result to obtain a predicted semantics; otherwise, do not perform semantic prediction on the current online speech recognition result.

[0141] Exemplarily, if the current online speech recognition result is "今" and does not hit the preset database, then no semantic prediction is performed on the current online speech recognition result. If the current online speech recognition result is "今天" and hits the preset database, then semantic prediction is performed on the current online speech recognition result.

[0142] In the embodiments of the present disclosure, when the current online speech recognition result hits the preset database, trigger the execution of performing semantic prediction on the current online speech recognition result to obtain a predicted semantics, without performing semantic prediction on all current online speech recognition results, reducing the workload of the server and also reducing the workload of the client to match the predicted semantics, thereby saving the computing resources of the server and the client.

[0143] Exemplarily, as Figure 5 shown, an interaction process of the speech processing method of the present disclosure includes:

[0144] S501, the client continuously obtains speech data to be recognized;

[0145] S502, the client sends the current speech segment of the obtained speech data to be recognized to the server;

[0146] S503, the server receives the current voice segment sent by the client, performs speech recognition on the current voice segment to obtain the current online speech recognition result, determines whether the current online speech recognition result hits the preset database, and if the current online speech recognition result hits the preset database, performs semantic prediction on the current online speech recognition result to obtain the predicted semantics, and generates the current response information corresponding to the current voice segment based on the predicted semantics.

[0147] S504, the server sends the predicted semantics and current response information to the client;

[0148] S505: The client receives and caches the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment. After receiving the complete speech data to be recognized, the client performs speech recognition on the complete speech data to be recognized and obtains the target offline speech recognition result.

[0149] S506: The client sends the complete voice data to be recognized to the server.

[0150] S507: The server receives the complete speech data to be recognized sent by the client, performs speech recognition on the complete speech data to be recognized to obtain the target online speech recognition result, performs semantic analysis on the target online speech recognition result to obtain online semantics, and generates online response information corresponding to the complete speech data to be recognized based on the online semantics.

[0151] S508, the client matches the target offline speech recognition result with all received predicted semantics. If a predicted semantic matches the target offline speech recognition result, the current response information corresponding to the predicted semantic matching the target offline speech recognition result is determined as the target response information of the complete speech data to be recognized. Otherwise, the client performs semantic parsing on the target offline speech recognition result to obtain the target semantics and determines whether the target semantics matches the preset whitelist. If it matches, the preset response information corresponding to the matched preset semantics is determined as the target response information of the complete speech data to be recognized. The preset whitelist contains at least one set of preset semantics and the preset response information corresponding to the preset semantics.

[0152] S509, the server sends the online response information to the client;

[0153] S510: If the target semantics does not match the preset whitelist, the client receives and caches the online response information sent by the server after performing speech recognition and semantic parsing on the complete speech data to be recognized, and determines the online response information as the target response information of the complete speech data to be recognized.

[0154] S511. The client displays or plays the information to be output, or executes an instruction. Among them, the target response information of the complete voice data to be recognized includes: the information to be output or the instruction.

[0155] Exemplarily, as Figure 6 shown, Figure 6 is a schematic diagram showing a voice processing process according to the present disclosure, including:

[0156] The application program (APP) of the client continuously obtains the voice data to be recognized, sends the current voice segment of the obtained voice data to be recognized to the server, and uses an offline recognition engine to perform voice recognition on the current voice segment. At time point t1, "today" is recognized. The server synchronously uses an online recognition engine to perform voice recognition on the current voice segment. At time point t2, "today" is recognized. The client recognizes "today" at time point t3, and the server recognizes "today" at time point t4, and sends the predicted semantics "today's weather" and the corresponding current response information to the client. The client recognizes "today's day" at time point t5, and the server recognizes "today's day" at time point t6, and sends the predicted semantics "how about today's weather" and the corresponding current response information to the client. The client completes the reception of the voice data to be recognized at time point t7 and recognizes "today's weather". This recognition result "today's weather" matches the predicted semantics "today's weather". At this time, directly determine the current response information corresponding to the predicted semantics "today's weather" as the target response information of the complete voice data to be recognized, and thus there is no need to wait for the server to recognize "today's weather" at time point t8 and perform online semantic analysis on "today's weather", and send the final online semantic result (i.e., the online response information of the complete voice data to be recognized) to the client at time point t9. In other words, after the client recognizes "today's weather" at time point t7, it determines the current response information corresponding to the predicted semantics "today's weather" as the target response information of the complete voice data to be recognized, and responds to the voice data to be recognized. The improved voice response speed time is t9 - t7.

[0157] Correspondingly, applying the voice processing method provided in the embodiments of the present disclosure, when the client has a predicted semantics that matches the target offline voice recognition result, the current response information corresponding to the predicted semantics that matches the target offline voice recognition result is determined as the target response information of the complete voice data to be recognized. The improved voice response speed time = t9 - t7 = (server online final recognition recognition recognition information, and the duration for the server to send the final response information to the client.

[0158] The embodiments of the present disclosure also provide a voice processing device, which is applied to the client. Refer to Figure 7 and this device includes:

[0159] The first sending module 701 is used to continuously acquire the speech data to be recognized and send the current speech segment of the acquired speech data to be recognized to the server.

[0160] The first receiving module 702 is used to receive and cache the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment;

[0161] The first recognition module 703 is used to perform speech recognition on the complete speech data to be recognized when it receives the complete speech data to be recognized, obtain the target offline speech recognition result, and send the complete speech data to be recognized to the server.

[0162] The matching module 704 is used to match the target offline speech recognition result with all received predicted semantics;

[0163] The first determining module 705 is used to determine the current response information corresponding to the predicted semantic that matches the target offline speech recognition result as the target response information of the complete speech data to be recognized when the matching module 704 finds that there is a predicted semantic that matches the target offline speech recognition result.

[0164] In this embodiment, while continuously receiving speech data to be recognized, the client sends the current speech segment from the currently received speech data to the server. The server performs speech recognition and semantic prediction on the current speech segment online and generates current response information corresponding to the current speech segment. When the client completes the reception of the complete speech data to be recognized and obtains the offline recognition result of the complete speech data to be recognized (i.e., the target offline speech recognition result), the target offline speech recognition result is matched with all received predicted semantics. If a match exists, the current response information corresponding to the predicted semantics that match the target offline speech recognition result is determined as the target response information of the complete speech data to be recognized. There is no need to wait for the server to return the online response information of the complete speech data to be recognized, thus improving the speed of speech processing.

[0165] In one possible implementation, the above-described apparatus further includes:

[0166] The first parsing module is used to perform semantic parsing on the target offline speech recognition result to obtain the target semantic when the matching module 704 finds that there is no predicted semantic that matches the target offline speech recognition result.

[0167] The second determining module is used to determine the target response information of the complete speech data to be recognized based on the target semantics and the preset strategy.

[0168] In one possible implementation, the second determining module includes:

[0169] The first determining unit is used to determine whether the target semantics matches a preset whitelist; the preset whitelist contains at least one set of preset semantics and preset response information corresponding to the preset semantics;

[0170] The second determining unit is used to determine the preset response information corresponding to the preset semantic that is hit as the target response information of the complete voice data to be recognized when the first determining unit determines that the target semantic hits the preset whitelist.

[0171] In one possible implementation, the second determining module further includes:

[0172] The receiving unit is configured to receive and cache the online response information sent by the server after performing speech recognition and semantic parsing on the complete speech data to be recognized, when the first determining unit determines that the target semantic does not match the preset whitelist.

[0173] The third determining unit is used to determine the online response information as the target response information of the complete voice data to be recognized.

[0174] In one possible implementation, the first receiving module 702 is specifically used to: after performing speech recognition on the current speech segment and obtaining the current offline speech recognition result, receive and cache the predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment.

[0175] In one possible implementation, the target response information of the complete speech data to be recognized includes: information or instructions to be output, and the device further includes:

[0176] The response module is used to display or play the information to be output, or to execute the instructions.

[0177] This disclosure also provides a voice processing device applied to a server, see [link to relevant documentation]. Figure 8 The device includes:

[0178] The second receiving module 801 is used to receive the current voice segment sent by the client while continuously acquiring voice data to be recognized;

[0179] The second recognition module 802 is used to perform speech recognition on the current speech segment to obtain the current online speech recognition result;

[0180] The semantic prediction module 803 is used to perform semantic prediction on the current online speech recognition result to obtain the predicted semantics;

[0181] The first generation module 804 is used to generate current response information corresponding to the current speech segment 5 based on the predicted semantics;

[0182] The second sending module 805 is used to send the predicted semantics and the current response information to the client.

[0183] In this embodiment of the disclosure, the server receives the client's continuously acquired speech data to be recognized.

[0184] In the case of receiving the current speech segment, the system performs speech recognition on the current speech segment and semantic prediction on the current online speech recognition result. Based on the predicted semantics, it generates the current response information corresponding to the current speech segment and sends the predicted semantics and current response information to the client. This allows the client to receive the complete speech data to be recognized and obtain the offline recognition result of the complete speech data (i.e., the target offline speech recognition result).

[0185] In this case, the target offline speech recognition result is matched with all received predicted semantics first. If a match exists, the current response information corresponding to the predicted semantics that match the target offline speech recognition result is determined as the target response information of the complete speech data to be recognized. There is no need to wait for the server to return the online response information of the complete speech data to be recognized, which improves the speed of speech processing and enables a faster response to user needs.

[0186] In one possible implementation, the above-described apparatus further includes: a third receiving module, configured to receive data from the client upon receiving complete voice data to be recognized.

[0187] The complete voice data to be recognized is sent below;

[0188] The third recognition module is used to perform speech recognition on the complete speech data to be recognized, and obtain the target online speech recognition result;

[0189] The second parsing module is used to perform semantic parsing on the target online speech recognition result to obtain 5 online semantics;

[0190] The second generation module is used to generate online response information corresponding to the complete speech data to be recognized based on the online semantics.

[0191] The third sending module is used to send the online response information to the client.

[0192] In one possible implementation, the above-mentioned device further includes: a third determining module, used to determine whether the current online speech recognition result hits a preset database;

[0193] The execution module is used to trigger the semantic prediction module to perform semantic prediction on the current online speech recognition result when the third determining module determines that the current online speech recognition result matches the preset database, thereby obtaining the predicted semantics.

[0194] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0195] This disclosure provides an electronic device, comprising:

[0196] At least one processor; and

[0197] A memory that is communicatively connected to at least one processor; wherein,

[0198] The memory stores instructions that can be executed by at least one processor to enable the at least one processor to perform any of the methods of this disclosure.

[0199] This disclosure provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform any of the methods described in this disclosure.

[0200] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements any of the methods described in this disclosure.

[0201] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information in this technical solution comply with relevant laws and regulations and do not violate public order and good morals. It should be noted that the head model in this embodiment is not a head model specific to any particular user and does not reflect the personal information of any particular user. It should also be noted that the two-dimensional face images in this embodiment are from publicly available datasets.

[0202] Figure 9 A schematic block diagram of an example electronic device 900 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0203] like Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 902 or a computer program loaded from storage unit 908 into random access memory (RAM) 903. RAM 903 may also store various programs and data required for the operation of device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.

[0204] Multiple components in device 900 are connected to I / O interface 905, including: input unit 906, such as keyboard, mouse, etc.; output unit 907, such as various types of monitors, speakers, etc.; storage unit 908, such as disk, optical disk, etc.; and communication unit 909, such as network card, modem, wireless transceiver, etc. Communication unit 909 allows device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0205] The computing unit 901 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above, such as speech processing methods. For example, in some embodiments, the speech processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on device 900 via ROM 902 and / or communication unit 909. When the computer program is loaded into RAM 903 and executed by the computing unit 901, one or more steps of the speech processing method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform speech processing methods by any other suitable means (e.g., by means of firmware).

[0206] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0207] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0208] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0209] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0210] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0211] A computer system can include clients and servers. Clients and servers are generally located far apart from each other.

[0212] They typically interact via communication networks. The client-server relationship is established by computer programs running on corresponding computers and having a client-server relationship with each other. The server can be a cloud server, a server in a distributed system, or a server incorporating blockchain technology.

[0213] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0214] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A voice processing method applied to a client, comprising: continuously obtaining to-be-recognized voice data, and sending a current voice segment of the obtained to-be-recognized voice data to a server; receiving and buffering predicted semantics and current response information sent by the server after performing voice recognition and semantic prediction on the current voice segment; the current response information is an operation instruction corresponding to the predicted semantics; in a case where complete to-be-recognized voice data is received, performing voice recognition on the complete to-be-recognized voice data to obtain a target offline voice recognition result, and sending the complete to-be-recognized voice data to the server; matching the target offline voice recognition result with all received predicted semantics; in a case where there is predicted semantics matching the target offline voice recognition result, determining the current response information corresponding to the predicted semantics matching the target offline voice recognition result as target response information of the complete to-be-recognized voice data; in a case where there is no predicted semantics matching the target offline voice recognition result, performing offline semantic analysis on the target offline voice recognition result to obtain a target semantic; the target semantic is used to represent an interaction purpose or intention corresponding to the complete to-be-recognized voice data; determining target response information of the complete to-be-recognized voice data based on the target semantic and a preset strategy.

2. The method of claim 1, wherein, The determining of the target response information of the complete to-be-recognized voice data based on the target semantic and the preset strategy comprises: determining whether the target semantic hits a preset white list; the preset white list contains at least one set of preset semantics and preset response information corresponding to the preset semantics; in a case where the target semantic hits the preset white list, determining the preset response information corresponding to the hit preset semantic as the target response information of the complete to-be-recognized voice data. 3.The method of claim 2, further comprising: in a case where the target semantic does not hit the preset white list, receiving and buffering online response information sent by the server after performing voice recognition and semantic analysis on the complete to-be-recognized voice data; determining the online response information as the target response information of the complete to-be-recognized voice data.

4. The method of claim 1, wherein, The receiving and buffering of the predicted semantics and the current response information sent by the server after performing voice recognition and semantic prediction on the current voice segment comprises: after obtaining a current offline voice recognition result by performing voice recognition on the current voice segment, receiving and buffering the predicted semantics and the current response information sent by the server after performing voice recognition and semantic prediction on the current voice segment.

5. The method of any of claims 1-4, wherein the target response information of the complete voice data to be recognized comprises: to-be-output information or an instruction, the method further comprising: displaying or playing the to-be-output information, or executing the instruction. 6.A voice processing method applied to a server, comprising: receiving a current voice segment sent by a client in a case where the client continuously obtains to-be-recognized voice data; performing voice recognition on the current voice segment to obtain a current online voice recognition result; performing semantic prediction on the current online voice recognition result to obtain predicted semantics; generating, based on the predicted semantics, current response information corresponding to the current voice segment; the current response information is an operation instruction corresponding to the predicted semantics; sending the predicted semantics and the current response information to the client, so that the client performs voice recognition on complete to-be-recognized voice data received by the client to obtain a target offline voice recognition result, and matches the target offline voice recognition result with all the predicted semantics; in a case where there is predicted semantics matching the target offline voice recognition result, current response information corresponding to the predicted semantics matching the target offline voice recognition result is determined as target response information of the complete to-be-recognized voice data; in a case where there is no predicted semantics matching the target offline voice recognition result, performing offline semantic analysis on the target offline voice recognition result to obtain target semantics; the target semantics are used to represent an interaction purpose or an intention corresponding to the complete to-be-recognized voice data; determining, based on the target semantics and a preset strategy, target response information of the complete to-be-recognized voice data.

7. The method of claim 6, further comprising: receiving complete to-be-recognized voice data sent by the client in a case where the client receives the complete to-be-recognized voice data; performing voice recognition on the complete to-be-recognized voice data to obtain a target online voice recognition result; performing semantic analysis on the target online voice recognition result to obtain online semantics; generating online response information corresponding to the complete to-be-recognized voice data based on the online semantics; sending the online response information to the client.

8. The method of claim 6, further comprising: determining whether the current online voice recognition result hits a preset database; in a case where the current online voice recognition result hits the preset database, triggering a step of performing semantic prediction on the current online voice recognition result to obtain predicted semantics.

9. A voice processing apparatus applied to a client, comprising: a first sending module configured to continuously acquire to-be-recognized voice data, and send a current voice segment of the acquired to-be-recognized voice data to a server; a first receiving module configured to receive and cache predicted semantics and current response information sent by the server after performing voice recognition and semantic prediction on the current voice segment; the current response information is an operation instruction corresponding to the predicted semantics; a first recognition module configured to, in a case where complete to-be-recognized voice data is received, perform voice recognition on the complete to-be-recognized voice data to obtain a target offline voice recognition result, and send the complete to-be-recognized voice data to the server; a matching module configured to match the target offline voice recognition result with all the predicted semantics received; a first determination module configured to, in a case where the matching module matches predicted semantics matching the target offline voice recognition result, determine current response information corresponding to the predicted semantics matching the target offline voice recognition result as target response information of the complete to-be-recognized voice data. The first analysis module is configured to perform offline semantic analysis on the target offline speech recognition result to obtain a target semantic in a case where the matching module matches a result that there is no predicted semantic matched with the target offline speech recognition result; the target semantic is used to represent an interactive purpose or an intent corresponding to the complete speech data to be recognized. The second determination module is configured to determine target response information of the complete speech data to be recognized based on the target semantic and a preset strategy.

10. The apparatus of claim 9, wherein, The second determination module includes: A first determination unit configured to determine whether the target semantic hits a preset white list; the preset white list includes at least one group of preset semantics and preset response information corresponding to the preset semantics; A second determination unit configured to determine, in a case where the first determination unit determines that the target semantic hits the preset white list, preset response information corresponding to the hit preset semantic as the target response information of the complete speech data to be recognized.

11. The apparatus of claim 10, wherein the second determination module further includes: A receiving unit configured to receive and cache online response information sent by the server after performing speech recognition and semantic analysis on the complete speech data to be recognized in a case where the first determination unit determines that the target semantic does not hit the preset white list; A third determination unit configured to determine the online response information as the target response information of the complete speech data to be recognized.

12. The apparatus of claim 9, wherein, The first receiving module is specifically configured to receive and cache predicted semantics and current response information sent by the server after performing speech recognition and semantic prediction on the current speech segment after performing speech recognition on the current speech segment to obtain a current offline speech recognition result.

13. An apparatus for speech processing, applied to a server, including: A second receiving module configured to receive a current speech segment sent by a client in a case where the client continuously acquires speech data to be recognized; A second recognition module configured to perform speech recognition on the current speech segment to obtain a current online speech recognition result; A semantic prediction module configured to perform semantic prediction on the current online speech recognition result to obtain predicted semantics; A first generation module configured to generate current response information corresponding to the current speech segment based on the predicted semantics; the current response information is an operation instruction corresponding to the predicted semantics. The second sending module is configured to send the predicted semantics and the current response information to the client, so that the client performs speech recognition on complete to-be-recognized speech data received by the client to obtain a target offline speech recognition result, and matches the target offline speech recognition result with all predicted semantics received by the client; in a case where there is predicted semantics matching the target offline speech recognition result, current response information corresponding to the predicted semantics matching the target offline speech recognition result is determined as target response information of the complete to-be-recognized speech data; in a case where there is no predicted semantics matching the target offline speech recognition result, the target offline speech recognition result is subjected to offline semantic analysis to obtain a target semantic meaning; and the target semantic meaning is used to represent an interaction purpose or an intention corresponding to the complete to-be-recognized speech data. The target response information of the complete to-be-recognized speech data is determined based on the target semantic meaning and a preset strategy.

14. The apparatus of claim 13, further comprising: a third receiving module configured to receive complete to-be-recognized speech data sent by the client in a case where the client receives the complete to-be-recognized speech data; a third recognizing module configured to perform speech recognition on the complete to-be-recognized speech data to obtain a target online speech recognition result; a second analyzing module configured to perform semantic analysis on the target online speech recognition result to obtain an online semantic meaning; a second generating module configured to generate online response information corresponding to the complete to-be-recognized speech data based on the online semantic meaning; a third sending module configured to send the online response information to the client.

15. The apparatus of claim 13, further comprising: a third determining module configured to determine whether the current online speech recognition result hits a preset database; an executing module configured to, in a case where the third determining module determines that the current online speech recognition result hits the preset database, trigger a semantic predicting module to perform semantic prediction on the current online speech recognition result to obtain predicted semantics.

16. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

17. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-8.

18. A computer program product comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Voice processing method and device, equipment, storage medium and computer program product

    CN112509580A

  • Vehicle voice interaction method, server and storage medium

    CN114822540A