Voice processing method and system
By extracting audio features and storing, identifying rising or falling signals, and combining speech processing methods of protocol layer, pre-processing layer, model inference layer, recognition layer and quick-revision layer, the existing speech recognition system has solved the problem of insufficient real-time, accuracy and flexibility in real time, achieving more accurate and stable speech recognition.
Patent Information
- Application Number
- CN202510402841.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-25
AI Technical Summary
The existing speech recognition systems have shortcomings in real-time and accuracy, which are difficult to adapt to complex interactive environments, and the system architecture is insufficient and the flexibility and scalability are sufficient, resulting in poor user experience.
By obtaining the to-process voice data, extracting audio features and storing, identifying rising or falling signals, performing a closed-microphone process, obtaining intermediate and final recognition results, and using the protocol layer, pre-processing layer, model inference layer, rejection layer, and quick-revision layer for voice processing, real-time response and accurate recognition are achieved.
It improves the robustness and adaptability in complex dialogue scenarios, provides more accurate final recognition results, meets users' diverse needs, and improves the stability and efficiency of voice interaction.
Smart Images

Figure CN120375869A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech processing, and particularly to a speech processing method and system. Background Art
[0002] As an increasingly popular device in modern families, smart speakers integrate functions of speech recognition, speech synthesis, artificial intelligence, and wireless network connection, interact with users through voice commands, and provide information and execute tasks. The core lies in being able to accurately understand users' voice instructions and respond in a natural and instant manner, which makes speech recognition technology a key factor in the user experience of smart speakers. With the development of speech processing technology, users have put forward higher requirements for the interaction experience of such devices.
[0003] The existing technology mainly adopts a single recognition service. The single recognition service focuses on converting audio signals into text, with the advantages of high accuracy and low latency, but has a single function and is difficult to meet the requirements of actual speaker interaction scenarios. Summary of the Invention
[0004] In view of this, the purpose of the embodiments of the present invention is to provide a speech processing method and system, which can achieve real-time response, provide more accurate and comprehensive final recognition results, improve the robustness and adaptability in complex dialogue scenarios, and thus better meet user needs.
[0005] In a first aspect, an embodiment of the present invention provides a speech processing method, which includes:
[0006] Obtain speech data to be processed;
[0007] Extract and store audio features from the speech data to be processed;
[0008] Identify the audio features to obtain an ascending signal or a descending signal;
[0009] In response to identifying a descending signal, execute a mute process;
[0010] In response to identifying an ascending signal, obtain an intermediate recognition result according to the audio features;
[0011] Execute a mute process according to the intermediate recognition result;
[0012] In response to the speech end completing the mute, obtain a final recognition result according to the audio features.
[0013] In a second aspect, an embodiment of the present invention provides a speech processing system, which includes:
[0014] A protocol layer for obtaining speech data to be processed;
[0015] The pre - processing layer is used to extract and store audio features from the to - be - processed speech data, identify the audio features to obtain a rising signal or a falling signal, and execute a mute process in response to identifying a falling signal;
[0016] The model inference layer is used to, in response to identifying a rising signal, obtain an intermediate recognition result based on the audio features, execute a mute process according to the intermediate recognition result, and obtain a final recognition result based on the audio features in response to the voice end completing the mute;
[0017] The rejection recognition layer is used to judge the validity of the final recognition result based on a predetermined rejection recognition rule;
[0018] The quick repair layer is used to correct the final recognition result.
[0019] The technical solution of the embodiment of the present invention obtains the to - be - processed speech data, extracts and stores audio features from the to - be - processed speech data, identifies the audio features to obtain a rising signal or a falling signal, executes a mute process in response to identifying a falling signal, obtains an intermediate recognition result based on the audio features in response to identifying a rising signal, executes a mute process according to the intermediate recognition result, and obtains a final recognition result based on the audio features in response to the voice end completing the mute. Thus, real - time response can be achieved, a more accurate and comprehensive final recognition result can be provided, the robustness and adaptability in complex dialogue scenarios can be improved, and thus the user needs can be better met. Description of the Drawings
[0020] Through the following description of the embodiments of the present invention with reference to the drawings, the above - mentioned and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0021] Figure 1 is a schematic diagram of the information interaction system of the embodiment of the present invention;
[0022] Figure 2 is a schematic diagram of the voice processing system of the embodiment of the present invention;
[0023] Figure 3 is a flowchart of the process flow and status control of the voice end in the embodiment of the present invention;
[0024] Figure 4 is a flowchart of the voice processing method in the embodiment of the present invention;
[0025] Figure 5 is a flowchart of obtaining audio features in the embodiment of the present invention;
[0026] Figure 6 is a flowchart of recognizing audio features in the embodiment of the present invention;
[0027] Figure 7 It is a flowchart for obtaining intermediate recognition results in an embodiment of the present invention;
[0028] Figure 8 It is a flowchart for executing the process of muting the microphone in an embodiment of the present invention;
[0029] Figure 9 It is a flowchart for obtaining final recognition results in an embodiment of the present invention;
[0030] Figure 10 It is a flowchart for the rejection recognition process in an embodiment of the present invention;
[0031] Figure 11 It is a schematic diagram of the stability guarantee strategy in an embodiment of the present invention;
[0032] Figure 12 It is a schematic diagram of the voice processing device in an embodiment of the present invention;
[0033] Figure 13 It is a schematic diagram of the electronic device in an embodiment of the present invention. Detailed implementation manners
[0034] The following describes the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. In order to avoid obscuring the essence of the present application, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0035] In addition, those of ordinary skill in the art should understand that the drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.
[0036] Unless the context clearly requires otherwise, words such as "including" and "comprising" in the entire application document should be interpreted as having an inclusive meaning rather than an exclusive or exhaustive meaning; that is, it is the meaning of "including but not limited to".
[0037] In the description of the present application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0038] For the solutions described in this specification and embodiments, if they involve personal information processing, they will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for the basic functions, it will not affect the user's use of the basic functions.
[0039] With the rapid development of artificial intelligence and speech processing technologies, speech recognition devices have also made great progress. Speech recognition devices are an important part of modern information technology. By converting human voices into machine-readable text information, they enable natural interaction between humans and machines. These devices include, but are not limited to, smartphones, in-vehicle systems, smart wearable devices, and smart speakers, etc. They not only greatly enrich people's daily lives but also promote the intelligentization process in all walks of life. Especially in the field of smart home, smart speakers are playing an increasingly important role as the home control center. Users can query information, play music, control household appliances, etc. through simple voice commands and enjoy a more convenient lifestyle.
[0040] Existing speech recognition mainly adopts a single recognition service or an integrated engine strategy.
[0041] Among them, the single recognition service focuses on converting the user's speech into text, with the advantages of high recognition accuracy and low latency. However, such services usually lack support for actual speaker interaction scenarios, such as the inability to handle functions like mute determination and end-of-sentence rejection determination. In addition, due to its relatively fixed process, it is difficult to adapt to the diverse access of upstream scenarios and the development of new features, which limits its application in complex interaction environments. Typical representatives include Alibaba Cloud Speech Recognition, iFlytek Recognition, etc.
[0042] The integrated engine strategy attempts to integrate all relevant functions, including but not limited to the Automatic Speech Recognition (ASR) system, mute determination, end-of-sentence rejection determination, etc., to provide a comprehensive speech processing solution. This method simplifies the system setup and process connection steps, improves the collaborative working ability between modules, and reduces the latency problem caused by independent services. However, traditional Automatic Speech Recognition (ASR) systems face many challenges when dealing with streaming speech input. First, the high real-time requirement means that the system must complete the conversion process from speech to text within an extremely short time, which poses strict requirements on the system's computing power and algorithm optimization. Second, the problem of unstable accuracy in dynamic environments is particularly prominent. Factors such as different background noises and speaker accent differences will all affect the recognition accuracy. In addition, to support multiple task scenarios, such as smart home control, online shopping, entertainment services, etc., the ASR system also needs to have good scalability and flexibility. Unfortunately, the current ASR systems, due to their fixed architectures, are difficult to quickly adapt to new algorithm features or expand new task functions, which directly leads to a poor user experience.
[0043] Therefore, the current technology faces several key problems and challenges. Real-time performance and accuracy: Traditional speech recognition systems have difficulty in ensuring both high real-time performance and accuracy in dynamic environments when processing streaming speech inputs. Flexibility of system architecture: Existing ASR systems, due to their fixed architectures, are difficult to quickly adapt to new task requirements and technological advancements, resulting in poor user experiences. Scalability and maintainability: Whether it is a single recognition service or an integrated engine strategy, there are deficiencies in supporting rapidly iterating new features and long-term maintenance. Market trends and user experiences: With the rapid development of the large model and smart speaker markets, users have put forward higher requirements for the voice interaction experience, which poses challenges to the brand's continuous improvement ability and user experience enhancement.
[0044] Therefore, the present invention proposes a speech processing method and system, aiming to overcome the limitations existing in existing speech recognition systems and meet the requirements of diverse application scenarios. At the same time, the technical solutions of the embodiments of the present invention are not only applicable to smart speakers but also can be extended to all devices that require speech recognition services.
[0045] Figure 1 It is a schematic diagram of the information interaction system according to an embodiment of the present invention. As Figure 1 shown, the information interaction system according to an embodiment of the present invention includes a voice terminal 1, a network 2, and a server 3. Among them, the voice terminal 1 is communicatively connected to the server 3 through the network 2.
[0046] The voice terminal 1 is the direct interface between the user and the information interaction system and can be implemented through a smart speaker or other devices that support voice interaction. It is mainly responsible for collecting the user's voice input and transmitting these audio signals to the server for processing. At the same time, it can receive the reply from the server 3 and play it.
[0047] In some embodiments, the voice terminal includes a microphone array, a preprocessing unit, a speaker, a communication unit, etc. Among them, the microphone array is used to capture the sound signals emitted by the user. Through the collaborative work of multiple microphones, the direction positioning of the sound source and noise suppression can be realized, thereby improving the accuracy of speech recognition. The preprocessing unit: performs preliminary processing on the captured audio signals, such as noise reduction, gain adjustment, etc., to optimize the subsequent speech recognition effect. The communication unit is used to communicate with the server, used to send the collected sound to the server and receive the reply from the server. The speaker is used to play the audio sent by the server received.
[0048] Network 2 can be used for the exchange of information and / or data. Among them, Network 2 can be any type of wired or wireless network, or a combination of them. In some embodiments, Network 2 may include a wired network, a wireless network, an optical fiber network, a telecommunications network, an intranet, the Internet, a Local Area Network (LAN), a Wide Area Network (WAN), a Wireless Local Area Networks (WLAN), a Metropolitan Area Network (MAN), a Wide Area Network (WAN), a Public Switched Telephone Network (PSTN), a Bluetooth network, a ZigBee network, or a Near Field Communication (NFC) network, etc., or any combination thereof. In some embodiments, Network 2 may include one or more network access points. For example, Network 2 may include a wired or wireless network access point, such as a base station and / or a network switching node, and one or more components of the information interaction system may be connected to the network through this access point to exchange data and / or information.
[0049] The server 3 is used to receive the audio data from the voice terminal 1, analyze and process it using a predetermined audio processing algorithm, and then generate a corresponding response and return it to the voice terminal. The server side usually consists of multiple subsystems, including but not limited to: Natural Language Processing (NLP) module: Parses the text content, understands the user's query or command, and makes an appropriate response according to the context. Knowledge base and service interface: Stores a large amount of information and can be docked with other service interfaces, such as weather forecast API, music playback service, etc., to meet the diverse needs of users. Feedback generator: Based on the results of ASR and NLP, generates a user-friendly reply and converts it into a voice form through Text-to-Speech (TTS) technology and sends it back to the voice terminal.
[0050] Figure 2 is a schematic diagram of the voice processing system of an embodiment of the present invention. Among them, the voice processing system is a part of the algorithm or module in the server for processing voice, such as Figure 2 As shown, the voice processing system includes a protocol layer L1, a preprocessing layer L2, a model inference layer L3, a quick repair layer L4, and a rejection recognition layer L5.
[0051] Among them, the protocol layer L1 is used to obtain the voice data to be processed. Specifically, the protocol layer L1 is used to communicate with the voice terminal to obtain the voice data to be processed from the voice terminal.
[0052] In some embodiments, the protocol layer L1 can be implemented through the WebSocket (Web Socket) protocol. Among them, WebSocket is a communication protocol that provides a channel for full-duplex communication over a single TCP (Transmission Control Protocol) connection. WebSocket uses the HTTP (Hypertext Transfer Protocol) protocol to initiate a handshake request. The voice side sends an HTTP request with specific header information to the server. If the server supports the WebSocket protocol, it will return a response indicating a switch from the HTTP protocol to the WebSocket protocol. WebSocket allows the server and the voice side to perform real-time data exchange. Through WebSocket, once the connection is established, both the voice side and the server side can actively send data to each other, thus achieving a more efficient and instant communication experience. When the communication ends, either party can initiate a request to close the connection. This usually involves sending a close frame to the other party, and the party receiving the close frame should close its connection end accordingly.
[0053] In some embodiments, most long-connection servers require the voice side to implement the long-connection establishment and message consumption logic by itself, and interact according to the corresponding protocols and processes. However, streaming and real-time voice interaction systems usually have high requirements for overall performance, process specifications, and usability. Implementing the logic, controlling the content structure, and process by the upstream will not only increase the repetitive work (among multiple upstreams), but also may cause service errors due to incorrect interaction actions of the upstream, increasing the stability risk. Therefore, the embodiments of the present invention provide a standard Java SDK (Java Software Development Kit) to the voice side, simplify the access process, standardize the protocols and error codes, and organize the documents. The behavior of the voice side is standardized through the Java SDK state machine, helping the upstream to achieve quick access within minutes, standardizing the behavior of the voice side while reducing the need for answering questions and providing assistance.
[0054] Specifically, the core of the state control of the voice side is the process flow and state management based on the state machine. Figure 3 It is the flowchart of the process flow and state management of the voice side in the embodiments of the present invention. As Figure 3 shown, the process flow and state management of the voice side include the following steps:
[0055] Step S101, initialize the state.
[0056] In this embodiment, when the request is initialized, the voice terminal is in the INIT (initialization state). In this state, the voice terminal is ready to start interacting with the server. Specifically, in this step, the voice terminal initializes environment variables, configuration parameters, etc.
[0057] Step S102, Instance creation failure state.
[0058] In this embodiment, the voice terminal attempts to create a client instance using the target service information provided by the service. If the target service is invalid or an error occurs during the creation process, it enters the CLIENT_CREATE_FAIL (instance creation failure state). At the same time, an error log is recorded, and the user or administrator may be notified about the failure of instance creation.
[0059] Step S103, Instance creation success state.
[0060] In this embodiment, if in step S102, the target service is valid and the client instance is successfully created, the system will enter the CLIENT_CREATE_SUCCESS (instance creation success state). At the same time, it is confirmed that the required resources and services have been correctly loaded and configured to prepare for the next step of connecting to the service.
[0061] Step S104, Failed to establish a connection with the ASR service.
[0062] In this embodiment, when the voice terminal is in the CLIENT_CREATE_SUCCESS state, it attempts to establish a connection with the Automatic Speech Recognition (ASR) service, that is, attempts to establish a connection with the server. If the connection fails due to network problems or other reasons, it enters the CONNECT_FAIL (failed to establish a connection with the service state). Handle the situation of connection failure, such as a retry mechanism or reporting the error to the user / administrative system.
[0063] Step S105, Successfully established a connection with the ASR service.
[0064] In this embodiment, if the connection with the ASR service is successful in step S104, it enters the CONNECT_SUCCESS (successfully established a connection with the service state). At the same time, verify the validity of the connection to ensure that data can be normally sent and received.
[0065] Step S106, ASR request has been sent.
[0066] In this embodiment, after successfully establishing a connection with the ASR service, in the CONNECT_SUCCESS state, an ASR request is initiated to the server side, and after filling the request content according to the SDK standard, it enters the REQUEST_SEND (requested state). That is, send the voice data stream to the ASR service and wait for the processing result.
[0067] Specifically, in a streaming and real-time voice interaction scenario, the voice end first establishes a connection with the server and initializes necessary settings. Then, it collects voice data from audio input devices such as microphones, and preprocesses (such as noise reduction) and frames these data to split the continuous voice stream into small data packets suitable for transmission. Subsequently, the voice end sends these processed voice data frames to the server one by one.
[0068] In some embodiments, when sending an ASR request, the voice end can pack multiple frames of voice data in one ASR request for sending. Thus, by sending multiple frames of data in batches, the additional overhead caused by frequently initiating network requests can be reduced. At the same time, batch processing and sending of data can generally make better use of network bandwidth and improve data transmission efficiency. Among them, the voice data to be processed obtained by protocol layer L1 is the voice data in the ASR request sent by the voice end.
[0069] Step S107, recognition / transmission of the end packet has stopped.
[0070] In this embodiment, in the REQUEST_SEND state, the voice data stream can be sent in real time, and the server can be notified to stop recognition at any time. At this time, it enters the STOP_SEND (state where recognition has been terminated). In addition, there are several other situations that will cause the process to end, including abnormal connection closure (CLOSE), error (EXCEPTION), and final result return (COMPLETED). In these situations, there is no subsequent state transition. Specifically, a signal is sent to the server indicating that the voice input has been completed, or the current session is interrupted for various reasons.
[0071] Among them, for state control, restrictions are mainly imposed on operations such as start (sending a voice request to start recognition), send (sending the voice stream in real time), and stop (notifying the server to stop recognition).
[0072] When the instance is in the INIT, CLIENT_CREATE_SUCCESS, CLIENT_CREATE_FAIL, CONNECT_FAIL, STOP_SEND, EXCEPTION, or CLOSE state, the above three operations cannot be performed;
[0073] When the instance is in the CONNECT_SUCCESS state, the start operation can be performed;
[0074] When the instance is in the REQUEST_SEND state, the send and stop operations can be performed;
[0075] When the instance is in the COMPLETED state, the start and stop operations can be performed;
[0076] Thus, through the above state machine-based design, different stages and their transitions in the speech recognition process can be effectively managed, ensuring the stability and response speed of the system.
[0077] In this embodiment, the preprocessing layer L2 is used to extract and store audio features from the to-be-processed speech data, identify the audio features to obtain a rising signal or a falling signal, and execute a mute process in response to the identified falling signal.
[0078] Among them, the preprocessing layer L2 includes a first decoding module 21, a feature extraction module 22, and a first determination module 23.
[0079] Among them, the first decoding module 21 is used to decode the to-be-processed speech data. Specifically, after the process receives the real-time speech data uplinked from the speaker end, due to differences between different devices, the audio formats sent may vary. Common formats include OGG, WAV, Opus, etc. Therefore, the first decoding module is used to uniformly decode these audio into a specified format. For example, it can be decoded into a single-channel WAV format based on PCM (Pulse Code Modulation) encoding for subsequent processing and analysis. In the request initialization stage, the first decoding module will initialize the corresponding local decoder according to the audio type reported by the speech end. For example, if the data reported by the speech end is in OGG format, the decoder for the OGG format will be initialized; if it is in Opus format, the corresponding Opus decoder will be initialized. Thus, it can be ensured that no matter what format the input audio is, a suitable decoding method can be found. In the subsequent process of streaming speech data, the locally configured JNI (Java Native Interface) decoder in the initialization stage is used to perform real-time decoding on the uplinked speech data. The JNI technology allows Java code to interact with applications or libraries written in other languages. Here, it is mainly used to call an efficient local decoding library to achieve fast and accurate audio decoding. In this way, various different formats of audio data can be effectively converted into a unified standard format (such as single-channel WAV encoded by PCM), thereby simplifying subsequent processing steps and improving the compatibility and efficiency of the entire system. This not only solves the problem of inconsistent audio formats caused by device differences but also ensures that speech recognition and other processing tasks can be executed efficiently and accurately.
[0080] The feature extraction module 22 is used to extract features from the decoded content to obtain audio features and store them. Among them, the feature extraction module 22 is a preset local feature extraction JNI (Java Native Interface). Specifically, it is first determined whether the first decoding module has successfully decoded valid content this time. This process includes verifying the integrity and format correctness of the decoded data to ensure the effectiveness of subsequent processing steps. If valid content is successfully decoded, the preset local feature extraction JNI (Java Native Interface) is used to extract features from the content decoded this time. The role of JNI here is to call efficient native libraries to perform complex mathematical operations and signal processing tasks, so as to extract useful feature information from audio data, such as Mel Frequency Cepstral Coefficients (MFCC), spectral features or other customized feature representations. After the feature extraction is completed, it is further determined whether audio features have been successfully extracted this time. The effectiveness and integrity of the extraction results can be checked to ensure that the extracted audio features can accurately reflect the key attributes of the original audio data and are suitable for subsequent processing or analysis tasks. If valid audio features are extracted, the content of these audio features is saved to the memory for subsequent processes to use. This can speed up data access and provide the necessary input data for the next task, thus ensuring the efficiency and coherence of the entire processing flow. In addition, the feature data saved in the memory can also facilitate real-time processing and quick response, meeting the low-latency requirements of the streaming voice interaction system.
[0081] The first determination module 23 is used to identify the audio features to obtain a rising signal or a falling signal, and in response to identifying a falling signal, execute the microphone mute process. Among them, the first determination module is a locally pre-set microphone mute determination (VAD, Voice Activity Detection) module. Specifically, it receives the audio characteristics transmitted from the feature extraction module in real time and performs microphone mute determination based on this audio feature information. In the request initialization stage, the server will provide basic configuration information, such as parameters like the longest mute duration and semantic microphone mute delay, to initialize the corresponding local VAD service. This step ensures that the VAD module can be adjusted and optimized according to the specific application scenario requirements. In the subsequent microphone mute determination process, the first determination module will obtain the audio features output in the previous stage, call the already initialized local VAD service, and parse the call result. First, it determines whether a rising signal (i.e., the start of speech) is recognized according to the result of the local VAD. If a rising signal is recognized, its core information, mainly the position of the rising signal (usually represented by the number of frames), is recorded for subsequent tracking of the starting point of the speech segment. Then, it continues to detect the audio stream to determine whether the service recognizes a falling signal (i.e., the end of speech). When a falling signal is detected, it means that a complete speech segment has ended. After confirming that a falling signal is recognized, the server will execute the microphone mute process. Specifically, the server sends a microphone mute instruction to notify the voice end to perform the microphone mute operation, and at the same time performs corresponding resource release (such as closing the audio acquisition channel, releasing memory, etc.), and records the relevant information of this event to provide data support for subsequent log analysis or problem troubleshooting. Thus, the server can effectively manage the detection and response of voice activities, ensure that the microphone is turned off in a timely manner when there is no voice activity, reduce unnecessary resource consumption, and improve the overall system efficiency and performance.
[0082] In this embodiment, the model inference layer L3 is used to, in response to identifying a rising signal, obtain an intermediate recognition result according to the audio features, execute the microphone mute process according to the intermediate recognition result, and in response to the voice end completing the microphone mute, obtain a final recognition result according to the audio features.
[0083] Specifically, the model inference layer L3 includes a first encoding module 31, a second determination module 32, a second encoding module 34, a pinyin cache module 35, and at least one second decoding module.
[0084] Among them, the first encoding module 31 is a streaming acoustic encoding module, which is used to determine the request information corresponding to the to-be-processed speech data, and the request information includes at least one of device identification, product identification, and request content; obtain hot words at the speech end according to the request information; segment the audio features according to the rising signal to obtain to-be-processed audio segments; encode the to-be-processed audio segments to obtain first encoding information; obtain and record the tail point information corresponding to the first encoding information; obtain the intermediate recognition result according to the first encoding information and the hot words, and compare the intermediate recognition result output in this round with the intermediate recognition result in the previous round; in response to the inconsistency between the intermediate recognition result output in this round and the intermediate recognition result in the previous round, synchronize the intermediate recognition result output in this round to the speech end; update the intermediate recognition result according to the intermediate recognition result output in this round.
[0085] Specifically, after the rising signal is recognized, start receiving the audio features extracted in real time and perform real-time streaming recognition on the content. First, initialize and call the codec recognition service to obtain unique identification information (such as the target service IP, sessionId, etc.), and perform subsequent operations based on this information, and this unique identification information is used to identify the unique streaming encoding service. Then, uniformly obtain the corresponding hot words based on the information such as the device identification, product label, and request content associated with the request information of the ASR request for subsequent decoding calls. When entering the real-time recognition stage, segment the audio features extracted in this round according to the rising signal information given by the first determination module 23, and call the preset streaming encoding service to process the acoustic signal according to the unique identification information to generate first encoding information, and the first encoding information is various feature representations. Subsequently, parse and record the tail point information in the encoding output, combine the encoding output content and the hot words, and call the decoding service according to the initial identification information. After obtaining the decoding information, select the result with the highest confidence as the intermediate recognition result, and the intermediate recognition result is the content of the current streaming recognition, and compare it with the previous recognition content. If the content is the same, synchronize the intermediate recognition result to the first determination module 23; if not, it is considered that the intermediate recognition result has changed, synchronize the latest intermediate recognition result to the speech end, and call the preset second determination module 32 to update the current semantic integrity, and then synchronize the intermediate recognition result to the first determination module 23. Thus, the server can efficiently process speech data, ensure the accuracy and timeliness of the recognition process, and optimize the user experience at the same time.
[0086] The second determination module 32 is a semantic-based mute determination module, which is used to update the intermediate recognition result according to the intermediate recognition result output in this round when the intermediate recognition result output in this round is inconsistent with that in the previous round. Among them, the second determination module 32 is used to perform semantic analysis on the newly recognized intermediate recognition result to determine whether this content is complete, coherent, and conforms to the expected logical structure or intention. Based on the result of semantic analysis, the second determination module 32 will evaluate and update the current semantic integrity. For example, it can judge whether the newly added information constitutes a complete sentence or command, and whether more input is needed to complete the current semantic unit, etc. Thereby, the recognition result can be made more accurate.
[0087] The second encoding module 34 is a non-streaming encoding acoustic encoding module, which is used to detect whether a mute instruction has been issued in response to the intermediate recognition result output in this round being consistent with that in the previous round, or updating the current intermediate recognition result; in response to no mute instruction having been issued, recognizing the tail point information and the current intermediate recognition result; in response to recognizing a descending signal, sending a mute instruction to the voice end, and releasing and recording resources.
[0088] Specifically, during the process of the server recognizing the entire sentence content at the corresponding time to obtain the final recognition result, the server first judges whether a mute instruction has been issued to the voice end or whether the voice end has synchronously sent an EOF (End of File) signal to the server. If either condition is met, it indicates that subsequent non-streaming recognition processing can be performed. The server obtains all the feature contents of this request and calls the preset non-streaming encoding service for processing. After obtaining the result of non-streaming encoding, the system will parse the pinyin modeling output in these results and perform corresponding vocabulary conversion to obtain the specific pinyin result. Based on the wake-up word segmentation logic and the locally preset pinyin cache, the system will check whether the cache content is hit. If it is hit, the content of the pinyin cache is directly used as the final recognition result; if it is not hit, the non-streaming decoding service is continuously called to obtain the final recognition result. The specific logic of this process is the same as the streaming decoding processing logic, that is, the result with the highest confidence in the decoding output is selected as the final recognition output. In this way, the system can efficiently and accurately complete the conversion from voice to text and feedback it to the user or the subsequent process in a timely manner while ensuring the integrity of the voice data. At the same time, in the case of detecting a descending signal and no mute instruction having been issued, the system will send a mute instruction to the voice end and perform necessary resource release and recording to optimize the system performance and user experience.
[0089] The Pinyin cache module 35 is a pre-set pinyin cache, which is used to store the pinyin results generated in the previous recognition process and their corresponding context information. Specifically, the content stored in this module may include pinyin sequences, context information, hit records, wake-up words or keyword information, etc.
[0090] Among them, the pinyin sequence: stores the specific pinyin sequences extracted and converted from the speech data. These pinyin sequences are the results obtained after processing the audio features based on the non-streaming encoding service and are formed after word table conversion.
[0091] Context information: In addition to the pinyin itself, it may also include context information related to these pinyins, such as the recognition timestamp, the session ID to which it belongs, the device identifier, etc. This information helps to provide more accurate matching and richer background support during subsequent processing or querying.
[0092] Hit records: Record which pinyin sequences have been successfully matched and used. This can improve the system's response speed. When the same speech segment appears again, the results in the cache can be directly utilized without having to perform complex processing procedures again.
[0093] Wake-up word or keyword information: In particular, if there are specific wake-up words or keywords in the system (for example, command words often said by the user), the pinyin cache module will also store the pinyin forms corresponding to these words and the related segmentation logic. In this way, during the real-time recognition process, once a matching item is detected, a quick response can be made.
[0094] In this way, the pinyin cache module can not only speed up the recognition speed, reduce the workload of repeated calculations, but also improve the recognition accuracy, especially when dealing with common phrases or command words. In addition, it also provides better stability and efficiency for the system, making the overall voice interaction experience more fluent and natural.
[0095] For at least one second decoding module, two examples are used in the embodiments of the present invention, namely 33a and 33b. Among them, the second decoding module 33a is used to obtain an intermediate recognition result according to the first encoding information and hot words during the streaming encoding process, and the second decoding module 33b is used to obtain the final recognition result through the second decoding module 33b if the pinyin result does not hit the content of the pinyin cache during the non-streaming encoding process.
[0096] The rejection layer L5 is used to judge the validity of the final recognition result based on a predetermined rejection rule. Among them, the rejection layer L5 includes a first rejection module 51 and a second rejection module 52.
[0097] Among them, the first rejection recognition module 51 is used to judge the validity of the final recognition result based on at least one of a blacklist, a whitelist, and voice parameters. Specifically, the first rejection recognition module 51 detects whether the final recognition result is in the blacklist; in response to the final recognition result being in the blacklist, it determines that the final recognition result is invalid; in response to the final recognition result not being in the blacklist, it detects whether the final recognition result is in the whitelist; in response to the final recognition result being in the whitelist, it determines that the final recognition result is valid; in response to the final recognition result not being in the whitelist, it obtains the voice parameters of the final recognition result, and the voice parameters include at least one of the number of words and the speech rate; in response to the voice parameters meeting a predetermined condition, it determines that the final recognition result is invalid; in response to the voice parameters not meeting the predetermined condition, it enters the second rejection recognition module 52 for the next judgment.
[0098] Among them, the second rejection recognition module 52 is used to judge the validity of the final recognition result based on at least one of a voice score and a text score. Specifically, in response to the voice parameters not meeting the predetermined condition, it determines the voice score corresponding to the final recognition result; in response to the voice score being less than a first predetermined threshold, it determines that the final recognition result is invalid; in response to the voice score being greater than or equal to the first predetermined threshold, it detects whether it is an active wake-up; in response to it being an active wake-up, it determines that the final recognition result is valid; in response to it not being an active wake-up, it determines the text score corresponding to the final recognition result; in response to the text score being less than a second predetermined threshold, it determines that the final recognition result is invalid; in response to the text score being greater than or equal to the second predetermined threshold, it determines that the final recognition result is valid.
[0099] The quick repair layer L4 is used to correct the final recognition result. Among them, the quick repair layer L4 includes a question answering module 41 and a quick repair module 42.
[0100] Among them, the question answering module 41 is a FAQ (Frequently Asked Questions), including a series of pre-prepared answers, which are aimed at questions frequently asked by users. The goal of the FAQ is to provide quick self-service so that users can find solutions by themselves without directly contacting the support team. The FAQ can significantly reduce the time for handling simple and repetitive questions, enabling support staff to concentrate on handling more complex problems.
[0101] The quick repair module 42 is a regular expression-based quick repair (RE, Regular Expression) module for correcting the final recognition result. Among them, regular expressions are a powerful tool for matching strings, which allows defining patterns for searching text. These patterns can describe single characters, groups of characters, numbers, word boundaries, etc., and support complex logics such as "or" conditions, "and" conditions, and repetition counts. Therefore, regular expressions are very suitable for finding, replacing, and parsing text data. The quick repair module 42 can automate some repetitive tasks that can be solved by pattern matching.
[0102] In some embodiments, the voice processing system of the embodiments of the present invention also supports configuration management to optimize its performance, adapt to specific application scenarios, or meet personalized needs.
[0103] For example, configurations can be made for the model name, model version, hot words, VAD, rejection recognition, etc. The models include all the models involved in the embodiments of the present invention, such as feature extraction models, streaming encoding models, streaming decoding models, non-streaming encoding models, non-streaming decoding models, etc. For the model name, specify the specific model name used for the speech recognition task. Different models may be optimized for different languages, dialects, or specific application fields. For the model version: select a specific version of the model for deployment. This is important for ensuring consistency and rolling back to a stable version, especially during the model update and iteration process. Hot words: These are user-defined keywords or phrases that the system will pay more attention to. For example, in smart home control, command words such as "turn on the light" and "close the window" can be set as hot words to improve the recognition accuracy. VAD (Voice Activity Detection) configuration is used to detect the voice activity in the audio stream and distinguish the speech segment from the non-speech segment. The configuration items may include the longest silence duration, sensitivity threshold, etc., so as to adjust the behavior of VAD according to the actual environment. Rejection recognition configuration is used to reject those inputs that do not conform to the expected pattern to prevent misrecognition. The configuration can include setting the words or phrases to be automatically ignored under certain specific conditions, or defining a minimum confidence threshold, and the results below this threshold will be rejected.
[0104] For another example, the dimensions can also be configured, including configurations for the global, request, device, tags, etc., further enhancing flexibility and customizability. Global configuration: Applies to all instances and requests of the entire system, providing a unified way to set default parameters. For example, default VAD parameters or default model versions can be set here. Request-level configuration: Configuration for a single request, allowing parameters to be dynamically adjusted when sending a request. For example, different model versions or hotword lists can be used in a special request. Device-level configuration: Configuration based on the device type or characteristics. For example, for devices with different hardware capabilities (such as high-end smartphones and low-end smart speakers), respective parameter sets can be configured. Tag configuration: Organizes and classifies configuration items through tags, making management and search easier. For example, tags can be created for specific types of voice interactions (such as customer service, home entertainment), and corresponding values can be set for all related configuration items under each tag.
[0105] In some embodiments, the voice processing system further includes general components, such as a frame assembly component, a logging component, a dimension conversion component, etc.
[0106] Among them, the frame assembly component is used to recombine segmented audio data into complete voice segments for subsequent processing or analysis. Specifically, when voice data is split into multiple small data frames (for example, to adapt to network transmission limitations or improve processing efficiency), the frame assembly component is responsible for recombining these frames in the correct order.
[0107] The logging component is used to record various events and status information during the operation of the system, which is crucial for detecting system performance, diagnosing problems, and optimizing operations. Specifically, the logging component is used to record various events that occur in the system, including but not limited to user requests, recognition results, error messages, etc., providing detailed debugging information to help developers track and solve potential problems. By viewing the logs, it is easier to locate the fault points and analyze the reasons. At the same time, the response time and resource usage of the system can also be detected, which helps to identify bottlenecks and optimize performance.
[0108] The dimension conversion component is used to convert the data representation form between different abstraction levels, enabling the seamless transfer and processing of data between different modules or systems. Specifically, the data is converted from one format to another to meet different processing requirements. For example, converting the original audio signal into a feature vector suitable for input to a machine learning model. When transmitting information between different modules, it may be necessary to adjust the context or perspective of the data. The dimension conversion component can convert the data specific to a certain module into the format required by another module while preserving the necessary semantic information. It supports multi-dimensional data analysis, allowing the system to process complex data sets from different sources or types within the same framework. This helps to improve the flexibility and scalability of the system. Ensure the consistency and integrity of the data during the conversion process to prevent information loss or distortion.
[0109] In an embodiment of the present invention, by obtaining the speech data to be processed, extracting and storing the audio features from the speech data to be processed, identifying the audio features to obtain an ascending signal or a descending signal, in response to identifying a descending signal, performing a mute process, in response to identifying an ascending signal, obtaining an intermediate recognition result according to the audio features, performing a mute process according to the intermediate recognition result, and in response to the speech end completing the mute, obtaining a final recognition result according to the audio features. Thus, real-time response can be achieved, providing a more accurate and comprehensive final recognition result, enhancing the robustness and adaptability in complex dialogue scenarios, and thus better meeting user needs.
[0110] Figure 4 It is a flowchart of the speech processing method according to an embodiment of the present invention. Figure 4 The shown speech processing method is executed by a server, and specifically includes the following steps:
[0111] Step S210: Obtain the speech data to be processed.
[0112] In this embodiment, the server receives an ASR request sent by the speech end, and the ASR request includes the speech data to be processed. Specifically, the speech end collects speech data from audio input devices such as microphones, and preprocesses (such as noise reduction) and frames these data so as to split the continuous speech stream into small data packets suitable for transmission. When sending an ASR request, the speech end can pack multiple frames of speech data in one ASR request and send it.
[0113] Step S220: Extract and store the audio features from the speech data to be processed.
[0114] In this embodiment, for multiple frames of speech data to be processed in one ASR request, according to the sequence of the speech data to be processed, each frame of the speech data to be processed is processed separately to extract and store the audio features.
[0115] Specifically,Figure 5 is a flowchart for obtaining audio features according to an embodiment of the present invention. As Figure 5 shown, extracting audio features from the to-be-processed speech data and storing them includes the following steps:
[0116] Step S221, initialize the first decoding module.
[0117] In this embodiment, after obtaining the to-be-processed speech data, due to the differences between different speech terminals, the sent audio formats may vary, and common formats include OGG, WAV, Opus, etc. Therefore, in the request initialization stage, the server initializes the corresponding local decoder (the first decoding module) according to the audio type reported by the speech terminal. For example, if the data reported by the speech terminal is in OGG format, the decoder for OGG format is initialized; if it is in Opus format, the corresponding Opus decoder is initialized. Thus, it can be ensured that no matter what format the input audio is, a suitable decoding method can be found.
[0118] Step S222, decode the to-be-processed speech data.
[0119] In this embodiment, the decoder configured in step S221 is used to perform real-time decoding on the to-be-processed speech data. In this way, various different formats of audio data can be effectively converted into a unified standard format (such as single-channel WAV encoded in PCM), thereby simplifying subsequent processing steps and improving the compatibility and efficiency of the entire system. This not only solves the problem of inconsistent audio formats caused by device differences but also ensures that speech recognition and other processing tasks can be executed efficiently and accurately.
[0120] Step S223, whether valid content is decoded.
[0121] In this embodiment, it is detected whether valid content is decoded from the to-be-processed speech data. Specifically, the integrity and format correctness of the decoded data can be verified to determine whether valid content is decoded to ensure the effectiveness of subsequent processing steps.
[0122] In response to decoding valid content, enter step S224.
[0123] In response to not decoding valid content, enter step S227.
[0124] Step S224, extract features from the decoded content.
[0125] In this embodiment, in response to decoding valid content, feature extraction is performed on the decoded content. Specifically, feature extraction is performed on the content obtained by this decoding to obtain audio features, where the audio features are useful feature information extracted, such as Mel Frequency Cepstral Coefficients (MFCCs), spectral features, or other customized feature representations.
[0126] Step S225: Audio features are extracted.
[0127] In this embodiment, it is detected whether audio features are extracted. Specifically, it can be done by checking the validity and integrity of the extraction result to ensure that the extracted audio features can accurately reflect the key attributes of the original audio data and are suitable for subsequent processing or analysis tasks.
[0128] In response to extracting audio features, proceed to step S226.
[0129] In response to not extracting audio features, proceed to step S227.
[0130] Step S226: Store the audio features in memory.
[0131] In this embodiment, in response to extracting audio features, the content of these audio features is saved to memory for use in subsequent processes.
[0132] Step S227: A termination signal is received.
[0133] In this embodiment, in response to not decoding valid content, or decoding valid content but not extracting audio features, it is detected whether a termination signal is received. Among them, the termination signal is an instruction or flag sent from the voice side to the server, used to notify the server that the current voice interaction process should be terminated or paused.
[0134] In response to not receiving the termination signal, return to step S221 to process the next frame of voice data to be processed.
[0135] In response to receiving the termination signal, proceed to step S228.
[0136] Step S228: Terminate the recognition process of the voice data to be processed.
[0137] In this embodiment, if the termination signal is not received, it means there is a problem with this round of audio, and the recognition process of the voice data to be processed is terminated.
[0138] That is, in response to receiving the termination signal sent from the voice side and not decoding valid content, terminate the recognition process of the voice data to be processed. Or, in response to receiving the termination signal sent from the voice side and not extracting audio features, terminate the recognition process of the voice data to be processed.
[0139] It should be noted that in the embodiments of the present invention, when the voice terminal sends an ASR request to the server, the request includes multiple frames of voice data. In the embodiments of the present invention, the processing process of one frame of voice data is referred to as one round.
[0140] Step S230: Identify the audio feature to obtain a rising signal or a falling signal.
[0141] In this embodiment, the obtained audio feature is retrieved from the memory, and the audio feature is identified to obtain a rising signal or a falling signal.
[0142] Specifically, Figure 6 is a flowchart for identifying the audio feature in the embodiments of the present invention. As Figure 6 shown, identifying the audio feature to obtain a rising signal or a falling signal includes the following steps:
[0143] Step S231: Initialize the first determination module.
[0144] In this embodiment, during the request initialization phase, the server provides basic configuration information, such as parameters like the longest silence duration and semantic mute delay, to initialize the first determination module, that is, to initialize the local VAD service. Thus, it can be ensured that the VAD module can be adjusted and optimized according to the specific application scenario requirements.
[0145] Step S232: Identify the audio feature through the first determination module.
[0146] In this embodiment, the first determination module retrieves the audio feature obtained from the previous stage, calls the initialized local VAD service, and parses the call result.
[0147] Step S233: Whether a rising signal is identified.
[0148] In this embodiment, it is determined whether a rising signal (i.e., the start of speech) is identified based on the result of the local VAD.
[0149] If a rising signal is identified, proceed to step S234.
[0150] If a rising signal is not identified, proceed to step S235.
[0151] Step S234: Record the rising signal information.
[0152] In this embodiment, if a rising signal is identified, the core information of the rising signal, such as the position of the rising signal (usually represented by the number of frames), is recorded to facilitate subsequent tracking of the starting point of the speech segment. Then proceed to step S235.
[0153] Step S235: Determine whether a descending signal is recognized.
[0154] In this embodiment, continue to recognize the audio signal to determine whether the service has recognized a descending signal (i.e., the end of the speech).
[0155] If no descending signal is recognized, do not execute the mute process.
[0156] If a descending signal is recognized, proceed to step S240.
[0157] Step S240: In response to recognizing a descending signal, execute the mute process.
[0158] In this embodiment, if a descending signal is recognized in the audio features, execute the mute process. Specifically, execute the mute process by sending a mute command to the voice terminal and releasing resources and recording.
[0159] Step S250: In response to recognizing an ascending signal, obtain an intermediate recognition result based on the audio features.
[0160] In this embodiment, if an ascending signal is recognized in steps S233 - S234 and the information of the ascending signal is recorded, obtain an intermediate recognition result based on the audio features.
[0161] Figure 7 This is the flowchart for obtaining an intermediate recognition result in an embodiment of the present invention. As Figure 7 shown, obtaining an intermediate recognition result based on the audio features includes the following steps:
[0162] Step S2501: Initialize the first encoding module.
[0163] In this embodiment, the first encoding module is a streaming acoustic encoding module. Specifically, initialize and call the codec recognition service to obtain unique identification information (such as the target service IP, sessionId, etc.), and perform subsequent operations based on this information, and this unique identification information is used to identify the unique streaming encoding service.
[0164] Step S2502: Obtain hotwords.
[0165] In this embodiment, determine the request information corresponding to the speech data to be processed. The request information includes at least one of device identification, product identification, and request content. Obtain the hotwords of the voice terminal according to the request information. Specifically, uniformly obtain the corresponding hotwords based on the information such as device identification, product label, and request content associated with the ASR request for subsequent decoding calls.
[0166] Step S2503: Segment the audio features according to the rising signal to obtain the audio segments to be processed.
[0167] In this embodiment, the audio features extracted in this round are segmented according to the rising signal to obtain the audio segments to be processed. Specifically, when a rising signal is detected, it indicates the start of a new speech segment. The currently collected audio features will be segmented according to this signal to distinguish the new speech segment from other speech segments or background noise.
[0168] Step S2504: Encode the audio segments to be processed to obtain the first encoded information.
[0169] In this embodiment, the preset streaming encoding service is called according to the unique identification information in step S2501 above to process the audio segments to be processed, and the first encoded information is generated. The first encoded information is a variety of feature representations.
[0170] Step S2505: Obtain and record the end point information corresponding to the first encoded information.
[0171] In this embodiment, the end point information in the first encoded information is parsed and recorded. Among them, the end point information refers to the time stamp or frame number of the end position of a speech segment, indicating the termination point of speech activity.
[0172] Step S2506: Obtain the intermediate recognition result according to the first encoded information and the hot words.
[0173] In this embodiment, the decoding service is called by combining the first encoded information and the hot words. The decoding service obtains the intermediate recognition result according to the first encoded information and the hot words. Specifically, after obtaining the decoding information output by the decoding service, the result with the highest confidence is selected as the intermediate recognition result. The intermediate recognition result is the semantic content of the current streaming recognition.
[0174] Step S2507: Detect whether the intermediate recognition result has changed.
[0175] In this embodiment, the intermediate recognition result output in this round is compared with the intermediate recognition result in the previous round to detect whether the intermediate recognition result has changed.
[0176] In response to the change in the intermediate recognition result, that is, the intermediate recognition result output in this round is inconsistent with the intermediate recognition result in the previous round, go to step S2507.
[0177] In response to the fact that the intermediate recognition result has not changed, that is, the intermediate recognition result output in this round is consistent with the intermediate recognition result in the previous round, go to step S2510.
[0178] Step S2508: Synchronize the intermediate recognition result output in this round to the voice side.
[0179] In this embodiment, in response to a change in the intermediate recognition result, that is, the intermediate recognition result output in this round is inconsistent with the intermediate recognition result in the previous round, synchronize the intermediate recognition result output in this round to the voice side.
[0180] Step S2509: Update the intermediate recognition result according to the intermediate recognition result output in this round.
[0181] In this embodiment, in response to a change in the intermediate recognition result, that is, the intermediate recognition result output in this round is inconsistent with the intermediate recognition result in the previous round, call the preset second determination module to update the current semantic integrity.
[0182] Among them, the second determination module is a voice-off determination module based on semantics, which is used to update the intermediate recognition result according to the intermediate recognition result output in this round when the intermediate recognition result output in this round is inconsistent with the intermediate recognition result in the previous round. Among them, the second determination module is used to perform semantic analysis on the newly recognized intermediate recognition result to determine whether this content is complete, coherent, and conforms to the expected logical structure or intention. Based on the result of semantic analysis, the second determination module will evaluate and update the current semantic integrity. For example, it can judge whether the newly added information constitutes a complete sentence or command, and whether more input is needed to complete the current semantic unit, etc. Thus, the recognition result can be made more accurate.
[0183] Step S2510: Information synchronization.
[0184] In this embodiment, in response to the intermediate recognition result output in this round being consistent with the intermediate recognition result in the previous round, or after updating the current intermediate recognition result, synchronize the processing result of this round to the voice-off service.
[0185] Step S260: Execute the voice-off process according to the intermediate recognition result.
[0186] In this embodiment, in response to the intermediate recognition result output in this round being consistent with the intermediate recognition result in the previous round, or after updating the current intermediate recognition result, synchronize the processing result of this round to the voice-off service, and the voice-off service executes the voice-off process according to the intermediate recognition result.
[0187] Figure 8 It is a flowchart of executing the voice-off process in an embodiment of the present invention. As Figure 8 shown, the steps of executing the voice-off process according to the intermediate recognition result are as follows:
[0188] Step S261: Detect whether a voice-off instruction has been issued.
[0189] In this embodiment, in response to the intermediate recognition result output in this round being the same as the intermediate recognition result in the previous round, or updating the current intermediate recognition result, it is detected whether a mute instruction has been issued. If no mute instruction has been issued, it means that the local VAD service did not give a mute determination in step S240, and step S262 is entered.
[0190] If a mute instruction has already been issued, it means that the local VAD service has already given a mute determination in step S240, and the mute process is not executed anymore.
[0191] Step S262: Obtain the tail point information and the intermediate recognition result.
[0192] In this embodiment, the tail point information and the intermediate recognition result obtained in the above step S250 are obtained.
[0193] Step S263: Detect whether a descending signal is recognized.
[0194] In this embodiment, the tail point information and the current intermediate recognition result are recognized to detect whether a descending signal is recognized. Specifically, in the embodiment of the present invention, the first determination module recognizes the tail point information and the current intermediate recognition result, and the specific recognition method is similar to that in step S230, which will not be elaborated herein in the embodiment of the present invention.
[0195] In response to recognizing a descending signal, step S264 is entered.
[0196] In response to not recognizing a descending signal, the mute process is not executed anymore.
[0197] Step S264: Send a mute instruction to the voice end, and perform resource release and recording.
[0198] In this embodiment, in response to recognizing a descending signal, a mute instruction is sent to the voice end, and resource release and recording are performed.
[0199] Step S270: In response to the voice end completing muting, obtain the final recognition result according to the audio feature.
[0200] In this embodiment, when the server sends a mute instruction to the voice end or receives a termination instruction sent by the voice end, it indicates that the voice end has completed muting, and the server obtains the final recognition result according to the audio feature.
[0201] Specifically, Figure 9 is the flowchart of obtaining the final recognition result in the embodiment of the present invention. As Figure 9 shown, obtaining the final recognition result according to the audio feature includes the following steps:
[0202] Step S271: Obtain the audio feature of this request.
[0203] In this embodiment, in response to the voice end completing muting the microphone, all audio features corresponding to the current request are obtained. Among them, all audio features corresponding to the current request include all audio features between the moment when the rising signal is recognized and the moment when the voice end completes muting the microphone. Among them, the moment when the rising signal is recognized is the most recent moment when the rising signal is recognized before the moment when the voice end completes muting the microphone.
[0204] Step S272: Encode the audio features corresponding to the current request to obtain second encoded information.
[0205] In this embodiment, the second encoding module encodes the audio features corresponding to the current request to obtain second encoded information. Among them, the second encoding module is a non-streaming encoding acoustic encoding module. Specifically, in the process of the server recognizing the entire sentence content at the corresponding time to obtain the final recognition result, the server first determines whether a muting instruction has been sent to the voice end or whether the voice end has synchronously sent an EOF (End of File) signal to the server. If either condition is satisfied, it indicates that subsequent non-streaming recognition processing can be performed. The server obtains all the feature contents of this request and invokes the preset non-streaming encoding service for processing to obtain the second encoded information.
[0206] Step S273: Perform vocabulary conversion on the second encoded information to obtain a pinyin result.
[0207] In this embodiment, after the server obtains the second encoded information obtained by non-streaming encoding, it parses the pinyin modeling output in the second encoded information and performs corresponding vocabulary conversion to obtain a specific pinyin result.
[0208] Step S274: Hit the pinyin cache.
[0209] In this embodiment, based on the wake-up word segmentation logic and the locally preset pinyin cache, the server checks whether the preset pinyin cache is hit.
[0210] In response to the pinyin result not hitting the preset pinyin cache, proceed to step S275.
[0211] In response to the pinyin result hitting the preset pinyin cache, proceed to step S276.
[0212] Step S275: Determine the final recognition result according to the second encoded information and the hot words.
[0213] In this embodiment, in response to the pinyin result not hitting the preset pinyin cache, the final recognition result is determined according to the second coding information and the hot words. Specifically, in response to the pinyin result not hitting the preset pinyin cache, the server invokes a non-streaming decoding service to obtain the final recognition result. The specific logic of this process is the same as the streaming decoding processing logic, that is, the result with the highest confidence in the decoding output is selected as the final recognition output.
[0214] Step S276: Determine the content in the pinyin cache as the final recognition result.
[0215] In this embodiment, in response to the pinyin result hitting the preset pinyin cache, the content in the pinyin cache is determined as the final recognition result.
[0216] In the embodiment of the present invention, by obtaining the voice data to be processed, extracting and storing the audio features from the voice data to be processed, recognizing the audio features to obtain a rising signal or a falling signal, in response to recognizing the falling signal, executing a mute process, in response to recognizing the rising signal, obtaining an intermediate recognition result according to the audio features, executing a mute process according to the intermediate recognition result, and in response to the voice end completing the mute, obtaining the final recognition result according to the audio features. Thus, real-time response can be achieved, a more accurate and comprehensive final recognition result can be provided, the robustness and adaptability in complex dialogue scenarios can be improved, and thus the user needs can be better met.
[0217] In some embodiments, the voice processing method further includes:
[0218] Step S280: Judge the validity of the final recognition result based on a predetermined rejection rule.
[0219] Wherein, the predetermined rejection rule includes at least one of a blacklist, a white list, voice parameters, a voice score, and a text score.
[0220] Specifically, Figure 10 is a flowchart of the rejection process in the embodiment of the present invention. As Figure 10 shown, judging the validity of the final recognition result based on a predetermined rejection rule includes the following steps:
[0221] Step S2801: Obtain the final recognition result.
[0222] In this embodiment, before non-streaming recognition, an asynchronous pre-call to a pre-voice rejection service that only depends on audio information is made in advance to reduce the overall response duration of the service. Among them, the pre-voice rejection service is implemented by a first rejection module. At the same time, the final recognition result output from the above steps is obtained.
[0223] Step S2802: Detect whether the blacklist is hit.
[0224] In this embodiment, the front-end voice rejection recognition service is used to determine whether the final recognition result hits the preset rejection blacklist.
[0225] In response to the final recognition result being in the blacklist, go to step S2812 to determine that the final recognition result is invalid.
[0226] In response to the final recognition result not being in the blacklist, go to step S2803.
[0227] Step S2803: Detect whether the whitelist is hit.
[0228] In this embodiment, in response to the final recognition result not being in the blacklist, detect whether the final recognition result is in the whitelist.
[0229] In response to the final recognition result being in the whitelist, go to step S2811 to determine that the final recognition result is valid.
[0230] In response to the final recognition result not being in the whitelist, go to step S2804.
[0231] Step S2804: Obtain the voice parameters of the final recognition result.
[0232] In this embodiment, in response to the final recognition result not being in the whitelist, obtain the voice parameters of the final recognition result, where the voice parameters include at least one of the number of words and the speech rate.
[0233] Step S2805: Detect whether the predetermined conditions are met.
[0234] In this embodiment, corresponding intervals of the number of words and the speech rate are preset in advance. If the number of words and the speech rate are within the intervals, it means that the predetermined conditions are met. If the number of words or the speech rate is not within the intervals, it means that the predetermined conditions are not met.
[0235] In response to the voice parameters meeting the predetermined conditions, go to step S2806.
[0236] In response to the voice parameters not meeting the predetermined conditions, go to step S2812 to determine that the final recognition result is invalid.
[0237] Step S2806: Determine the voice score corresponding to the final recognition result.
[0238] In this embodiment, in response to the voice parameters meeting the predetermined conditions, call the back-end voice rejection recognition service. Among them, the back-end voice rejection recognition service is implemented by the second rejection module. At the same time, obtain the final recognition result output by the above steps and determine the voice score corresponding to the final recognition result.
[0239] In some embodiments, the server scores the input audio features based on an acoustic model to obtain the speech score. Each segment of audio features is assigned an acoustic score, which represents the degree of its match with the model prediction. The higher the score, the more the audio features conform to the sound pattern expected by the model.
[0240] Step S2807: Is it less than the first predetermined threshold?
[0241] In this embodiment, the first predetermined threshold is set in advance to detect whether the speech score is less than the first predetermined threshold.
[0242] In response to the speech score being less than the first predetermined threshold, go to step S2812 to determine that the final recognition result is invalid.
[0243] In response to the speech score being greater than or equal to the first predetermined threshold, go to step S2808.
[0244] Step S2808: Is it an active wake-up?
[0245] In this embodiment, it is determined whether the wake-up type reported by the current speech terminal is an active wake-up by the user (i.e., the user actively says a predetermined wake-up word for interaction).
[0246] In response to it being an active wake-up, go to step S2811 to determine that the final recognition result is valid.
[0247] In response to it not being an active wake-up, go to step S2809.
[0248] Step S2809: Determine the text score corresponding to the final recognition result.
[0249] In this embodiment, in response to the speech score being greater than or equal to the first predetermined threshold, the text score corresponding to the final recognition result is determined. Specifically, the server integrates information such as audio features, the final recognition result, and the speech rejection recognition result, and calls a preset text rejection service to score the text corresponding to the final recognition result to obtain the text score. The text score reflects the rationality of the generated text in terms of grammar and semantics. The higher the score, the more the text conforms to the structure and habits of natural language.
[0250] Step S2810: Is it less than the second predetermined threshold?
[0251] In this embodiment, the second predetermined threshold is set in advance, and it is detected whether the text score is less than the second predetermined threshold.
[0252] In response to the text score being less than the second predetermined threshold, go to step S2812 to determine that the final recognition result is invalid.
[0253] In response to the text score being greater than or equal to a second predetermined threshold, proceed to step S2811 to determine that the final recognition result is valid.
[0254] Step S2811: Determine that the final recognition result is valid.
[0255] Step S2812: Determine that the final recognition result is invalid.
[0256] Thus, the quality and reliability of the final recognition result can be improved through the rejection service.
[0257] In some embodiments, the speech processing method further includes:
[0258] Step S290: Correct the final recognition result.
[0259] In this embodiment, the server corrects the final recognition result through a quick repair module. The quick repair module is a Regular Expression (RE) - based fast repair module for correcting the final recognition result. Among them, a regular expression is a powerful tool for matching strings, which allows defining patterns for searching text. These patterns can describe single characters, groups of characters, numbers, word boundaries, etc., and support complex logics such as "or" conditions, "and" conditions, and repetition counts. Therefore, regular expressions are very suitable for finding, replacing, and parsing text data. The quick repair module can automate some repetitive tasks that can be solved through pattern matching.
[0260] In some embodiments, the speech processing method further includes:
[0261] Step S310: Store the final recognition result in a cache.
[0262] In some embodiments, the speech processing method further includes:
[0263] Step S320: Before receiving the termination signal sent by the speech terminal, detect whether there is a final recognition result in the cache.
[0264] Step S320: In response to there being no final recognition result in the cache, pre - obtain the final recognition result and store it in the cache.
[0265] Step S320: In response to receiving the termination signal sent by the speech terminal, obtain the final recognition result from the cache and send it to the speech terminal.
[0266] Specifically, the real-time voice interaction scenario has high requirements for latency. One important measurement metric is the tail packet rt (the time interval between receiving the client eof signal and sending out the final result). The process includes submitting the calculation logic in advance after the microphone is muted, caching the final recognition result, and directly returning it after receiving the client eof signal, which greatly reduces the tail packet rt.
[0267] Furthermore, the real-time voice recognition task has even higher requirements for the stability, reliability, and robustness of the overall system. Especially for devices such as intelligent voice speakers, the user's conversation with the speaker strongly depends on a correct and stable ASR system. If there is a problem during the real-time voice recognition process, the entire voice interaction ability of the speaker will be unavailable. In this case, even if quick positioning and response are carried out through components such as logs, the handling and repair of online problems still require a certain amount of time, thereby causing failures and task losses.
[0268] Therefore, the embodiments of the present invention also provide a stability guarantee strategy to make the link have higher availability. There are alarms for exceptions, backups for the core, contingency plans for emergencies, and switches for modules.
[0269] Specifically, Figure 11 is a schematic diagram of the stability guarantee strategy of the embodiments of the present invention. As Figure 11 shown, the stability guarantee strategy of the embodiments of the present invention includes an emergency plan 111, a backup plan 112, a degradation plan 113, a stability plan 114, and an engineering switch 115.
[0270] In this embodiment, the emergency plan 111 includes a large-capacity pinyin cache, a degraded non-streaming encoder, a degraded streaming decoder, etc.
[0271] Among them, the large-capacity pinyin cache can avoid the situation of insufficient pinyin capacity in the cache, reduce the computational pressure on the system, and ensure the stability of the core service.
[0272] The streaming decoder and streaming encoder adopted in the embodiments of the present invention usually focus on real-time performance and accuracy, and may use complex models and algorithms to achieve a good user experience. In the case of setting ordinary streaming decoders and streaming encoders, setting a degraded non-streaming encoder and a degraded streaming decoder can give priority to the stability and real-time performance of the system under resource constraints or high load, reduce the computational overhead by simplifying the models and algorithms, and ensure that the system can continue to provide services under limited resources.
[0273] In the fallback solution 112, a four-level automatic fallback strategy is formulated, namely non-streaming decoding, non-streaming encoding, streaming decoding, and streaming encoding. When the result of the previous level is missing due to service errors / timeouts, etc., the result of the next level is automatically adopted. For example, when non-streaming decoding fails, the non-streaming encoding result is automatically used as a fallback.
[0274] The downgrade solution 113 includes streaming decoding, voice semantic rejection recognition, semantic judgment, quick repair, streaming VAD, etc. In the case of resource constraints or high load, automatic downgrading reduces the computational overhead by simplifying the model and algorithm to ensure that the system can continue to provide services under limited resources.
[0275] The stability solution 114 includes downtime self-start, alarm, and regular cleaning, which can guarantee the stability of the system.
[0276] The engineering switches 115 include rejection recognition switch, hot word switch, quick repair switch, pinyin cache switch, VAD switch, etc. In the case of resource constraints or high load, or when any module encounters an abnormal error, the corresponding function is turned off through the engineering switch to ensure that the system can continue to run.
[0277] In traditional speech recognition systems, it is usually divided into a single recognition service and an integrated engine strategy as described in Chapter 5. Among them, the single recognition service specifically provides the speech recognition (ASR) function, directly inputs the audio signal into the ASR model and outputs the corresponding text result. This service has characteristics such as high accuracy and low latency, and is easy to integrate. However, its function is single, it is difficult to meet the requirements of complex speaker interaction scenarios, and the process is relatively fixed, which is not conducive to diverse access and the development of new features. Facing real-time speaker interaction scenarios such as Tmall Genie, more accurate mute determination, more precise recognition effects, and rejection recognition capabilities more suitable for the interaction scenario are required, and the single recognition service has obvious deficiencies. The integrated engine strategy integrates all relevant functions such as ASR, mute determination, and end-of-sentence rejection recognition, provides a comprehensive speech processing solution, simplifies the system setup and process connection steps, can be quickly launched, and the modules can work better together to reduce latency issues. However, this method has deficiencies in scalability and maintainability. Due to the high coupling between functions, the update of sub-functions usually requires the entire package to be repackaged and released, which increases the complexity of compatibility and debugging. In addition, the instability of a single sub-function may affect the performance of the entire system. For large services that need to handle high-concurrency requests, this increases the cost of maintenance and iteration as well as the stability risk.
[0278] In the embodiment of the present invention, through a specially designed process scheduling system, the capabilities of each sub-module are connected in series. With an orchestration, scheduling, and highly scalable architecture design, the core features of multiple recognition services (such as mute determination, speech recognition, and rejection judgment) are integrated into a unified system. At the same time, the system stability is taken into account, forming a streaming voice request processing process and system with alarms for anomalies, backups for cores, contingency plans for emergencies, and switches for modules. This approach not only improves the overall maintainability and stability but also enhances the independence and update flexibility of each module's functions, ensuring better performance and user experience in a high-concurrency environment.
[0279] In the embodiment of the present invention, by obtaining the voice data to be processed, extracting and storing the audio features from the voice data to be processed, identifying the audio features to obtain an ascending signal or a descending signal, in response to identifying a descending signal, executing a mute process, in response to identifying an ascending signal, obtaining an intermediate recognition result according to the audio features, executing a mute process according to the intermediate recognition result, and in response to the voice end completing the mute, obtaining a final recognition result according to the audio features. Thus, real-time response can be achieved, providing a more accurate and comprehensive final recognition result, enhancing the robustness and adaptability in complex dialogue scenarios, and better meeting the user needs.
[0280] Figure 12 It is a schematic diagram of the voice processing device in the embodiment of the present invention. As Figure 12 shown, the voice processing device in the embodiment of the present invention includes a data acquisition unit 121, an extraction unit 122, a signal acquisition unit 123, a first mute unit 124, an identification unit 125, a second mute unit 126, and a result acquisition unit 127. Among them, the data acquisition unit 121 is used to acquire the voice data to be processed. The extraction unit 122 is used to extract and store the audio features from the voice data to be processed. The signal acquisition unit 123 is used to identify the audio features to obtain an ascending signal or a descending signal. The first mute unit 124 is used to execute a mute process in response to identifying a descending signal. The identification unit 125 is used to obtain an intermediate recognition result according to the audio features in response to identifying an ascending signal. The second mute unit 126 is used to execute a mute process according to the intermediate recognition result. The result acquisition unit 127 is used to obtain a final recognition result according to the audio features in response to the voice end completing the mute.
[0281] In an embodiment of the present invention, by obtaining the voice data to be processed, extracting and storing the audio features from the voice data to be processed, identifying the audio features to obtain an ascending signal or a descending signal, in response to identifying a descending signal, performing a mute process, in response to identifying an ascending signal, obtaining an intermediate recognition result according to the audio features, performing a mute process according to the intermediate recognition result, and in response to the voice end completing the mute, obtaining a final recognition result according to the audio features. Thus, real-time response can be achieved, a more accurate and comprehensive final recognition result can be provided, the robustness and adaptability in complex dialogue scenarios can be improved, and thus the user needs can be better met.
[0282] Figure 13 It is a schematic diagram of an electronic device according to an embodiment of the present invention. In this embodiment, the electronic device 13 includes a server, a terminal, etc. As Figure 13 shown, the electronic device 13: includes at least one processor 131; and, a memory 132 communicatively connected to at least one processor 131; and, a communication component 133 communicatively connected to the scanning device, and the communication component 133 receives and sends data under the control of the processor 131; wherein, the memory 132 stores instructions executable by at least one processor 131, and the instructions are executed by at least one processor 131 to implement the above-mentioned voice processing method.
[0283] Specifically, the electronic device includes: one or more processors 131 and a memory 132, Figure 13 wherein one processor 131 is taken as an example. The processor 131 and the memory 132 can be connected by a bus or other means, Figure 13 wherein taking the connection by a bus as an example. The memory 132, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The processor 131 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in the memory 132, that is, implements the above-mentioned voice processing method.
[0284] The memory 132 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store an option list, etc. In addition, the memory 132 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 132 can optionally include a memory remotely set relative to the processor 131, and these remote memories can be connected to an external device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise internal network, a local area network, a mobile communication network, and combinations thereof.
[0285] One or more modules are stored in the memory 132 and, when executed by one or more processors 131, perform the voice processing method in any of the above method embodiments.
[0286] The above product can execute the method provided in the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference may be made to the method provided in the embodiments of the present application.
[0287] In an embodiment of the present invention, by obtaining voice data to be processed, extracting and storing audio features from the voice data to be processed, identifying the audio features to obtain a rising signal or a falling signal, in response to identifying a falling signal, performing a mute process, in response to identifying a rising signal, obtaining an intermediate recognition result according to the audio features, performing a mute process according to the intermediate recognition result, and in response to the voice end completing the mute, obtaining a final recognition result according to the audio features. Thus, real-time response can be achieved, a more accurate and comprehensive final recognition result can be provided, and the robustness and adaptability in complex dialogue scenarios can be improved, so as to better meet the user's needs.
[0288] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, and the computer-readable program is used for a computer to execute some or all of the above method embodiments.
[0289] That is, those skilled in the art can understand that all or part of the steps in implementing the above method embodiments can be completed by instructing relevant hardware through a program, and the program is stored in a storage medium, including several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the method described in the embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0290] The foregoing are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A voice processing method, characterized in that, The method includes: Obtain the voice data to be processed; Extract and store audio features from the voice data to be processed; Identify the audio features to obtain an ascending signal or a descending signal; In response to identifying a descending signal, execute the mute process; In response to identifying an ascending signal, obtain an intermediate recognition result according to the audio features; Execute the mute process according to the intermediate recognition result; In response to the voice end completing the mute, obtain a final recognition result according to the audio features.
2. The method according to claim 1, wherein The extracting and storing audio features from the voice data to be processed includes: Decode the voice data to be processed; In response to decoding valid content, extract features from the decoded content; In response to extracting audio features, store the audio features in memory.
3. The method according to claim 2, wherein The method further includes: In response to receiving a termination signal sent by the voice end and no valid content being decoded, terminate the recognition process of the voice data to be processed; or In response to receiving a termination signal sent by the voice end and no audio features being extracted, terminate the recognition process of the voice data to be processed.
4. The method according to claim 1, characterized in that, The executing the mute process in response to identifying a descending signal includes: Send a mute instruction to the voice end, and release and record resources.
5. The method according to claim 1, characterized in that, The obtaining an intermediate recognition result according to the audio features in response to identifying an ascending signal includes: Determine the request information corresponding to the voice data to be processed, where the request information includes at least one of a device identifier, a product identifier, and a request content; Obtain the hot words of the voice end according to the request information; Segment the audio features according to the ascending signal to obtain audio segments to be processed; Encode the audio segments to be processed to obtain first encoding information; Obtain and record the tail point information corresponding to the first encoding information; Obtain the intermediate recognition result according to the first encoding information and the hot words.
6. The method according to claim 5, wherein The obtaining an intermediate recognition result according to the audio features in response to identifying an ascending signal further includes: Compare the intermediate recognition result output in this round with the intermediate recognition result of the previous round; In response to the intermediate recognition result output in this round being inconsistent with the intermediate recognition result of the previous round, synchronize the intermediate recognition result output in this round to the voice end; Update the intermediate recognition result according to the intermediate recognition result output in this round.
7. The method according to claim 6, characterized in that, The method further includes: In response to the intermediate recognition result output in this round being consistent with the intermediate recognition result of the previous round, or updating the current intermediate recognition result, detect whether a mute instruction has been issued; In response to no mute instruction having been issued, recognize the tail point information and the current intermediate recognition result; In response to identifying a descending signal, send a mute instruction to the voice end, and release and record resources.
8. The method according to claim 1, wherein The obtaining a final recognition result according to the audio features in response to the voice end completing the mute includes: Obtain the audio features corresponding to this request; Encode the audio features corresponding to this request to obtain second encoding information; Convert the second encoding information into a word list to obtain a pinyin result; Compare the pinyin result with a preset pinyin cache; In response to the pinyin result hitting the preset pinyin cache, determine the content in the pinyin cache as the final recognition result; In response to the pinyin result not hitting the preset pinyin cache, determine the final recognition result according to the second coding information and hot words.
9. The method according to claim 1, wherein The method further includes: Judging the validity of the final recognition result based on a predetermined rejection rule; Wherein, the predetermined rejection rule includes at least one of a blacklist, a whitelist, voice parameters, voice scores, and text scores.
10. The method according to claim 9, wherein The judging the validity of the final recognition result based on a predetermined rejection rule includes: Detecting whether the final recognition result is in the blacklist; In response to the final recognition result being in the blacklist, determining that the final recognition result is invalid; In response to the final recognition result not being in the blacklist, detecting whether the final recognition result is in the whitelist; In response to the final recognition result being in the whitelist, determining that the final recognition result is valid; In response to the final recognition result not being in the whitelist, obtaining the voice parameters of the final recognition result, where the voice parameters include at least one of the number of words and the speech rate; In response to the voice parameters meeting a predetermined condition, determining that the final recognition result is invalid; In response to the voice parameters not meeting the predetermined condition, determining the voice score corresponding to the final recognition result; In response to the voice score being less than a first predetermined threshold, determining that the final recognition result is invalid; In response to the voice score being greater than or equal to the first predetermined threshold, detecting whether it is an active wake-up; In response to it being an active wake-up, determining that the final recognition result is valid; In response to it not being an active wake-up, determining the text score corresponding to the final recognition result; In response to the text score being less than a second predetermined threshold, determining that the final recognition result is invalid; In response to the text score being greater than or equal to the second predetermined threshold, determining that the final recognition result is valid.
11. The method according to claim 1, wherein The method further includes: Correcting the final recognition result.
12. The method according to claim 1, wherein The method further includes: Storing the final recognition result in a cache.
13. The method according to claim 12, characterized in that, The method further includes: Before receiving the termination signal sent by the voice side, detecting whether there is a final recognition result in the cache; In response to there being no final recognition result in the cache, obtaining the final recognition result in advance and storing it in the cache; In response to receiving the termination signal sent by the voice side, obtaining the final recognition result from the cache and sending it to the voice side.
14. A voice processing system, characterized in that, The system includes: A protocol layer for obtaining voice data to be processed; A preprocessing layer for extracting and storing audio features from the voice data to be processed, recognizing the audio features to obtain a rising signal or a falling signal, and in response to recognizing a falling signal, performing a mute process; A model inference layer for, in response to recognizing a rising signal, obtaining an intermediate recognition result according to the audio features, performing a mute process according to the intermediate recognition result, and in response to the voice side completing the mute, obtaining a final recognition result according to the audio features; A rejection layer for judging the validity of the final recognition result based on a predetermined rejection rule; A quick repair layer for correcting the final recognition result.