Audio processing method and device
By segmenting the audio stream and performing speech recognition, combined with silence detection and semantic analysis, the time limitation and resource consumption problems of long speech recognition in existing technologies are solved, and real-time and accurate long speech recognition is achieved.
Patent Information
- Application Number
- CN202511254826.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing speech recognition systems have problems in long speech recognition services, such as short audio recognition time limit, high inference resource consumption, and easy word omission in boundary scenarios, which cannot meet the needs of continuous recognition.
By traversing the sentence segmentation list, the audio stream is intercepted based on the silence detection results, the endpoint detection results of the sentence segmentation are determined, and speech recognition is performed based on the start and end frames of the sentence segmentation. The truncation position is determined by combining semantic and contextual information to achieve continuous and accurate long speech recognition.
It achieves real-time, continuous long speech recognition, breaks through the time limitations of traditional audio recognition, improves recognition accuracy, and reduces hardware resource consumption and recognition delay.
Smart Images

Figure CN120748375A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and more particularly, to an audio processing method and device. Background Art
[0002] In some scenarios, the client needs to have the ability to keep the microphone open and recognize speech in real time, that is, to provide long speech recognition services.
[0003] However, the long speech recognition service of the existing speech recognition system has problems such as short audio recognition time limit, high inference resource consumption, and easy word omission in boundary scenarios, which cannot meet the needs of continuous recognition well. Summary of the Invention
[0004] In view of this, embodiments of the present invention aim to provide an audio processing method and apparatus to achieve continuous and accurate long speech recognition.
[0005] In a first aspect, an embodiment of the present invention provides an audio processing method, the method comprising: Traversing a sentence list to determine a current sentence, the sentence list including at least one sentence, the sentence being determined by intercepting the audio stream based on a silence detection result; Determine an endpoint detection result of the current segmentation; In response to the presence of a sentence starting point in the current sentence segment, determining the starting frame of the current sentence segment as the starting frame of the latest sentence to be recognized; In response to the existence of a sentence end in the current sentence, determining the end frame of the current sentence as the end frame of the latest sentence to be recognized; Determining the latest sentence to be recognized based on the start frame and the end frame of the latest sentence to be recognized; Perform speech recognition on the latest sentence to be recognized to determine a corresponding speech recognition result.
[0006] Furthermore, the method further comprises: Receive audio stream; Extracting features from the audio stream to determine a corresponding audio feature sequence; Performing silence detection on the audio feature sequence; In response to detecting that the continuous silence duration reaches a configured duration, determining a truncation position; The audio feature sequence is truncated at the truncation position to determine the corresponding sentence segment.
[0007] Furthermore, in response to detecting that the continuous silence duration reaches a configured duration, determining the truncation position includes: In response to detecting that the continuous silence duration reaches a configured duration, determining scene information of a non-silent audio segment closest to the current audio frame, the scene information including semantic information and / or context information; In response to the scene information being complete, a truncation position is determined.
[0008] Furthermore, the truncation position is any audio frame within the silence period corresponding to the continuous silence duration.
[0009] Furthermore, the extracting features from the audio stream to determine the corresponding audio feature sequence includes: Feature extraction is performed on the audio stream to determine a corresponding Mel spectrum feature sequence.
[0010] Furthermore, the method further comprises: Performing a rejection determination on a speech recognition result corresponding to the audio stream, and outputting speech interaction content based on the rejection determination result.
[0011] In a second aspect, an embodiment of the present invention is directed to providing an audio processing device, the device comprising: A traversal unit, configured to traverse a sentence list and determine a current sentence, wherein the sentence list includes at least one sentence, and the sentence is determined by intercepting the audio stream based on the silence detection result; a detection unit, configured to determine an endpoint detection result of the current sentence; in response to a sentence start point in the current sentence, determine the start frame of the current sentence as the start frame of the latest sentence to be recognized; in response to a sentence end point in the current sentence, determine the end frame of the current sentence as the end frame of the latest sentence to be recognized; and determine the latest sentence to be recognized based on the start frame and end frame of the latest sentence to be recognized; The recognition unit is used to perform speech recognition on the latest sentence to be recognized and determine a corresponding speech recognition result.
[0012] In a third aspect, an embodiment of the present invention aims to provide a computer program product, which includes a computer program / instructions, and when the computer program / instructions are executed by a processor, implements the method as described in any one of the above items.
[0013] In a fourth aspect, an embodiment of the present invention aims to provide an electronic device comprising a memory and a processor, wherein the memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of the above items.
[0014] In a fifth aspect, an embodiment of the present invention aims to provide a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method described in any one of the above items.
[0015] The technical solution of the embodiment of the present invention intercepts the audio stream based on the audio detection result to determine the sentence list corresponding to the audio stream, traverses the sentences in the sentence list to determine the sentences to be recognized and performs voice recognition on the sentences to be recognized to determine the voice recognition results. It can achieve real-time and continuous long speech recognition, breaking through the time limit of traditional audio recognition. Moreover, since the sentences to be recognized determined by the sentence start point and sentence end point in the endpoint detection results of the sentence segmentation include complete content, the audio processing method in this embodiment can also improve the accuracy of long speech recognition. At the same time, the long speech recognition method in this embodiment does not need to rely on frequent microphone opening and closing operations and end-side backtracking logic, which reduces hardware resource consumption and recognition delay, and further improves the processing capability of long speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The above and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which: Figure 1 It is a schematic diagram of a long speech recognition mode in the prior art; Figure 2 is a flowchart of an audio processing method according to an embodiment of the present invention; Figure 3 This is a flowchart of determining sentence segmentation according to an embodiment of the present invention; Figure 4 is a schematic diagram of determining a truncation position according to an embodiment of the present invention; Figure 5 is a flow chart of determining a truncation position according to an embodiment of the present invention; Figure 6 is a schematic diagram of determining a truncation position according to an embodiment of the present invention; Figure 7 is an overall flow chart of audio processing according to an embodiment of the present invention; Figure 8 is a schematic diagram of audio recognition according to an embodiment of the present invention; Figure 9 is a schematic diagram of an audio processing device according to an embodiment of the present invention; Figure 10 is a schematic diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The present application is described below based on the following embodiments, but the present application is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these details. To avoid obscuring the essence of the present application, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0018] Furthermore, persons of ordinary skill in the art will appreciate that the figures provided herein are for illustration purposes only and are not necessarily drawn to scale.
[0019] Unless the context clearly requires otherwise, words like “include”, “comprising” and the like throughout this application should be interpreted as including rather than exclusive or exhaustive; that is, as meaning “including but not limited to”.
[0020] In the description of this application, it should be understood that the terms "first", "second", etc. are used for descriptive purposes only and should not be understood to indicate or imply relative importance. In addition, in the description of this application, unless otherwise specified, "plurality" means two or more.
[0021] Any solutions described in this specification and in the examples that involve the processing of personal information will be processed only with a legitimate basis (such as with the consent of the personal information subject or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect their use of these functions.
[0022] Existing technologies, constrained by audio processing models and numerous upstream and downstream system services, limit the duration of a single voice interaction, making it difficult to provide long-duration speech recognition services. At the same time, some strategies based on the "open mic, close mic, and then open mic again" cycle, combined with client-side audio backtracking mechanisms, simulate continuous audio recognition to provide a near-continuous speech recognition experience.
[0023] Figure 1 This is a schematic diagram of the long speech recognition mode in the prior art. In the prior art, there is a solution that guides the client to continuously open and close the microphone, uses multiple rounds of recognition operations to achieve continuous real-time audio recognition, and combines a backtracking mechanism to prevent audio signal loss. However, since this solution relies on the client to implement caching and backtracking logic, the client cache capacity is limited, and long-term operation will result in large system resource consumption, thereby affecting the end-side performance and even causing data loss. At the same time, the actual recognition of the backtracked audio needs to wait until the next time the microphone is opened to complete. Such backtracking will cause a significant time delay in the recognition result; and there is no intrinsic connection between the multiple recognition processes. If the microphone is closed at an inappropriate time or the backtracking is inaccurate, problems such as extra words or missing words will occur, thereby affecting the accuracy of the speech recognition results.
[0024] In view of this, an embodiment of the present invention provides an audio processing method for achieving continuous and accurate long speech recognition. While this embodiment uses the long speech recognition process in a scenario where the client is smart glasses as an example to illustrate the long speech recognition method, it should be understood that the long speech recognition method in this embodiment is also applicable to other application scenarios requiring long speech recognition, such as smart speakers and wearable smart devices, and the application scenario of this method is not limited here.
[0025] When smart glasses support new application features such as translation, flash memos, and teleprompters, they must be capable of continuous microphone interaction and real-time speech recognition. Furthermore, to accommodate the lightweight form factor of smart glasses and provide a comfortable wearing experience for the wearer, and given the limited cache and processing capacity of smart glasses, after receiving the audio stream, smart glasses typically transmit it to a server. The server can be a processor, a server (including a single server or a cluster of multiple servers), or a smart terminal connected to the smart glasses. The server can be deployed locally or remotely from the smart glasses. The server then performs speech recognition on the audio stream and returns the results to the smart glasses. This enables long-duration speech recognition in real-time and continuous interaction scenarios when the smart glasses are in use. This provides long-duration speech recognition support for downstream business scenarios (such as industry and education). It also lays the foundation for future upgrades to full-duplex interaction scenarios (i.e., two-way, real-time, and natural voice or information exchange between the user and smart glasses, similar to the synchronous communication mode of face-to-face conversation).
[0026] Figure 2 FIG. 1 is a flow chart of an audio processing method according to an embodiment of the present invention. Figure 2 As shown, the audio processing method in this embodiment includes: In step S210, the segmentation list is traversed to determine the current segmentation, wherein the segmentation list includes at least one segmentation, and each segmentation is determined by intercepting the audio stream based on the silence detection result.
[0027] In this embodiment, after receiving the audio stream to be recognized, a silence detection is performed on the audio stream to determine the silence detection result, and the audio stream is intercepted based on the silence detection result. Each time a segment is intercepted, a segment is obtained. Assuming that the most recently intercepted segment is defined as the current segment, since the segment before the current segment that has been intercepted may not be fully recognized, in this embodiment, each time a segment is intercepted, the segment is added to the segment list. The earlier it is added to the segment list, the higher the processing order, so that the segments in the segment list can be processed later, ensuring that all voice content in the audio stream is processed in an orderly manner, and improving the accuracy of audio recognition.
[0028] In other optional implementations, in this embodiment, when adding each segment to the segment list, the traversal order of the segment in all the segments in the segment list can be represented by a unique identifier, and each segment can be processed in parallel or serially after being traversed.
[0029] Further, Figure 3 This is a flow chart of determining sentence segmentation according to an embodiment of the present invention. Figure 3 As shown, in this embodiment, the sentence segmentation is determined by the following method, which specifically includes the following steps.
[0030] In step S310, an audio stream is received.
[0031] In step S320, feature extraction is performed on the audio stream to determine a corresponding audio feature sequence.
[0032] In this embodiment, in order to realize the recognition of audio content in the audio stream, it is necessary to perform feature extraction on the audio stream to obtain a corresponding audio feature sequence, so that subsequent processing of the audio feature sequence can obtain a speech recognition result.
[0033] Optionally, the audio feature sequence in this embodiment can adopt time-domain features, frequency-domain features, Mel-spectrogram or other audio feature representation forms. Furthermore, since Mel-spectrogram is more in line with the auditory perception characteristics of the human ear than other audio feature representations, it can also filter out some noise and redundant information, while compressing redundant high-frequency details, thereby improving subsequent recognition accuracy and efficiency, this embodiment adopts Mel-spectrogram for audio feature representation, and extracts features from the audio stream through Mel Filter Banks to determine the corresponding Mel-spectrogram feature sequence (also known as Fbank feature sequence). This not only realizes the audio feature extraction of the audio stream, but also helps to improve the recognition efficiency and accuracy of subsequent audio recognition.
[0034] In step S330, silence detection is performed on the audio feature sequence.
[0035] Optionally, in this embodiment, the silence period in the audio feature sequence can be detected based on a silence detection method such as energy detection or spectrum detection. The silence period is a continuous audio frame in which a preset sound type (such as human voice) does not exist. When detecting the silence period based on energy, the energy of the audio signal is calculated to determine whether the audio frame is in a silent state. When detecting the silence period based on the spectrum, the audio frame in a silent state is identified by analyzing the spectral characteristics of the audio signal. Therefore, by performing silence detection on the audio stream and determining the segmentation based on the silence detection result, the silent audio frames in the audio stream can be eliminated, reducing the data transmission and processing volume of subsequent audio recognition, and improving the efficiency of audio recognition.
[0036] In step S340 , in response to detecting that the continuous silence duration reaches a configured duration, a truncation position is determined.
[0037] In this embodiment, the configuration duration can be set according to the actual application scenario, for example, it can be set to 300ms, 500ms, etc. Detecting that the continuous silence duration reaches the configured duration is a triggering operation for intercepting the audio feature sequence, and each continuous silence duration has a corresponding silence period.
[0038] Optionally, the duration of continuous silence in this embodiment can be determined by single continuous timing or multiple continuous timing. If single continuous timing is used, when a silent audio frame is detected, a silent period is determined to have occurred and timing begins, until the end of silence is detected or other silence detection termination operations (such as audio interception) are stopped, and the continuous silence duration of the corresponding silent period is the duration from the start timing to the end timing. For example, silence is detected at audio frame 1700ms and timing begins, and the end of silence is detected at audio frame 2100ms (i.e., the timing reaches 400ms), then the continuous silence duration of the corresponding silent period (i.e., audio frame 1700ms-2100ms) is 400ms.
[0039] If multiple continuous timings are used, each timing has a corresponding predetermined duration. When the predetermined duration is reached, the timing ends, the number of times the timing is recorded, and the next timing begins. Timing continues until the end of silence is detected or another silence detection termination operation (such as audio interception) is performed, at which point the timing stops. The duration of continuous silence for the corresponding silence period is determined based on the predetermined duration, the number of times the timing is performed, and the duration of the last timing. For example, if the predetermined duration of each timing is 150ms, and silence is detected at audio frame 1700ms, the first timing begins. The second timing begins at audio frame 1850ms (i.e., the first timing reaches 150ms), the third timing begins at audio frame 2000ms (i.e., the second timing reaches 150ms), and the end of silence is detected at audio frame 2100ms (i.e., the third timing reaches 100ms). In this case, the duration of continuous silence for the corresponding silence period (i.e., audio frame 1700ms - 2100ms) is (150 × 2 + 100) ms, or 400ms. Therefore, in this embodiment, by providing different continuous silence duration detection methods, it is convenient for specific application scenarios to select an appropriate method to determine the continuous silence duration of the silence period, thereby improving the convenience of detecting the continuous silence duration.
[0040] Furthermore, in this embodiment, a single continuous timing method is used to determine the continuous silence duration. When determining the truncation position, in an optional implementation method, in this embodiment, the truncation position can be immediately determined when it is detected that the continuous silence duration reaches the configured duration. The truncation position can be any audio frame within the silence period corresponding to the continuous silence duration. For example, it can be an audio frame at the beginning of the silence period, an audio frame when the silence duration in the silence period reaches the configured duration, or an audio frame between the audio frame at the beginning of the silence period and the audio frame when the silence duration reaches the configured duration.
[0041] Figure 4 FIG. 1 is a schematic diagram of determining the truncation position according to an embodiment of the present invention. Figure 4 As shown in , assuming a configured duration of 500ms and the audio content of the audio stream is "Helper, I want to listen to music, singer A's new album" (the audio content here is the speech recognition result to be recognized), the start frame of the audio feature sequence corresponding to the audio stream is recorded as audio frame 0ms. Silence is detected at audio frame 1700ms. The corresponding silence period is recorded as silence period ① and timing begins. When the timing duration (i.e., the continuous silence duration) reaches 400ms (i.e., audio frame 2100ms), the silence ends. At this time, the continuous silence duration of silence period ① is stopped and the continuous silence duration of silence period ① is determined to be 400ms. Since the continuous silence duration of 400ms in silence period ① does not reach the configured duration of 500ms, interception is not triggered and silence detection continues on the audio stream.
[0042] At 4600ms of the audio frame, silence is detected again. The corresponding silence period is recorded as silence period ② and timing begins. When the timing reaches the configured duration of 500ms (that is, audio frame 5100ms), since the audio stream is still silent, interception is triggered and audio frame 5100ms is determined as the truncation position.
[0043] Later, silence is detected again at audio frame 8700ms. The corresponding silence period is recorded as silence period ③ and timing begins. When the timing duration (i.e., the continuous silence duration) reaches the configured duration of 500ms (i.e., audio frame 9200ms), since the audio stream is still silent, interception is triggered and audio frame 9200ms is determined as the truncation position.
[0044] In another optional implementation, in this embodiment, after detecting that the continuous silence duration reaches the configured duration, the audio interception operation can be confirmed through audio-related scene information, such as the integrity of the audio semantics, the integrity of the audio context information, semantic emotions, etc., and the truncation position is determined when it is confirmed that the current audio information meets the interception conditions. If the interception conditions are not met, the interception time is extended until the interception conditions are met. Among them, the scene information used to determine whether the interception conditions are met can be selected according to the actual application scenario. For example, in specific scenarios such as real-time translation, meetings, and speeches, the semantic and context information will be combined to decide whether the interception conditions are met, so that the audio feature sequence is intercepted to determine the corresponding sentence segmentation when the interception conditions are met.
[0045] Figure 5 FIG. 1 is a flow chart of determining the truncation position according to an embodiment of the present invention. Figure 5 As shown, in this embodiment, the truncation position is determined by the following method.
[0046] In step S510 , in response to detecting that the continuous silence duration reaches a configured duration, scene information of a non-silent audio segment closest to the current audio frame is determined, where the scene information includes semantic information and / or context information.
[0047] In this embodiment, when the duration of continuous silence is detected to have reached a configured duration, the audio frame at which the duration of continuous silence reaches the configured duration is used as the current audio frame. The nearest non-silent audio segment immediately preceding the current audio frame and the scene information for that non-silent audio segment are then retrieved. The semantic and contextual information in the scene information can be determined by combining methods such as speech recognition technology and natural language processing analysis.
[0048] In step S520, it is determined whether the scene information is complete.
[0049] In this embodiment, after determining the scene information of the non-silent audio segment closest to the current audio frame, the integrity of the semantic information and / or contextual information of the non-silent audio segment is determined to determine whether the corresponding scene information is complete. For example, when the scene information includes only semantic information or contextual information, if either the semantic information is complete or the contextual information is complete, the corresponding scene information is determined to be complete. When the scene information includes both semantic information and contextual information, if both the semantic information and the contextual information are complete, the corresponding scene information is determined to be complete.
[0050] Specifically, the integrity of the semantic information, contextual information, and corresponding scene information in the scene information in this embodiment can be determined by combining methods such as speech recognition technology and natural language processing analysis. Furthermore, in this embodiment, if the scene information is determined to be complete, step S530 is continued; if the scene information is incomplete, step S540 is continued.
[0051] In step S530 , in response to the scene information being complete, a truncation position is determined.
[0052] In this embodiment, when it is determined that the scene information is complete, it indicates that the current clipping condition is met. At this time, the clipping position can be determined to clip the audio feature sequence, and the clipping position can be any audio frame within the silent period corresponding to the continuous silent duration.
[0053] In step S540 , in response to the scene information being incomplete, silence detection is continued.
[0054] In this embodiment, when it is determined that the scene information is incomplete, it indicates that the current interception conditions are not met, the truncation position is not determined within the silent period corresponding to the current continuous silence duration, and the audio features of the current non-silent audio segment are retained, and silence detection is continued. When it is detected again that the continuous duration reaches the configured duration, the above processing steps are repeated until the current audio stream ends.
[0055] Figure 6 FIG. 1 is a schematic diagram of determining the truncation position according to an embodiment of the present invention. Figure 6As shown in the figure, assume the configured duration is 500ms, semantic information is used to determine the truncation position, the audio content of the audio stream is "Helper, I want to listen to music, singer A's new album" (the audio content here is the speech recognition result to be recognized), and the start frame of the audio feature sequence corresponding to the audio stream is recorded as audio frame 0ms. Silence is detected at audio frame 1700ms. The corresponding silence period is recorded as silence period ① and timing begins. When the timing duration (i.e., the continuous silence duration) reaches 200ms (i.e., audio frame 1900ms), the silence ends. At this time, the continuous silence duration of silence period ① is stopped and the continuous silence duration of silence period ① is set to 200ms. Since the continuous silence duration of 200ms in silence period ① does not reach the configured duration of 500ms, truncation is not triggered and silence detection continues on the audio stream.
[0056] At 3300ms of the audio frame, silence is detected again. At this time, the corresponding silent period is recorded as silent period ② and the timing begins. When the timing reaches the configured duration of 500ms (i.e., 3800ms of the audio frame), since the audio stream is still in a silent state, it will be triggered to determine whether the interception conditions are met. Since the current non-silent audio segment (i.e., the audio feature sequence portion corresponding to the "Little Assistant, I want to listen" audio content) is detected to be semantically incomplete, the truncation position is not determined within the silent period ②, that is, the interception operation is not triggered. The interception time needs to be extended to intercept after the audio feature sequence portion corresponding to the semantically complete audio content, further improving the accuracy of audio recognition.
[0057] Afterwards, at 6700ms of the audio frame, silence is detected again. At this time, the corresponding silent period is recorded as silent period ③ and the timing begins. When the timing reaches the configured duration of 500ms (i.e., 7200ms of the audio frame), since the audio stream is still in a silent state, it will be triggered to determine whether the interception condition is met. Since the current non-silent audio segment (i.e., the audio feature sequence portion corresponding to the audio content "Assistant, I want to listen to singer A's new album") is detected to be semantically complete, the truncation position is determined within the silent period ③, which triggers the interception operation. The truncation position during interception can be at 7200ms of the audio frame. This ensures that interception is performed after the audio feature sequence portion corresponding to the semantically complete audio content, further improving the accuracy of audio recognition.
[0058] In step S350, the audio feature sequence is truncated at the truncation position to determine the corresponding sentence segment.
[0059] In this embodiment, after the truncation position is determined using the above method, the audio feature sequence is truncated at the truncation position to determine the corresponding sentence segment. Specifically, for the current truncation position, after the audio feature sequence is truncated at the current truncation position, the sentence segment corresponding to the current truncation position is the portion of the audio feature sequence formed by the audio features between the previous truncation position and the current truncation position.
[0060] For example, in Figure 4 In the audio feature sequence, after the first truncation position (i.e., audio frame 5100ms) is cut off, the first sentence can be obtained. The audio feature sequence corresponding to the first sentence is the audio feature sequence portion corresponding to the audio content "Assistant, I want to listen to music"; after the audio feature sequence is cut off at the second truncation position (i.e., audio frame 9200ms), the second sentence can be obtained. The audio feature sequence corresponding to the second sentence is the audio feature sequence portion corresponding to the audio content "Singer A's new album". For another example, in Figure 6 In the example, after the audio feature sequence is intercepted at the truncation position of the audio frame 7200ms, the first sentence can be obtained. The audio feature sequence corresponding to the first sentence is the audio feature sequence part corresponding to the audio content "Assistant, I want to listen to the new album of singer A".
[0061] Furthermore, after the audio feature sequence corresponding to the audio stream is intercepted by the above method to obtain the sentence segmentation list corresponding to the audio stream, this embodiment will traverse the sentence segmentation list to recognize the voice content in the audio stream based on the sentences in the sentence segmentation list.
[0062] In step S220, the endpoint detection result of the current segment is determined.
[0063] In this embodiment, after traversing to the current segment, the endpoint detection result of the current segment is determined. The endpoint detection result of the current segment can be predetermined, such as when determining the segment; or it can be determined when traversing the segment.
[0064] Optionally, in this embodiment, after determining the segmentation, the endpoint detection result of the segmentation is determined by endpoint detection, and the endpoint detection result of the segmentation is stored. During audio processing, when traversing to the segmentation, the endpoint detection result of the corresponding segmentation is determined by obtaining the stored endpoint detection result of the segmentation. Among them, endpoint detection is also called voice activity detection (VAD), which is used to detect the speech segment (i.e., the part containing actual speech information) and non-speech segment (i.e., the silence or background noise part) in the audio signal, so as to determine the sentence starting point and sentence ending point corresponding to the speech segment. The endpoint detection result is used to characterize whether there is a sentence starting point and sentence ending point in the segmentation. The sentence starting point refers to the time point when a meaningful speech or sentence begins, and the sentence ending point is the time point when the speech or sentence ends. By performing endpoint detection on the segmentation, it is possible to identify which parts of the audio feature sequence corresponding to the segmentation contain actual speech information and which parts are silence or background noise.
[0065] Optionally, in this embodiment, a combination of one or more of energy-based methods, zero-crossing rate, frequency domain features, and deep learning methods can be used to implement endpoint detection. For example, when performing endpoint detection on sentence segments based on an energy-based method, the starting point and ending point of a sentence in a speech segment can be distinguished by calculating the energy level of the audio signal. When the energy of the audio signal exceeds a preset energy threshold, it is considered that the speech begins (starting point), and when it is lower than the energy threshold, it is considered that the speech ends (end point). For another example, when performing endpoint detection on sentence segments based on frequency domain features, it can be determined whether it is a speech segment based on the change in the Mel spectrum features. Specifically, a machine learning model can be used to train to identify speech patterns under different features, thereby determining the starting and ending points of the speech.
[0066] In step S230 , in response to the presence of a sentence start point in the current sentence segment, the start frame of the current sentence segment is determined as the start frame of the latest sentence to be recognized.
[0067] In this embodiment, if the current segment is determined to have a sentence start point, this indicates that the speech content corresponding to the current segment is the beginning of a new sentence. In this case, the starting frame of the current segment is determined to be the starting frame of the latest sentence to be recognized. However, if the current segment is determined to have no sentence start point, this indicates that the speech content corresponding to the current segment is a continuation of the audio content of the previous segment. The starting frame of the latest sentence to be recognized (i.e., the most recent sentence to be recognized) is the audio frame determined based on the segment immediately preceding the current segment. It should also be understood that the frames mentioned in this embodiment (including various frames such as the starting frame and the ending frame) are all audio frames used in audio processing.
[0068] In step S240 , in response to the existence of a sentence end point in the current sentence segment, the end frame of the current sentence segment is determined as the end frame of the latest sentence to be recognized.
[0069] In this embodiment, if the current segment is determined to have a sentence end, it indicates that the speech content corresponding to the current segment is the end of a new sentence to be recognized. In this case, the end frame of the current segment is determined as the end frame of the new sentence to be recognized. However, if the current segment is determined to have no sentence end, it indicates that the audio content of the new sentence to be recognized is still continuing. It is necessary to determine the end frame of the new sentence to be recognized by combining the sentence after the current segment.
[0070] In step S250 , the latest sentence to be recognized is determined based on the start frame and the end frame of the latest sentence to be recognized.
[0071] In this embodiment, after determining the starting frame and ending frame of the latest sentence to be recognized, the audio feature sequence portion between the starting frame and the ending frame is read according to the starting frame and the ending frame of the latest sentence to be recognized, and the audio feature sequence portion is determined as the audio feature sequence portion corresponding to the latest sentence to be recognized.
[0072] In step S260, speech recognition is performed on the latest sentence to be recognized to determine a corresponding speech recognition result.
[0073] In this embodiment, after determining the most recent sentence to be recognized, the portion of the audio feature sequence corresponding to the sentence to be recognized is sent to the speech recognition module for speech recognition to determine the speech recognition result for the sentence to be recognized. Furthermore, because the portion of the audio feature sequence corresponding to the most recent sentence to be recognized covers the entire content from the start to the end of the sentence, speech recognition of the most recent sentence to be recognized preserves the complete speech content of the sentence to be recognized, thereby improving the accuracy of the speech recognition result for the sentence to be recognized.
[0074] At the same time, in order to improve the overall audio recognition efficiency of the audio stream, in this embodiment, after the latest sentence to be recognized is sent to the speech recognition module for speech recognition, the latest sentence to be recognized will be determined as the previous sentence to be recognized, and the sentence segmentation list will continue to be traversed based on the same method as mentioned above to determine the corresponding latest sentence to be recognized until all the sentence segments corresponding to the audio stream are traversed and all the sentences to be recognized in the audio stream are determined to be completed.
[0075] Optionally, considering that after voice recognition, it is usually necessary to display the voice recognition results and make corresponding replies based on the voice recognition results, in order to improve the reliability of subsequent replies and further enhance the reliability of client functions and user experience, in this embodiment, after obtaining the voice recognition results of all sentences to be recognized corresponding to the audio stream based on the aforementioned method, that is, the voice recognition results corresponding to the audio stream, the voice recognition results corresponding to the audio stream will also be rejected, so as to output voice interaction content based on the rejection judgment results.
[0076] Specifically, the rejection decision of speech recognition results is used to assess the credibility of speech recognition results. Rejection decisions can be made based on available methods such as statistics, rules, machine learning, and deep learning. For example, confidence can be estimated by calculating the confidence probability distribution of the speech recognition results output by the speech recognition model; using a predefined set of rules to determine the validity of the speech recognition results; and using a trained classifier to predict whether a given input should be accepted.
[0077] Furthermore, when the credibility of the speech recognition result is high (e.g., above a preset credibility threshold), it indicates that the speech recognition result is acceptable, and the speech recognition result and corresponding response content based on the speech recognition result can be output. On the other hand, when the credibility of the speech recognition result is low (e.g., below a preset credibility threshold), it indicates that the speech recognition result is unacceptable, and neither the speech recognition result nor the response content related to the speech recognition result can be output. Therefore, by performing a rejection judgment on the speech recognition result corresponding to the audio stream in this embodiment, the effectiveness of the speech recognition result display and the reliability of the voice interaction content corresponding to the speech recognition result can be improved, thereby enhancing the reliability of client functions and the user experience.
[0078] The technical solution of this embodiment intercepts the audio stream based on the audio detection results to determine the sentence list corresponding to the audio stream, traverses the sentences in the sentence list to determine the sentences to be recognized, and performs voice recognition on the sentences to be recognized to determine the voice recognition results. It can achieve real-time and continuous long speech recognition, breaking through the time limit of traditional audio recognition. Moreover, since the sentences to be recognized determined by the sentence start point and sentence end point in the endpoint detection results of the sentence segmentation include complete content, the audio processing method in this embodiment can also improve the accuracy of long speech recognition. At the same time, the long speech recognition method in this embodiment does not need to rely on frequent microphone opening and closing operations and end-side backtracking logic, reducing hardware resource consumption and recognition delay, and further improving the processing capability of long speech recognition.
[0079] To facilitate a clearer understanding of the audio processing method in this embodiment, the overall process of audio processing is described below in conjunction with the following explanations in this embodiment.
[0080] Figure 7 This is the overall flow chart of the audio processing of the embodiment of the present invention. Figure 7 The audio processing process in this embodiment includes the following processing operations.
[0081] In step S710, an audio stream is received.
[0082] In step S720, feature extraction is performed on the audio stream to determine the corresponding Fbank feature sequence. The Fbank feature is the Filter Bank feature, which is obtained by processing the speech signal corresponding to the audio stream through a set of triangular filters with Mel-scale distribution.
[0083] In step S730, the Fbank feature sequence is intercepted to determine a sentence segmentation list.
[0084] In this embodiment, the Fbank feature sequence is truncated according to the order in which the audio frames appear in the feature sequence. Each truncation generates a sentence segment, which is then added to the segmentation list. When truncating the Fbank feature sequence, this embodiment can use a pre-set local speech segmentation module to implement a request-response-based sentence segmentation protocol according to the aforementioned method to determine the truncation positions and corresponding sentences.
[0085] Moreover, for the determined segmentation, in this embodiment, it can be represented by a one-dimensional list, which includes the following contents: the start position (begin_feature_index), which indicates that the speech segmentation has recognized the beginning of a new sentence (that is, there is a sentence start point in the segmentation), and the start position corresponding to the start position (that is, the audio frame corresponding to the sentence start point in the segmentation) will also be returned; the stop position (end_feature_index), which indicates that the speech segmentation has recognized the end of the sentence (that is, there is a sentence end point in the segmentation), and the end position corresponding to the stop position (that is, the audio frame corresponding to the sentence end point in the segmentation) will also be returned. In addition, if the start and end are recognized at the same time in the segmentation, the start position and the end position are returned; if the start is not recognized in the segmentation and only the end is recognized, the start position is returned corresponding to -1 (or not returned); if only the start is recognized in the segmentation and the end is not recognized, the end position is returned corresponding to -1 (or not returned).
[0086] It should be noted that the audio frames in this embodiment are arranged in the order of their appearance in the audio stream. The positions corresponding to the above-mentioned start and stop positions include (but are not limited to) the fbank feature frame index / original audio / PCM corresponding byte order index. Their main purpose is to transmit the specific position information of the start / stop judgment to the upstream. At the same time, this embodiment will use the example of returning -1 at the start position beg_inx if the start is not recognized, and returning -1 at the end position end_inx if the end is not recognized.
[0087] In step S740, speech recognition is performed on the sentences to be recognized corresponding to the sentences in the sentence list to determine the speech recognition results of the audio stream.
[0088] In this embodiment, after determining the sentence list corresponding to the audio stream, the sentences in the sentence list are traversed to determine the sentences to be recognized in the audio stream. Each time a sentence to be recognized is determined, the determined sentence to be recognized is sent to a preset speech recognition module for speech recognition, thereby performing real-time speech recognition on the audio stream and ultimately determining the speech recognition result of the audio stream. Optionally, in this embodiment, different processing cores can be used to respectively determine the sentence list corresponding to the audio stream and perform speech recognition on the sentences to be recognized in the audio stream in real time, which can further improve the real-time performance of audio processing and the user experience.
[0089] In step S750 , the speech recognition result of the audio stream is evaluated.
[0090] In this embodiment, the speech recognition result of the audio stream is evaluated by performing a rejection judgment on the speech recognition result corresponding to the audio stream. The specific processing method can refer to the above content and will not be repeated here.
[0091] At the same time, in order to further facilitate the understanding of the audio processing method of the audio stream, this embodiment will be combined with Figure 8 This section describes a method for performing speech recognition on the sentences to be recognized corresponding to the segments in the segment list and determining the speech recognition results of the audio stream.
[0092] Figure 8 FIG is a schematic diagram of audio recognition according to an embodiment of the present invention. Figure 8 As shown, in this embodiment, the following processing method is used to achieve the determination of the sentence to be recognized in the audio stream and speech recognition.
[0093] In step S810, the sentence segmentation list is traversed.
[0094] In step S820, it is determined whether the sentence segmentation list is empty.
[0095] In this embodiment, when there are unprocessed segments in the segment list, step S830 is continued to be executed; if the segment list is empty, indicating that the traversal is completed, step S8A0 is executed to end the traversal.
[0096] In step S830, the current segment is obtained.
[0097] In step S840, it is determined whether there is a sentence starting point in the current segment.
[0098] In this embodiment, the endpoint detection result of the current segment is obtained to determine whether the current segment has a sentence starting point, that is, the starting position beg_dex returned by the current segment is determined to be -1. If the starting position beg_dex returned by the current segment is not -1, it indicates that the current segment has a sentence starting point, and step S850 is continued at this time; if the starting position beg_dex returned by the current segment is -1, it indicates that the current segment does not have a sentence starting point, and step S850' is continued at this time.
[0099] In step S850 , in response to the existence of a sentence start point in the current segment, the start frame of the current segment is output. The start frame of the current segment is the start position beg_dex of the current segment.
[0100] In step S850 ′, in response to the absence of a sentence start point in the current segment, the start frame of the latest sentence to be recognized is output.
[0101] In step S860, it is determined whether there is a sentence end point in the current segment.
[0102] In this embodiment, whether the current sentence has a sentence endpoint is determined by obtaining the endpoint detection result of the current sentence, that is, by determining whether the end position end_dex returned by the current sentence is -1. If the end position end_dex returned by the current sentence is not -1, it indicates that the current sentence has a sentence endpoint, and step S870 is continued. If the end position end_dex returned by the current sentence is -1, it indicates that the current sentence does not have a sentence endpoint, and step S870' is continued.
[0103] In step S870 , in response to the existence of a sentence end point in the current sentence, the end frame of the current sentence is output. The end frame of the current sentence is the end position end_dex of the current sentence.
[0104] In step S880, the latest state of the statement to be recognized is recorded as the end state.
[0105] In step S870', in response to the current segment not having a sentence ending point, the end frame of the previous sentence to be recognized is output.
[0106] In step S880', the state of the latest statement to be recognized is recorded as an unfinished state.
[0107] In step S890, the latest sentence to be recognized is determined or sent.
[0108] In this embodiment, when the latest sentence to be recognized is in an unfinished state, the latest sentence to be recognized is continuously determined until the latest sentence to be recognized is in a finished state. When the latest sentence to be recognized is in a finished state, the audio from the start frame to the end frame of the latest sentence to be recognized is sent to the speech recognition module to perform speech recognition on the latest sentence to be recognized. At the same time, the process returns to step S810 and continues until all sentences to be recognized in the audio stream are recognized.
[0109] The technical solution of this embodiment determines the sentence to be recognized by traversing the sentence segments in the sentence segmentation list and performs voice recognition on the sentence to be recognized to determine the voice recognition result. It can realize real-time and continuous long voice recognition, breaking through the time limit of traditional audio recognition. Moreover, since the sentence to be recognized determined by the sentence starting point and sentence end point in the endpoint detection result based on the sentence segmentation includes the complete content, the audio processing method in this embodiment can also improve the accuracy of long voice recognition.
[0110] Figure 9 Schematic diagram of an audio processing device according to an embodiment of the present invention. Figure 9 As shown. The audio processing device in this embodiment includes a traversal unit 91, a detection unit 92 and a recognition unit 93. The traversal unit 91 is used to traverse the sentence list and determine the current sentence, wherein the sentence list includes at least one sentence, and the sentence is determined by intercepting the audio stream based on the silence detection result. The detection unit 92 is used to determine the endpoint detection result of the current sentence; in response to the existence of a sentence starting point in the current sentence, the starting frame of the current sentence is determined as the starting frame of the latest sentence to be recognized; in response to the existence of a sentence ending point in the current sentence, the ending frame of the current sentence is determined as the ending frame of the latest sentence to be recognized; and the latest sentence to be recognized is determined according to the starting frame and the ending frame of the latest sentence to be recognized. The recognition unit 93 is used to perform speech recognition on the latest sentence to be recognized and determine the corresponding speech recognition result.
[0111] Figure 10 Schematic diagram of an electronic device according to an embodiment of the present invention. In this embodiment, the electronic device includes a server, a terminal, etc. Figure 10 As shown, the electronic device includes: at least one processor 101; a memory 102 communicatively connected to the at least one processor 101; and a communication component 103 communicatively connected to the scanning device, wherein the communication component 103 receives and sends data under the control of the processor 101; wherein the memory 102 stores instructions that can be executed by the at least one processor 101, and the instructions are executed by the at least one processor 101 to implement the above-mentioned audio processing method.
[0112] Specifically, the electronic device includes: one or more processors 101 and a memory 102, Figure 10A processor 101 is taken as an example. The processor 101 and the memory 102 may be connected via a bus or other means. Figure 10 In the example above, a bus connection is used. Memory 102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer executable programs, and modules. Processor 101 executes the non-volatile software programs, instructions, and modules stored in memory 102 to execute various functional applications and data processing of the device, thereby implementing the aforementioned audio processing method.
[0113] The memory 102 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store a list of options, etc. In addition, the memory 102 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 102 may optionally include a memory remotely located relative to the processor 101, and these remote memories may be connected to an external device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0114] One or more modules are stored in the memory 102 , and when executed by one or more processors 101 , perform the audio processing method in any of the above method embodiments.
[0115] The above-mentioned product can execute the method provided in the embodiment of this application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in this embodiment, please refer to the method provided in the embodiment of this application.
[0116] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, wherein the computer-readable program is used to enable a computer to execute part or all of the above method embodiments.
[0117] That is, those skilled in the art will understand that all or part of the steps in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a program. The program is stored in a storage medium and includes a number of instructions for causing a device (which may be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps in the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., various media that can store program code.
[0118] The foregoing is merely a preferred embodiment of the present application and is not intended to limit the present application. Persons skilled in the art will readily appreciate that the present application is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. An audio processing method, characterized in that: The method comprises: Traversing a sentence list to determine a current sentence, the sentence list including at least one sentence, the sentence being determined by intercepting the audio stream based on a silence detection result; Determine an endpoint detection result of the current segmentation; In response to the presence of a sentence starting point in the current sentence segment, determining the starting frame of the current sentence segment as the starting frame of the latest sentence to be recognized; In response to the existence of a sentence end in the current sentence, determining the end frame of the current sentence as the end frame of the latest sentence to be recognized; Determining the latest sentence to be recognized based on the start frame and the end frame of the latest sentence to be recognized; Perform speech recognition on the latest sentence to be recognized to determine a corresponding speech recognition result.
2. The method according to claim 1, characterized in that The method further comprises: Receive audio stream; Extracting features from the audio stream to determine a corresponding audio feature sequence; Performing silence detection on the audio feature sequence; In response to detecting that the continuous silence duration reaches a configured duration, determining a truncation position; The audio feature sequence is truncated at the truncation position to determine the corresponding sentence segment.
3. The method according to claim 2, characterized in that In response to detecting that the continuous silence duration reaches a configured duration, determining the truncation position includes: In response to detecting that the continuous silence duration reaches a configured duration, determining scene information of a non-silent audio segment closest to the current audio frame, the scene information including semantic information and / or context information; In response to the scene information being complete, a truncation position is determined.
4. The method according to claim 2 or 3, characterized in that The truncation position is any audio frame within the silence period corresponding to the continuous silence duration.
5. The method according to claim 2, characterized in that Extracting features from the audio stream to determine a corresponding audio feature sequence includes: Feature extraction is performed on the audio stream to determine a corresponding Mel spectrum feature sequence.
6. The method according to claim 2, characterized in that The method further comprises: Performing a rejection determination on a speech recognition result corresponding to the audio stream, and outputting speech interaction content based on the rejection determination result.
7. An audio processing device, characterized in that: The device comprises: A traversal unit, configured to traverse a sentence list and determine a current sentence, wherein the sentence list includes at least one sentence, and the sentence is determined by intercepting the audio stream based on the silence detection result; a detection unit, configured to determine an endpoint detection result of the current sentence; in response to a sentence start point in the current sentence, determine the start frame of the current sentence as the start frame of the latest sentence to be recognized; in response to a sentence end point in the current sentence, determine the end frame of the current sentence as the end frame of the latest sentence to be recognized; and determine the latest sentence to be recognized based on the start frame and end frame of the latest sentence to be recognized; The recognition unit is used to perform speech recognition on the latest sentence to be recognized and determine a corresponding speech recognition result.
8. A computer program product, characterized in that The computer program product comprises a computer program / instruction, and when the computer program / instruction is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
9. An electronic device comprising a memory and a processor, characterized in that: The memory is configured to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Continuous and long voice recognition method and system and hardware equipment
CN105719642A
Voice recognizing method and device
CN110767236A
Speech VAD tail point determination method and device, electronic equipment and computer readable medium
CN111627463A
Voice segmentation method and device and computer-readable storage medium
CN112466287A
Audio signal processing method, model training method and device, equipment and medium
CN113380238A
Cited By
Asynchronous alignment and accurate dynamic truncation method for audio frame and streaming recognition text
CN122177119A
Asynchronous alignment of audio frames with streaming recognized text and accurate dynamic truncation method
CN122177119B