Offline real-time voice transcription optimization method and device for domestic operating system

By segmenting and identifying the original audio signal, correcting word errors in real-time and post-processing, the accuracy and response speed of speech transcription technology in complex environments is solved, and high-quality text generation and user experience is achieved, which is suitable for network-free environments.

CN119943054AActive Publication Date: 2025-05-06NAT UNIV OF DEFENSE TECH

Patent Information

Application Number
CN202510050378.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-06
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

The existing speech transcription technology has low recognition accuracy and slow response speed in complex environments, and is difficult to be compatible with multilingual and flexible terms such as new words and slang, and the text punctuation and mark configuration is not ideal.

Method used

It provides an offline real-time speech transcription optimization method for domestic operating systems. By segmenting and identifying the original audio signal, correcting words in real time, and post-processing to judge potential pauses and sentence boundaries, automatically adding punctuation marks, and dynamically adjusting the speech crop length and output speed.

Benefits of technology

It improves the accuracy and response speed of speech transcription in complex environments, enhances the naturalness of text generation, ensures user experience, and operates stably in a network-free environment, ensuring user privacy and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943054A_ABST
    Figure CN119943054A_ABST
Patent Text Reader

Abstract

The invention discloses a domestic operating system-oriented offline real-time voice transcription optimization method and device, and belongs to the field of voice transcription. According to the off-line real-time voice transcription optimization method oriented to the domestic operating system, the voice cutting length and the output speed are dynamically adjusted by performing segmentation and recognition operation on the original audio signal, the transcription accuracy is guaranteed, meanwhile, the real-time performance of voice transcription is enhanced, and the user experience is improved; after the preliminary transcription text is obtained, real-time word error correction is carried out, so that the accuracy of a transcription result is further improved; after error correction, the transcription text is post-processed to obtain an off-line real-time voice transcription result, proper punctuation marks are automatically added for the transcription result according to semantics, the readability of the text is improved, the voice transcription process is automatically stopped, and the transcription process is more in line with the intention and use habits of a user. The offline real-time voice transcription optimization method facing the domestic operating system provided by the invention does not need to depend on the Internet, can be stably operated in a network-free environment, is suitable for various application scenes, effectively avoids the risks of data leakage and privacy intrusion, and ensures the privacy security of a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of speech transcription, and in particular to an offline real-time speech transcription optimization method and device for a domestic operating system. Background Art

[0002] In today's digital age, real-time speech transcription technology, as an important tool for information exchange, has been widely used in a variety of scenarios, including meeting records, real-time subtitle generation, virtual assistants, voice input, etc. With the continuous development of artificial intelligence and natural language processing technology, speech transcription technology has made significant progress. However, with the complexity of demand scenarios, such technology still faces many challenges in terms of accuracy, response speed and applicability.

[0003] First, when processing long-term continuous voice input, the instability of voice signals and the interference of environmental noise often lead to a decrease in recognition accuracy. Especially in complex acoustic environments, such as noisy background noise and multi-party conversations, the performance of traditional voice recognition systems is difficult to be satisfactory. Most of the current solutions rely on complex voice processing algorithms. Although these algorithms have improved in accuracy, they still have shortcomings in terms of computing resource consumption and operational convenience. At the same time, in some scenarios with high real-time requirements, such as real-time conference translation or subtitle generation, too slow a response speed directly affects the user experience, and blindly pursuing response speed may lead to a decrease in recognition accuracy.

[0004] Secondly, the compatibility issue of multi-language and multi-regional environments is also one of the bottlenecks of existing speech transcription technology. For example, in the conversion between traditional and simplified Chinese characters, current technology is difficult to fully meet the needs of users in different regions. In addition, in the process of real-time speech transcription, the emergence of flexible expressions such as new words, slang or slips of the tongue often exceeds the coverage of the system dictionary and grammatical rules, thus affecting the transcription quality of the system.

[0005] Furthermore, the text generated by speech transcription usually lacks reasonable punctuation configuration, which not only affects the readability of the text, but also reduces its value in subsequent use. The punctuation addition method in traditional systems relies on static rules, which makes it difficult to generate text that conforms to the expression logic according to the context. By introducing automatic punctuation addition technology based on context analysis, not only can the text quality be significantly improved, but also the user's reading experience can be improved.

[0006] Therefore, how to effectively improve the accuracy, response speed and naturalness of text generation of real-time speech transcription technology in complex environments has become a key research direction. This requires the combination of advanced speech recognition algorithms, multi-language compatible design and intelligent text processing technology to promote the practical application of speech transcription technology in more scenarios. Summary of the invention

[0007] The embodiments of the present application provide an offline real-time speech transcription optimization method and device for a domestic operating system.

[0008] On the one hand, an offline real-time speech transcription optimization method for a domestic operating system is provided, the method comprising:

[0009] Acquire an original audio signal for segmentation and recognition operations to obtain a preliminary transcription text, wherein the original audio signal includes a custom long audio signal and a custom short audio signal;

[0010] Performing real-time word error correction on the preliminary transcribed text to obtain a transcribed text after error correction;

[0011] Post-processing the corrected transcribed text to obtain an offline real-time speech transcription result, wherein the post-processing includes potential pause and sentence boundary judgment and speech transcription process stop judgment;

[0012] In response to detecting that the original audio signal is the custom short audio signal, segmentation and recognition operations are performed on the custom short audio signal through a preset audio sampling rate and a preset buffer.

[0013] Optionally, the obtaining of the original audio signal for segmentation and recognition operations to obtain a preliminary transcription text includes:

[0014] Receive the original audio signal collected by the collection end;

[0015] Preprocessing the original audio signal to obtain an enhanced audio signal;

[0016] According to the voice activity detection result, the enhanced audio signal is segmented and then spliced ​​to obtain an audio signal to be recognized;

[0017] Speech recognition is performed on the audio signal to be recognized and converted into the preliminary transcription text.

[0018] Optionally, the segmenting and splicing of the enhanced audio signal according to the voice activity detection result to obtain the audio signal to be recognized includes:

[0019] Segmenting the enhanced audio signal into a plurality of speech segments according to the speech activity detection result, wherein the segmentation is based on natural pauses or customized time lengths;

[0020] Adjacent speech segments are spliced ​​to obtain the audio signal to be recognized, wherein the adjacent speech segments have an overlapping portion of a custom length at the splicing position.

[0021] Optionally, performing speech recognition on the audio signal to be recognized and converting it into the preliminary transcription text includes:

[0022] Identifying the overlapping parts of the audio signal to be identified in a sliding window mode;

[0023] The overlapping parts are removed to obtain the preliminary transcription text.

[0024] Optionally, performing real-time word error correction on the preliminary transcribed text to obtain a transcribed text after error correction includes:

[0025] Performing a word-by-word grammar and spelling check on the preliminary transcription using word segmentation technology;

[0026] Use context information to identify words that are not grammatical or semantically correct and mark them as uncertain words;

[0027] Calculating the edit distance of the uncertain words and generating a set of candidate correction words in combination with context information;

[0028] The candidate correction words are sorted, and the correction words with the best matching items are selected and replaced into the preliminary transcription text to obtain the error-corrected transcription text.

[0029] Optionally, the post-processing of the error-corrected transcribed text to obtain an offline real-time speech transcription result includes:

[0030] Analyzing the grammatical structure, topic and contextual semantics of the corrected transcribed text to identify potential pauses and sentence boundaries;

[0031] Analyzing in real time the response level of the original audio signal during speech input, and identifying significant pauses or emphasis changes and rhythm and pitch changes in the speech stream corresponding to the original audio signal;

[0032] Assisting in determining the potential pauses and sentence boundaries based on the identified significant pauses or changes in emphasis and changes in rhythm and pitch;

[0033] Generating punctuation suggestions for the corrected transcribed text according to the potential pauses and sentence boundaries using natural language technology;

[0034] According to a pause length threshold specified by a user, real-time monitoring of the silence duration in the original audio signal;

[0035] In response to the silence duration in the original audio signal exceeding the silence duration defined by the user, an automatic stop function of the speech transcription process is triggered.

[0036] Optionally, after obtaining the original audio signal for segmentation and recognition operations to obtain a preliminary transcription text, the method further includes:

[0037] Calculate the confidence score C(x) of each speech segment and perform repeated recognition adjustment on each speech segment, wherein the calculation formula of the confidence score C(x) is: x represents the recognized speech segment, P i (x) represents the confidence probability of the i-th model for x, and N is the total number of evaluation methods;

[0038] In response to detecting a segment of speech having a confidence score C(x) below a target threshold θ, an automatic adjustment processing strategy is performed.

[0039] Optionally, in response to detecting a speech segment whose confidence score C(x) is lower than a target threshold θ, executing an automatic adjustment processing strategy comprises:

[0040] For speech segments whose confidence score C(x) is lower than the target threshold θ, increase the audio length L during repeated recognition repeat , where L repeat =L base +α×(L complex -L base ), L base is the basic processing length, L complex is the ideal processing length after complexity judgment, and α is an adjustment coefficient ranging from 0 to 1.

[0041] On the other hand, an offline real-time speech transcription optimization device for a domestic operating system is provided, the device comprising:

[0042] A signal operation module, used to obtain an original audio signal for segmentation and recognition operations to obtain a preliminary transcription text, wherein the original audio signal includes a custom long audio signal and a custom short audio signal;

[0043] A word error correction module, used for performing real-time word error correction on the preliminary transcribed text to obtain a transcribed text after error correction;

[0044] A post-processing module, used for post-processing the transcribed text after error correction to obtain an offline real-time speech transcription result, wherein the post-processing includes potential pause and sentence boundary judgment and speech transcription process stop judgment;

[0045] The short audio processing module is used for performing segmentation and recognition operations on the custom short audio signal by using a preset audio sampling rate and a preset buffer in response to detecting that the original audio signal is the custom short audio signal.

[0046] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the offline real-time speech transcription optimization method for domestic operating systems as described in the above aspects.

[0047] On the other hand, a computer program product is also provided, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the offline real-time speech transcription optimization method for domestic operating systems described in the above aspect.

[0048] The technical effects brought about by this application are at least as follows.

[0049] The present application provides an offline real-time speech transcription optimization method for domestic operating systems. By segmenting and recognizing the original audio signal, the dynamic adjustment of the speech clipping length and output speed is achieved. While ensuring the transcription accuracy, the real-time nature of speech transcription is enhanced, and the user experience is improved. After obtaining the preliminary transcribed text, real-time word error correction is performed to further improve the accuracy of the transcription result. The transcribed text after error correction is post-processed to obtain the offline real-time speech transcription result, and appropriate punctuation marks are automatically added to the transcription result according to the semantics to improve the readability of the text, and the speech transcription process is automatically stopped, so that the transcription process is more in line with the user's intentions and usage habits. The offline real-time speech transcription optimization method for domestic operating systems provided by the present application does not need to rely on the Internet, can run stably in a network-free environment, is suitable for a variety of application scenarios, effectively avoids the risk of data leakage and privacy invasion, and ensures the privacy and security of users.

[0050] In other embodiments, the number of voice repetition recognition times and the output speed are adjusted according to different user audio input densities, thereby further improving the user experience.

[0051] In other embodiments, a sliding window speech transcription mode is also used to perform repeated recognition and correct text errors in real time. At the same time, traditional Chinese characters in the transcribed text are automatically recognized and converted into simplified Chinese characters, further improving the accuracy of the transcription results.

[0052] In other embodiments, the loudness of the voice input is monitored in real time and the voice context is analyzed to more accurately determine the demarcation point of the audio segment, effectively filter the background noise, adapt to complex application scenarios, and make the transcription result closer to the actual semantics of the user. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A flowchart of an offline real-time speech transcription optimization method for a domestic operating system provided by an exemplary embodiment of the present application is shown;

[0054] Figure 2 Shows the corresponding Figure 1 Schematic diagram of the interaction process between users and AI assistants;

[0055] Figure 3 The following is a schematic diagram showing the structure of each module corresponding to the embodiment of the present application;

[0056] Figure 4 The following are test results in quiet and noisy environments;

[0057] Figure 5 The following are test results at different speaking speeds. DETAILED DESCRIPTION

[0058] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.

[0059] The term "multiple" as used herein refers to two or more than two. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the related objects are in an "or" relationship.

[0060] Example 1

[0061] Please refer to Figure 1 , which shows a flow chart of an offline real-time speech transcription optimization method for a domestic operating system provided by an exemplary embodiment of the present application, Figure 2 Shows the corresponding Figure 1 A user interacts with an AI assistant through a flowchart, the method comprising:

[0062] Step 101, obtaining the original audio signal for segmentation and recognition operations to obtain a preliminary transcription text.

[0063] In a possible implementation, step 101 includes the following steps.

[0064] Step S11, receiving the original audio signal collected by the collection end.

[0065] The original audio signal includes a customized long audio signal and a customized short audio signal, and the length of the original audio signal is not limited to 10 seconds.

[0066] Step S12: pre-process the original audio signal to obtain an enhanced audio signal.

[0067] Among them, the process of preprocessing the original audio signal can be as follows: the processing system first receives the original audio signal input by the acquisition end such as a microphone, and analyzes the amplitude distribution and energy density of the input original audio signal through fast Fourier transform (FFT) to provide data support for subsequent amplitude adjustment.

[0068] Furthermore, the dynamic range of the input signal is calculated, the peak amplitude is identified, and the overall power level of the signal is evaluated as a reference for determining the uniform amplitude range. Further, based on the calculated dynamic range, the signal amplitude is adjusted by linear scaling or dynamic compression technology to map the signal amplitude into the standard range.

[0069] The above processing ensures that all audio signal data is within the range that the system can handle, which helps to reduce distortion and enhance the signal-to-noise ratio.

[0070] Step S13, segmenting and splicing the enhanced audio signal according to the voice activity detection result to obtain an audio signal to be recognized.

[0071] The enhanced audio signal is segmented into several speech segments according to the speech activity detection result, wherein the segmentation is based on natural pauses or customized time lengths.

[0072] The activity process of the voice activity detection result is as follows. After the microphone at the acquisition end receives the voice signal, it will save the voice for synchronous detection. The detection process mainly includes fast Fourier transform (FFT) of the input voice, analyzing the amplitude distribution and energy density of the input audio signal, calculating the peak amplitude, and evaluating the overall power level; adjusting the overall signal amplitude through linear transformation to adjust the signal to a certain range.

[0073] It should be noted that the preset long segment speech is generally defined as speech of more than 1 minute according to actual needs, such as scenes such as meetings and speeches. Of course, technical personnel in this field can define the duration according to actual needs and there is no limitation on this.

[0074] Adjacent speech segments are concatenated to obtain an audio signal to be recognized, wherein adjacent speech segments have an overlapping portion of a custom length at the concatenation position.

[0075] In one example, for a preset long segment of speech, the back end of the processing will intercept 10 seconds of the previous segment of speech for repeated detection, so the length of the repeated speech is 10 seconds. Of course, those skilled in the art can define the detection time of the overlapping part according to actual needs, which is not limited here.

[0076] Step S14, performing speech recognition on the audio signal to be recognized and converting it into a preliminary transcription text.

[0077] In one possible implementation, the audio signal to be recognized is recognized and converted into text, the text is integrated, and the overlapping parts between the two voice texts are removed. Specifically, the overlapping parts in the audio signal to be recognized are recognized in a sliding window mode, and then the overlapping parts are removed to obtain a preliminary transcription text. This operation can enhance the accuracy of audio recognition at the end of each voice segment while better describing the context of the entire voice segment, which serves as a basis for adding appropriate punctuation marks.

[0078] Step 102, performing real-time word error correction on the preliminary transcribed text to obtain a transcribed text after error correction.

[0079] In a possible implementation, step 102 includes the following steps.

[0080] Step S21, using word segmentation technology to perform word-by-word grammar and spelling checks on the preliminary transcription text.

[0081] Step S22, using context information to identify words that do not conform to grammar or semantics and mark them as uncertain words;

[0082] Step S23, calculating the edit distance of the uncertain words and generating a set of candidate correction words in combination with the context information.

[0083] In one possible implementation, a basic grammar and spelling check is performed on the preliminary transcription text in real time, such as using word segmentation technology to decompose the text into independent words and phrases to facilitate error detection of each word. Context information is used to identify words that are not grammatical or semantically correct, and these words are marked as uncertain words, i.e., possible errors. The edit distance of the uncertain words is calculated, and a set of candidate correction words are generated in combination with the context, which can be in the form of words or phrases, wherein the context is used to predict the candidate correction words, which can improve the overall semantic consistency.

[0084] Step S24, sorting the candidate correction words, selecting the correction words with the best match and replacing them into the preliminary transcription text, so as to obtain the error-corrected transcription text.

[0085] Furthermore, considering common error patterns in speech recognition, candidate correction words are sorted, and options that are most likely to match the context are displayed first. High-confidence correction suggestions are automatically applied and directly replaced into the transcribed text. That is, the correction word of the best match is selected and replaced into the preliminary transcribed text to obtain the corrected transcribed text.

[0086] In addition, you can also add content that checks traditional Chinese characters in text and converts them into simplified Chinese characters based on different regions and personal usage habits.

[0087] Step 103, post-processing the transcribed text after error correction to obtain an offline real-time speech transcription result, wherein the post-processing includes potential pauses and sentence boundaries judgment and speech transcription process stop judgment.

[0088] In a possible implementation, step 103 includes the following steps.

[0089] Step S31, analyzing the grammatical structure, topic and contextual semantics in the transcribed text after error correction, and identifying potential pauses and sentence boundaries.

[0090] Step S32, analyzing the response level of the original audio signal during speech input in real time, and identifying significant pauses or emphasis changes and rhythm and pitch changes in the original audio signal corresponding to the speech stream.

[0091] Step S33, assisting in determining potential pauses and sentence boundaries based on the identified significant pauses or emphasis changes and rhythm and tone changes.

[0092] This step is for speech feature analysis. Specifically, after the initial completion of speech transcription and recognition, the corrected transcribed text is obtained, and post-processing is performed to analyze the grammatical structure, theme and contextual semantics in the corrected transcribed text, and identify potential pauses and sentence boundaries. Furthermore, the loudness level of the speech input is analyzed in real time, significant pauses or emphasis changes in the speech stream are identified, and changes in speech rhythm and pitch are detected to assist in judging potential pauses and sentence boundaries, that is, to judge potential punctuation positions in preparation for the next step, such as predicting interrogative sentences based on rising pitch.

[0093] Step S34, generating punctuation suggestions for the corrected transcribed text based on potential pauses and sentence boundaries using natural language technology.

[0094] In one possible implementation, natural language processing technology is used in combination with contextual information to generate punctuation suggestions for text segmentation points (i.e., potential pauses and sentence boundaries), and appropriate punctuation, such as periods, commas, question marks, and exclamation marks, is added to common sentence patterns and semantic patterns.

[0095] Step S35 , monitoring the silence duration in the original audio signal in real time according to the pause length threshold specified by the user.

[0096] Step S36, in response to the silence duration in the original audio signal exceeding the silence duration defined by the user, triggering the automatic stop function of the speech transcription process.

[0097] In a possible implementation, the system sends a stop signal to the speech transcription process to pause transcription and save the current offline real-time speech transcription result in a timely manner.

[0098] In addition, considering the need for timely response to short voice, the embodiment of the present application also provides a processing method for short audio signals.

[0099] Step 104 , in response to detecting that the original audio signal is a custom short audio signal, segmenting and identifying the custom short audio signal by using a preset audio sampling rate and a preset buffer.

[0100] In one example, fast speech recognition is performed on the original audio signal with a preset audio sampling rate of 16 kHz and a preset buffer of 0.5 seconds.

[0101] In response to detecting that the original audio signal is a short audio signal, rapid speech recognition is performed on the original audio signal using a preset 16 kHz audio sampling rate and a 0.5 second buffer.

[0102] In one possible implementation, when user input is detected, a shorter audio signal is intercepted for speech recognition, and fast speech recognition is performed on the original audio signal using a preset 16kHz audio sampling rate and a 0.5-second buffer, i.e., FFT processing is performed, wherein the 16kHz audio sampling rate and the 0.5-second buffer can be user settings or system default settings.

[0103] It should be noted that since the length of the short audio signal is also determined by the input from the acquisition end, in order to improve the response rate of short voice, in a possible implementation, the segmentation and recognition operations on the custom short audio signal mentioned in step 104 also require the segmentation operations, word correction, and real-time output after processing such as steps 101 to 103 for short voice. For example, within 10 seconds, the first 4 seconds will be cut first for rapid recognition and output of the recognition result, and then repeated recognition and correction of the output result will be performed after splicing to ensure the immediacy of user input to output.

[0104] Therefore, the embodiment of the present application ensures that the system can quickly analyze short audio signals, and the recognized text will be output in real time through a ring buffer mechanism so as to be quickly updated to the user interface. For example, if the initial response time is set to 2 seconds, the user can see the preliminary transcription results within 3 seconds.

[0105] The present application provides an offline real-time speech transcription optimization method for domestic operating systems. By segmenting and recognizing the original audio signal, the dynamic adjustment of the speech clipping length and output speed is achieved. While ensuring the transcription accuracy, the real-time nature of speech transcription is enhanced, and the user experience is improved. After obtaining the preliminary transcribed text, real-time word error correction is performed to further improve the accuracy of the transcription result. The transcribed text after error correction is post-processed to obtain the offline real-time speech transcription result, and appropriate punctuation marks are automatically added to the transcription result according to the semantics to improve the readability of the text, and the speech transcription process is automatically stopped, so that the transcription process is more in line with the user's intentions and usage habits. The offline real-time speech transcription optimization method for domestic operating systems provided by the present application does not need to rely on the Internet, can run stably in a network-free environment, is suitable for a variety of application scenarios, effectively avoids the risk of data leakage and privacy invasion, and ensures the privacy and security of users.

[0106] In addition, it also includes adjusting the number of voice repetition recognition and output speed according to different user audio input densities, further improving the user experience.

[0107] In addition, it also includes a sliding window voice transcription mode to perform repeated recognition and correct text errors in real time. At the same time, it automatically recognizes and converts traditional Chinese characters in the transcribed text into simplified Chinese, further improving the accuracy of the transcription results.

[0108] In addition, it also includes real-time monitoring of the loudness of voice input and contextual analysis of voice to more accurately determine the demarcation points of audio segments, effectively filter background noise, adapt to complex application scenarios, and make the transcription results closer to the user's actual semantics.

[0109] Example 2

[0110] In addition, in order to further improve the accuracy of the speech transcription text, the embodiment of the present application also provides a repeated recognition adjustment mechanism, and the content after the method step 101 corresponding to the mechanism is as follows.

[0111] Calculate the confidence score C(x) of each speech segment and perform repeated recognition adjustment on each speech segment. The calculation formula of the confidence score C(x) is: x represents the recognized speech segment, P i (x) represents the confidence probability of the i-th model for x, N is the total number of evaluation methods, and in response to detecting a speech segment whose confidence score C(x) is lower than the target threshold θ, an automatic adjustment processing strategy is executed. In one example, the model uses the openAI open source speech recognition model whisper.

[0112] Furthermore, for the speech segments whose confidence score C(x) is lower than the target threshold θ, the audio length L for repeated recognition is increased.repeat , where L repeat =L base +α×(L complex -L base ), L base is the basic processing length, L complex is the ideal processing length after complexity judgment, and α is an adjustment coefficient ranging from 0 to l. When a text segment with a confidence lower than a certain threshold θ is detected (i.e. C(x)<<θ), the processing strategy will be automatically adjusted. The value range of α is dynamically adjusted according to the complexity evaluation.

[0113] Furthermore, the text output speed is adjusted. Based on the general speaking speed, the standard speaking speed (v standard ) is 160 words per minute (WPM), and for fast-speech scenarios, a buffer time Δt is set base The basic L is set to 200 milliseconds to ensure stable text output in most environments; the adjustment coefficient γ is set to 0.2 seconds to fine-tune the sensitivity of the delay to different speech speed changes. For low-speed scenarios, the basic L is set to base 0.5 seconds; Set the adjustment coefficient β to 0.5 seconds to determine the extent to which the recognition length is increased when the speech speed is lower than the standard. user ) is greater than v standard When the buffer time (Δt output ) drops to Δt base 70%-80%; when v user Less than v standard When the recognition length is increased by the calculation formula, the update time of the text output is slowed down to ensure high accuracy of speech recognition:

[0114] Δt output =Δt base +γ×(v standard -v user );

[0115] L repeat =L base +β×(v standard -v user ).

[0116] Among them, the situation below the threshold is further explained. Below the threshold means that the recognition effect of the paragraph is not good and further repeated recognition is required.

[0117] The occurrence of low confidence includes but is not limited to the presence of unfiltered background noise, insufficient speech clarity, and high word error rate. The main approach to low-confidence speech segments is to increase the length of repeated recognition audio, and to improve the accuracy by repeatedly recognizing longer segments of audio, without taking different response strategies for different situations.

[0118] The process of complexity judgment is explained below.

[0119] In this application, complexity judgment is mainly reflected in the processing and repeated recognition mechanism of speech segments with low confidence scores. The specific steps are as follows.

[0120] When the system calculates the confidence score for each speech segment, if the confidence score is lower than the target threshold (indicating that there is a large uncertainty in the transcription result), further processing is required.

[0121] For low-confidence segments, the following two steps are used for optimization: increasing the audio length and adjusting the text output speed.

[0122] As the audio length increases, when repeated recognition is performed, the audio processing length is increased to improve the recognition accuracy. The specific calculation formula is: L repeat =L base +α×(L complex -L base ).

[0123] Text output speed adjustment: dynamically adjust the text output speed based on the user's input speed. For example, if the user's input speed is higher than the standard value (160WPM), reduce the buffer time and speed up the output. If the user's input speed is lower than the standard value, increase the buffer time and slow down the output speed to improve recognition accuracy.

[0124] Through the complexity judgment and processing mechanism, low-confidence segments are effectively optimized, the overall accuracy and real-time performance of speech transcription are improved, and the needs of users with different input speeds are adapted. This mechanism not only takes into account the dynamic complexity of speech input, but also improves the performance of the system in complex scenarios through adaptive adjustment.

[0125] The following describes how the value range of α is dynamically adjusted according to the complexity evaluation.

[0126] In the present application, "the value range of α is dynamically adjusted according to the complexity evaluation" means that the system will dynamically adjust the value range of the parameter α according to the complexity of the current voice input to adapt to different scenarios, thereby optimizing the audio processing length and recognition results.

[0127] First of all, complexity assessment is mainly based on the following indicators.

[0128] Signal quality: background noise intensity, audio clarity, silence intervals, etc.

[0129] Recognition accuracy: judged by confidence score.

[0130] Input speaking rate: Information density in the audio segment (number of words / audio length).

[0131] Number of repeated recognitions: the degree of convergence of multiple recognition results for a certain audio segment.

[0132] When the complexity is high (such as low confidence, high noise, and fast speaking speed), the system needs to dynamically adjust the parameter α to expand the audio processing length and enhance the capture of context, thereby improving the recognition accuracy.

[0133] In one example, for simple speech scenarios (such as clear, single-person speech, low noise), α takes a lower value, such as 0.1-0.3; for complex speech scenarios (such as large background noise, multiple conversations, unclear accents), α takes a higher value, such as 0.7-0.9. Of course, the value of α is positively correlated with complexity. In this application, the system dynamically calculates α based on indicators such as confidence and noise level.

[0134] In summary, by dynamically adjusting the value range of α, the system can flexibly optimize the processing length according to the actual speech environment. For simple scenarios, it can maintain real-time performance and reduce computing resource consumption; for complex scenarios, it can enhance recognition accuracy and ensure result quality. This mechanism ensures that the system can achieve the best speech transcription effect in environments of different complexity.

[0135] Example 3

[0136] It should be noted that segmenting and splicing the enhanced audio signal according to the voice activity detection result is a real-time operation. Each segment of the speech is sequential and the next step is performed in chronological order. Let's take an example to illustrate.

[0137] like Figure 2 As shown, when the user interacts with the AI ​​assistant, they can choose between keyboard input and voice input, and the audio is pre-processed after receiving the voice input.

[0138] The first step is to perform audio preprocessing, including segmentation and recognition operations.

[0139] Set the initial audio length of the speech segment (e.g. 2 seconds). When 2 seconds of non-silent audio is received, segmentation is performed to obtain 2 seconds of speech segments. As long as the non-silent speech segments are obtained continuously in real time, the system will transcribe the audio into text for each speech segment and store it, and output it to the user interface at the same time. The output process is also real-time and continuous due to the real-time transcription, that is, the subsequent collected original audio signal continues to be segmented and spliced ​​with the recognized audio. The initially transcribed audio will also be repeatedly recognized to improve accuracy.

[0140] The second is the content of repetition recognition and text integration.

[0141] The repeatedly recognized text is integrated with the stored audio, and the integrated text is semantically analyzed to segment the sentence boundaries; the density of the user input content is analyzed, and the ratio of the number of transcribed text words to the audio length is used as the evaluation index; when this index exceeds the normal speaking speed range, it is determined that the user input speed is fast. The system feeds this signal back to the text preprocessing and text output links, increases the length of the speech splicing, and speeds up the output speed of the text; when this index is lower than the normal speaking speed range, the length of the speech splicing is appropriately shortened, and the output speed of the text is reduced. In this way, the speech can be recognized repeatedly, strengthen the connection with the context, and improve the recognition accuracy.

[0142] The second step is to initially transcribe the text for real-time word correction and optimization.

[0143] For the integrated repeated recognition text, instant error correction is performed in combination with the context to identify and correct possible errors; traditional Chinese characters in the transcribed text are identified and converted into simplified Chinese characters instantly; based on the analysis of semantics and context in the text analysis phase, appropriate punctuation marks are added at the text boundaries to improve the readability of the text.

[0144] The last part is post-processing, i.e. automatic stop and result saving.

[0145] According to the user-specified pause threshold, the silence duration in the voice signal is monitored in real time to determine the end point of the voice input and the pause intended by the user; if the user-defined silence duration is detected, the automatic stop function is triggered. A stop signal is sent to the transcription process to pause the transcription and save the current text result in time.

[0146] Example 4

[0147] For the actual application of this application, the device module is explained by taking the Qizhi-Lingyu AI assistant application running on the domestic operating system openKylin as an example.

[0148] like Figure 3As shown, a schematic diagram of the module structure corresponding to the embodiment of the present application is shown. In order to become a more complete system module, the functional modules can be increased or decreased according to the actual situation, and there is no limitation on this. Figure 3 is a possible improved system module structure.

[0149] In one example, the system mainly includes a front-end page module, an audio processing module, and a text analysis and processing module.

[0150] The front-end page module includes a user audio input unit, a text output unit and a text editing unit; the audio processing module includes an audio noise reduction filter unit, an audio cropping and splicing unit, an audio transcription unit and an audio input monitoring unit; the text analysis and processing module includes a first functional unit and a second functional unit, wherein the first functional unit is used for text error correction and adding punctuation marks, and the second functional unit is used for analyzing the text context and calculating the user information input density.

[0151] Among them, the user audio input unit can be used to execute step S11, the audio noise reduction filter unit can be used to execute step S12, the audio cropping and splicing unit can be used to execute step S13, the audio transcription unit can be used to execute step S14, and the first functional unit can be used to execute step S21, step S22, step S23, step S24, step S31, step S32, step S33, step S34, step S35 and step S36.

[0152] The second functional unit may calculate the user information input density and may also be used to process the short audio signal. For example, the basis for calculating the user information input density is the short audio signal judgment.

[0153] Among them, the text editing unit can be used to execute the repeated recognition adjustment mechanism, such as executing the repeated recognition adjustment mechanism on the continuously obtained transcription results, and performing text editing adjustments and real-time output.

[0154] In a possible implementation, when the front-end page module detects the presence of user audio input, it obtains the initial audio signal and enters the audio processing module for processing in each unit. The audio processing module also monitors the silent state of the user input in real time to determine when to stop transcribing. After the audio transcription of each segment is completed, it will also be sent to the text analysis and processing module for further analysis of the text context. The user information input density is calculated to perform audio trimming and splicing operations, and the audio transcription of the segmented segments continues until the transcription is stopped.

[0155] When the density of user information input is too high, the text output speed will also be adjusted, ultimately achieving real-time transcription text output. The output results will also be processed for text editing, i.e. word correction.

[0156] Example 6

[0157] In order to illustrate the optimization effect of the offline real-time speech transcription optimization method for domestic operating systems provided by this application, the applicant conducted experimental tests.

[0158] The experimental test uses aishelltech's AISHELL-4 multi-channel Chinese conference speech dataset, which is an eight-channel Chinese Mandarin conference scene recorded by a microphone array. It contains a total of 211 meetings, with 4 to 8 people in each meeting, and the dataset is about 120 hours in total. This dataset aims to promote the research on multi-speaker processing in practical application scenarios, and includes various important features in actual meeting scenarios, such as pauses, overlaps, speaker turns, noise, etc. The test voices are selected from 20 interviews and conference recordings in quiet and noisy environments. The selected voices include low speech speed, normal speech speed and high speech speed. The speech transcription effect is quantified by the accuracy (A), which is expressed as follows: Where N correct Represents the correct number of words in the transcribed text, N total Represents the total number of words in the transcribed text. Transcribed text errors include typos, traditional Chinese characters, etc.

[0159] The experimental test was evaluated from three aspects: 1) the presence or absence of environmental noise, and the selected recordings were all at normal speaking speed; 2) different speaking speeds, and the selected voices were all recorded in a quiet environment; 3) the overall optimization effect.

[0160] Figure 4 The following are the test results in a quiet environment and a noisy environment. Figure 5 The test results at different speaking speeds are shown in the figure. The experimental test results show that in a quiet environment, the average accuracy of the optimized transcription is 91.67%, which is 3.93% higher than before optimization; in a noisy environment, the average accuracy of the optimized transcription is 90.09%, which is 3.73% higher than before optimization. The average accuracy after optimization under low speed is 92.14%, which is 1.81% higher than before optimization; the average accuracy after optimization under normal speaking speed is 92.07%, which is 3.58% higher than before optimization; the average accuracy after optimization under high speaking speed is 90.84%, which is 5.04% higher than before optimization. There are no traditional Chinese characters in the optimized transcription text, and traditional Chinese characters can be correctly converted into simplified Chinese. The experimental results show that the method improves the accuracy of text transcription under different environments and different speaking speeds.

[0161] On the other hand, a computer-readable storage medium is provided, wherein the storage medium stores at least one instruction, and the at least one instruction is used to be executed by a processor to implement the offline real-time speech transcription optimization method for domestic operating systems as described in the above aspects.

[0162] On the other hand, a computer program product is also provided, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the offline real-time speech transcription optimization method for domestic operating systems described in the above aspect.

[0163] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0164] A person skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware or by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0165] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. An offline real-time speech transcription optimization method for domestic operating systems, characterized in that: The method comprises: Acquire an original audio signal for segmentation and recognition operations to obtain a preliminary transcription text, wherein the original audio signal includes a custom long audio signal and a custom short audio signal; Performing real-time word error correction on the preliminary transcribed text to obtain a transcribed text after error correction; Post-processing the corrected transcribed text to obtain an offline real-time speech transcription result, wherein the post-processing includes potential pause and sentence boundary judgment and speech transcription process stop judgment; In response to detecting that the original audio signal is the custom short audio signal, segmentation and recognition operations are performed on the custom short audio signal through a preset audio sampling rate and a preset buffer.

2. The method according to claim 1, characterized in that The method of obtaining the original audio signal for segmentation and recognition operations to obtain a preliminary transcription text includes: Receive the original audio signal collected by the collection end; Preprocessing the original audio signal to obtain an enhanced audio signal; According to the voice activity detection result, the enhanced audio signal is segmented and then spliced ​​to obtain an audio signal to be recognized; Speech recognition is performed on the audio signal to be recognized and converted into the preliminary transcription text.

3. The method according to claim 2, characterized in that The step of segmenting and splicing the enhanced audio signal according to the voice activity detection result to obtain the audio signal to be recognized includes: Segmenting the enhanced audio signal into a plurality of speech segments according to the speech activity detection result, wherein the segmentation is based on natural pauses or customized time lengths; Adjacent speech segments are spliced ​​to obtain the audio signal to be recognized, wherein the adjacent speech segments have an overlapping portion of a custom length at the splicing position.

4. The method according to claim 3, characterized in that The performing speech recognition on the audio signal to be recognized and converting it into the preliminary transcription text comprises: Identifying the overlapping parts of the audio signal to be identified in a sliding window mode; The overlapping parts are removed to obtain the preliminary transcription text.

5. The method according to claim 1, characterized in that The performing real-time word error correction on the preliminary transcribed text to obtain a transcribed text after error correction includes: Performing a word-by-word grammar and spelling check on the preliminary transcription using word segmentation technology; Use context information to identify words that are not grammatical or semantically correct and mark them as uncertain words; Calculating the edit distance of the uncertain words and generating a set of candidate correction words in combination with context information; The candidate correction words are sorted, and the correction words with the best matching items are selected and replaced into the preliminary transcription text to obtain the error-corrected transcription text.

6. The method according to claim 1, characterized in that The post-processing of the error-corrected transcribed text to obtain an offline real-time speech transcription result includes: Analyzing the grammatical structure, topic and contextual semantics of the corrected transcribed text to identify potential pauses and sentence boundaries; Analyzing in real time the response level of the original audio signal during speech input, and identifying significant pauses or emphasis changes and rhythm and pitch changes in the speech stream corresponding to the original audio signal; Assisting in determining the potential pauses and sentence boundaries based on the identified significant pauses or emphasis changes and rhythm and tone changes; Generating punctuation suggestions for the corrected transcribed text according to the potential pauses and sentence boundaries using natural language technology; According to a pause length threshold specified by a user, real-time monitoring of the silence duration in the original audio signal; In response to the silence duration in the original audio signal exceeding the silence duration defined by the user, an automatic stop function of the speech transcription process is triggered.

7. The method according to claim 3, characterized in that After obtaining the original audio signal for segmentation and recognition operations to obtain a preliminary transcription text, the method further includes: Calculate the confidence score C(x) of each speech segment and perform repeated recognition adjustment on each speech segment, wherein the calculation formula of the confidence score C(x) is: x represents the recognized speech segment, P i (x) represents the confidence probability of the i-th model for x, and N is the total number of evaluation methods; In response to detecting a segment of speech having a confidence score C(x) below a target threshold θ, an automatic adjustment processing strategy is performed.

8. The method according to claim 7, characterized in that In response to detecting a speech segment whose confidence score C(x) is lower than a target threshold θ, executing an automatic adjustment processing strategy comprises: For speech segments whose confidence score C(x) is lower than the target threshold θ, increase the audio length L during repeated recognition repeat , where L repeat =L base +α×(L complex -L base ), L base is the basic processing length, L complex is the ideal processing length after complexity judgment, and α is an adjustment coefficient ranging from 0 to 1.

9. An offline real-time speech transcription optimization device for domestic operating systems, characterized in that: The device comprises: A signal operation module, used to obtain an original audio signal for segmentation and recognition operations to obtain a preliminary transcription text, wherein the original audio signal includes a custom long audio signal and a custom short audio signal; A word error correction module, used for performing real-time word error correction on the preliminary transcribed text to obtain a transcribed text after error correction; A post-processing module, used for post-processing the transcribed text after error correction to obtain an offline real-time speech transcription result, wherein the post-processing includes potential pause and sentence boundary judgment and speech transcription process stop judgment; The short audio processing module is used for performing segmentation and recognition operations on the custom short audio signal by using a preset audio sampling rate and a preset buffer in response to detecting that the original audio signal is the custom short audio signal.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the offline real-time speech transcription optimization method for domestic operating systems as described in any one of claims 1 to 8 above.

Citation Information

Patent Citations

  • Depressed mood phone automatic speech recognition screening system

    CN102339606A

  • Voice converting method and apparatus

    CN106920547A

  • Streaming speech recognition method

    CN110942764A

  • Method and device for improving quality of speech recognition text

    CN112447172A

  • Error correction method and device for voice text

    CN113012705A

Cited By

  • Two-process error correction method and device for real-time speech transcription

    CN120690225A

  • TTS audio generation system and method based on sound cloning

    CN121171202A