Speech input method, apparatus, device, and storage medium

CN122799861APending Publication Date: 2026-09-22GOERTEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611309155.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]本申请的主要目的在于提供一种语音输入方法、装置、设备及存储介质,旨在解决现有的在进行语音输入时,每当生成新的识别结果时屏幕上已有的整段内容均全部擦除,用户体验较差的技术问题

Benefits of technology

[0015]本申请提供了一种语音输入方法、装置、设备及存储介质,所述方法包括:获取用户口述的当前语音,并确定上一语音与所述当前语音之间是否存在自然断句点,所述自然断句点表征所述上一语音与所述当前语音处于不同断句;若不存在所述自然断句点的情况下,确定所述当前语音所在的当前断句的已显示文本,将所述当前断句中的已显示文本设置为允许修正状态,基于所述当前语音对所述当前断句的已显示文本进行修正,将所述当前断句的修正后的已显示文本以及所述当前语音的文本作为所述当前断句的新的已显示文本并显示,并返回执行所述获取用户口述的当前语音的步骤;若存在所述自然断句点的情况下,确定上一语音所在的上一断句的已显示文本,将所述上一断句的已显示文本设置为禁止修正状态,将所述当前语音的文本作为所述当前断句的已显示文本并显示,并返回执行所述获取用户口述的当前语音的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122799861A_ABST
    Figure CN122799861A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of voice processing, and discloses a voice input method, device and equipment and a storage medium, the method comprising the following steps: acquiring a current voice spoken by a user, and determining whether a natural punctuation point exists between a previous voice and the current voice; if no natural punctuation point exists, determining displayed text of a current punctuation where the current voice is located, setting the displayed text in the current punctuation to an allowed correction state, correcting the displayed text of the current punctuation based on the current voice, and displaying the corrected displayed text of the current punctuation and the text of the current voice as new displayed text of the current punctuation; and if a natural punctuation point exists, determining displayed text of a previous punctuation where a previous voice is located, setting the displayed text of the previous punctuation to a prohibited correction state, and displaying the text of the current voice as displayed text of a current punctuation. The application does not need to repeatedly reposition a visual line, and user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a speech input method, apparatus, device and storage medium. Background Technology

[0002] Currently, speech recognition technology has been widely deployed in devices such as smartphones, in-vehicle information systems, smart headphones, and conference recording terminals. Especially in scenarios that require the continuous generation of long texts (such as dictating meeting minutes, long instant messaging messages, and inputting navigation addresses and commands in vehicles), voice interaction has become a key way to improve the efficiency of digital content production.

[0003] Existing text presentation methods for continuous speech input typically employ "full-segment streaming transcription and refresh," treating the entire spoken content as a single text block. Relying on the streaming intermediate and final results output by the speech recognition engine, the system continuously generates new recognized text and corrects any previously output words as the user progresses. Consequently, each time a new recognition result is generated, the system erases the entire existing text on the screen and re-displays the latest version of the complete text, causing the text to constantly jump and flicker. The user's gaze is repeatedly interrupted, requiring them to reposition and review the changed content after each jump. This makes it impossible to reliably reread spoken content during the build phase, resulting in a poor user experience. Summary of the Invention

[0004] The main purpose of this application is to provide a voice input method, apparatus, device, and storage medium, which aims to solve the technical problem that when voice input is performed, the entire content on the screen is erased whenever a new recognition result is generated, resulting in a poor user experience.

[0005] To achieve the above objectives, this application proposes a voice input method, the method comprising: The system acquires the current speech dictated by the user and determines whether there is a natural punctuation point between the previous speech and the current speech. The natural punctuation point indicates that the previous speech and the current speech are at different punctuation points. If no natural punctuation point exists, determine the displayed text of the current punctuation where the current speech is located, set the displayed text of the current punctuation to the correction-allowed state, correct the displayed text of the current punctuation based on the current speech, and display the corrected displayed text of the current punctuation and the text of the current speech as the new displayed text of the current punctuation, and return to the step of obtaining the current speech dictated by the user. If the natural punctuation point exists, determine the displayed text of the previous punctuation point where the previous speech is located, set the displayed text of the previous punctuation point to a state where correction is prohibited, display the text of the current speech as the displayed text of the current punctuation point, and return to the step of obtaining the current speech dictated by the user.

[0006] In one embodiment, the step of determining whether there is a natural punctuation mark between the previous speech and the current speech includes: Feature extraction is performed on the current speech to obtain cue features, which include at least two of the following: speech pause features, intonation features, punctuation assumption features, and semantic boundary features. The clue features are weighted and fused according to preset weights to obtain a comprehensive score, and the existence of a natural punctuation point between the previous speech and the current speech is determined based on the comprehensive score.

[0007] In one embodiment, the cue features include speech pause features, intonation features, punctuation assumption features, and semantic boundary features of the current speech. The step of extracting features from the current speech to obtain cue features includes: Speech activity detection is performed on the current speech, the silence duration of the silent segment in the current speech is detected, and the speech pause feature is generated based on the silence duration; Extract the fundamental frequency of the speech frame preceding the silence segment in the current speech, and generate the intonation trend feature based on the fundamental frequency of the speech frame; Determine the punctuation classification probability vector of the current speech, and generate the punctuation hypothesis features based on the punctuation classification probability vector; Determine the semantic vector of the current sentence segment where the current speech is located, determine the similarity between the semantic vector and the preset complete statement template vector, and generate the semantic boundary feature based on the similarity.

[0008] In one embodiment, the step of setting the displayed text of the previous sentence segment to a state where correction is disabled includes: Set the displayed text of the previous sentence to a stable state, and determine the triggering conditions satisfied by the displayed text of the previous sentence. When the trigger condition is a modification condition, the displayed text of the previous sentence is modified and displayed based on the modification requirements to obtain the new displayed text of the previous sentence, and the process returns to the step of determining the trigger condition satisfied by the displayed text of the previous sentence. When the trigger condition is a locking condition, the displayed text of the previous sentence is set to a state where correction is prohibited.

[0009] In one embodiment, the step of setting the displayed text of the previous sentence segment to a stable state includes: Obtain the semantic completeness of the previous sentence segment and obtain the previous recognition result of the previous sentence segment. The previous recognition result includes at least: the text of the previous sentence segment, the confidence score, and the stability indicator. Based on the semantic completeness and each of the previous recognition results, it is determined whether the previous sentence segment meets the preset stability conditions. The preset stability conditions include at least one of the following: at least one of the stability flags is a preset stability flag, the text of the previous sentence segment remains unchanged for a preset number of consecutive segments, at least one of the confidence scores reaches a preset confidence threshold, and the semantic completeness reaches a preset completeness threshold. If the previous sentence segment meets the preset stability condition, the displayed text of the previous sentence segment is set to a stable state.

[0010] In one embodiment, the step of determining the triggering condition satisfied by the displayed text of the previous sentence segment includes: Upon receiving the modification request triggered by the user, the triggering condition satisfied by the displayed text of the previous sentence is determined to be the modification condition. If the displayed text of the previous sentence is continuously maintained for a preset configuration duration and / or a confirmation operation triggered by the user is received, the trigger condition satisfied by the displayed text of the previous sentence is determined to be a locking condition.

[0011] In one embodiment, the step of correcting the displayed text of the current sentence segment based on the current speech includes: Based on the current speech, generate hypothetical text and determine the sentence segment where the hypothetical text is located; If the given sentence segment is the current sentence segment and the displayed text of the current sentence segment is in the allowed correction state, determine the character position of the hypothetical text in the current sentence segment; The text corresponding to the character position in the currently displayed text of the sentence segment is corrected to the hypothetical text to obtain the corrected displayed text of the current sentence segment.

[0012] Furthermore, to achieve the above objectives, this application also proposes a voice input device, the device comprising: The acquisition module is used to acquire the current speech dictated by the user and determine whether there is a natural punctuation point between the previous speech and the current speech. The natural punctuation point indicates that the previous speech and the current speech are at different punctuation points. The first determining module is used to determine the displayed text of the current segment where the current speech is located if the natural segment break point does not exist, set the displayed text of the current segment to a state where correction is allowed, correct the displayed text of the current segment based on the current speech, display the corrected displayed text of the current segment and the text of the current speech as the new displayed text of the current segment, and return to execute the step of obtaining the current speech dictated by the user. The second determining module is used to determine the displayed text of the previous segment where the previous speech is located if the natural segment break point exists, set the displayed text of the previous segment break point to a state where correction is prohibited, display the text of the current speech as the displayed text of the current segment break point, and return to execute the step of obtaining the current speech dictated by the user.

[0013] In addition, to achieve the above objectives, this application also proposes a voice input device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the voice input method as described above.

[0014] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and which, when executed by a processor, implements the steps of the voice input method described above.

[0015] This application provides a voice input method, apparatus, device, and storage medium. The method includes: acquiring the current voice spoken by a user, and determining whether there is a natural punctuation point between the previous voice and the current voice, wherein the natural punctuation point indicates that the previous voice and the current voice are at different punctuation points; if there is no natural punctuation point, determining the displayed text of the current punctuation point in which the current voice is located, setting the displayed text of the current punctuation point to a state where correction is allowed, correcting the displayed text of the current punctuation point based on the current voice, displaying the corrected displayed text of the current punctuation point and the text of the current voice as the new displayed text of the current punctuation point, and returning to the step of acquiring the current voice spoken by the user; if there is a natural punctuation point, determining the displayed text of the previous punctuation point in which the previous voice is located, setting the displayed text of the previous punctuation point to a state where correction is prohibited, displaying the text of the current voice as the displayed text of the current punctuation point, and returning to the step of acquiring the current voice spoken by the user.

[0016] This application first continuously acquires the user's current speech while the user is speaking, then determines whether the current speech and the previous speech are in different sentences, i.e., whether there is a natural punctuation mark. If there is no natural punctuation mark, it determines the displayed text of the current punctuation mark to which the current speech belongs, sets the displayed text of the current punctuation mark to an editable state, corrects the displayed text of the current punctuation mark according to the current speech, and displays the corrected displayed text and the text of the current speech mark together as the new displayed text of the current punctuation mark, and returns to step one to continue acquiring the next speech segment. If there is a natural punctuation mark, it determines the displayed text of the previous punctuation mark to which the previous speech belongs, sets the displayed text of the previous punctuation mark to an editable state, displays the text of the current speech mark as the displayed text of the current punctuation mark, and returns to step one to continue acquiring the next speech segment. This application detects natural punctuation points between preceding and following speech in real time during voice input. It only sets the currently written punctuation point to a state where correction is allowed, and once a punctuation point is detected, it switches the previous punctuation point to a state where correction is prohibited and fixes it. Subsequent speech recognition results will not alter the fixed punctuation point. Compared to existing technologies that erase and re-render the entire text whenever a new recognition result arrives, this application can prohibit correction of the previous punctuation point when there is a different punctuation point between the current speech and the previous speech. This allows users to stably review the content they have spoken during continuous dictation without repeatedly repositioning their gaze, thus improving the user experience. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the voice input device structure in the hardware operating environment involved in the embodiments of this application; Figure 2 This is a flowchart illustrating the first embodiment of the voice input method of this application; Figure 3 This is a flowchart illustrating the second embodiment of the voice input method of this application; Figure 4 This is a flowchart illustrating the third embodiment of the voice input method of this application; Figure 5This is a structural block diagram of the voice input device of this application.

[0020] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0021] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0022] Reference Figure 1 , Figure 1 This is a schematic diagram of the voice input device structure in the hardware operating environment involved in the embodiments of this application.

[0023] like Figure 1 As shown, the voice input device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may be connected to a display screen. Optionally, the user interface 1003 may include a standard wired interface or a wireless interface; in this application, the wired interface of the user interface 1003 may be a USB interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be high-speed random access memory (RAM) or non-volatile memory (NVM), such as a disk storage device. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0024] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on the voice input device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0025] like Figure 1 As shown, the memory 1005, which is identified as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a voice input program.

[0026] exist Figure 1In the voice input device shown, the network interface 1004 is mainly used to connect to the backend server and communicate data with the backend server; the user interface 1003 is mainly used to connect to the user equipment; the voice input device calls the voice input program stored in the memory 1005 through the processor 1001 and executes the steps of the voice input method provided in the embodiments of this application.

[0027] It should be noted that speech recognition technology is currently widely deployed in devices such as smartphones, in-vehicle information systems, smart headphones, and conference recording terminals. Especially in scenarios that require the continuous generation of long texts (such as dictating meeting minutes, long instant messaging messages, and inputting navigation addresses and commands in vehicles), voice interaction has become a key way to improve the efficiency of digital content production.

[0028] Existing text presentation methods for continuous speech input generally employ "full-segment streaming transcription and refresh," treating the entire content spoken by the user as a single text block. Relying on the streaming intermediate and final results output by the speech recognition engine, the system continuously generates new recognized text and corrects any previously output words as the user progresses. Then, whenever a new recognition result is generated, the system erases the existing entire segment on the screen and re-displays the latest version of the complete text.

[0029] For example, a user first says "tomorrow afternoon at 3 PM," then adds, "We'll discuss this in the conference room." Because the speech recognition engine can only make immediate guesses based on limited forward acoustic signals at the initial stage of the user's speech, its output of "3 PM" is only a preliminary assumption. As the engine continues to receive subsequent speech signals such as "discuss in the conference room," it dynamically re-scores based on a more complete speech context and language model. It finds that "discussing in the conference room at 4 PM" has higher confidence in semantics and temporal logic. Therefore, the engine determines that the previously output "3 PM" was a misjudgment and actively corrects it to "4 PM." In this correction process, the system doesn't just replace the words "3" and "4," but erases the entire text in the display window containing "tomorrow afternoon at 3 PM" and re-renders a complete new text block containing "tomorrow afternoon at 4 PM." Because each subsequent round of speech input may trigger a global rewrite of the engine's assumptions about several preceding words, the text already displayed to the user on the screen, located in the previous position, frequently shifts in position and changes in content, causing the text to constantly jump and flicker. The user's gaze is repeatedly interrupted, and after each jump, the user needs to reposition and review the previously changed content. It is impossible to reread the content that has been said during the build phase, resulting in a poor user experience.

[0030] Therefore, to address the aforementioned shortcomings, this embodiment provides a voice input method, apparatus, device, and storage medium. This embodiment first continuously acquires the user's current voice while the user speaks, then determines whether the current voice and the previous voice are in different sentences, i.e., whether a natural punctuation point exists. If no natural punctuation point exists, the displayed text of the current punctuation point to which the current voice belongs is determined, and the displayed text of the current punctuation point is set to a correction-allowed state. The displayed text of the current punctuation point is corrected based on the current voice, and the corrected displayed text and the text of the current voice are displayed together as the new displayed text of the current punctuation point. The process then returns to step one to continue acquiring the next voice segment. If the natural punctuation point exists, the displayed text of the previous punctuation point to which the previous voice belongs is determined, and the displayed text of the previous punctuation point is set to a correction-disallowed state. The text of the current voice is displayed as the displayed text of the current punctuation point, and the process then returns to step one to continue acquiring the next voice segment. This embodiment detects natural punctuation points between preceding and following speech in real time during voice input. It only sets the currently being written punctuation point to a state where correction is allowed. Once a punctuation point is detected, the previous punctuation point is switched to a state where correction is prohibited and fixed. Subsequent speech recognition results do not alter the fixed punctuation point. Compared to existing technologies that erase and re-render the entire text whenever a new recognition result arrives, this embodiment prohibits correction of the previous punctuation point when there is a different punctuation point between the current and previous speech. This allows users to stably review what they have said during continuous dictation without repeatedly repositioning their gaze, thus improving the user experience.

[0031] For ease of understanding, the following is combined with Figures 2 to 5 The voice input method provided in the embodiments of this application will be described in detail.

[0032] Reference Figure 2 , Figure 2 This is a flowchart illustrating the first embodiment of the voice input method of this application. The first embodiment of the voice input method of this application is presented as follows: Figure 2 As shown, in this embodiment, the specific method includes: Step S10: Obtain the current speech dictated by the user, and determine whether there is a natural punctuation point between the previous speech and the current speech. The natural punctuation point indicates that the previous speech and the current speech are at different punctuation points.

[0033] It is understood that the method of this embodiment can be applied to any electronic device with data processing, program execution and voice input functions, such as smart glasses, smart bracelets, etc. This embodiment takes a voice input device as an example (hereinafter referred to as the device) to describe this embodiment and the following embodiments.

[0034] It should be noted that the "current speech" mentioned above refers to a segment of speech signal captured by the device at the current moment during the user's continuous dictation. The "previous speech" mentioned above refers to a segment of speech signal captured earlier in time than the current speech. The current speech and the previous speech are adjacent in time.

[0035] It should also be noted that the aforementioned natural punctuation points can refer to the boundary positions in a continuous speech stream used to indicate that two consecutive speech segments belong to different sentences. In this embodiment, natural punctuation points can represent that the previous speech segment and the current speech segment do not belong to the same semantically complete sentence.

[0036] It should be emphasized that the criteria for determining natural punctuation points in this embodiment may include speech pauses, intonation and rhythm boundaries, punctuation assumptions output by the recognition system, and relatively complete semantic boundaries, etc. This embodiment does not impose any limitations on these criteria.

[0037] For example, if a user says "3 PM today" followed by a long period of silence, and then says "We're discussing this in the meeting room," that silence can be used as one of the criteria for determining a natural punctuation point. Similarly, if a user says "The plan has been finalized" with a noticeable drop in tone, and then says "Next, we'll discuss the budget," that drop in tone can be used as one of the criteria for determining a natural punctuation point.

[0038] In its implementation, the device captures the user's voice signal in real time via a microphone during continuous dictation. The captured voice signal is then streamed, and the currently captured voice segment is taken as the current voice. Simultaneously, the device can acquire the preceding voice segment that is chronologically adjacent to the current voice segment.

[0039] Next, after acquiring the current speech, the device compares it with the previous speech to determine if there is a natural punctuation mark between them. If a natural punctuation mark is found, the current and previous speech belong to different sentences. If no natural punctuation mark is found, the current and previous speech belong to the same sentence.

[0040] To facilitate understanding, the following examples are provided for illustration, but they do not impose specific limitations on this embodiment. For instance, in a meeting minutes speaking scenario, a user repeatedly says, "We believe this solution is feasible (pause 0.8 seconds). Next, we need to discuss the budget issue." The device can capture the previous speech "We believe this solution is feasible" and the current speech "Next, we need to discuss the budget issue." The device can detect whether there is a natural pause between the two speech segments.

[0041] Step S20: If there is no natural punctuation point, determine the displayed text of the current punctuation where the current speech is located, set the displayed text of the current punctuation to the correction-allowed state, correct the displayed text of the current punctuation based on the current speech, use the corrected displayed text of the current punctuation and the text of the current speech as the new displayed text of the current punctuation and display it, and return to the step of obtaining the current speech dictated by the user.

[0042] Understandably, the aforementioned current sentence segment can refer to a sentence unit to which the current speech belongs that has not yet been determined to be completely finished. In terms of text composition, the aforementioned current sentence segment can include previously displayed text on the screen as well as newly identified text.

[0043] It is also understood that the aforementioned displayed text can refer to the text content that has been output by the device and presented on the display interface in the current sentence segmentation. Furthermore, in this embodiment, the displayed text can be the recognition result that the device has already shown to the user before the current moment.

[0044] It should be understood that the aforementioned "allowed correction state" refers to a state in which the text content of a sentence is allowed to be corrected by subsequent speech recognition results. In this embodiment, the text content of a sentence in the "allowed correction state" can be modified by newly arrived speech signals, including but not limited to the replacement, deletion, or addition of existing words.

[0045] It should also be understood that the above-mentioned correction may refer to the operation of updating the already displayed text in the current sentence segment based on the recognition result of the newly arrived current speech. Similarly, in this embodiment, the correction may include replacing previously recognized words, such as correcting "three dots" to "four dots"; it may also include appending newly recognized text to the existing text.

[0046] It should be noted that the aforementioned new displayed text can refer to the entire text content of the updated version obtained after the current sentence segmentation has been corrected. The new displayed text may include the corrected old text and the newly identified text, constituting the complete text representation of the current sentence segmentation at the current moment.

[0047] In its implementation, the device determines the current sentence segment of the current speech if there is no natural punctuation between the previous and current speech segments. The device can then retrieve the displayed text of the current sentence segment already shown on the screen at the current moment. This displayed text can then be set to a state where correction is permitted.

[0048] When correction is enabled, the device can correct the displayed text of the current sentence based on the recognition result of the current speech, and then perform recognition based on the acoustic signal and language model of the current speech to obtain the recognized text corresponding to the current speech. The recognized text of the current speech can then be combined with the existing displayed text of the current sentence. The text in correction-enabled mode may flash or jump to prompt the user.

[0049] During the combination process, the device can use the current speech recognition results to correct previous assumptions with low confidence in the existing displayed text, and replace the corresponding parts of the existing displayed text with the corrected content. The corrected complete text is then used as the new displayed text for the current sentence segment. Finally, the new displayed text can be output to the display interface for display. This display replaces the original displayed text for the current sentence segment, allowing the display interface to present the latest complete text content for the current sentence segment. After completing the above display operation, the device returns to the step of acquiring the user's current dictated speech and continues to collect subsequent speech signals from the user for the next round of processing.

[0050] To facilitate understanding, the following examples are provided for illustration, but they do not impose specific limitations on this embodiment. Continuing from the previous example, in a scenario where a user dictates meeting minutes, they first say, "We believe this solution is feasible." The device can then display this text as the previously punctuated text and set it to a state where correction is disabled.

[0051] The user then continues, "Next, we need to discuss the budget issue." The device captures this segment of speech as the current speech and determines that there is a natural punctuation point between the previous speech, "We believe this solution is feasible," and the current speech, "Next, we need to discuss the budget issue." The user then continues, "The budget issue mainly includes three aspects." The device captures this phrase as the new current speech and determines that there is no natural punctuation point between the current speech, "Next, we need to discuss the budget issue," and the new current speech, "The budget issue mainly includes three aspects," indicating that both belong to the same punctuation point.

[0052] The device can then determine that the current sentence segment is "Next, we need to discuss the budget issue, which mainly includes three aspects", obtain the currently displayed text "Next, we need to discuss the budget issue", and set the displayed text to a state where correction is allowed.

[0053] After receiving the subsequent speech "mainly includes three aspects," the speech recognition engine dynamically re-scored its previous recognition assumptions based on a more complete speech context. Finding that "budget" in "discussing budget issues" had low confidence, it corrected it to "cost." The device can then correct the already displayed text "Next, we need to discuss budget issues" based on the current speech, replacing "budget" with "cost," resulting in the corrected displayed text "Next, we need to discuss cost issues." This corrected text can then be merged with the current speech text "mainly includes three aspects," resulting in the new displayed text "Next, we need to discuss cost issues, mainly including three aspects," which is then displayed in the current sentence segment's display area. The device then returns to the step of acquiring the user's current spoken speech.

[0054] Step S30: If the natural punctuation point exists, determine the displayed text of the previous punctuation point where the previous speech is located, set the displayed text of the previous punctuation point to a state where correction is prohibited, display the text of the current speech as the displayed text of the current punctuation point, and return to the step of obtaining the current speech dictated by the user.

[0055] It should be noted that the aforementioned "previous sentence segment" can refer to the sentence to which the previous speech belonged. In this implementation, the previous sentence segment can be located before the current sentence segment in chronological order. The aforementioned "no correction" state can refer to the state of the displayed text of the previous sentence segment. In this state, the displayed text can no longer accept any corrections or rewritings from subsequent speech recognition results; that is, the displayed text is locked as the current content. The aforementioned "displayed text of the current sentence segment" can refer to the text content that the current sentence segment is first presented to the user on the display interface, and this text content comes from the recognition result corresponding to the current speech.

[0056] In its implementation, the device determines that the previous and current speech segments belong to different sentence units when a natural punctuation point exists between them. Furthermore, the device can identify the previous sentence segment to which the previous speech segment belongs and retrieve the text of that segment that has already been displayed to the user on the screen.

[0057] Next, the device can set the displayed text of the previous sentence to a state where correction is disabled. In this state, the displayed text of the previous sentence will no longer accept any corrections or rewriting from subsequent speech recognition results, and subsequent recognition results will not affect the content of the previous sentence. The text in the state of disabling correction will remain displayed continuously, flashing or jumping.

[0058] Subsequently, the device can use the recognition result corresponding to the current speech as the text content, determine this text content as the displayed text of the current sentence segment, and display it in the display area corresponding to the current sentence segment on the display interface. During the above display process, the displayed text of the previous sentence segment remains in its original display position with a fixed style, and its content and position do not change. After the display is completed, the device returns to the step of obtaining the user's current speech and continues to process the next segment of collected speech.

[0059] To facilitate understanding, the following explanation uses examples, but does not impose specific limitations on this embodiment. Continuing the previous example, the device has determined that the current sentence segment is "Next, we need to discuss the cost issues, which mainly include three aspects," and there is no natural punctuation point between the previous speech and the current speech. Subsequently, the user continues to utter "The first point is personnel configuration," and the device can then acquire "The first point is personnel configuration" as the new current speech, and determine that there is a natural punctuation point between the current speech "Next, we need to discuss the cost issues, which mainly include three aspects" and the new current speech "The first point is personnel configuration," because the semantics of "three aspects" are complete and the new topic "the first point" appears subsequently. The device can then determine the previous sentence segment to which the previous speech "Next, we need to discuss the cost issues, which mainly include three aspects" belongs, and acquire the displayed text of that previous sentence segment, "Next, we need to discuss the cost issues, which mainly include three aspects."

[0060] Next, the device can set the displayed text of the previous sentence to a non-correction state; the text is locked and will no longer accept any subsequent corrections. Then, the device will display the text of the current speech, "The first point is personnel configuration," as the displayed text of the current sentence, starting on a new line in the display interface. At this time, in the display interface, the previous sentence, "Next, we need to discuss cost issues, mainly including three aspects," is displayed stably in a fixed style, with its position and content remaining unchanged; the new current sentence, "The first point is personnel configuration," is displayed on the next line in a writing style. The device then returns to the step of obtaining the user's current spoken speech.

[0061] This embodiment detects natural punctuation points between preceding and following speech in real time during voice input. It only sets the currently being written punctuation point to a state where correction is allowed. Once a punctuation point is detected, the previous punctuation point is switched to a state where correction is prohibited and fixed. Subsequent speech recognition results do not alter the fixed punctuation point. Compared to existing technologies that erase and re-render the entire text whenever a new recognition result arrives, this embodiment prohibits correction of the previous punctuation point when there is a different punctuation point between the current and previous speech. This allows users to stably review what they have said during continuous dictation without repeatedly repositioning their gaze, thus improving the user experience.

[0062] Furthermore, considering that during the correction process, if the entire sentence of the current punctuation is modified again, some text that does not need to be modified will still be erased and re-displayed, affecting the user's viewing experience, therefore, in this embodiment, the step of correcting the displayed text of the current punctuation based on the current speech includes: Step S21: Generate hypothetical text based on the current speech and determine the sentence segment where the hypothetical text is located.

[0063] It should be noted that the hypothetical text mentioned above can refer to the candidate recognized text result obtained by the device after performing speech recognition processing based on the current speech. During the speech recognition process, due to the ambiguity of the speech signal and the uncertainty of the context, the initial recognition result output by the device after acquiring the current speech through the speech recognition engine is the hypothetical text.

[0064] It is important to emphasize that the hypothetical text mentioned above can be the optimal result among one or more candidate recognition results obtained by the speech recognition engine after decoding the current speech based on acoustic and language models. The sentence segment mentioned above refers to the sentence to which the hypothetical text belongs in the continuous speech stream; that is, which sentence segment the hypothetical text should be assigned to. During the speech input construction phase, the device can divide the continuous speech stream into multiple sentence-level structural units using natural sentence break points as boundaries. The sentence segment to which the hypothetical text belongs can be used to indicate which sentence-level unit the hypothetical text should be included in for display and management.

[0065] In its implementation, the device captures the user's spoken voice through a microphone and then sends the voice to a speech recognition engine for processing. The speech recognition engine decodes the voice based on acoustic and language models to generate corresponding candidate text results, which is the hypothetical text mentioned above.

[0066] After acquiring the hypothetical text, the device can further determine the sentence segment to which the hypothetical text belongs in the continuous speech stream. The process of determining the sentence segment to which the hypothetical text belongs may include: the device acquiring the position information of the current speech in the time sequence and the boundary information of the existing sentence-level structural units; based on whether the preceding speech of the current speech has been determined to have a natural sentence break and the boundary position of the current sentence break, it determines whether the hypothetical text should belong to an existing current sentence break or a newly generated sentence break. If there is no natural sentence break between the previous speech and the current speech, the hypothetical text belongs to an existing current sentence break; if there is a natural sentence break between the previous speech and the current speech, the hypothetical text belongs to a newly generated sentence break.

[0067] To facilitate understanding, the following example illustrates the concept, but does not impose specific limitations on this embodiment. Continuing from the previous example, the user continues to utter "The second point is the schedule." The device acquires the current speech "The second point is the schedule," sends it to the speech recognition engine, and the engine decodes the speech and outputs the hypothetical text "The second point is the schedule." Simultaneously, the device determines whether there is a natural punctuation point between the previous speech "The first point is personnel configuration" and the current speech "The second point is the schedule." Since there is a clear semantic boundary after "The first point is personnel configuration" and "The second point" is a new parallel topic, the device determines that there is a natural punctuation point between them. Based on the above determination, the device determines that the punctuation point of the hypothetical text "The second point is the schedule" is a newly generated punctuation point, rather than an existing previous punctuation point. Subsequently, the device displays this hypothetical text as the displayed text of the new punctuation point on a new line.

[0068] Step S22: If the current sentence is the segmented sentence and the displayed text of the current sentence is in the allowed correction state, determine the character position of the hypothetical text in the current sentence. Step S23: Correct the text corresponding to the character position in the currently displayed text of the sentence segment to the hypothetical text, and obtain the corrected displayed text of the current sentence segment.

[0069] Understandably, the aforementioned character position can refer to the specific position or range that each character in the hypothetical text should occupy in the displayed text of the corresponding sentence. The character position can be used to indicate the insertion point, replacement start point, or replacement end point of the hypothetical text in the displayed text of the current sentence; that is, the device can determine, based on the character position, where the hypothetical text should be placed in the displayed text of the current sentence.

[0070] In this embodiment, the character position is related to the order of the characters and can be represented by a character number or offset. For example, the character position of a hypothetical text is "after the 5th character" or "between the 3rd and 6th characters".

[0071] In its implementation, the device, upon determining that the segment containing the hypothetical text is the current segment, can further obtain the current state of the displayed text within that current segment. The device then determines whether the displayed text of the current segment is in a state where correction is permitted. When the device determines that the displayed text of the current segment is in a state where correction is permitted, it can determine the character position of the hypothetical text within the displayed text of the current segment.

[0072] Specifically, the foregoing process of determining a character position may include: a device may determine, according to a timestamp position of current speech in a time sequence relative to existing displayed text, and an alignment relationship between hypothesis text output by a speech recognition engine and the displayed text, a specific position where the hypothesis text should be placed in the displayed text. The foregoing character position may include a start position and an end position, or may include a target position to be inserted. For example, if the hypothesis text is a replacement result for some words in the displayed text, the character position is a start character order and an end character order of the part of words in the displayed text; if the hypothesis text is newly added content, the character position is a character order of a to-be-inserted position of the newly added content in the displayed text. After determining the character position, the device performs corresponding correction processing on the currently segmented displayed text according to the character position, including replacement, insertion or deletion operations.

[0073] Considering that the hypothesis text output by the speech recognition engine may involve correction of words at any position in the currently segmented displayed text, in the process of determining the character position, the character position corresponding to the hypothesis text may be a position for replacing existing words in the displayed text, or may be a position for inserting between existing words in the displayed text.

[0074] For ease of understanding, the following description is provided through an example, which does not impose specific limitations on this embodiment. Continuing the previous example, the device has determined that the sentence where the hypothesis text "the second point is the schedule" is located is a newly generated sentence, and uses the hypothesis text as the displayed text of the new sentence. Subsequently, the user continues to dictate "the schedule is on Wednesday morning", after obtaining the current speech "the schedule is on Wednesday morning", the device generates the hypothesis text "the schedule is on Wednesday morning", the device determines that the sentence where the speech is located is the current sentence where the previous sentence "the second point is the schedule" is located, and the displayed text "the second point is the schedule" of the current sentence is in an allowable correction state. The device aligns the hypothesis text "the schedule is on Wednesday morning" with the displayed text "the second point is the schedule", and determines that the hypothesis text should be inserted at the end of the displayed text. The device determines that the character position is a position after the 10th character (that is, the character "安排" (schedule)) of the displayed text "the second point is the schedule". Then the hypothesis text "the schedule is on Wednesday morning" can be inserted after the character "排" (arrange) in the displayed text "the second point is the schedule" of the current sentence, to obtain the corrected displayed text "the second point is the schedule is on Wednesday morning".

[0075] If the hypothesis text is for replacing an existing word, for example, the device generates the hypothesis text "Thursday" and determines that the corresponding character position is the position of "Wednesday" in the displayed text "the second point is the schedule is on Wednesday morning", then the device replaces "Wednesday" with "Thursday", to obtain the corrected displayed text "the second point is the schedule is on Thursday morning".

[0076] Since this embodiment only corrects and displays the text corresponding to the character position, other text outside the character position in the currently displayed text of the sentence segment remains unchanged and does not need to be erased and redisplayed. Therefore, during the voice input construction period, only the corrected portion of the displayed content of the current sentence segment changes, while the unmodified text remains stable, reducing text flickering and jitter, and further improving the user's viewing experience during continuous speaking.

[0077] Reference Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the voice input method of this application. Based on the first embodiment described above, a second embodiment of the voice input method of this application is proposed.

[0078] To accurately determine whether there is a natural punctuation point between the previous and current speech, such as Figure 3 As shown, in this embodiment, the step of determining whether there is a natural punctuation mark between the previous speech and the current speech includes: Step S11: Extract features from the current speech to obtain cue features. The cue features include at least two of the following: speech pause features, intonation features, punctuation assumption features, and semantic boundary features.

[0079] It should be noted that the aforementioned speech pause features refer to the acoustic features related to the speaker's pauses extracted from the current speech. These speech pause features may include silence duration and pause pattern. Silence duration refers to the length of time there is no speech signal between the current speech and the previous speech; pause pattern refers to the regularity and distribution of the location of pauses, such as the distinction between sentence-ending pauses and sentence-intermediate pauses.

[0080] It should also be noted that the aforementioned intonation trend features can refer to the prosodic features related to the speaker's pitch changes extracted from the current speech. These intonation trend features include the rise and fall of intonation and sentence-end prosodic features. For example, a drop in intonation at the end of a sentence usually indicates the end of a sentence.

[0081] Understandably, the aforementioned punctuation hypothesis features can refer to the hypothesis information related to punctuation marks output by the speech recognition system during the recognition process, such as the confidence level or candidate punctuation types of punctuation marks such as periods, commas, and question marks inserted by the recognition system in the recognition results.

[0082] It is also understood that the aforementioned semantic boundary features can refer to features related to semantic integrity extracted from the recognized text of the current speech. These semantic boundary features can be used to characterize whether the text content of the current speech has reached a relatively complete boundary at the semantic level, such as whether the subject, predicate, and object of a sentence are complete.

[0083] Step S12: The cue features are weighted and fused according to preset weights to obtain a comprehensive score, and the existence of a natural punctuation point between the previous speech and the current speech is determined based on the comprehensive score.

[0084] It should be understood that the aforementioned preset weights may refer to the combination coefficients pre-assigned to each type of clue feature. These preset weights can be used to reflect the different degrees of influence of different clue features on the sentence segmentation judgment results during the fusion process.

[0085] It should also be understood that the above comprehensive score may refer to the numerical result obtained by the device after weighting and fusing multiple cue features according to preset weights. The above comprehensive score is used to quantitatively represent the comprehensive confidence that there is a natural punctuation point between the current speech and the previous speech.

[0086] In its implementation, after acquiring the current speech, the device extracts features from the speech to obtain cue features for sentence segmentation determination. These cue features include at least two of the following: speech pause features, intonation features, punctuation assumption features, and semantic boundary features.

[0087] The device can extract acoustic parameters from the current speech, obtaining silence duration and pause patterns as speech pause features; it can analyze the fundamental frequency changes of the current speech, obtaining the rise and fall trend of intonation and the prosodic features at the end of sentences as intonation direction features; it can obtain the punctuation candidate and their confidence level output by the speech recognition system during the recognition of the current speech as punctuation hypothesis features; and it can perform semantic analysis on the recognized text of the current speech, determining whether the text has grammatical and semantic integrity as semantic boundary features.

[0088] After extracting the aforementioned cue features, the device can obtain pre-stored preset weights corresponding to various cue features. The extracted cue features are then weighted and fused according to these preset weights. The quantized value of each cue feature is multiplied by its corresponding preset weight, and the products are summed to obtain a comprehensive score. This comprehensive score is then compared with a pre-configured sentence segmentation threshold. Based on the comparison result, it is determined whether a natural break exists between the previous and current speech. If the comprehensive score reaches or exceeds the sentence segmentation threshold, the device determines that a natural break exists between the previous and current speech; if the comprehensive score is below the sentence segmentation threshold, the device determines that no natural break exists between the previous and current speech.

[0089] To facilitate understanding, the following example is used for explanation, but it does not impose specific limitations on this embodiment. Continuing from the previous example, after the user utters "The second point is that the time is scheduled for Thursday morning," there is a 0.6-second pause, followed by "The third point is the budget." The device can acquire the current speech "The third point is the budget" and extract features from it. The extracted features include a 0.6-second silence duration as the pause feature, a drop in intonation at the end of the previous speech, a punctuation hypothesis feature (the recognition system outputs a period candidate with a confidence level of 0.85), and a semantic boundary feature (the recognized text of the previous speech, "The second point is that the time is scheduled for Thursday morning," has a complete subject-verb-object structure).

[0090] Next, preset weights are obtained, as follows: the weight of the pause feature can be 0.3, the weight of the intonation feature can be 0.2, the weight of the punctuation assumption feature can be 0.25, and the weight of the semantic boundary feature can be 0.25. Then, weighted fusion can be performed according to the above preset weights, and the comprehensive score can be calculated as 0.6×0.3+0.85×0.2+0.85×0.25+0.9×0.25, where the quantization value of the semantic boundary feature is 0.9, and the comprehensive score is calculated as 0.18+0.17+0.2125+0.225=0.7875. The preset sentence segmentation threshold is 0.7. When the comprehensive score of 0.7875 reaches the sentence segmentation threshold, the device determines that there is a natural break between the previous speech "The second point is that the time is scheduled for Thursday morning" and the current speech "The third point is the budget".

[0091] Considering that the weights of each clue feature and the sentence segmentation gate threshold in the multi-clue fusion process affect the sensitivity and accuracy of sentence segmentation determination, the aforementioned preset weights and sentence segmentation gate thresholds are configurable parameters. The device can configure different preset weights and sentence segmentation gate thresholds for the aforementioned clue features according to different application scenarios. For example, in an in-vehicle voice scenario, due to higher environmental noise and reduced reliability of silence duration, the weight of the speaking pause feature can be reduced and the sentence segmentation gate threshold adjusted accordingly; in a quiet meeting recording scenario, the reliability of each clue feature is higher, and the default weight configuration can be used. This embodiment does not impose any limitations on this.

[0092] Furthermore, in order to improve the accuracy of natural punctuation point judgment, the cue features in this embodiment include the speech pause features, intonation features, punctuation assumption features, and semantic boundary features of the current speech. The step of extracting features from the current speech to obtain cue features includes: Step S111: Perform speech activity detection on the current speech, detect the silence duration of the silent segment in the current speech, and generate the speech pause feature based on the silence duration.

[0093] It should be explained that the aforementioned voice activity detection refers to the device's processing of the acquired voice signal to determine whether or not there is voice activity, which can be used to distinguish between voice segments and non-voice segments (i.e., silence segments). In this embodiment, the purpose of voice activity detection is to identify from a continuous audio stream which time periods contain the user's voice signal and which time periods do not.

[0094] It should also be noted that the aforementioned silent segment can refer to the time interval determined by speech activity detection to be free of speech activity, that is, the time segment in a continuous speech stream where the speech signal energy is lower than a preset threshold.

[0095] The aforementioned silence duration refers to the duration of the silence segment on the time axis, which can be measured in milliseconds or seconds in this embodiment. During continuous voice input, silence segments typically appear during pauses in the user's speech, and their duration is one of the important criteria for judging speech pauses.

[0096] It is understood that the aforementioned speech pause features can refer to feature information generated based on the silence duration to characterize the pauses a user makes during speech. These speech pause features may include the silence duration itself and pause pattern information derived from the silence duration.

[0097] In its specific implementation, after acquiring the current speech, the aforementioned device can perform speech activity detection. The speech activity detection process may include: the device performing frame-segmentation on the audio signal of the current speech, dividing the continuous audio signal into multiple short frames; then calculating acoustic parameters such as signal energy or zero-crossing rate for each short frame; and then comparing the acoustic parameters of each short frame with a preset speech activity detection threshold. If the acoustic parameters reach or exceed the speech activity detection threshold, the short frame is determined to be a speech frame; if the acoustic parameters are below the speech activity detection threshold, the short frame is determined to be a silence frame.

[0098] After detecting all short frames, the device identifies the time interval formed by consecutive silence frames in the current speech based on the detection results. This time interval is the silence segment. The device then obtains the start and end times of the silence segment on the time axis and calculates its duration based on these times. This duration is the silence duration. Finally, a speech pause feature is generated based on the calculated silence duration.

[0099] The process of generating speech pause features based on silence duration described above may include: the device directly using the numerical value of silence duration as the quantified value of the speech pause feature; or, the device matching the silence duration with multiple preset duration intervals, determining the corresponding pause type (e.g., short pause, medium pause, long pause) based on the duration interval to which the silence duration belongs, and using the pause type as a component of the speech pause feature. This embodiment does not impose any limitations on this.

[0100] To facilitate understanding, the following example is used for explanation, but it does not impose specific limitations on this embodiment. Continuing from the previous example, after the user says "The second point is that the time is scheduled for Thursday morning," there is a pause, followed by "The third point is the budget." The device acquires the current speech "The third point is the budget," divides the current speech into multiple short frames, and calculates the energy value of each short frame. The energy value of each short frame is compared with a preset speech activity detection threshold, detecting a continuous silence frame before the current speech. The time interval formed by the above continuous silence frames is identified as a silence segment, and the start and end times of the silence segment are acquired, calculating the duration of the silence segment to be 0.6 seconds. 0.6 seconds is used as the silence duration, and a speech pause feature is generated based on this silence duration. Alternatively, the device can use the 0.6-second silence duration as the quantization value of the speech pause feature, and simultaneously match 0.6 seconds with a preset duration interval to determine that the silence duration belongs to a long pause interval, generating a speech pause feature containing "long pause" type information.

[0101] Step S112: Extract the fundamental frequency of the speech frame preceding the silence segment in the current speech, and generate the intonation trend feature based on the fundamental frequency of the speech frame.

[0102] It should be understood that the fundamental frequency of the aforementioned speech frame can refer to the fundamental frequency parameter extracted from the speech frame preceding the silence segment in the current speech. The fundamental frequency can refer to the frequency of vocal cord vibration, and in this embodiment, it can be measured in Hertz. Since the fundamental frequency is the physical quantity corresponding to pitch in human auditory perception, the fundamental frequency change pattern at the end of a sentence (such as a decrease or increase in fundamental frequency) can be used to determine the tone direction.

[0103] It should also be understood that the aforementioned speech frame can refer to each speech signal unit obtained after framing the current speech. The speech frame preceding the silence segment refers to one or more speech frames that are immediately adjacent to the silence segment in time sequence, because the speech content corresponding to this speech frame can be located at the end of the sentence before the user pauses.

[0104] It is understood that the above-mentioned intonation trend features may refer to the feature information extracted from the fundamental frequency of the speech frame to characterize the speaker's intonation change trend. In this embodiment, the intonation trend features may include the rise and fall trend of intonation and the prosodic features at the end of the sentence. For example, a downward trend in the fundamental frequency at the end of the sentence usually indicates the end of a declarative sentence, while an upward trend in the fundamental frequency at the end of the sentence usually indicates an interrogative sentence or an incomplete semantic.

[0105] In its implementation, after detecting a silence segment in the current speech, the device can determine the speech signal portion preceding the silence segment. The device can then extract the fundamental frequency of the speech frame from this preceding speech signal portion.

[0106] The process of extracting the fundamental frequency of a speech frame described above may include: the device performing frame segmentation on the speech signal portion preceding the silence segment to obtain multiple speech frames; performing fundamental frequency detection on each speech frame, calculating the fundamental frequency value of each speech frame using fundamental frequency detection algorithms such as autocorrelation or cepstral method; then smoothing the extracted fundamental frequency values ​​of multiple speech frames to remove abnormal fundamental frequency transition points. After completing the extraction of the fundamental frequency of the speech frames, intonation characteristics are generated based on the aforementioned fundamental frequency of the speech frames.

[0107] The process of generating intonation trend features based on the fundamental frequency of speech frames described above may include: the device can perform trend analysis on the fundamental frequency values ​​of multiple speech frames preceding the silence segment to determine the direction of change of the fundamental frequency in the time series. If the fundamental frequency shows a decreasing trend from high to low, an intonation trend feature indicating a decreasing intonation at the end of the sentence is generated; if the fundamental frequency shows an increasing trend from low to high, an intonation trend feature indicating an increasing intonation at the end of the sentence is generated; if the fundamental frequency change is not obvious, an intonation trend feature indicating a flat intonation is generated. Finally, the device can also calculate the difference between the fundamental frequency value of the last speech frame preceding the silence segment and the average fundamental frequency value, and use this difference as part of the intonation trend feature.

[0108] To facilitate understanding, the following example is used for explanation, but it does not impose specific limitations on this embodiment. Continuing from the previous example, after detecting the silence segment (0.6 seconds) before the current speech "The third point is the budget," the device extracts the fundamental frequency of the speech frame preceding that silence segment, that is, extracts the fundamental frequency of the speech frame from the end of the previous speech "The second point is that the time is scheduled for Thursday morning." The speech signal portion preceding the silence segment is divided into frames, and fundamental frequency detection is performed on each frame to obtain a series of fundamental frequency values. The device performs trend analysis on this series of fundamental frequency values ​​and finds that the fundamental frequency gradually decreases from 210Hz at the beginning to 165Hz at the end, showing a clear downward trend. Based on the above fundamental frequency downward trend, the device generates intonation characteristics, indicating that the sentence-ending intonation of the previous speech is falling, that is, the intonation at "Thursday morning" is the ending intonation of a declarative sentence.

[0109] Step S113: Determine the punctuation classification probability vector of the current speech, and generate the punctuation hypothesis features based on the punctuation classification probability vector.

[0110] It should be explained that the above-mentioned punctuation classification probability vector can refer to the probability distribution vector obtained by the device after performing punctuation classification processing on the current speech or the recognized text corresponding to the current speech. This probability distribution vector can be used to represent the probability distribution of various punctuation marks corresponding to the text boundary positions of the current speech.

[0111] In this embodiment, each dimension of the punctuation classification probability vector can correspond to a punctuation mark type (e.g., period, comma, question mark, exclamation mark, or empty punctuation). The value of each dimension can represent the probability value of that punctuation mark type being true, and the sum of the probability values ​​of all dimensions is 1.

[0112] It should also be explained that the aforementioned punctuation hypothesis features can refer to information generated based on punctuation classification probability vectors, used to characterize the speech recognition system's tendency towards punctuation marks in the current speech. These punctuation hypothesis features may include punctuation types and their corresponding confidence levels, such as the period hypothesis output by the recognition system and its confidence level.

[0113] In its implementation, after acquiring the current speech or its corresponding recognized text, the device can perform punctuation classification processing on the current speech or its recognized text. This punctuation classification processing can be completed locally on the device or on a cloud server, with the results then sent back to the device.

[0114] The above-mentioned punctuation classification process may include: the device inputs the recognized text corresponding to the current speech into a pre-trained punctuation classification model, which can be built based on a deep neural network and used to classify punctuation types at the boundary positions of the recognized text; the punctuation classification model encodes the recognized text, extracting semantic and contextual features; the punctuation classification model calculates the classification probability of each type of punctuation mark corresponding to the boundary positions of the recognized text based on the extracted features, and outputs a punctuation classification probability vector; or, the device directly performs punctuation classification on the acoustic features of the current speech, and outputs a punctuation classification probability vector through an acoustic-punctuation joint model. After obtaining the punctuation classification probability vector, the punctuation type corresponding to the maximum probability value can be extracted from the punctuation classification probability vector as the target punctuation type, and the maximum probability value can be used as the confidence level of the punctuation type. Then, punctuation hypothesis features are generated based on the target punctuation type and its confidence level. The punctuation hypothesis features may include information in the form of "period, confidence level 0.85" or "comma, confidence level 0.70".

[0115] To facilitate understanding, the following example is used for explanation, but it does not impose specific limitations on this embodiment. Continuing from the previous example, the device can obtain the recognized text corresponding to the current speech "The third point is the budget" as "The third point is the budget". Then, the recognized text is input into a pre-trained punctuation classification model. The punctuation classification model encodes the recognized text, calculates the classification probability of various punctuation marks at the text boundary positions (i.e., the positions after "budget"), and outputs a punctuation classification probability vector, whose dimensions correspond to five categories: period, comma, question mark, exclamation mark, and empty punctuation. For example, this punctuation classification probability vector is [period: 0.88, comma: 0.05, question mark: 0.02, exclamation mark: 0.03, empty punctuation: 0.02]. Then, the maximum probability value of 0.88 can be extracted from the punctuation classification probability vector to determine the corresponding punctuation type as a period. This maximum probability value of 0.88 is the confidence level of the period hypothesis. Finally, based on the above period type and its confidence level of 0.88, the punctuation hypothesis feature is generated, resulting in the punctuation hypothesis feature "period, confidence level 0.88".

[0116] Step S114: Determine the semantic vector of the current sentence segment where the current speech is located, determine the similarity between the semantic vector and the preset complete statement template vector, and generate the semantic boundary feature based on the similarity.

[0117] It should be noted that the aforementioned semantic vector can refer to a numerical vector obtained by vectorizing the text content of the current sentence segment containing the current speech. This semantic vector can be used to represent the semantic information of the text in mathematical space, and in this embodiment, it can exist in the form of a fixed-dimensional floating-point vector.

[0118] It should also be noted that the aforementioned preset complete statement template vector can refer to a pre-constructed template vector used to represent the semantic features of a complete statement. The preset complete statement template vector can be obtained statistically based on the semantic vectors of complete statements in a large-scale corpus. For example, after encoding a large number of complete statement texts, the average vector can be taken as the template vector. Of course, it can also be obtained in other ways, and this embodiment does not limit it.

[0119] Understandably, the aforementioned similarity can refer to the degree of similarity between a semantic vector and a preset complete statement template vector, and can be measured by the distance between vectors (e.g., cosine distance or Euclidean distance) or the cosine value of the angle between vectors. In this embodiment, the numerical range of the similarity can be between 0 and 1, with a higher value indicating that the semantics of the current sentence segment is closer to the semantics of the complete statement.

[0120] It is also understood that the aforementioned semantic boundary features can refer to feature information generated based on similarity, used to characterize whether the current sentence segment has reached a relatively complete boundary at the semantic level. Semantic boundary features can be used to indicate the degree of semantic integrity of the current sentence segment.

[0121] In its specific implementation, after obtaining the displayed text of the current sentence in which the current speech is located, the device performs semantic encoding on the displayed text of the current sentence to generate a semantic vector of the current sentence.

[0122] The aforementioned semantic encoding process may include: the device inputting the currently displayed text of the segmented sentence into a pre-trained semantic encoding model (e.g., an encoder based on the Transformer architecture). The semantic encoding model performs word segmentation and sequence encoding on the input text, extracts semantic information from the text, and outputs a fixed-dimensional semantic vector. After generating the semantic vector of the current segmented sentence, the device can obtain a pre-stored preset complete statement template vector. This preset complete statement template vector is stored in the device's internal memory and can be pre-calculated by averaging all encoding results after semantically encoding a large number of complete statement texts.

[0123] Next, the device can calculate the similarity between the semantic vector of the current sentence segment and the preset complete statement template vector. This similarity can be calculated by taking the cosine similarity between the two vectors, which is obtained by dividing the dot product of the two vectors by the product of their magnitudes. Finally, semantic boundary features are generated based on this similarity. For example, the calculated similarity value can be directly used as the quantified value of the semantic boundary feature; a higher similarity indicates that the semantics of the current sentence segment is closer to a complete statement, meaning it is more likely to form a natural sentence boundary at the semantic level.

[0124] To facilitate understanding, the following explanation uses examples, but does not impose specific limitations on this embodiment. Continuing from the previous example, the device acquires the displayed text of the currently segmented sentence "The third point is the budget." The device can input the displayed text into a pre-trained semantic encoding model. The semantic encoding model encodes the text "The third point is the budget" and outputs a semantic vector of the text. This semantic vector can be a 256-dimensional floating-point vector. Next, a pre-stored preset complete statement template vector can be acquired. This template vector is also a 256-dimensional floating-point vector, representing the average semantic features of the complete statement.

[0125] The cosine similarity between the semantic vector and the preset complete statement template vector is then calculated, yielding a similarity of 0.87. Based on this similarity of 0.87, semantic boundary features are generated. This similarity value of 0.87 is used as the quantification value of the semantic boundary features. This quantification value indicates that the current sentence segment "The third point is the budget" has high semantic completeness, meaning it is semantically close to a complete statement. If the current sentence segment is an incomplete expression, such as "The third point is," its semantic vector may only have a similarity of 0.45 with the preset complete statement template vector, indicating that the sentence segment is semantically incomplete and does not belong to a complete sentence boundary.

[0126] Reference Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the voice input method of this application. Based on the above embodiments, the third embodiment of the voice input method of this application is proposed.

[0127] Furthermore, considering that if natural punctuation marks exist, directly setting the displayed text of the previous punctuation mark to a state where correction is disabled would prevent users from making changes. Therefore, if... Figure 4 As shown, in this embodiment, the step of setting the displayed text of the previous sentence segment to a state where correction is disabled includes: Step S31: Set the displayed text of the previous sentence to a stable state, and determine the triggering conditions satisfied by the displayed text of the previous sentence.

[0128] It should be noted that the aforementioned stable state can refer to the intermediate state of the displayed text of the previous sentence segment, which lies between the state where correction is allowed and the state where correction is prohibited. Displayed text in a stable state indicates that the text content of the sentence segment has become stable at the current moment; that is, the speech recognition engine's recognition result for the sentence segment no longer changes significantly in several consecutive rounds of recognition output, but it has not yet been ultimately locked as unmodifiable. Therefore, in this embodiment, the stable state is the intermediate stage in the transition from the writing state to the state where correction is prohibited.

[0129] It should also be noted that the triggering conditions mentioned above can refer to the conditions that the displayed text used to determine the previous sentence segment must meet.

[0130] In its implementation, after determining that there is a natural punctuation point between the previous speech and the current speech, the device identifies the previous punctuation point to which the previous speech belongs and obtains the displayed text of the previous punctuation point. The displayed text of the previous punctuation point is then set to a stable state.

[0131] After setting the displayed text of the previous segment to a stable state, the device can further determine the triggering conditions satisfied by the displayed text of the previous segment.

[0132] Step S32: When the trigger condition is a modification condition, modify and display the displayed text of the previous sentence based on the modification requirements, obtain the new displayed text of the previous sentence, and return to the step of determining the trigger condition satisfied by the displayed text of the previous sentence. Step S33: When the trigger condition is a locking condition, set the displayed text of the previous sentence segment to a state where correction is prohibited.

[0133] Understandably, the aforementioned modification conditions can refer to specific conditions used to trigger modification operations on the displayed text of the previous sentence that is in a stable state. In this embodiment, the modification conditions can represent that the user has a need to edit or correct the displayed text of the previous sentence. The aforementioned locking conditions can refer to specific conditions used to trigger setting the displayed text of the previous sentence that is in a stable state to a state where modification is prohibited. In this embodiment, the locking conditions can represent that the displayed text of the previous sentence should be finally locked and no longer accept any subsequent modifications.

[0134] It is also understood that the aforementioned modification request can refer to the user's specific intention or instruction to modify the previously displayed text. Modification requests may include replacing, deleting, inserting specific words or phrases in the displayed text, or rewriting the entire sentence. The aforementioned new displayed text can refer to the updated complete text content obtained after the device performs modification operations on the previously displayed text.

[0135] In its implementation, after determining the triggering conditions satisfied by the displayed text of the previous segment, the device executes the corresponding processing branch according to the type of the triggering conditions.

[0136] When the trigger condition is a modification condition, the device can obtain the modification request input by the user. The acquisition of this modification request can include: the device receiving a verbal modification instruction from the user via a voice interaction interface, such as the user saying "change Thursday to Friday," and the device parsing the specific modification request after speech recognition and semantic understanding; or, the device receiving operations performed by the user on the previously displayed text, such as cursor positioning, text selection, and keyboard input, via a touch interaction interface, and determining the modification request based on these operations. The device then modifies the previously displayed text according to the modification request, including replacing, deleting, or inserting characters at specific positions in the displayed text. After completing the modification operation, the device uses the modified complete text as the new displayed text for the previous sentence and displays it on the display interface. After display, the device returns to the step of determining the trigger condition satisfied by the previously displayed text, and continues to monitor and determine the trigger condition for the modified previously displayed text.

[0137] When the trigger condition is a locking condition, the device sets the displayed text of the previous sentence to a state where correction is prohibited. In the state where correction is prohibited, the displayed text of the previous sentence will no longer accept any correction or rewriting from subsequent speech recognition results, and the text content of that sentence is ultimately locked. The aforementioned locking condition can be any one of the timeout trigger condition, the next sentence start trigger condition, or the explicit confirmation operation trigger condition.

[0138] To facilitate understanding, the following example is used for explanation, but it does not impose specific limitations on this embodiment. Continuing from the previous example, after the device sets the displayed text of the previous sentence "The second point is that the time is scheduled for Thursday morning" to a stable state, it detects that the user speaks the modification command "Change Thursday to Friday". The device determines that the trigger condition is a modification condition and obtains the modification request as replacing "Thursday" with "Friday". The device modifies the displayed text of the previous sentence "The second point is that the time is scheduled for Thursday morning" according to the modification request, replacing "Thursday" with "Friday", obtaining the new displayed text "The second point is that the time is scheduled for Friday morning", and displays it on the display interface. Subsequently, the device returns to the step of determining the trigger condition satisfied by the displayed text of the previous sentence, and re-monitors the trigger condition for the modified displayed text. If the modified displayed text does not change again within 2 seconds, the device detects a timeout, determines that the trigger condition is a locking condition, sets the displayed text of the previous sentence "The second point is that the time is scheduled for Friday morning" to a prohibited correction state, and the sentence is finally locked.

[0139] Furthermore, considering that not all cases can be directly set to a stable state if natural punctuation breaks exist, in this embodiment, the step of setting the displayed text of the previous punctuation break to a stable state includes: Step S311: Obtain the semantic completeness of the previous sentence segment and obtain the previous recognition result of the previous sentence segment. The previous recognition result includes at least: the text of the previous sentence segment, the confidence score, and the stability indicator.

[0140] It should be noted that the aforementioned semantic completeness can be a quantitative indicator used to measure whether the displayed text of the previous sentence has complete expressive power at the semantic level. In this embodiment, semantic completeness can characterize whether the text content of the previous sentence contains a complete subject-verb-object structure or whether it expresses a complete semantic unit.

[0141] It should also be noted that the aforementioned previous recognition result can refer to the recognition result information output by the speech recognition engine for the previous sentence segment. The previous recognition result can be the recognition output obtained by the speech recognition engine after recognizing the previous speech before the current speech arrives.

[0142] It is understood that the confidence score mentioned above can refer to the quantitative score of the credibility of the text content in the previous recognition result by the speech recognition engine. In this embodiment, the confidence score can be expressed in the form of probability value or numerical value. The higher the value, the more confident the recognition engine is in the recognition result.

[0143] It is also understood that the aforementioned stability flag can refer to the flag information output by the speech recognition engine to indicate the stability of the previous recognition result. The stability flag can include a flag to indicate whether the recognition result has converged, such as the final result flag (final flag) output by the recognition engine.

[0144] In its specific implementation, the device can obtain the semantic completeness of the previous sentence after setting the displayed text of the previous sentence to a stable state.

[0145] The process of obtaining the semantic completeness of the previous sentence segment may include: the device can obtain the displayed text of the previous sentence segment, input the displayed text of the previous sentence segment into a pre-trained semantic completeness evaluation model, the semantic completeness evaluation model performs semantic analysis on the displayed text, determines whether the text has a complete semantic structure, and outputs a quantitative score of semantic completeness, the higher the score, the more complete the semantics.

[0146] The semantic completeness assessment model described above can be built on a deep neural network and trained using a large-scale text corpus labeled with semantic completeness as training data. Of course, it can also be built in other ways, and this embodiment does not limit it.

[0147] In addition to the above, the method of obtaining semantic completeness may also include: the device uses a rule-based method to perform syntactic structure analysis on the displayed text of the previous sentence segment, and determines whether the text contains complete core syntactic components such as subject, predicate and object. If the syntactic structure is complete, a higher semantic completeness is assigned; otherwise, a lower semantic completeness is assigned.

[0148] Simultaneously, the device can obtain the previous recognition result of the previous sentence segment; that is, the device can read the previous recognition result output by the speech recognition engine when the previous speech was recognized from its internal memory. The aforementioned previous recognition result includes at least the text of the previous sentence segment, a confidence score, and a stability flag. The text of the previous sentence segment can be the recognized text content from the previous recognition result; this content can be the same as or different from the displayed text of the previous sentence segment. The confidence score can be the credibility rating given by the speech recognition engine for the recognition result of the previous sentence segment. The stability flag can be the flag information output by the speech recognition engine indicating whether the recognition result of the previous sentence segment has reached a stable state.

[0149] To facilitate understanding, the following explanation uses examples, but does not impose specific limitations on this embodiment. Continuing from the previous example, after the device sets the displayed text of the previous sentence "The second point is that the time is scheduled for Friday morning" to a stable state, it can obtain the semantic completeness of the previous sentence. The device inputs "The second point is that the time is scheduled for Friday morning" into the semantic completeness evaluation model, which outputs a semantic completeness score of 0.92, indicating that the text has a high semantic completeness. The device retrieves the previous recognition result from its internal memory, which includes: the text of the previous sentence "The second point is that the time is scheduled for Friday morning", a confidence score of 0.89, and a stability flag of "final" (indicating that the recognition result has converged).

[0150] Step S312: Based on the semantic completeness and each of the previous recognition results, determine whether the previous sentence segment meets the preset stability conditions. The preset stability conditions include at least one of the following: at least one stability flag is a preset stability flag, the text of the previous sentence segment remains unchanged for a preset number of consecutive segments, at least one confidence score reaches a preset confidence threshold, and the semantic completeness reaches a preset completeness threshold. Step S313: If the previous sentence segment meets the preset stability condition, set the displayed text of the previous sentence segment to a stable state.

[0151] It should be noted that the aforementioned preset stability conditions may refer to a set of pre-set judgment conditions used to determine whether the displayed text of the previous sentence segment has reached a stable state.

[0152] In this embodiment, the aforementioned preset stability conditions include at least one of the following: at least one stability flag is a preset stability flag; the text of the previous sentence remains unchanged for a preset number of consecutive sentences; at least one confidence score reaches a preset confidence threshold; and the semantic integrity reaches a preset integrity threshold.

[0153] It is important to emphasize that the aforementioned preset stability flag can refer to a pre-specified stability flag type used to indicate that the recognition result has converged, such as the final result flag (final flag) output by the speech recognition engine. The aforementioned preset number of consecutive occurrences can refer to a pre-set quantity representing the number of times the text of the previous sentence segment needs to remain unchanged. The aforementioned preset reliability threshold can refer to a pre-set confidence score threshold. The aforementioned preset completeness threshold can refer to a pre-set semantic completeness score threshold.

[0154] In a specific implementation, after obtaining the semantic completeness of the previous sentence segment and each previous recognition result, the device can determine whether the previous sentence segment meets the preset stability conditions based on the information in the semantic completeness and each previous recognition result.

[0155] The above judgment process may include: the device comparing each sub-condition in the preset stability conditions one by one, and determining whether the previous sentence segment meets the preset stability conditions based on the comparison results. Specifically, it may be as follows: The device can detect stability flags in each previous recognition result and determine whether at least one stability flag is a preset stability flag. A preset stability flag could be, for example, the "final" flag output by the speech recognition engine. If the device detects that at least one stability flag in a previous recognition result is a "final" flag, then this sub-condition is satisfied.

[0156] The device can acquire the text of the previous sentence for a preset number of consecutive consecutive times. This preset number of consecutive previous sentence texts refers to the recognized text corresponding to the previous sentence received by the device in multiple consecutive recognition output cycles. The device determines whether all the texts of the preset number of consecutive previous sentences are identical. If they are all identical, the sub-condition is satisfied. The preset number of consecutive sentences can be, for example, three, meaning that the text of the previous sentence remains unchanged in three consecutive recognition outputs.

[0157] The device can detect the confidence scores in each previous identification result and determine whether at least one confidence score reaches a preset confidence threshold. The preset confidence threshold can be set to, for example, 0.85. If the device detects that at least one confidence score in a previous identification result reaches or exceeds the preset confidence threshold, then the sub-condition is satisfied.

[0158] The device can detect whether the semantic completeness of the acquired previous sentence segment reaches a preset completeness threshold. The preset completeness threshold can be set to, for example, 0.8. If the semantic completeness reaches or exceeds the preset completeness threshold, the sub-condition is satisfied.

[0159] If the device determines that the previous sentence segment meets the preset stability conditions, the displayed text of the previous sentence segment can be set to a stable state.

[0160] Considering that multiple sub-conditions in the preset stability conditions are related as "at least one", the device only needs to determine that the previous sentence satisfies any one or more sub-conditions included in the preset stability conditions to determine that the previous sentence satisfies the preset stability conditions. All sub-conditions in the above preset stability conditions can be enabled, or some sub-conditions can be selectively enabled according to the application scenario. The above preset reliability threshold and preset completeness threshold are both configurable parameters and can be set according to different application scenarios.

[0161] To facilitate understanding, the following example is used for explanation, but it does not impose specific limitations on this embodiment. Continuing from the previous example, the device obtains a semantic completeness of 0.92 for the previous sentence segment "The second point is that the time is scheduled for Friday morning". The previous recognition results include: the text of the first recognition result is "The second point is that the time is scheduled for Friday morning", with a confidence score of 0.82 and a stability flag of empty; the text of the second recognition result is "The second point is that the time is scheduled for Friday morning", with a confidence score of 0.89 and a stability flag of "final". The preset stability conditions include: at least one stability flag is the "final" flag, the text of three consecutive previous sentences remains unchanged, at least one confidence score reaches a preset confidence threshold of 0.85, and the semantic completeness reaches a preset completeness threshold of 0.8. The device detects that the stability flag in the second recognition result is the "final" flag, satisfying the sub-condition of "at least one stability flag is a preset stability flag". The device also detects that the confidence score of 0.89 reaches the preset confidence threshold of 0.85, satisfying the sub-condition of "at least one confidence score reaches the preset confidence threshold". The device detected that the semantic integrity score of 0.92 reached the preset integrity threshold of 0.8, satisfying the sub-condition of "semantic integrity score reaches the preset integrity threshold". Therefore, the device determined that the previous sentence segment met the preset stability condition and set the displayed text of the previous sentence segment, "The second point is that the time is scheduled for Friday morning", to a stable state.

[0162] Furthermore, in order to determine the specific triggering conditions that are met, in this embodiment, the step of determining the triggering conditions met by the displayed text of the previous sentence segment includes: Step S314: Upon receiving the modification request triggered by the user, determine that the triggering condition satisfied by the displayed text of the previous sentence segment is the modification condition; Step S315: If it is detected that the displayed text of the previous sentence continues to maintain the preset configuration duration and / or a confirmation operation triggered by the user is received, determine that the trigger condition satisfied by the displayed text of the previous sentence is a locking condition.

[0163] It should be noted that the aforementioned user-triggered modification requests can refer to specific requests initiated by users through the interactive interface to edit or correct the displayed text of the previous sentence. These modification requests can be triggered by users inputting modification commands via voice (e.g., saying "change [something] to [something]"), or by users performing touch operations (e.g., positioning the cursor, selecting text, and inputting text using the keyboard on the display interface).

[0164] It should also be noted that the aforementioned preset configuration duration can refer to a pre-configured time period for a sentence segment in a stable state. The preset configuration duration can be used to determine whether the displayed text of a sentence segment in a stable state remains unchanged within that time period, in order to decide whether to trigger the locking condition. The aforementioned user-triggered confirmation operation can refer to an operation initiated by the user through an interactive interface to confirm that the displayed text of the previous sentence segment is correct and to agree to lock it. The triggering method for the confirmation operation can include the user clicking the confirmation button on the display interface, issuing a voice confirmation command (e.g., saying "confirm" or "okay"), or performing a preset confirmation gesture, etc., which are not limited in this embodiment.

[0165] In its implementation, after setting the previously displayed text of the previous sentence to a stable state, the device continuously monitors the reception of user-triggered modification requests, the duration for which the previously displayed text of the previous sentence remains unchanged, and the reception of user-triggered confirmation operations. Based on these monitoring results, the triggering conditions satisfied by the previously displayed text of the previous sentence are determined.

[0166] When the device receives a modification request triggered by the user, it can determine that the triggering condition met by the displayed text of the previous segment is the modification condition, perform the modification operation on the displayed text of the previous segment according to the modification request, and return to re-monitor the triggering condition after the modification is completed.

[0167] If the device detects that the displayed text of the previous segment remains unchanged for a preset configuration duration, the device determines that the trigger condition met by the displayed text of the previous segment is a locking condition. The aforementioned "remaining unchanged for a preset configuration duration" means that the timer starts from the moment the displayed text of the previous segment is set to a stable state, or from the moment the displayed text of the previous segment last changed, and the content of the displayed text does not change within the preset configuration duration. After determining that the trigger condition is a locking condition, the device sets the displayed text of the previous segment to a state where correction is prohibited.

[0168] When the device receives a confirmation operation triggered by the user, it determines that the trigger condition met by the displayed text of the previous segment is a locking condition. The confirmation operation can be triggered by the user at any time after the previous segment has entered a stable state. Once the confirmation operation is triggered, the device determines that the trigger condition is a locking condition and sets the displayed text of the previous segment to a state where correction is prohibited.

[0169] In this embodiment, the device can determine the trigger condition as a locking condition if both or either of the following conditions are met: "the displayed text of the previous sentence is continuously maintained for a preset configuration duration" and "a user-triggered confirmation operation is received". The preset configuration duration can be configured according to different application scenarios. For example, in a mobile phone input method scenario, the preset configuration duration can be configured to 2 seconds; in an in-vehicle voice scenario, the preset configuration duration can be configured to 1.5 seconds; and in a meeting recording scenario, the preset configuration duration can be configured to 3 seconds.

[0170] To facilitate understanding, the following explanation uses examples, but does not impose specific limitations on this embodiment. Continuing from the previous example, the device sets the displayed text of the previous sentence, "The second point is that the time is scheduled for Friday morning," to a stable state and begins monitoring trigger conditions. The device is configured with a preset duration of 2 seconds. During the stable state, if the device receives a voice modification command from the user, "Change Friday to Saturday," the device determines the trigger condition as a modification condition, enters the modification processing branch, and modifies the displayed text. If the device does not receive any modification request and detects that the displayed text remains unchanged for 2 seconds, the device determines the trigger condition as a locking condition and sets the displayed text of the previous sentence to a state where modification is prohibited. If, during the stable state, the device receives a confirmation operation triggered by the user clicking the confirmation button on the display interface, regardless of whether the displayed text has remained stable within the preset duration, the device determines the trigger condition as a locking condition and sets the displayed text of the previous sentence to a state where modification is prohibited.

[0171] Furthermore, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the voice input method described above.

[0172] In addition, refer to Figure 5 , Figure 5 This is a structural block diagram of the voice input device of this application. Figure 5 As shown in the figure, this application also proposes a voice input device, which includes: The acquisition module is used to acquire the current speech dictated by the user and determine whether there is a natural punctuation point between the previous speech and the current speech. The natural punctuation point indicates that the previous speech and the current speech are at different punctuation points. The first determining module is used to determine the displayed text of the current segment where the current speech is located if the natural segment break point does not exist, set the displayed text of the current segment to a state where correction is allowed, correct the displayed text of the current segment based on the current speech, display the corrected displayed text of the current segment and the text of the current speech as the new displayed text of the current segment, and return to execute the step of obtaining the current speech dictated by the user. The second determining module is used to determine the displayed text of the previous segment where the previous speech is located if the natural segment break point exists, set the displayed text of the previous segment break point to a state where correction is prohibited, display the text of the current speech as the displayed text of the current segment break point, and return to execute the step of obtaining the current speech dictated by the user.

[0173] This embodiment detects natural punctuation points between preceding and following speech in real time during voice input. It only sets the currently being written punctuation point to a state where correction is allowed. Once a punctuation point is detected, the previous punctuation point is switched to a state where correction is prohibited and fixed. Subsequent speech recognition results will not alter the fixed punctuation point. Compared to existing technologies that erase and re-render the entire text whenever a new recognition result arrives, this embodiment prohibits correction of the previous punctuation point when there is a different punctuation point between the current and previous speech. This allows users to stably review the spoken content during continuous dictation without repeatedly repositioning their gaze, thus improving the user experience.

[0174] Other embodiments or specific implementations of the voice input device described in this application can be found in the above-described method embodiments, and will not be repeated here.

[0175] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0176] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0177] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as a read-only memory image (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0178] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A voice input method, characterized in that, The method includes: The system acquires the current speech dictated by the user and determines whether there is a natural punctuation point between the previous speech and the current speech. The natural punctuation point indicates that the previous speech and the current speech are at different punctuation points. If no natural punctuation point exists, determine the displayed text of the current punctuation where the current speech is located, set the displayed text of the current punctuation to the correction-allowed state, correct the displayed text of the current punctuation based on the current speech, and display the corrected displayed text of the current punctuation and the text of the current speech as the new displayed text of the current punctuation, and return to the step of obtaining the current speech dictated by the user. If the natural punctuation point exists, determine the displayed text of the previous punctuation point where the previous speech is located, set the displayed text of the previous punctuation point to a state where correction is prohibited, display the text of the current speech as the displayed text of the current punctuation point, and return to the step of obtaining the current speech dictated by the user.

2. The method as described in claim 1, characterized in that, The step of determining whether there is a natural punctuation mark between the previous speech and the current speech includes: Feature extraction is performed on the current speech to obtain cue features, which include at least two of the following: speech pause features, intonation features, punctuation assumption features, and semantic boundary features. The clue features are weighted and fused according to preset weights to obtain a comprehensive score, and the existence of a natural punctuation point between the previous speech and the current speech is determined based on the comprehensive score.

3. The method as described in claim 2, characterized in that, The cue features include the speech pause features, intonation features, punctuation assumption features, and semantic boundary features of the current speech. The step of extracting features from the current speech to obtain cue features includes: Speech activity detection is performed on the current speech, the silence duration of the silent segment in the current speech is detected, and the speech pause feature is generated based on the silence duration; Extract the fundamental frequency of the speech frame preceding the silence segment in the current speech, and generate the intonation trend feature based on the fundamental frequency of the speech frame; Determine the punctuation classification probability vector of the current speech, and generate the punctuation hypothesis features based on the punctuation classification probability vector; Determine the semantic vector of the current sentence segment where the current speech is located, determine the similarity between the semantic vector and the preset complete statement template vector, and generate the semantic boundary feature based on the similarity.

4. The method as described in claim 1, characterized in that, The step of setting the displayed text of the previous sentence segment to a state where correction is disabled includes: Set the displayed text of the previous sentence to a stable state, and determine the triggering conditions satisfied by the displayed text of the previous sentence. When the trigger condition is a modification condition, the displayed text of the previous sentence is modified and displayed based on the modification requirements to obtain the new displayed text of the previous sentence, and the process returns to the step of determining the trigger condition satisfied by the displayed text of the previous sentence. When the trigger condition is a locking condition, the displayed text of the previous sentence is set to a state where correction is prohibited.

5. The method as described in claim 4, characterized in that, The step of setting the displayed text of the previous sentence segment to a stable state includes: Obtain the semantic completeness of the previous sentence segment and obtain the previous recognition result of the previous sentence segment. The previous recognition result includes at least: the text of the previous sentence segment, the confidence score, and the stability indicator. Based on the semantic completeness and each of the previous recognition results, it is determined whether the previous sentence segment meets the preset stability conditions. The preset stability conditions include at least one of the following: at least one of the stability flags is a preset stability flag, the text of the previous sentence segment remains unchanged for a preset number of consecutive segments, at least one of the confidence scores reaches a preset confidence threshold, and the semantic completeness reaches a preset completeness threshold. If the previous sentence segment meets the preset stability condition, the displayed text of the previous sentence segment is set to a stable state.

6. The method as described in claim 4, characterized in that, The step of determining the triggering conditions satisfied by the displayed text of the previous sentence segment includes: Upon receiving the modification request triggered by the user, the triggering condition satisfied by the displayed text of the previous sentence is determined to be the modification condition. If the displayed text of the previous sentence is continuously maintained for a preset configuration duration and / or a confirmation operation triggered by the user is received, the trigger condition satisfied by the displayed text of the previous sentence is determined to be a locking condition.

7. The method as described in claim 1, characterized in that, The step of correcting the displayed text of the current sentence segment based on the current speech includes: Based on the current speech, generate hypothetical text and determine the sentence segment where the hypothetical text is located; If the given sentence segment is the current sentence segment and the displayed text of the current sentence segment is in the allowed correction state, determine the character position of the hypothetical text in the current sentence segment; The text corresponding to the character position in the currently displayed text of the sentence segment is corrected to the hypothetical text to obtain the corrected displayed text of the current sentence segment.

8. A voice input device, characterized in that, The device includes: The acquisition module is used to acquire the current speech dictated by the user and determine whether there is a natural punctuation point between the previous speech and the current speech. The natural punctuation point indicates that the previous speech and the current speech are at different punctuation points. The first determining module is used to determine the displayed text of the current segment where the current speech is located if the natural segment break point does not exist, set the displayed text of the current segment to a state where correction is allowed, correct the displayed text of the current segment based on the current speech, display the corrected displayed text of the current segment and the text of the current speech as the new displayed text of the current segment, and return to execute the step of obtaining the current speech dictated by the user. The second determining module is used to determine the displayed text of the previous segment where the previous speech is located if the natural segment break point exists, set the displayed text of the previous segment break point to a state where correction is prohibited, display the text of the current speech as the displayed text of the current segment break point, and return to execute the step of obtaining the current speech dictated by the user.

9. A voice input device, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the voice input method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the voice input method as described in any one of claims 1 to 7.