Voice input method, terminal device and storage medium based on gesture interaction
By introducing gesture interaction on terminal devices and combining voice input with gesture area operation, real-time modification of voice text and recommendation assistance are achieved, which solves the problems of cumbersome operation and low efficiency in existing technologies and improves the convenience and accuracy of voice input.
Patent Information
- Application Number
- CN202510955265.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-11
AI Technical Summary
Existing voice input methods are cumbersome when users need to modify input content, and voice input is separated from text modification methods, resulting in low input efficiency and poor user experience.
Gesture interaction is introduced to display a mobile icon in a specific area of the terminal device, which follows the user's finger touch position, enabling voice input and real-time modification, and providing recommended word auxiliary correction.
It improves the fluency and convenience of voice input, enhances the immediacy and accuracy of voice text, avoids the inconvenience of frequent switching input modes, and improves overall input efficiency and user experience.
Smart Images

Figure CN120447861B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of gesture interaction technology, and in particular to a voice input method, terminal device and storage medium based on gesture interaction. Background Art
[0002] Voice input is a technical means of converting information into text or instructions through the user's voice. Simply put, the user speaks to the device, and the device can recognize what is said and convert it into text or perform corresponding operations. When the user needs to modify the input content, the existing voice input method usually switches to keyboard mode to implement text correction. When using this method, users often encounter the problem of not being able to make changes immediately during the voice input process. At the same time, the interaction methods of keyboard input and voice input do not match, making the operation process cumbersome, thereby affecting the overall input efficiency and user experience. Summary of the Invention
[0003] In view of this, the embodiments of the present application provide a voice input method, terminal device and storage medium based on gesture interaction, which can effectively solve the problems of cumbersome and low efficiency of voice input operations.
[0004] In a first aspect, an embodiment of the present application provides a voice input method based on gesture interaction, comprising:
[0005] In response to a voice input command received by the terminal device, entering the voice input interface and displaying the first gesture area;
[0006] In response to a user's movement operation instruction in the first gesture area, controlling the display of a movement icon in the first gesture area and controlling the movement icon to follow the touch position of the user's finger;
[0007] Controlling voice input and performing a first modification operation on the input voice text according to a movement state of the movement icon in the first gesture area;
[0008] Under the condition that the first modification operation is performed on the voice text, corresponding recommended words of the retained voice text are obtained and displayed, and the recommended words are used for the user to correct the retained voice text.
[0009] In a first possible embodiment of the first aspect, the gesture interaction-based voice input method further includes:
[0010] After responding to the movement operation instruction of the first gesture area, displaying the second gesture area;
[0011] In response to a user's movement operation instruction in the second gesture area, controlling the display of a movement icon in the second gesture area and controlling the movement icon to follow the touch position of the user's finger;
[0012] According to the movement state of the moving icon in the second gesture area, the voice input is controlled to continue, and a second modification operation is performed on the input voice text.
[0013] In a second possible embodiment of the first aspect, the first gesture area includes a first function area, a second function area, and a third function area, and controlling voice input and performing a first modification operation on the input voice text according to a movement state of the moving icon in the first gesture area includes:
[0014] controlling voice input according to a moving state of the moving icon in the first functional area to obtain the input voice text;
[0015] controlling cancellation of voice input according to a movement state of the moving icon in the second functional area;
[0016] According to the moving state of the moving icon in the third function area, a first modification operation is performed on the input voice text, wherein the first modification operation includes a partial deletion and restoration operation of the voice text.
[0017] In a third possible embodiment of the first aspect, the moving state of the mobile icon in the third functional area includes a sliding interval and a sliding direction of a user's finger;
[0018] The finger sliding interval and the finger sliding direction are used to select the target deleted text or restored text in the voice text;
[0019] The finger sliding interval is also used to control the text deletion or restoration speed of the voice text.
[0020] In a fourth possible embodiment of the first aspect, obtaining and displaying the corresponding recommended words of the retained voice text includes:
[0021] Selecting a plurality of to-be-processed words from the retained speech text according to a preset direction;
[0022] Selecting a plurality of candidate words as the recommended words based on semantic matching results between the word to be processed and a plurality of candidate words;
[0023] The recommended words are sorted and displayed according to their usage frequency and semantic matching degree.
[0024] In a fifth possible embodiment of the first aspect, the second gesture area includes a fourth function area and a fifth function area, and controlling the continuation of voice input and performing a second modification operation on the input voice text according to a movement state of the move icon in the second gesture area includes:
[0025] continuing voice input or inputting a voice modification instruction according to the movement state of the moving icon in the fourth functional area;
[0026] Under the condition that the voice modification instruction is input, the second modification operation is controlled to be performed according to the moving state of the moving icon in the fifth functional area.
[0027] In a sixth possible embodiment of the first aspect, the second gesture area further includes a sixth function area, and the method further includes:
[0028] When the mobile icon moves to the fifth functional area, displaying the sixth functional area;
[0029] According to the moving state of the moving icon in the sixth function area, control is performed to cancel the second modifying operation.
[0030] In a seventh possible embodiment of the first aspect, the controlling the second modification operation includes:
[0031] determining the target word phonetics in the phonetically modified text;
[0032] Mark all words in the speech text that are phonetically identical to the target word and are to be changed;
[0033] Marking and displaying all modified words in the phonetically modified text that are phonetically identical to the target word;
[0034] According to the moving state of the moving icon in the fifth function area, control is performed to modify the selected target word to be changed into the target word to be changed.
[0035] In a second aspect, an embodiment of the present application provides a terminal device, which includes a processor and a memory, wherein the memory stores a computer program, and the processor is used to execute the computer program to implement the above-mentioned voice input method based on gesture interaction.
[0036] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed on a processor, implements the above-mentioned voice input method based on gesture interaction.
[0037] The embodiments of the present application have the following beneficial effects:
[0038] A speech input method based on gesture interaction in an embodiment of the present application includes: responding to a speech input instruction received by a terminal device, entering a speech input interface and displaying a first gesture area; responding to a user's movement operation instruction in the first gesture area, controlling the display of a moving icon in the first gesture area and controlling the moving icon to follow the touch position of the user's finger; controlling speech input according to the movement state of the moving icon in the first gesture area, and performing a first modification operation on the input speech text; under the condition of performing the first modification operation on the speech text, obtaining and displaying the corresponding recommended words of the retained speech text, the recommended words are used for the user to correct the retained speech text. The present application introduces gesture interaction during the speech input process, allowing the user to make real-time adjustments to the speech text without interrupting the speech input, and introduces gesture interaction after the speech input, allowing the user to modify the speech text without switching to the keyboard input method. By combining gesture interaction with speech input, the present application effectively avoids the inconvenience caused by frequent switching of input modes, improves the fluency of speech input and the overall interaction convenience, thereby significantly improving the speech input efficiency and the accuracy of speech text. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.
[0040] Figure 1 A schematic diagram of the structure of a terminal device according to an embodiment of the present application is shown;
[0041] Figure 2 A first flow chart of a voice input method based on gesture interaction according to an embodiment of the present application is shown;
[0042] Figure 3 A second flow chart of the voice input method based on gesture interaction according to an embodiment of the present application is shown;
[0043] Figure 4 A first interaction diagram of the first gesture area in an embodiment of the present application is shown;
[0044] Figure 5 A schematic diagram of a second interaction in the first gesture area according to an embodiment of the present application is shown;
[0045] Figure 6 A schematic diagram of the sliding area in the voice deletion and restoration area in an embodiment of the present application is shown;
[0046] Figure 7A third interaction diagram of the first gesture area in an embodiment of the present application is shown;
[0047] Figure 8 A third flow chart of the voice input method based on gesture interaction according to an embodiment of the present application is shown;
[0048] Figure 9 A fourth flow chart of the voice input method based on gesture interaction according to an embodiment of the present application is shown;
[0049] Figure 10 A schematic diagram of a first interaction in the second gesture area according to an embodiment of the present application is shown;
[0050] Figure 11 A second interaction diagram of the second gesture area in an embodiment of the present application is shown;
[0051] Figure 12 A third interaction schematic diagram of the second gesture area in an embodiment of the present application is shown.
[0052] Description of main component symbols:
[0053] 100 - terminal device; 110 - memory; 120 - processor; 130 - human-computer interaction interface. DETAILED DESCRIPTION
[0054] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments.
[0055] The components of the embodiments of the present application generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed application, but rather merely represents selected embodiments of the present application. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative effort are within the scope of protection of the present application.
[0056] Hereinafter, the terms "including", "having" and their cognates used in various embodiments of the present application are intended only to indicate specific features, numbers, steps, operations, elements, components or combinations of the aforementioned items, and should not be understood as excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the aforementioned items or adding the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the aforementioned items. In addition, the terms "first", "second", "third" and the like are only used to distinguish descriptions and should not be understood as indicating or implying relative importance.
[0057] Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those skilled in the art to which the various embodiments of the present application belong. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meaning as in the context of the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in the various embodiments of the present application.
[0058] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.
[0059] Existing voice input technology has the following defects and deficiencies in voice-to-text modification: Insufficient instant modification capabilities: When the current voice input system recognizes errors, slips of the tongue, or redundant input, users usually need to wait for the input to be completed before manually modifying it, which makes it difficult to meet the demand for real-time adjustment of text during the voice input process. Fragmented interaction modes: Voice input and text modification methods are usually independent of each other. Users need to frequently switch between voice input and keyboard editing, resulting in cumbersome operation processes, affecting the smoothness of input and user experience. The accuracy of voice modification: Text modification based on voice commands may affect the accuracy of the modification results due to factors such as homophones and contextual ambiguity. Existing technologies lack an effective mechanism to assist users in choosing the correct modification plan.
[0060] To address the above issues, this application proposes a voice input method, terminal device, and storage medium based on gesture interaction. This method allows for input, deletion, restoration, and modification of voice based on the user's interaction information in the gesture interaction area. This application optimizes the voice input editing process, improving the immediacy, convenience, and accuracy of voice text modification, thereby enhancing overall voice input efficiency and user experience.
[0061] First, the present application embodiment provides a terminal device 100, please refer to Figure 1 , which is a structural block diagram of the terminal device 100 provided in an embodiment of the present application. The terminal device 100 may include a memory 110 and a processor 120. The memory 110 and the processor 120 may be electrically connected directly or indirectly to the human-computer interaction interface 130 to realize voice input and interaction.
[0062] The processor 120 can process information and / or data related to the voice input method based on gesture interaction to perform one or more functions described in this application. For example, the processor 120 can perform voice input operations based on the user's gesture interaction results in the first gesture area of the interactive interface, and partially delete and / or restore the input voice text; under the condition of partially deleting the voice text, obtain recommended words for the undeleted voice text to select recommended words to supplement the deleted voice text; continue to perform voice input operations based on the user's gesture interaction results in the second gesture area, or input voice modification instructions to modify the voice text according to the voice modification instructions. This enables the terminal device 100 to perform operations such as deletion, restoration and modification on the input voice text in combination with gesture interaction, thereby improving the efficiency and accuracy of voice input.
[0063] The processor 120 may be an integrated circuit chip with signal processing capabilities. The processor 120 may be a general-purpose processor 120, including at least one of a central processing unit 120 (CPU), a graphics processing unit 120 (GPU), a network processor 120 (NP), a digital signal processor 120 (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor 120 may be a microprocessor 120 or any conventional processor 120, and may implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0064] It is understandable that Figure 1 The structure of the terminal device 100 shown is only a schematic structure. The terminal device 100 may also include Figure 1 The structure shown in FIG may have more or fewer components or modules, or may have Figure 1 The structure shown in FIG is configured or constructed differently. Figure 1 Each component shown in the figure may be implemented by hardware, software, or a combination of both.
[0065] Furthermore, it should be understood that the terminal device 100 provided herein may employ different configurations or structures depending on the requirements of the actual application. For example, the terminal device 100 provided herein may be an electronic device with voice recognition, computing, and storage functions (e.g., a smartphone, computer, smart home device, VR device, etc.).
[0066] For ease of understanding, the following examples of this application will be combined with Figure 1 The terminal device 100 shown illustrates the voice input method based on gesture interaction provided in an embodiment of the present application.
[0067] Figure 2 A flowchart of a voice input method based on gesture interaction according to an embodiment of the present application is shown. Exemplarily, the voice input method based on gesture interaction includes the following steps:
[0068] S210, in response to a voice input instruction received by the terminal device, enter the voice input interface and display the first gesture area.
[0069] Exemplarily, the gesture area refers to a specific interaction area defined by the terminal device 100 for the user in the voice input interface. In the gesture area, the mobile icon has different movement states following the touch position of the user's finger, such as sliding, clicking, long pressing, etc., which can trigger the voice input function or operation.
[0070] S220 , in response to a user's movement operation instruction in the first gesture area, controlling the display of a movement icon in the first gesture area and controlling the movement icon to follow the touch position of the user's finger.
[0071] In this embodiment, when a user's finger presses down and begins touching the gesture area, the terminal device 100 detects the touch and displays a move icon at the touch location. As the user's finger slides within the gesture area, the terminal device 100 continuously updates the movement status of the move icon, ensuring that the move icon always follows the user's finger. The movement status of the move icon within the gesture area indicates different operational intentions.
[0072] S230: Control the voice input according to the movement state of the moving icon in the first gesture area, and perform a first modification operation on the input voice text.
[0073] In the embodiment of the present application, the first gesture area includes a first function area, a second function area, and a third function area. Exemplarily, the first function area is located in the middle area of the first gesture area, the second function area is located below the first gesture area, and the third function area is located above the first gesture area.
[0074] For example, Figure 3 As shown, this application performs corresponding voice operations based on the user's interaction results in the first gesture area, specifically including the following steps:
[0075] S231 : Controlling voice input according to the moving state of the moving icon in the first function area to obtain input voice text.
[0076] In one embodiment, during the voice input process, the user controls the mobile icon to long press the first function area of the first gesture area to input voice, and continues to input voice without releasing the finger, so that the terminal device 100 converts the user's voice into voice text through voice recognition. Figure 4 As shown, the user long presses the first function area with his finger to input the corresponding voice text.
[0077] S232: Controlling cancellation of the voice input according to the moving state of the moving icon in the second function area.
[0078] S233: Perform a first modification operation on the input voice text according to the movement state of the movement icon in the third function area, wherein the first modification operation includes a partial deletion and restoration operation of the voice text.
[0079] For example, when a user finds an error during voice input, the user can slide his finger downward from the first function area to the second function area without releasing his finger to cancel the input voice. The user can also control the mobile icon to slide upward from the first function area to the third function area to delete and restore the voice text immediately. Figure 5 As shown, the user slides in the third function area to delete part of the voice text.
[0080] In one embodiment, the movement state of the mobile icon in the third functional area includes the sliding range and sliding direction of the user's finger. The sliding range and sliding direction are used to select the target text to be deleted or restored in the voice text. For example, sliding the finger to the left can delete the selected voice text, and sliding the finger to the right can restore the selected voice text.
[0081] In another embodiment, the finger sliding range is also used to control the speed of text deletion or restoration in voice text. The finger sliding range refers to the sliding distance of the mobile icon on the third area, which determines the length of the target text to be deleted or restored in the voice text. For example, a short sliding range can delete a small amount of voice text (such as a single word or phrase); a long sliding range can delete a larger range of voice text (such as an entire sentence or paragraph). The finger sliding range controls the speed of character deletion and restoration. For example, the initial position of the mobile icon is the origin. The farther you slide to the left, the faster the deletion speed. For example, the farther you slide to the right, the faster the restoration speed.
[0082] In one embodiment, when the mobile icon follows the touch position of the user's finger into the third function area, the "delete and restore mode" is activated. In this mode, the sliding of the mobile icon in the sliding area can be detected (such as Figure 6), until you release your finger or move out of the sliding area to exit this mode. In this mode, the movement distance and direction of the move icon will be mapped to the movement of the cursor. After releasing your finger, all characters after the cursor will be deleted.
[0083] S240 , under the condition that the first modification operation is performed on the voice text, corresponding recommended words of the retained voice text are obtained and displayed, where the recommended words are used for the user to correct the retained voice text.
[0084] For example, after the user partially deletes the voice text, the delete key in the third functional area moves to the right, and the left area in the third functional area provides the user with recommended words. In this application, the recommended words are associated words based on the undeleted voice text. The function is to directly select the target recommended word to modify the deleted voice text without the need for subsequent voice modification operations when making a slip of the tongue. Figure 7 As shown, the left area of the third function area displays corresponding recommended words.
[0085] In one embodiment, the present application selects multiple words to be processed from the retained spoken text according to a preset direction. In this embodiment, the range of words to be processed can be set to a certain number of characters or words before the end of the spoken text (for example, the first 5 characters or 2 words). Based on the semantic matching results between the word to be processed and multiple candidate words, the present application selects several candidate words as recommended words; the recommended words are sorted and displayed based on their frequency of use and degree of semantic matching.
[0086] In one embodiment, the present application selects multiple words to be processed from a speech text in a back-to-front manner. Based on the extracted words to be processed, the present application converts the extracted words to be processed into vector representations using a pre-trained word vector model (such as Word2Vec, BERT, etc.); calculates the semantic similarity between the vector of the word to be processed and the vectors of each candidate word using cosine similarity; and after comparing the semantic similarities, selects the candidate words that are most similar to the word to be processed as recommended words.
[0087] This application counts the historical usage frequency of each candidate word (i.e., the number of times the candidate word has been selected by users in the past). Frequently used candidate words are more likely to be commonly used expressions by users and will be recommended and displayed first. This application prioritizes the recommended words with the highest semantic similarity to the word to be processed based on the semantic similarity between the recommended word and the word to be processed.
[0088] It is understandable that when users are inputting voice, they may produce redundant or incorrect words due to slips of the tongue, recognition errors, or pauses in thinking. To address this problem, this application introduces gesture interaction during the voice input process to enable instant deletion and modification of text, thereby maximizing the accuracy of text output without affecting the input process.
[0089] Exemplarily, after the user finishes inputting voice through the first gesture area, the second gesture area is displayed in the voice input interface, and the user can continue to input voice and further modify the input voice text.
[0090] In one embodiment, if Figure 8 As shown, the speech input method based on gesture interaction of the present application also includes:
[0091] S250: After responding to the movement operation instruction of the first gesture area, display the second gesture area.
[0092] S260 , in response to a movement operation instruction of the user in the second gesture area, controlling the display of a movement icon in the second gesture area and controlling the movement icon to follow the touch position of the user's finger.
[0093] S270 , controlling the continuation of the voice input and performing a second modification operation on the input voice text according to the movement state of the moving icon in the second gesture area.
[0094] In one embodiment, the initial area distribution of the second gesture area includes a fourth functional area and a fifth functional area. This application continues voice input based on the movement state of the mobile icon in the fourth functional area. If the user confirms that the input voice text is correct, there is no need to modify the voice text. The user can long press the mobile icon to control it and continue voice input. If the user confirms that the input voice text is incorrect, the user can enter a voice modification instruction. Under the condition that the voice modification instruction is entered, the voice modification text is obtained and displayed based on the movement state of the mobile icon in the fifth functional area, and the second modification operation is controlled.
[0095] In another embodiment, when the mobile icon moves to the fifth functional area, the sixth functional area is displayed; and based on the movement state of the mobile icon in the sixth functional area, the second modification operation is canceled. In this embodiment, when the user controls the mobile icon to long-press the fifth functional area, a cancel modification area is displayed below the second gesture area. The user can control the mobile icon to move from the fifth functional area to the sixth functional area to cancel the voice text modification operation, and the second gesture area returns to the initial area layout.
[0096] In one embodiment, in the initial area distribution of the second gesture area, the fourth functional area is located in the lower area of the second gesture area, and the fifth functional area is located in the upper area of the second gesture area. After the sixth functional area appears, the fourth functional area is located in the middle area of the second gesture area, the fifth functional area is located in the upper area of the second gesture area, and the sixth functional area is located in the lower area of the second gesture area.
[0097] In one embodiment, the terminal device 100 has voice recognition to transcribe the user's voice modification instruction into a text instruction in real time, and preprocess the instruction text, including stop word removal processing, standardized synonymous expression processing, etc.
[0098] Stop words refer to words that frequently appear in the instruction text but contribute little to the semantics, such as "de", "le", "shi", "he", etc. Stop word removal processing is to delete the stop words in the instruction text. Standardized synonymous expression processing is used to unify different expression forms with the same or similar meanings in the instruction text into a standard format. For example, expressions such as "huancheng", "tihuaiwei" can be unified into "gaiwei". Standardized synonymous expression processing can simplify the instruction text structure and facilitate subsequent structural semantic matching.
[0099] In the embodiment of the present application, as Figure 9 shown, the present application controls to perform a second modification operation, which specifically includes the following steps:
[0100] S271, determine the target word voice in the voice modification text.
[0101] In one embodiment, the present application performs semantic structure parsing on the text instruction based on a preset instruction template, and extracts the target word voice in the text instruction. The preset instruction template is a set of predefined structured patterns used to match the user input text instruction. The preset instruction template is represented in a specific syntax form and can capture the "word to be modified" and "modified word" in the text instruction.
[0102] For example, by mapping the expression of the user's text instruction to the preset instruction template, the present application can quickly extract the core content of the text instruction. For example, the user input text instruction is: change "yuanyinsu" to "yuansu", and the preset instruction template that can be matched is: "ba + word to be modified + gaicheng + modified word", and the corresponding word to be modified "yuanyinsu" and modified word "yuansu" in the text instruction can be automatically extracted through template recognition. When the user's text instruction is: add "shoushi" before "shou shi" and delete "shou shi", the preset instruction template that can be matched is: "zai + word to be modified + qian + jiashang + modified word, jiang + word to be modified + shanchu", and the corresponding word to be modified "shou shi" and modified word "shoushi" in the text instruction can be automatically extracted through template recognition. [[ID=第十九]]
[0103] In one embodiment, the present application uses the "word to be modified" and "modified word" voices in the captured text instruction as the target word voice. When the target word voice is determined, the present application determines all the words to be modified in the voice text that are the same as the target word voice according to the homophone library, and determines all possible homophone modified words.
[0104] It can be understood that if there are multiple meanings or uncertainties in the recognition of the modified word (such as the recognition of "comma" and ",", the recognition of "gesture" and "jewelry"), this application determines and displays all possible homophonic modified words for the user to click and confirm, ensuring the accuracy and controllability of the modification operation.
[0105] In another embodiment, on the one hand, this application can determine the homophonic modified word through the homophone of the target word pronunciation. On the other hand, a prompt statement can be added to the voice modification instruction to determine the modified word, and the modified word can be determined through the prompt statement of the glyph plus pronunciation. For example, the voice modification instruction is: change jun to "jun with a single-person radical", and the modified word can be determined as "俊". This application can also determine the modified word through the prompt statement of finding a character in the context. For example, the voice modification instruction is: change jun to "jun with a single-person radical", and the modified word can be determined as "俊". This application can also determine the modified word through the prompt statement of finding a character by radical. For example, change mang to "盲 with wang on the top and mu on the bottom", and the modified word can be determined as "盲".
[0106] S272, mark all the words to be modified in the voice text that have the same pronunciation as the target word.
[0107] In one embodiment, this application matches the word to be modified according to the target word pronunciation in the voice text, marks and displays all the words to be modified in the voice text, so that the user can select the word to be modified from all the marked words to be modified. For example, when the number of marked words to be modified is 1, there is no need to select, but when the number of marked words to be modified is greater than 1, the user needs to select independently.
[0108] In one implementation manner, this application can add an order prompt to the voice modification instruction to determine and mark the position information of the word to be modified. For example, in combination with the word order prompt such as "the nth" word to be modified or "all" words to be modified in the voice modification instruction, judge and mark the position of the word to be modified in the voice text.
[0109] S273, mark and display all the modified words in the voice modification text that have the same pronunciation as the target word.
[0110] In this embodiment, the user can select between the modified words with the same pronunciation as the target word to avoid modification errors. When the number of marked modified words is 1, there is no need to select, but if the number of marked modified words is greater than 1, the user needs to select independently.
[0111] [[ID=二十一]]S274, according to the moving state of the moving icon in the fifth function area, control to modify the selected target word to be modified into the target modified word.
[0112] In another implementation manner, after this application determines the target word pronunciation and the modified word of the text instruction, it combines the voice modification text to display the target word pronunciation and multiple possible homophonic modified words above the second gesture area.
[0113] Such as Figure 10 As shown, the user controls the mobile icon to long press the fourth function area, enters the voice modification command "change shoushi to gesture", and then controls the mobile icon to move to the fifth gesture area. "Change shoushi to gesture / hand ornaments" is displayed above the fifth gesture area, and the target word and the modified word are grayed out. Figure 11 As shown, the user controls the mobile icon to select the target word to be changed "jewelry"; then selects the target word to be changed "gesture" to change the word to be changed in the voice text "jewelry" to "gesture"; finally, the user controls the mobile icon to the send button to send the modified voice text. Figure 12 As shown, after the user clicks the send button, the target voice text can be sent.
[0114] It is understood that this application modifies the voice text through voice modification instructions after the voice input is completed, ensuring the consistency of the voice input. During the voice modification process, the user is provided with all possible word changes, ensuring that the user can select the most appropriate modification result from reasonable options, thereby improving the accuracy and controllability of text modification.
[0115] The present application also provides a computer-readable storage medium for storing a computer program used in the terminal device 100. For example, the computer-readable storage medium may include, but is not limited to, various media capable of storing program codes, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0116] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or flowchart, and the combination of boxes in the structure diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.
[0117] In addition, the functional modules or units in the various embodiments of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0118] If a function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a smart phone, personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of this application.
[0119] The above is only a specific implementation method of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the protection scope of the present application.
Claims
1. A speech input method based on gesture interaction, characterized in that: include: In response to a voice input command received by the terminal device, entering the voice input interface and displaying the first gesture area; In response to a user's movement operation instruction in the first gesture area, controlling the display of a movement icon in the first gesture area and controlling the movement icon to follow the touch position of the user's finger; Controlling voice input and performing a first modification operation on the input voice text according to a movement state of the movement icon in the first gesture area; Under the condition that the first modification operation is performed on the voice text, obtaining and displaying corresponding recommended words of the retained voice text, the recommended words are used for the user to correct the retained voice text; wherein the first modification operation includes partial deletion and restoration of the voice text; After responding to the movement operation instruction of the first gesture area, displaying the second gesture area; In response to a user's movement operation instruction in the second gesture area, controlling the display of a movement icon in the second gesture area and controlling the movement icon to follow the touch position of the user's finger; controlling the continuation of voice input and performing a second modification operation on the input voice text according to the movement state of the movement icon in the second gesture area; The second gesture area includes a fourth function area and a fifth function area, and the controlling of continuing the voice input and performing a second modification operation on the input voice text according to the movement state of the moving icon in the second gesture area includes: continuing voice input or inputting a voice modification instruction according to the movement state of the moving icon in the fourth functional area; Under the condition that the voice modification instruction is input, controlling the second modification operation according to the movement state of the moving icon in the fifth functional area; and obtaining and displaying the corresponding recommended words of the retained voice text, including: Selecting a plurality of to-be-processed words from the retained speech text according to a preset direction; Selecting a plurality of candidate words as the recommended words based on semantic matching results between the word to be processed and a plurality of candidate words; The recommended words are sorted and displayed according to their usage frequency and semantic matching degree.
2. The speech input method based on gesture interaction according to claim 1, characterized in that: The first gesture area includes a first function area, a second function area, and a third function area. Controlling voice input and performing a first modification operation on the input voice text according to the movement state of the moving icon in the first gesture area include: controlling voice input according to a moving state of the moving icon in the first functional area to obtain the input voice text; controlling cancellation of voice input according to a movement state of the moving icon in the second functional area; A first modification operation is performed on the input voice text according to the moving state of the moving icon in the third function area.
3. The speech input method based on gesture interaction according to claim 2, characterized in that: The movement state of the mobile icon in the third functional area includes the user's finger sliding interval and finger sliding direction; The finger sliding interval and the finger sliding direction are used to select the target deleted text or restored text in the voice text; The finger sliding interval is also used to control the text deletion or restoration speed of the voice text.
4. The speech input method based on gesture interaction according to claim 1, characterized in that: The second gesture area also includes a sixth function area, and the method further includes: When the mobile icon moves to the fifth functional area, displaying the sixth functional area; According to the moving state of the moving icon in the sixth function area, control is performed to cancel the second modifying operation.
5. The speech input method based on gesture interaction according to claim 1, characterized in that: The controlling the second modification operation includes: determining the target word phonetics in the phonetically modified text; Mark all words in the speech text that are phonetically identical to the target word and are to be changed; Marking and displaying all modified words in the phonetically modified text that are phonetically identical to the target word; According to the moving state of the moving icon in the fifth function area, control is performed to modify the selected target word to be changed into the target word to be changed.
6. A terminal device, characterized in that: The terminal device includes a processor and a memory, the memory stores a computer program, and the processor is used to execute the voice input method based on gesture interaction as described in any one of claims 1 to 5.
7. A computer-readable storage medium, characterized in that It stores a computer program, which, when executed on a processor, implements the speech input method based on gesture interaction according to any one of claims 1-5.
Citation Information
Patent Citations
Voice input error correction method and system
CN103366741A
Information correction method and device and electronic equipment
CN112684913A