Text editing method, device, equipment and storage medium based on voice input

By combining voice and image acquisition instructions in the voice input interface, and receiving voice and user action instructions, the problem of unknown meaning of voice input is solved, and precise text content display and editing is achieved.

CN113590073BActive Publication Date: 2025-09-02TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110137739.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-01
Publication Date
2025-09-02
Estimated Expiration
2041-02-01

AI Technical Summary

Technical Problem

The prior art cannot accurately understand the meaning of the voice information expressed by the voice input, resulting in the information displayed during text editing that does not match the user's intention.

Method used

By presenting the voice input interface of the voice acquisition area and the text display area, combining image acquisition instructions, voice and user action instructions are received, accurate text content display and editing operations are achieved.

Benefits of technology

It realizes the precise display of text content of voice input and the editing operation that is different from text input, and text editing that conforms to user's intentions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113590073B_ABST
    Figure CN113590073B_ABST
Patent Text Reader

Abstract

The present application provides a text editing method, apparatus, device and computer-readable storage medium based on voice input; the method comprises: presenting a voice input interface including a voice acquisition area and a text display area, and presenting image acquisition indication information indicating that image acquisition is in progress in the voice input interface; receiving a text operation instruction based on the voice input interface; when the text operation instruction is a voice instruction triggered based on the voice acquisition area, presenting the text content input indicated by the voice instruction in the text display area; when the text operation instruction is a user action instruction triggered based on the image acquisition, executing the text editing operation indicated by the user action instruction in the text display area that is different from the text input. Through the present application, the text content of the voice input can be accurately displayed, and a text editing operation different from the text input can be executed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of Internet technology, and in particular to a text editing method, apparatus, device, and computer-readable storage medium based on voice input. Background Art

[0002] With the popularization of terminal devices and the development of voice recognition technology, users can quickly edit text through voice input in terminal devices; however, during the text editing process, related technologies cannot accurately know the meaning expressed by the voice information of the voice input, and can only display the text content of the voice input, but cannot perform text editing operations that are different from text input, resulting in the displayed text information being contrary to the user's intention. Summary of the Invention

[0003] Embodiments of the present application provide a text editing method, apparatus, device, and computer-readable storage medium based on voice input, which can accurately display the text content of voice input and perform text editing operations that are different from text input.

[0004] The technical solution of the embodiment of the present application is implemented as follows:

[0005] The present invention provides a text editing method based on voice input, including:

[0006] Presenting a voice input interface including a voice collection area and a text display area, and presenting image collection indication information indicating that image collection is in progress in the voice input interface;

[0007] Based on the voice input interface, receiving text operation instructions;

[0008] When the text operation instruction is a voice instruction triggered based on the voice collection area, presenting the text content input indicated by the voice instruction in the text display area;

[0009] When the text operation instruction is a user action instruction triggered based on the image acquisition, a text editing operation indicated by the user action instruction, which is different from text input, is performed in the text display area.

[0010] In the above solution, before presenting the voice input interface including the voice collection area and the text display area, the method further includes:

[0011] In the voice input interface, a voice input function item is presented;

[0012] In response to a trigger operation on the voice input function item, when the trigger operation is the first trigger operation on the voice input function item, presenting first permission prompt information indicating obtaining control permission for the image acquisition component;

[0013] In response to a determination operation on the first permission prompt information, second permission prompt information is presented for indicating that the control permission of the sound acquisition component and the image acquisition component has been obtained.

[0014] The present invention provides a text editing device based on voice input, comprising:

[0015] A first presentation module is configured to present a voice input interface including a voice collection area and a text display area, and to present image collection indication information indicating that image collection is in progress in the voice input interface;

[0016] An instruction receiving module, configured to receive text operation instructions based on the voice input interface;

[0017] A second presentation module is configured to present, in the text display area, text content input as indicated by the voice instruction when the text operation instruction is a voice instruction triggered based on the voice collection area;

[0018] An operation execution module is used to execute a text editing operation indicated by the user action instruction, which is different from text input, in the text display area when the text operation instruction is a user action instruction triggered based on the image acquisition.

[0019] In the above solution, the text operation instruction includes a voice instruction and a user action instruction, and the instruction receiving module is further used to present prompt information in the voice input interface;

[0020] When the prompt information is used to indicate the presence of voice input, in response to a determination operation on the prompt information, a voice instruction triggered based on the voice collection area is received;

[0021] When the prompt information is used to indicate that there is a user action of image acquisition, in response to a determination operation on the prompt information, a user action instruction triggered by the image acquisition is received.

[0022] In the above solution, the first presentation module is further used to present the voice input function item in the voice input interface;

[0023] In response to the triggering operation for the voice input function item, in the voice collection area, voice collection indication information for indicating that voice collection is in progress is presented, and

[0024] In the voice collection area, image collection indication information for indicating that image collection is in progress is presented.

[0025] In the above solution, the first presentation module is further configured to present a dynamically changing sequence of expressions in the voice collection area;

[0026] The expression sequence is used as image acquisition indication information indicating that image acquisition is in progress.

[0027] In the above solution, the device further comprises:

[0028] A cancel module is used to present a function item for ending voice input during the voice collection process;

[0029] In response to the triggering operation for the end voice input function item, the presented voice collection instruction information and the image collection instruction information are canceled.

[0030] In the above solution, when the text operation instruction is the user action instruction, after executing the text editing operation indicated by the user action instruction that is different from text input, the device further includes:

[0031] An operation prompt module, used for presenting operation instruction information corresponding to the text editing operation;

[0032] The operation instruction information includes at least one of graphics and text, which is used to indicate the operation content corresponding to the editing operation.

[0033] In the above solution, before presenting the voice input interface including the voice collection area and the text display area, the device further includes:

[0034] An authority verification module, configured to present a voice input function item in the voice input interface;

[0035] In response to a trigger operation on the voice input function item, when the trigger operation is the first trigger operation on the voice input function item, presenting first permission prompt information indicating obtaining control permission for the image acquisition component;

[0036] In response to a determination operation on the first permission prompt information, second permission prompt information is presented for indicating that the control permission of the sound acquisition component and the image acquisition component has been obtained.

[0037] In the above solution, the first presentation module is further used to send an authorization request for the sound acquisition component and the image acquisition component, and the authorization request carries verification information;

[0038] Wherein, the sound acquisition component is used for voice acquisition, and the image acquisition component is used for image acquisition;

[0039] After receiving the verification result of the legitimacy of the authorization request based on the verification information, the verification result is returned;

[0040] When the verification result indicates that the authorization request is legal, a voice input interface including a voice collection area and a text display area is presented.

[0041] In the above solution, the device further comprises:

[0042] Verification information generation module, used to generate a timestamp and obtain the currently logged-in user account;

[0043] generating a signature based on the timestamp and the user account;

[0044] An application identifier is obtained, and the application identifier and the signature are combined to obtain the verification information.

[0045] In the above solution, the operation execution module is used to obtain the text editing operation indicated by the user action instruction, and the text editing operation includes at least one of the following: text format editing, text style editing and text layout editing;

[0046] The text editing operation is performed in the text display area based on the cursor or the selected text content in the text display area.

[0047] In the above solution, the text operation instruction includes a voice instruction and a user action instruction, and the instruction receiving module is further used to obtain the input voice information in real time based on the voice input interface and to capture the user image in real time;

[0048] When voice information is acquired and the voice information meets the triggering condition of the voice instruction, determining the voice instruction corresponding to the voice information to receive the voice instruction triggered based on the voice collection area;

[0049] When it is determined that a corresponding user action instruction exists based on the collected user image, the user action instruction corresponding to the collected user image is acquired, so as to receive the user action instruction triggered based on the image collection.

[0050] In the above solution, before determining the voice instruction corresponding to the voice information, the instruction receiving module is further used to

[0051] Acquiring the volume intensity of the voice information, and when the volume intensity reaches a target intensity, determining that the voice information meets the triggering condition of the voice command; or,

[0052] Acquiring the semantic content of the voice information, and when the semantic content is the target content, determining that the voice information meets the triggering condition of the voice instruction; or,

[0053] The volume intensity of the voice information is obtained, and the user's mouth shape feature is extracted from the user image corresponding to the input time of the voice information. When the volume intensity reaches the target intensity and the user's mouth shape feature matches the target mouth shape feature, it is determined that the voice information meets the triggering condition of the voice command.

[0054] In the above solution, before obtaining the user action instruction corresponding to the collected user image, the instruction receiving module is further used to perform action recognition on the collected user image to obtain the corresponding user action;

[0055] When the user action is a target action, it is determined that there is a corresponding user action instruction.

[0056] In the above solution, before obtaining the user action instruction corresponding to the collected user image, the instruction receiving module is further configured to extract a target number of target images from the multiple user images when there are multiple user images collected;

[0057] Performing action recognition on each of the target images to obtain the user posture in each of the target images;

[0058] Determining a corresponding user action based on a user posture in each of the target images, and determining that a corresponding user action instruction exists when the user action is a target action;

[0059] The target image is a user image captured from a first time point when the voice information does not meet the triggering condition of the voice command and stops capturing at a second time point when the voice information meets the triggering condition of the voice command again.

[0060] In the above solution, the device further comprises:

[0061] an image capture module, configured to select, from the collected user images, a target number of user images within a target time period before a first time point as reference images, where the first time point is a time point at which the voice information does not meet a triggering condition of the voice command;

[0062] Performing action recognition on the reference image to obtain a corresponding action trend;

[0063] The instruction receiving module is further configured to determine that a corresponding user action instruction exists when the action trend matches the action trend of the target action and the user action is the target action.

[0064] In the above solution, before presenting the text content inputted by the voice instruction in the text display area, the device further includes:

[0065] A content acquisition module, configured to acquire voice information input based on the voice input interface;

[0066] Taking the time point when the voice information meets the triggering condition of the voice command as the starting point and the time point when the voice information does not meet the triggering condition of the voice command as the end point, intercepting the voice content between the starting point and the end point in the voice information;

[0067] The intercepted voice content is converted into text to obtain the text content input indicated by the voice instruction.

[0068] In the above solution, the content acquisition module is used to perform noise reduction processing on the voice content to obtain the noise-reduced voice content;

[0069] The de-noised voice content is converted into text to obtain the text content indicated by the voice instruction.

[0070] In the above solution, before the text content indicated by the voice instruction is presented in the text display area, the content acquisition module is also used to

[0071] Acquiring voice information input based on the voice input interface;

[0072] Taking the time point when the voice information meets the triggering condition of the voice command as the starting point, continuously intercepting the voice content of the voice information from the starting point onward, and stopping intercepting when the voice information no longer meets the triggering condition of the voice command;

[0073] As the voice input continues, the voice content of the voice information from the new starting point is continuously captured, starting from the time when the voice information again meets the triggering condition of the voice command, and the capture is stopped when the voice information again fails to meet the triggering condition of the voice command;

[0074] The intercepted voice content is converted into text to obtain the text content input indicated by the voice instruction.

[0075] An embodiment of the present application provides an electronic device, including:

[0076] a memory for storing executable instructions;

[0077] The processor is used to implement the text editing method based on voice input provided in the embodiment of the present application when executing the executable instructions stored in the memory.

[0078] An embodiment of the present application provides a computer-readable storage medium storing executable instructions for causing a processor to execute instructions to implement the text editing method based on voice input provided in the embodiment of the present application.

[0079] The embodiments of the present application have the following beneficial effects:

[0080] By presenting a voice input interface including a voice collection area and a text display area, and presenting image collection indication information indicating that image collection is in progress in the voice input interface; receiving text operation instructions based on the voice input interface; when the text operation instruction is a voice instruction triggered based on the voice collection area, presenting the text content input indicated by the voice instruction in the text display area; when the text operation instruction is a user action instruction triggered based on image collection, executing the text editing operation indicated by the user action instruction, which is different from the text input, in the text display area; in this way, the user's text operation instruction can be detected through the dual channels of voice collection and image collection, and when the text operation instruction is a voice instruction, the text input is controlled by the voice instruction, and the text content of the voice input is accurately displayed; when the text operation instruction is a user action instruction, the execution of the text editing operation, which is different from the text input, is controlled by the user action instruction, so that the text content and text editing operation that meet the user's intention can be accurately displayed and executed. BRIEF DESCRIPTION OF THE DRAWINGS

[0081] Figures 1A-1B A schematic diagram of a text editing interface based on voice input provided in an embodiment of the present application;

[0082] Figure 2 A schematic diagram of an optional architecture of a text editing system 100 based on voice input provided in an embodiment of the present application;

[0083] Figure 3 An optional structural diagram of an electronic device 500 provided in an embodiment of the present application;

[0084] Figure 4 An optional flowchart of a text editing method based on voice input provided in an embodiment of the present application;

[0085] Figure 5 A schematic diagram of the authorization interface for the voice input function provided in an embodiment of the present application;

[0086] Figure 6 A schematic diagram of an indication information display interface provided in an embodiment of the present application;

[0087] Figure 7 A schematic diagram of an indication information display interface provided in an embodiment of the present application;

[0088] Figure 8A schematic diagram of an indication information display interface provided in an embodiment of the present application;

[0089] Figure 9 Schematic diagram of the text operation instruction receiving interface provided in an embodiment of the present application;

[0090] Figure 10 Schematic diagram of the text operation instruction receiving interface provided in an embodiment of the present application;

[0091] Figure 11 A schematic diagram of the execution interface for text editing operations provided in an embodiment of the present application;

[0092] Figure 12 A schematic diagram of the execution interface for text editing operations provided in an embodiment of the present application;

[0093] Figure 13 A schematic diagram of a prompt information display interface provided in an embodiment of the present application;

[0094] Figure 14 A flowchart of a text editing method based on voice input provided in an embodiment of the present application;

[0095] Figure 15 A flow chart of the speech recognition method provided in an embodiment of the present application;

[0096] Figure 16 A schematic diagram of the structure of a text editing device based on voice input provided in an embodiment of the present application. DETAILED DESCRIPTION

[0097] In order to make the purpose, technical solutions and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting this application. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0098] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.

[0099] In the following description, the terms "first\second..." are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understandable that "first\second..." can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0100] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0101] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.

[0102] 1) Client: An application running in a terminal to provide various services, such as a video player client, a game client, etc.

[0103] 2) In response, it is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.

[0104] See also Figures 1A-1B , Figures 1A-1B This is a diagram of a text editing interface based on voice input provided by an embodiment of the present application. In the voice input process, when the user enters the voice information of "line break", Figure 1A In the process of text conversion, the terminal converts the voice information into text, obtains and displays the text content of "line break", but cannot perform the text editing operation of "line break"; Figure 1B In the embodiment of the present invention, the terminal recognizes the voice information and determines that the voice information is used to instruct the execution of the text editing operation of "line break", that is, the cursor wraps the line, but the text content of "line break" cannot be input by voice. It can be seen that the relevant technology cannot distinguish whether the voice information such as "line break" is used to instruct the presentation of the corresponding text content or the execution of the corresponding text editing operation of line break, that is, it is impossible to accurately know the meaning expressed by the voice information of the voice input, which may cause the displayed text information to be contrary to the user's intention.

[0105] In view of this, embodiments of the present application provide a text editing method, apparatus, device, and computer-readable storage medium based on voice input to at least solve the above-mentioned problems.

[0106] See also Figure 2 , Figure 2 An optional architectural diagram of a text editing system 100 based on voice input provided in an embodiment of the present application. To support an exemplary application, terminals (terminal 400-1 and terminal 400-2 are shown as examples) are connected to a server 200 via a network 300. The network 300 can be a wide area network or a local area network, or a combination of the two, and data transmission is achieved using a wireless link.

[0107] In actual applications, the terminal can be various types of user terminals such as smart phones, tablet computers, laptops, etc., and can also be desktop computers, game consoles, televisions, or a combination of any two or more of these data processing devices; the server 200 can be a separately configured server that supports various services, or it can be configured as a server cluster, or it can be a cloud server, etc.

[0108] In actual implementation, a client is provided on the terminal, such as a video playback client, an instant messaging client, a live broadcast client, a shopping client or other social clients. When the user opens the client on the terminal to edit text through voice input, the terminal is used to present a voice input interface including a voice collection area and a text display area, and in the voice input interface, presents image collection indication information for indicating that image collection is in progress; based on the voice input interface, receives text operation instructions, and sends an analysis request for the text operation instructions to the server 200; the server 200 determines an analysis result for representing the meaning indicated by the text operation instruction based on the analysis request, and returns the analysis result to the terminal; the terminal receives the analysis result, and when the analysis result represents that the text operation instruction is a voice instruction triggered based on the voice collection area, presents the text content indicated by the voice instruction in the text display area; when the analysis result represents that the text operation instruction is a user action instruction triggered based on image collection, executes the text editing operation indicated by the user action instruction in the text display area, which is different from the text input.

[0109] See also Figure 3 , Figure 3 This is an optional structural diagram of the electronic device 500 provided in the embodiment of the present application. In practical applications, the electronic device 500 may be Figure 2 The terminal or server 200 in the embodiment of the present invention is an electronic device. Figure 2 Taking the terminal shown as an example, an electronic device that implements the text editing method based on voice input according to an embodiment of the present application is described. Figure 3 The electronic device 500 shown includes: at least one processor 510, a memory 550, at least one network interface 520 and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to achieve connection and communication between these components. In addition to including a data bus, the bus system 540 also includes a power bus, a control bus and a status signal bus. However, for the sake of clarity, the bus system 540 is not shown in FIG. Figure 3 Various buses are labeled as bus system 540 .

[0110] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.

[0111] The user interface 530 includes one or more output devices 531 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.

[0112] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 550 may optionally include one or more storage devices that are physically remote from the processor 510.

[0113] The memory 550 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.

[0114] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplified below.

[0115] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0116] A network communication module 552 for reaching other computing devices via one or more (wired or wireless) network interfaces 520 , exemplary network interfaces 520 including Bluetooth, WiFi, and USB;

[0117] a presentation module 553 for enabling presentation of information via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with the user interface 530 (e.g., a user interface for operating peripheral devices and displaying content and information);

[0118] The input processing module 554 is configured to detect one or more user inputs or interactions from one of the one or more input devices 532 and to translate the detected inputs or interactions.

[0119] In some embodiments, the text editing device based on voice input provided by the embodiments of the present application can be implemented in software. Figure 3 A text editing device 555 based on voice input stored in a memory 550 is shown, which can be software in the form of a program and plug-in, etc., including the following software modules: a first presentation module 5551, an instruction receiving module 5552, a second presentation module 5553 and an operation execution module 5554. These modules are logical, and therefore can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.

[0120] In other embodiments, the text editing device based on voice input provided in the embodiments of the present application can be implemented in hardware. As an example, the text editing device based on voice input provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the information display method based on graphic identification provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic components.

[0121] Next, the text editing method based on voice input provided by the embodiment of the present application is described. In actual implementation, the text editing method based on voice input provided by the embodiment of the present application can be Figure 2 The server or terminal shown is implemented separately, and can also be implemented by Figure 2 The server and terminal shown in the figure are implemented in coordination. Figure 2 ,by Figure 2 The terminal shown implements the text editing method based on voice input provided in the embodiment of the present application as an example for illustration.

[0122] See also Figure 4 , Figure 4 An optional flow chart of a text editing method based on voice input provided in the embodiment of the present application is combined with Figure 4 The steps shown are explained.

[0123] Step 101: The terminal presents a voice input interface including a voice collection area and a text display area, and presents image collection indication information indicating that image collection is in progress in the voice input interface.

[0124] In actual applications, when the terminal is equipped with an instant messaging client, live broadcast client, shopping client or other social clients, when the user opens the client on the terminal to edit text through voice input, the terminal presents a voice collection area for collecting voice information, and a text display area for displaying the text content corresponding to the voice information. In addition, the text display area can also display the text editing results corresponding to the text editing operation triggered by the image collection indication information.

[0125] In some embodiments, when a user uses the voice input function for the first time, the terminal may authorize the voice input function in the following manner before presenting the voice input interface including the voice collection area and the text display area:

[0126] In the voice input interface, a voice input function item is presented; in response to a trigger operation for the voice input function item, when the trigger operation is the first trigger operation for the voice input function item, first permission prompt information indicating that the control permission of the image acquisition component is obtained is presented; in response to a confirmation operation for the first permission prompt information, second permission prompt information indicating that the control permission of the sound acquisition component and the image acquisition component has been obtained is presented.

[0127] Here, when a user uses the voice input function for the first time, the voice input state needs to be initialized. That is, it is necessary to obtain control permissions (such as access permissions) for both the image acquisition component (such as a camera) for image acquisition and the sound acquisition component (such as a microphone) for sound acquisition. After the control permissions for the image acquisition component and the sound acquisition component are granted, the sound acquisition component can be used to collect the voice information of the user's voice input, and the image acquisition component can be used to collect the user's image in real time to capture the user's movements in real time, thereby obtaining a user image sequence sufficient to recognize the user's movements.

[0128] In some embodiments, the terminal can determine whether the control authority of the image acquisition component and the sound acquisition component is permitted in the following manner: sending an authorization request for the sound acquisition component and the image acquisition component, wherein the authorization request carries verification information; receiving a verification result returned after the legitimacy of the authorization request is verified based on the verification information; accordingly, when the verification result indicates that the authorization request is legal, a voice input interface including a voice acquisition area and a text display area is presented.

[0129] In some embodiments, the terminal may obtain verification information by: generating a timestamp and obtaining the currently logged-in user account; generating a signature based on the generated timestamp and user account; obtaining an application identifier, and combining the application identifier and signature to obtain verification information.

[0130] Among them, the timestamp is the system time corresponding to when the user opens the client on the terminal to edit the text, and the user account is the user account currently logged in to the client. The signature is generated based on the system time and user account, and the application identifier corresponding to the sound acquisition component and the image acquisition component are spliced ​​or compressed with the generated signature to generate verification information for verifying the legitimacy of the authorization request.

[0131] Here, when the current user is a legitimate user and the sound acquisition component and the image acquisition component can operate normally, it is determined that the control authority over the image acquisition component and the sound acquisition component is permitted.

[0132] When displaying the first permission prompt information or the second permission prompt information, it can be presented through a pop-up window or a sub-interface independent of the voice input interface, wherein the size and position of the pop-up window or sub-interface can change with the user's drag operation.

[0133] See also Figure 5 , Figure 5 Schematic diagram of the authorization interface of the voice input function provided in an embodiment of the present application, in which a voice input function item 501 is presented in the voice input interface. When the user triggers the voice input function item 501 for the first time, the terminal responds to the triggering operation and presents a prompt interface 502 indicating the acquisition of control permission for the image acquisition component. In the prompt interface 502, a first permission prompt information 503 such as "Voice input wants to access your camera" is presented, as well as a confirmation function item 504 and a cancel function item 505 corresponding to the first permission prompt information 503. When the user does not allow the voice input function, the cancel function item 505 can be triggered. When the user allows the voice input function, the confirmation function item 504 can be triggered. In response to the triggering operation, the terminal presents a second permission prompt information 506 for indicating that the control permission for the sound acquisition component and the image acquisition component has been acquired.

[0134] In some embodiments, the terminal may present image acquisition indication information indicating that image acquisition is in progress in the voice input interface in the following manner:

[0135] In the voice input interface, a voice input function item is presented; in response to a trigger operation on the voice input function item, in the voice collection area, voice collection indication information for indicating that voice collection is in progress is presented, and in the voice collection area, image collection indication information for indicating that image collection is in progress is presented.

[0136] Here, when the voice input function mode is entered by triggering the voice input function item, the terminal presents a voice collection area, and presents voice collection indication information, such as voice ripple waves, in the voice collection area, and presents image collection indication information, such as an image capture indicator, in the voice collection area to indicate that user images are being captured in real time. The number of user images is multiple, that is, multiple user images constitute an image sequence to identify user mouth features or user actions.

[0137] See also Figure 6 , Figure 6 Schematic diagram of the indication information display interface provided in an embodiment of the present application. When the user triggers the voice input function item 601, the terminal responds to the trigger operation and presents voice collection indication information 603 in the voice collection area 602 to indicate that voice collection is in progress, and presents image collection indication information 604 in the voice collection area 602 to indicate that image collection is in progress.

[0138] In some embodiments, the terminal may present image acquisition indication information for indicating that image acquisition is in progress in the voice acquisition area in the following manner: presenting a dynamically changing sequence of expressions in the voice acquisition area; and using the dynamically changing sequence of expressions as image acquisition indication information indicating that image acquisition is in progress.

[0139] Here, the dynamically changing expression sequence can be a fixed number of expression cycles, see Figure 7 , Figure 7 The schematic diagram of the indication information display interface provided in the embodiment of the present application uses a dynamically changing expression sequence 701 as image acquisition indication information indicating that image acquisition is in progress.

[0140] In some embodiments, during the voice collection process, the terminal may also present an end voice input function item; in response to a trigger operation for the end voice input function item, the presented voice collection indication information and image collection indication information are canceled.

[0141] Here, when the user triggers the end voice input function item, the terminal responds to the trigger operation and exits the voice input function mode, that is, cancels the acquisition instruction information and the image acquisition instruction information.

[0142] See also Figure 8 , Figure 8 This is a schematic diagram of the indication information display interface provided in an embodiment of the present application. When the user triggers the end voice input function item, the terminal responds to the trigger operation and cancels the presented voice collection indication information and image collection indication information.

[0143] Step 102: Receive text operation instructions based on the voice input interface.

[0144] In some embodiments, the text operation instruction includes a voice instruction and a user action instruction. The terminal can receive the text operation instruction based on the voice input interface in the following manner:

[0145] Prompt information is presented in the voice input interface; when the prompt information is used to indicate the presence of voice input, a voice instruction triggered based on the voice collection area is received in response to a confirmation operation on the prompt information; when the prompt information is used to indicate the presence of a user action of image collection, a user action instruction triggered based on image collection is received in response to a confirmation operation on the prompt information.

[0146] See also Figure 9 , Figure 9 A schematic diagram of the text operation instruction receiving interface provided in an embodiment of the present application. When the user triggers the voice input function item 901 and inputs voice information, the terminal receives the voice information input by the user, and presents a prompt information 902 for indicating the existence of voice information input, and presents a cancel function item 903 and a confirmation function item 904 corresponding to the prompt information 902. When the user triggers the confirmation function item 904, the terminal responds to the trigger operation and determines that a voice instruction triggered based on the voice collection area has been received.

[0147] See also Figure 10 , Figure 10 A schematic diagram of the text operation instruction receiving interface provided in an embodiment of the present application, when the user triggers the voice input function item 1001, the terminal performs image capture on the user image, and when a user action is recognized in the user image, a prompt message 1002 is presented for indicating the presence of voice input and the presence of a user action for image capture, and a cancel function item 1003 and a confirmation function 1004 corresponding to the prompt message 1002 are presented. When the user triggers the confirmation function item 1004, the terminal responds to the trigger operation and determines that a user action instruction triggered based on image capture has been received.

[0148] In some embodiments, text operation instructions include voice instructions and user action instructions. The terminal can receive text operation instructions based on the voice input interface in the following manner: based on the voice input interface, obtain input voice information in real time and collect user images in real time; when voice information is obtained and the voice information meets the triggering conditions of the voice instruction, determine the voice instruction corresponding to the voice information to receive the voice instruction triggered based on the voice collection area; when it is determined that there is a corresponding user action instruction based on the collected user image, obtain the user action instruction corresponding to the collected user image to receive the user action instruction triggered based on the image collection.

[0149] Here, in the process of acquiring voice information input by voice, the acquired voice information and the collected user image are recognized in real time. In actual applications, in the same time sequence (i.e., within the same time period), the terminal either receives a voice command and displays the text content input as indicated by the voice command in the text display area, or receives a user action command and performs a text editing operation indicated by the user action command that is different from the text input. Therefore, it is necessary to recognize the real-time collected voice information and user image to determine whether the user is currently inputting voice information or performing a user action. In actual implementation, if the voice information meets the triggering conditions of the voice command, the corresponding voice command can be triggered. After the user stops entering voice information, if the corresponding user action command exists in the user image, the user action command corresponding to the collected user image can be obtained.

[0150] In some embodiments, before determining the voice instruction corresponding to the voice information, the terminal may determine whether the acquired voice information meets the triggering condition of the voice instruction by:

[0151] Obtain the volume intensity of the voice information, and when the volume intensity reaches the target intensity, determine that the voice information meets the triggering conditions of the voice command; or, obtain the semantic content of the voice information, and when the semantic content is the target content, determine that the voice information meets the triggering conditions of the voice command; or, obtain the volume intensity of the voice information, and extract the user's mouth shape features from the user image corresponding to the input time of the voice information, and when the volume intensity reaches the target intensity and the user's mouth shape features match the target mouth shape features, determine that the voice information meets the triggering conditions of the voice command.

[0152] Here, the terminal performs volume intensity detection on the voice information of each audio frame acquired in real time, and when the volume intensity of the voice information of the current audio frame reaches the target intensity, it is considered that the voice information is valid voice information being input by the user, and the valid voice information can trigger the corresponding voice command. Alternatively, the terminal performs semantic analysis on the voice information acquired in real time to obtain the semantic content corresponding to the voice information. When the semantic content is the pre-stored target content for indicating the triggering condition for triggering the voice command, or the semantic content matches the target content, it is determined that the voice information meets the triggering condition of the voice command. For example, assuming that the pre-stored target content is "start voice input", when the semantic content corresponding to the voice information is "start voice input" or "start voice input, please present the corresponding text content", etc., it is determined that the voice information can trigger the voice command, that is, when the user enters the voice information, the terminal can receive the corresponding voice command.

[0153] In actual implementation, in addition to obtaining the volume intensity of the voice information, the user's mouth shape features can also be extracted from the user image corresponding to the input time of the voice information. When the volume intensity reaches the target intensity and the user's mouth shape features match the target mouth shape features, it is determined that the obtained voice information meets the triggering conditions of the voice command. In this way, through the dual judgment of volume intensity and user mouth shape features, the accuracy of voice recognition can be greatly improved.

[0154] In some embodiments, before obtaining the user action instruction corresponding to the collected user image, the terminal may determine whether the corresponding user action instruction exists in the collected user image in the following manner: performing action recognition on the collected user image to obtain the corresponding user action; when the user action is a target action, determining whether the corresponding user action instruction exists.

[0155] In some embodiments, before obtaining the user action instruction corresponding to the collected user image, the terminal may determine whether the corresponding user action instruction exists in the collected user image in the following manner: when the number of collected user images is multiple, extracting a target number of target images from the multiple user images; performing action recognition on each target image separately to obtain the user posture in each target image; determining the corresponding user action based on the user posture in each target image, and when the user action is a target action, determining the existence of the corresponding user action instruction.

[0156] Among them, the target image is a user image captured from the first time point when the acquired voice information does not meet the triggering conditions of the voice command, and stops capturing at the second time point when the acquired voice information meets the triggering conditions of the voice command again; that is, the target image used to identify user actions is the user image captured by the image acquisition component after the user stops entering voice information, that is, the voice input end point (first time point) is used as the image acquisition starting point, and the new voice input start point (second time point) located after the aforementioned voice input end point is used as the image acquisition end point, and the user image between the image acquisition starting point and the image acquisition end point is captured for identifying user action instructions.

[0157] In some embodiments, the terminal may also select a target number of user images within a target time period before a first time point from the collected user images as reference images, wherein the first time point is the time point when the voice information does not meet the triggering condition of the voice command; perform action recognition on the reference image to obtain the corresponding action trend; when the action trend meets the action trend of the target action and the user action is the target action, it is determined that there is a corresponding user action instruction

[0158] Here, from the captured user images, user images between a third time point that is a target time length before the first time point and the first time point are captured as reference images; wherein the third time point is before the first time point, that is, the user enters voice information at the third time point. In addition to analyzing the voice information corresponding to the third time point, action recognition is also performed on the user images captured at the third time point and a period of time after the third time point. This is because it is taken into account that the user may have a tendency to start making actions before the voice input ends. For example, the user action of "nodding" represents the "enter symbol", which is used to control the execution of the enter operation. Before making the user action of "nodding", the user may have an unconscious tendency to raise his head; for example, the user action of "quotation mark gesture" represents the "double quotation mark", which is used to control the presentation of the "double quotation mark" in the text display area. Before making the user action of "quotation mark gesture", the user often raises his hand. Therefore, multiple frames before and after the sequence frame interval between the image capture start point and the image capture end point can be captured to assist in the judgment, so as to further improve the accuracy of action recognition.

[0159] Step 103: When the text operation instruction is a voice instruction triggered based on the voice collection area, the text content indicated by the voice instruction input is presented in the text display area.

[0160] In some embodiments, before the terminal presents the text content indicated by the voice instruction in the text display area, the terminal may obtain the text content indicated by the voice instruction in the following manner:

[0161] Acquire voice information input based on a voice input interface; take the time point when the acquired voice information meets the triggering condition of the voice instruction as the starting point, and the time point when the acquired voice information does not meet the triggering condition of the voice instruction as the end point, and intercept the voice content between the starting point and the end point shown in the acquired voice information; convert the intercepted voice content into text to obtain the text content indicated by the voice instruction.

[0162] Here, in actual implementation, the terminal determines whether the voice information of each audio frame acquired in real time meets the triggering conditions of the voice command. The specific determination method can be referred to the determination method described above. For example, the voice information of each audio frame is detected for volume intensity. When the volume intensity of the voice information of the current audio frame reaches the target intensity, it is further determined whether the volume intensity of the voice information of the audio frame immediately preceding the current audio frame is less than the target intensity. If the volume intensity of the voice information of the audio frame immediately preceding the current audio frame is less than the target intensity, the time point corresponding to the current audio frame is determined as the voice input start point. As the voice information is input, starting from the voice input start point, it is determined whether the volume intensity of subsequent voice information is less than the target intensity. If a voice information with a volume intensity less than the target intensity subsequently appears, the time point corresponding to the audio frame immediately preceding the audio frame with a volume intensity less than the target intensity is determined as the voice input end point. After determining the voice input start point and voice input end point of the user's actual voice input, the voice information between the voice input start point and the voice input end point is intercepted to eliminate voice information that does not represent the user's actual input. The intercepted voice information is then converted into text to obtain corresponding text content, which is then presented in the text display area.

[0163] In some embodiments, the terminal may perform text conversion on the intercepted voice content to obtain the text content indicated by the voice instruction in the following manner: perform noise reduction processing on the voice content to obtain the noise-reduced voice content; perform text conversion on the noise-reduced voice content to obtain the text content indicated by the voice instruction.

[0164] Here, the intercepted voice information is denoised and then converted into text to obtain the corresponding text content, which is beneficial to improving the accuracy of voice recognition. At the same time, some invalid and meaningless voice data is eliminated and filtered out, which can reduce user traffic expenditure and reduce the functional consumption of terminal equipment.

[0165] In some embodiments, before the terminal presents the text content indicated by the voice instruction in the text display area, the terminal may obtain the text content indicated by the voice instruction in the following manner:

[0166] Acquire voice information input based on a voice input interface; take the time point when the acquired voice information meets the triggering condition of the voice instruction as the starting point, continuously intercept the voice content from the starting point onwards in the acquired voice information, and stop intercepting when the acquired voice information does not meet the triggering condition of the voice instruction; as the voice input proceeds, take the time point when the acquired voice information meets the triggering condition of the voice instruction again as a new starting point, continue to intercept the voice content from the new starting point onwards in the acquired voice information, and stop intercepting when the acquired voice information does not meet the triggering condition of the voice instruction again; convert the intercepted voice content into text to obtain the text content indicated by the voice instruction.

[0167] For example, the time point when collection is stopped is the end point, and the voice information between the starting point and the end point can form a voice segment. Repeat the above steps to determine the new starting point and end point, and multiple real-time voice segments can be generated. The real-time input voice information between the end point of each starting point and the starting point can be used to generate a real-time voice segment. The repeatedly determined new starting point and new end point are used to generate another real-time voice segment, and each voice segment is used as the collected voice content.

[0168] Step 104: When the text operation instruction is a user action instruction triggered based on image acquisition, a text editing operation indicated by the user action instruction, which is different from text input, is performed in the text display area.

[0169] In some embodiments, the terminal can execute a text editing operation indicated by the user action instruction, which is different from text input, in the text display area in the following manner: obtain the text editing operation indicated by the user action instruction, wherein the text editing operation includes at least one of the following: text format editing, text style editing, and text layout editing; based on the cursor or selected text content in the text display area, execute the text editing operation in the text display area.

[0170] In actual applications, for text format editing, such as switching the font of the presented text content, such as the user action of "blinking", the corresponding user action instruction is "switch font". Before the user performs the "blinking" action, if the user has selected part of the content in the text display area, the object of font switching is that part of the content. If the user has not selected any text content, the font of the entire text is switched; for text style editing, such as underlining, such as the user action of "shaking his head", the corresponding user action instruction is "underlining". Before the user performs the "shaking his head", if the user has selected part of the content in the text display area, the object of underlining is that part of the content. If the user has not selected any text content, the entire text is underlined; for text layout editing, such as line break, if the user performs a gesture and determines that a line break operation needs to be performed, the object of the operation is the content after the cursor.

[0171] See also Figure 11-12 , Figure 11-12 This is a schematic diagram of the execution interface of the text editing operation provided in the embodiment of the present application. Figure 11 In the example, before the user performs the "blink" action corresponding to switching the font, part of the content 1101 in the text display area has been selected, and the font is switched for the part of the content 1101, that is, only the font of the part of the content 1101 is switched; Figure 12 In FIG, before the user performs the “shaking head” action corresponding to switching the font, no text content is selected, and the entire text 1202 is underlined.

[0172] In some embodiments, when the text operation instruction is a user action instruction, the terminal may also present operation instruction information corresponding to the text editing operation after executing the text editing operation indicated by the user action instruction that is different from text input; wherein the operation instruction information includes at least one of graphics and text, which is used to indicate the operation content corresponding to the editing operation.

[0173] See also Figure 13 , Figure 13 This is a schematic diagram of the prompt information display interface provided in an embodiment of the present application. After the terminal executes the text editing operation indicated by the user action instruction, which is different from the text input, the terminal displays a text such as "Execute Line Break" and The operation instruction information 1301 of this graphic simultaneously presents the corresponding operation content 1302 on the text editing interface, that is, the operation content of the cursor wrapping is presented.

[0174] The following describes an exemplary application of the embodiments of the present application in a practical application scenario.

[0175] See also Figure 14 , Figure 14 This is a flow chart of a text editing method based on voice input provided in an embodiment of the present application. The method is implemented by a terminal and a server in collaboration, wherein a client is provided on the terminal. The server is a backend server corresponding to the client on the terminal. If the client is an instant messaging client, the server is a backend server corresponding to the instant messaging client. The user action recognition software development kit (SDK) and the voice recognition SDK can be provided on the terminal or in the server. Figure 14 The steps shown are explained.

[0176] Step 201: The terminal sends a request to obtain a symbol configuration table corresponding to a user action instruction to a server.

[0177] Here, the server stores a symbol configuration table corresponding to user action instructions. The symbol configuration table is used to store the correspondence between user actions and user action instructions. For example, if the user action is "blinking", the corresponding user action instruction is "switch font"; if the user action is "slide gesture to the right", the corresponding user action instruction is "adjust input cursor position", and so on.

[0178] When the user uses the voice input function for the first time, that is, when the user opens the client on the terminal and runs the voice input function on the client, the terminal responds to the user's opening operation and requests the server to obtain the symbol configuration table corresponding to the user action instruction for subsequent judgment of the user action instruction.

[0179] It should be noted that, in actual applications, the symbol configuration table corresponding to the user action instruction may also be stored in the terminal. In this case, the symbol configuration table corresponding to the user action instruction may be directly obtained locally to determine the user action instruction.

[0180] Step 202: Based on the acquisition request, the server obtains and returns a symbol configuration table corresponding to the user action instruction to the terminal.

[0181] Step 203: The terminal receives the symbol configuration table corresponding to the user action instruction returned by the server.

[0182] Step 204: The terminal presents first permission prompt information indicating obtaining the control permission of the image acquisition component and the voice acquisition component.

[0183] Here, when the client runs the voice input function, the terminal needs to determine whether the user has granted control permissions (such as access permissions) to the image acquisition component (such as a camera) and the voice acquisition component (such as a microphone). If not authorized, the terminal renders an authorization prompt pop-up window, and presents the first permission prompt information in the authorization prompt pop-up window for indicating the acquisition of control permissions for the image acquisition component and the voice acquisition component.

[0184] Step 205: In response to the confirmation operation on the first permission prompt information, the terminal presents second permission prompt information indicating that the control permission of the sound acquisition component and the image acquisition component has been obtained.

[0185] Here, when the user confirms to grant the control permission of the camera and microphone based on the first permission prompt information, it can be achieved by a confirmation operation on the first permission prompt information. In response to the user's confirmation operation, the terminal renders and presents a pop-up window prompting a successful authorization, and presents the second permission prompt information indicating that the control permission of the camera and microphone has been obtained in the pop-up window prompting a successful authorization.

[0186] Step 206: The terminal activates the image acquisition component and the voice acquisition component in response to the triggering operation on the voice input function item.

[0187] Here, the terminal presents a voice input function item in the voice input interface. When the user triggers the voice input function item, the terminal responds to the trigger operation, activates the voice acquisition component (such as a microphone), turns on the voice recording mode, and turns on the image acquisition component (such as a camera). That is, the terminal presents voice acquisition indication information for indicating that voice acquisition is in progress, and presents image acquisition indication information for indicating that image acquisition is in progress.

[0188] Step 207: In response to the user's voice input operation, the terminal obtains the input voice information in real time based on the voice input interface, collects the user image in real time, and sends the obtained voice information and user image to the user action recognition SDK.

[0189] Step 208: When the user action recognition SDK determines that the voice information meets the triggering condition of the voice command, the voice information is sent to the voice recognition SDK.

[0190] Here, in actual implementation, the user action SDK obtains the volume intensity of the voice information and extracts the user's mouth shape features from the user image corresponding to the input time of the voice information. When the volume intensity reaches the target intensity and the user's mouth shape features match the target mouth shape features, it is determined that the obtained voice information meets the trigger conditions of the voice command, and the voice information that meets the trigger conditions of the voice command is sent to the voice recognition SDK for voice recognition.

[0191] Before sending voice information, a communication connection (such as a long link) must be established between the user action recognition SDK and the voice recognition SDK. Through the established communication connection, the user action recognition SDK will send the voice information that meets the trigger conditions of the voice command to the voice recognition SDK for voice recognition.

[0192] Step 209: The voice recognition SDK performs content recognition on the voice information and returns the recognition result to the user action recognition SDK.

[0193] Here, the speech recognition SDK performs content recognition on the speech information, that is, converts the speech information into text, obtains the corresponding text content, and uses the obtained text content as the recognition result.

[0194] Step 210: The user action recognition SDK returns the recognition result to the terminal.

[0195] Step 211: The terminal presents the recognition result in the text display area.

[0196] Step 212: When the user stops voice input, the terminal sends the user image collected after the voice input stops to the motion recognition SDK.

[0197] Here, the voice input end point is used as the image acquisition starting point, and the new voice input starting point located after the previous voice input end point is used as the image acquisition end point. The user image between the image acquisition starting point and the image acquisition end point is captured for identifying user action instructions.

[0198] In actual applications, it is considered that the user may have a tendency to start making movements before the voice input is completed. For example, the user action of "nodding" represents the "enter symbol", which is used to control the execution of the enter operation. The user may have an unconscious tendency to raise his head before making the user action of "nodding"; for another example, the user action of "quotation mark gesture" represents "double quotation marks", which is used to control the presentation of "double quotation marks" in the text display area. The user often raises his hand before making the user action of "quotation mark gesture". Therefore, the method of capturing the first and last frames of the sequence frame interval between the image acquisition starting point and the image acquisition end point can be used to assist in judgment, so as to further improve the accuracy of action recognition.

[0199] Step 213: The action recognition SDK performs action recognition on the user image, obtains and returns the corresponding user action to the terminal.

[0200] Step 214: The terminal obtains a corresponding action instruction from the symbol configuration table based on the user action returned by the action recognition SDK, and executes a text editing operation indicated by the user action instruction that is different from text input.

[0201] Step 215: In response to the triggering operation for ending the voice input function item, the terminal cancels the presented voice collection instruction information and image collection instruction information.

[0202] Here, when the user triggers the end voice input function item to exit the voice input, the terminal cancels the presented voice collection instruction information and image collection instruction information, and disconnects the long link with the server.

[0203] In actual applications, in the same timing, the terminal either receives a voice command and presents the text content indicated by the voice command in the text display area, or receives a user action command and executes a text editing operation indicated by the user action command that is different from text input. The above steps 207 to 210 are used to determine whether the user is inputting voice information, recognize the voice information being input by the user, and return the recognized results to the terminal.

[0204] In actual implementation, it can also be combined with Figure 15 Determine whether the user is inputting voice information. Figure 15 The flow chart of the speech recognition method provided in the embodiment of the present application will be combined with Figure 15The steps shown illustrate the speech recognition method (ie, steps 207 to 210 ) provided in the embodiment of the present application.

[0205] Step 301: In response to a trigger operation on a voice input function item, the terminal obtains input voice information in real time and captures a user image in real time.

[0206] Here, when the user triggers the voice input function item, the terminal activates the voice acquisition component (such as a microphone), turns on the voice input mode to collect voice information, and turns on the image acquisition component (such as a camera) to collect user images.

[0207] Step 302: When the volume intensity of the current audio frame reaches the target intensity, the terminal determines whether the volume intensity of the previous audio frame is less than the target intensity.

[0208] Here, the terminal performs volume intensity detection on the voice information of each audio frame obtained in real time. When the volume intensity of the voice information of the current audio frame reaches the target intensity, it further determines whether the volume intensity of the voice information of the previous adjacent audio frame of the current audio frame is less than the target intensity to determine whether the current audio frame is the starting point of voice input.

[0209] If the volume intensity of the voice information of the previous adjacent audio frame of the current audio frame is less than the target intensity, step 303 is executed; if the volume intensity of the voice information of the previous adjacent audio frame of the current audio frame is not less than the target intensity, step 304 is executed.

[0210] Step 303: Mark the time point corresponding to the current audio frame as the speech input starting point.

[0211] Step 304: If the volume level of the voice information of the adjacent audio frame before the current audio frame is not less than the target level, determine whether the volume level of the voice information of the adjacent audio frame after the current audio frame reaches the target level.

[0212] Here, as the voice information is input, starting from the voice input starting point, it is determined whether the volume intensity of the subsequent voice information is lower than the target intensity to determine the voice input ending point.

[0213] If the volume level of the voice information of the subsequent adjacent audio frame of the current audio frame does not reach the target level, step 305 is executed; otherwise, step 301 is executed.

[0214] Step 305: Mark the time point corresponding to the current audio frame as the end point of the speech input.

[0215] Here, when voice information with a volume intensity lower than the target intensity appears subsequently, the time point corresponding to the previous adjacent frame of the audio frame with a volume intensity lower than the target intensity is considered to be the voice input end point.

[0216] Step 306: The terminal intercepts the voice information and user image between the voice input start point and the voice input end point.

[0217] Step 307: The terminal sends the intercepted voice information and user image to the recognition SDK for recognition.

[0218] Here, the terminal may further perform noise reduction processing on the intercepted voice information, and send the noise-reduced voice information and the user image to the recognition SDK for recognition, wherein the recognition SDK includes the above-mentioned user action recognition SDK and voice recognition SDK.

[0219] The terminal sends the intercepted voice information and user image to the user action recognition SDK. The user action recognition SDK can further determine the voice information actually input by the user based on the voice information and user image to further intercept the voice information actually input by the user to eliminate the voice information that fails to represent the actual input of the user. The user action SDK obtains the volume intensity of the voice information and extracts the user's mouth shape features from the user image corresponding to the input time of the voice information. When the volume intensity reaches the target intensity and the user's mouth shape features match the target mouth shape features, it is determined that the acquired voice information meets the triggering conditions of the voice command (that is, the voice information actually input by the user), and the voice information that meets the triggering conditions of the voice command is sent to the voice recognition SDK for voice recognition. Then, the voice recognition SDK performs content recognition on the voice information actually input by the user, that is, converts the voice information into text, obtains the corresponding text content, and returns the obtained text content to the terminal as the recognition result.

[0220] Step 308: The terminal receives the recognition result returned by the recognition SDK.

[0221] Through the above method, when determining the voice information actually input by the user, the volume intensity of the voice information and the user's mouth shape features in the corresponding time sequence of the user image are combined to improve the accuracy of the determination; before performing voice recognition, the voice information is subjected to noise reduction processing, which is conducive to improving the accuracy of voice recognition, and at the same time, some invalid and meaningless voice data are eliminated and filtered out, which can reduce user traffic expenditure and reduce the functional consumption of the terminal device; when performing user action recognition, a few frames before and after the starting point or end point of the image acquisition are captured to assist in the judgment, thereby improving the accuracy of action recognition; finally, the voice recognition results and action recognition results are arranged in chronological order, and the text editing information that meets the user's intention is drawn and rendered on the terminal interface, thereby improving the user experience.

[0222] The following continues to describe the exemplary structure of the text editing device 555 based on voice input provided by the embodiment of the present application as a software module. In some embodiments, such as Figure 16 As shown, Figure 16 This is a schematic diagram of the structure of a text editing device based on voice input provided in an embodiment of the present application. The software modules stored in the text editing device based on voice input 555 in the memory 550 may include:

[0223] A first presentation module 5551 is configured to present a voice input interface including a voice collection area and a text display area, and to present image collection indication information indicating that image collection is in progress in the voice input interface;

[0224] An instruction receiving module 5552 is configured to receive text operation instructions based on the voice input interface;

[0225] A second presentation module 5553 is configured to present, in the text display area, the text content input as indicated by the voice instruction when the text operation instruction is a voice instruction triggered based on the voice collection area;

[0226] The operation execution module 5554 is used to execute the text editing operation indicated by the user action instruction, which is different from text input, in the text display area when the text operation instruction is a user action instruction triggered based on the image acquisition.

[0227] In some embodiments, the text operation instruction includes a voice instruction and a user action instruction, and the instruction receiving module is further used to present prompt information in the voice input interface;

[0228] When the prompt information is used to indicate the presence of voice input, in response to a determination operation on the prompt information, a voice instruction triggered based on the voice collection area is received;

[0229] When the prompt information is used to indicate that there is a user action of image acquisition, in response to a determination operation on the prompt information, a user action instruction triggered by the image acquisition is received.

[0230] In some embodiments, the first presentation module is further configured to present a voice input function item in the voice input interface;

[0231] In response to the triggering operation for the voice input function item, in the voice collection area, voice collection indication information for indicating that voice collection is in progress is presented, and

[0232] In the voice collection area, image collection indication information for indicating that image collection is in progress is presented.

[0233] In some embodiments, the first presentation module is further configured to present a dynamically changing sequence of expressions in the voice collection area;

[0234] The expression sequence is used as image acquisition indication information indicating that image acquisition is in progress.

[0235] In some embodiments, the apparatus further comprises:

[0236] A cancel module is used to present a function item for ending voice input during the voice collection process;

[0237] In response to the triggering operation for the end voice input function item, the presented voice collection instruction information and the image collection instruction information are canceled.

[0238] In some embodiments, when the text operation instruction is the user action instruction, after performing the text editing operation indicated by the user action instruction that is different from text input, the device further includes:

[0239] An operation prompt module, used for presenting operation instruction information corresponding to the text editing operation;

[0240] The operation instruction information includes at least one of graphics and text, which is used to indicate the operation content corresponding to the editing operation.

[0241] In some embodiments, before presenting the image acquisition indication information indicating that image acquisition is in progress, the apparatus further comprises:

[0242] An authority verification module, configured to present a voice input function item in the voice input interface;

[0243] In response to a trigger operation on the voice input function item, when the trigger operation is the first trigger operation on the voice input function item, presenting first permission prompt information indicating obtaining control permission for the image acquisition component;

[0244] In response to a determination operation on the permission prompt information, second permission prompt information is presented for indicating that the control permission of the sound acquisition component and the image acquisition component has been obtained.

[0245] In some embodiments, the first presentation module is further configured to send an authorization request for the sound acquisition component and the image acquisition component, wherein the authorization request carries verification information;

[0246] Wherein, the sound acquisition component is used for voice acquisition, and the image acquisition component is used for image acquisition;

[0247] After receiving the verification result of the legitimacy of the authorization request based on the verification information, the verification result is returned;

[0248] When the verification result indicates that the authorization request is legal, a voice input interface including a voice collection area and a text display area is presented.

[0249] In some embodiments, the apparatus further comprises:

[0250] Verification information generation module, used to generate a timestamp and obtain the currently logged-in user account;

[0251] generating a signature based on the timestamp and the user account;

[0252] An application identifier is obtained, and the application identifier and the signature are combined to obtain the verification information.

[0253] In some embodiments, the operation execution module is used to obtain the text editing operation indicated by the user action instruction, and the text editing operation includes at least one of the following: text format editing, text style editing, and text layout editing;

[0254] The text editing operation is performed in the text display area based on the cursor or the selected text content in the text display area.

[0255] In some embodiments, the text operation instruction includes a voice instruction and a user action instruction, and the instruction receiving module is further used to obtain input voice information in real time based on the voice input interface and to capture a user image in real time;

[0256] When voice information is acquired and the voice information meets the triggering condition of the voice instruction, determining the voice instruction corresponding to the voice information to receive the voice instruction triggered based on the voice collection area;

[0257] When it is determined that a corresponding user action instruction exists based on the collected user image, the user action instruction corresponding to the collected user image is acquired, so as to receive the user action instruction triggered based on the image collection.

[0258] In some embodiments, before determining the voice instruction corresponding to the voice information, the instruction receiving module is further configured to:

[0259] Acquiring the volume intensity of the voice information, and when the volume intensity reaches a target intensity, determining that the voice information meets the triggering condition of the voice command; or,

[0260] Acquiring the semantic content of the voice information, and when the semantic content is the target content, determining that the voice information meets the triggering condition of the voice instruction; or,

[0261] The volume intensity of the voice information is obtained, and the user's mouth shape feature is extracted from the user image corresponding to the input time of the voice information. When the volume intensity reaches the target intensity and the user's mouth shape feature matches the target mouth shape feature, it is determined that the voice information meets the triggering condition of the voice command.

[0262] In some embodiments, before obtaining the user action instruction corresponding to the collected user image, the instruction receiving module is further configured to perform action recognition on the collected user image to obtain the corresponding user action;

[0263] When the user action is a target action, it is determined that there is a corresponding user action instruction.

[0264] In some embodiments, before acquiring the user action instruction corresponding to the collected user image, the instruction receiving module is further configured to extract a target number of target images from the plurality of user images when there are multiple collected user images;

[0265] Performing action recognition on each of the target images to obtain the user posture in each of the target images;

[0266] Based on the user posture in each of the target images, a corresponding user action is determined, and when the user action is a target action, it is determined that a corresponding user action instruction exists.

[0267] The target image is a user image captured from the first time point when the voice information does not meet the triggering condition of the voice command and stops capturing at the second time point when the voice information meets the triggering condition of the voice command again.

[0268] In some embodiments, the apparatus further comprises:

[0269] an image interception module, configured to intercept, from the collected user images, a user image between a third time point before the first time point by a target time length and the first time point as a reference image;

[0270] Performing action recognition on the reference image to obtain a corresponding action trend;

[0271] The instruction receiving module is further configured to determine that a corresponding user action instruction exists when the action trend matches the action trend of the target action and the user action is the target action.

[0272] In the above solution, before presenting the text content inputted by the voice instruction in the text display area, the device further includes:

[0273] A content acquisition module, configured to acquire voice information input based on the voice input interface;

[0274] Taking the time point when the voice information meets the triggering condition of the voice command as the starting point and the time point when the voice information does not meet the triggering condition of the voice command as the end point, intercepting the voice content between the starting point and the end point in the voice information;

[0275] The intercepted voice content is converted into text to obtain the text content input indicated by the voice instruction.

[0276] In some embodiments, the content acquisition module is configured to perform noise reduction processing on the voice content to obtain noise-reduced voice content;

[0277] The de-noised voice content is converted into text to obtain the text content indicated by the voice instruction.

[0278] In some embodiments, before the text content indicated by the voice instruction is presented in the text display area, the content acquisition module is further configured to:

[0279] Acquiring voice information input based on the voice input interface;

[0280] Taking the time point when the voice information meets the triggering condition of the voice command as the starting point, continuously intercepting the voice content of the voice information from the starting point onward, and stopping intercepting when the voice information no longer meets the triggering condition of the voice command;

[0281] As the voice input continues, the voice content of the voice information from the new starting point is continuously captured, starting from the time when the voice information again meets the triggering condition of the voice command, and the capture is stopped when the voice information again fails to meet the triggering condition of the voice command;

[0282] The intercepted voice content is converted into text to obtain the text content input indicated by the voice instruction.

[0283] The present invention provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described method for text editing based on voice input according to the present invention.

[0284] An embodiment of the present application provides a computer-readable storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will execute the text editing method based on voice input provided by the embodiment of the present application.

[0285] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or various devices including one or any combination of the above memories.

[0286] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0287] As an example, executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions).

[0288] By way of example, executable instructions may be deployed to be executed on one computing device, or on multiple computing devices at one site, or on multiple computing devices distributed across multiple sites and interconnected by a communication network.

[0289] The above description is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the scope of protection of the present application.

Claims

1. A text editing method based on voice input, characterized in that: The method comprises: Presenting a voice input interface including a voice collection area and a text display area, and presenting image collection indication information indicating that image collection is in progress in the voice input interface; Based on the voice input interface, receiving text operation instructions; When the text operation instruction is a voice instruction triggered based on the voice collection area, presenting the text content input indicated by the voice instruction in the text display area; When the text operation instruction is a user action instruction triggered by the image acquisition, performing a text editing operation indicated by the user action instruction that is different from text input in the text display area; In a case where the text operation instruction includes a voice instruction and a user action instruction, the receiving the text operation instruction based on the voice input interface includes: Based on the voice input interface, the input voice information is acquired in real time, and the user image is captured in real time; When the voice information is acquired and the voice information meets the triggering condition of the voice instruction, determining the voice instruction corresponding to the voice information to receive the voice instruction triggered based on the voice collection area; When it is determined that a corresponding user action instruction exists based on the collected user image, the user action instruction corresponding to the collected user image is acquired, so as to receive the user action instruction triggered based on the image collection.

2. The method according to claim 1, wherein The image acquisition indication information for indicating that image acquisition is in progress is presented in the voice input interface, including: In the voice input interface, a voice input function item is presented; In response to the triggering operation for the voice input function item, in the voice collection area, voice collection indication information for indicating that voice collection is in progress is presented, and In the voice collection area, image collection indication information for indicating that image collection is in progress is presented.

3. The method according to claim 2, wherein Presenting image acquisition indication information in the voice acquisition area for indicating that image acquisition is in progress includes: In the voice collection area, a dynamically changing sequence of expressions is presented; The expression sequence is used as image acquisition indication information indicating that image acquisition is in progress.

4. The method according to claim 1, wherein The method further comprises: During the voice collection process, a function item for ending voice input is presented; In response to the triggering operation for the end voice input function item, the presented voice collection instruction information and the image collection instruction information are canceled.

5. The method according to claim 1, wherein When the text operation instruction is the user action instruction, after executing the text editing operation indicated by the user action instruction that is different from text input, the method further includes: Presenting operation instruction information corresponding to the text editing operation; The operation instruction information includes at least one of graphics and text, which is used to indicate the operation content corresponding to the editing operation.

6. The method according to claim 1, wherein The voice input interface including the voice collection area and the text display area is presented, including: Sending an authorization request for the sound acquisition component and the image acquisition component, wherein the authorization request carries verification information; Wherein, the sound acquisition component is used for voice acquisition, and the image acquisition component is used for image acquisition; After receiving the verification result of the legitimacy of the authorization request based on the verification information, the verification result is returned; When the verification result indicates that the authorization request is legal, a voice input interface including a voice collection area and a text display area is presented.

7. The method according to claim 6, wherein The method further comprises: Generate a timestamp and obtain the currently logged-in user account; generating a signature based on the timestamp and the user account; An application identifier is obtained, and the application identifier and the signature are combined to obtain the verification information.

8. The method according to claim 1, wherein The performing of the text editing operation indicated by the user action instruction in the text display area, which is different from text input, includes: Acquire a text editing operation indicated by the user action instruction, wherein the text editing operation includes at least one of the following: text format editing, text style editing, and text layout editing; The text editing operation is performed in the text display area based on the cursor or the selected text content in the text display area.

9. The method according to claim 1, wherein Before determining the voice instruction corresponding to the voice information, the method further includes: Acquiring the volume intensity of the voice information, and when the volume intensity reaches a target intensity, determining that the voice information meets the triggering condition of the voice command; or, Acquiring the semantic content of the voice information, and when the semantic content is the target content, determining that the voice information meets the triggering condition of the voice instruction; or, The volume intensity of the voice information is obtained, and the user's mouth shape feature is extracted from the user image corresponding to the input time of the voice information. When the volume intensity reaches the target intensity and the user's mouth shape feature matches the target mouth shape feature, it is determined that the voice information meets the triggering condition of the voice command.

10. The method according to claim 1, wherein Before obtaining the user action instruction corresponding to the collected user image, the method further includes: When there are multiple user images collected, extracting a target number of target images from the multiple user images; Performing action recognition on each of the target images to obtain the user posture in each of the target images; Based on the user posture in each of the target images, a corresponding user action is determined, and when the user action is a target action, it is determined that a corresponding user action instruction exists.

11. The method according to claim 1, wherein Before obtaining the user action instruction corresponding to the collected user image, the method further includes: Selecting, from the collected user images, a target number of user images within a target time period before a first time point as reference images, where the first time point is a time point at which the voice information does not meet a triggering condition of the voice command; Performing action recognition on the reference image to obtain a corresponding action trend; When the action trend is consistent with the action trend of the target action and the user action is the target action, it is determined that there is a corresponding user action instruction.

12. The method according to claim 1, wherein Before presenting the text content inputted by the voice instruction in the text display area, the method further includes: Acquiring voice information input based on the voice input interface; Taking the time point when the voice information meets the triggering condition of the voice command as the starting point and the time point when the voice information does not meet the triggering condition of the voice command as the end point, intercepting the voice content between the starting point and the end point in the voice information; The intercepted voice content is converted into text to obtain the text content input indicated by the voice instruction.

13. A text editing device based on voice input, characterized in that: The device comprises: A first presentation module is configured to present a voice input interface including a voice collection area and a text display area, and to present image collection indication information indicating that image collection is in progress in the voice input interface; An instruction receiving module, configured to receive text operation instructions based on the voice input interface; a second presentation module, configured to display, in the text display area, text content indicated by the voice instruction input when the text operation instruction is a voice instruction triggered based on the voice collection area; an operation execution module, configured to, when the text operation instruction is a user action instruction triggered based on the image acquisition, execute, in the text display area, a text editing operation different from text input indicated by the user action instruction; The instruction receiving module is further configured to, when the text operation instruction includes a voice instruction and a user action instruction, acquire input voice information in real time based on the voice input interface and capture a user image in real time; When the voice information is acquired and the voice information meets the triggering condition of the voice instruction, determining the voice instruction corresponding to the voice information to receive the voice instruction triggered based on the voice collection area; When it is determined that a corresponding user action instruction exists based on the collected user image, the user action instruction corresponding to the collected user image is acquired, so as to receive the user action instruction triggered based on the image collection.

14. An electronic device, characterized in that: include: a memory for storing executable instructions; The processor is configured to implement the text editing method based on voice input according to any one of claims 1 to 12 when executing the executable instructions stored in the memory.

15. A computer-readable storage medium, characterized in that Executable instructions are stored for implementing the text editing method based on voice input according to any one of claims 1 to 12 when executed by a processor.

16. A computer program product comprising computer executable instructions or a computer program, characterized in that When the computer executable instructions or computer program are executed by a processor, the text editing method based on voice input according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Method and system for text operation

    CN103176594A

  • Realization method, realization device and terminal for voice input

    CN104346127A