Method, apparatus, device, storage medium and program product for voice input

By automatically switching to the input method with voice recognition function after detecting a trigger event and confirming the connection, the problem of cumbersome input method switching and inaccurate recognition results in the prior art is solved, and the voice input effect of simplified operation and improved stability is achieved.

CN122493858APending Publication Date: 2026-07-31BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-06-25
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, the steps for users to switch input methods to use voice input are cumbersome, and the voice recognition results are difficult to accurately submit to the target input object, resulting in limited stability.

Method used

The system automatically switches to a second input method with voice recognition functionality by detecting trigger events, and activates the voice recognition function after confirming the connection with the target input object, ensuring that the voice recognition result is accurately submitted through the established connection.

Benefits of technology

It simplifies the voice input process, improves the stability and accuracy of voice recognition, and reduces invalid voice recognition sessions and client retries due to abnormalities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493858A_ABST
    Figure CN122493858A_ABST
Patent Text Reader

Abstract

A method, apparatus, device, storage medium, and program product for voice input are provided. The method includes: switching from a first input method to a second input method in response to a first triggering event, the first triggering event being used to trigger a voice recognition function of the second input method; activating the voice recognition function of the second input method in response to the establishment of a first connection between the second input method and a first input object; and inputting a voice recognition result into the first input object via the first connection in response to receiving voice content, the voice recognition result being obtained by the voice recognition function performing voice recognition on the voice content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The examples in this article generally relate to the field of computers, and in particular to methods, apparatuses, devices, computer-readable storage media, and computer program products for voice input. Background Technology

[0002] With the application of speech recognition capabilities in input methods, users can convert spoken content into text using the voice function provided by the input method, and then input the text into input objects such as text boxes. This improves input efficiency and convenience. Summary of the Invention

[0003] In a first aspect of this document, a method for voice input is provided, comprising: switching from a first input method to a second input method in response to a first triggering event, the first triggering event being used to trigger a voice recognition function of the second input method; activating the voice recognition function of the second input method in response to a first connection being established between the second input method and a first input object; and inputting a voice recognition result into the first input object through the first connection in response to receiving voice content, the voice recognition result being obtained by the voice recognition function performing voice recognition on the voice content.

[0004] In a second aspect of this document, an electronic device is provided, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform a method according to any one of the first aspects when executed by the at least one processor.

[0005] In a third aspect of this document, a computer-readable storage medium is provided having computer-executable instructions stored thereon, which can be executed by a processor to implement the method according to any one of the first aspects.

[0006] In a fourth aspect of this document, a computer program product is provided, which is tangibly stored in a computer storage medium and includes computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of the first aspects.

[0007] In this way, when a user is using another input method, the input method can be switched by triggering an event on the speech recognition function. Furthermore, once it is confirmed that the input method has taken over the input object (i.e., the first connection has been established), the speech recognition function is automatically activated. This simplifies the process of cross-input method voice input and ensures that the speech recognition result is accurately submitted to the first input object, thereby improving the stability of the speech recognition function.

[0008] It should be understood that the content described in this section is not intended to limit the key or important features of the examples in this article, nor is it intended to restrict the scope of the solution. Other features will become readily apparent from the following description. Attached Figure Description

[0009] The above and other features, advantages, and aspects of the various examples herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein: Figure 1 The diagram shows example environments for some scenarios; Figure 2 Flowcharts illustrating the process for voice input in several scenarios are shown; Figure 3 Flowcharts illustrating example procedures for voice input in several scenarios are shown; Figure 4A and Figure 4B Example interfaces for voice input are shown in some scenarios; Figure 5 Flowcharts of example procedures for voice input in other scenarios are shown; Figures 6A to 6D Example interfaces for voice input are shown in other scenarios; Figure 7 Flowcharts of example procedures for voice input are shown in several other scenarios; Figure 8 Flowcharts of example procedures for voice input in several other scenarios are shown; Figure 9 Block diagrams of devices for voice input in several scenarios are shown; and Figure 10 A block diagram of an electronic device capable of implementing multiple illustrative scenarios is shown. Detailed Implementation

[0010] The examples in the text will now be described in more detail with reference to the accompanying drawings. While some examples are shown in the drawings, it should be understood that solutions can be implemented in various forms and should not be construed as limited to the examples presented herein. Rather, these examples are provided to provide a more thorough and complete understanding of the solutions. It should be understood that the drawings and examples in this document are for illustrative purposes only and are not intended to limit the scope of protection of the solutions.

[0011] It should be noted that the headings of any section / subsection provided herein are not restrictive. Various examples are described throughout this document, and examples of any type may be included under any section / subsection. Furthermore, examples described in any section / subsection may be combined in any way with any other examples described in the same section / subsection and / or different sections / subsections.

[0012] In the description of the examples in this document, the term "including" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an example" or "the example" should be understood as "at least one example". The term "some examples" should be understood as "at least some examples". Other explicit and implicit definitions may also be included below. The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.

[0013] The examples in this document may involve user data, data acquisition, and / or use. All of these aspects comply with relevant laws, regulations, and provisions. In the examples presented herein, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, when implementing each example, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained through appropriate means, in accordance with relevant laws and regulations. The specific methods of notification and / or authorization can vary depending on the actual situation and application scenario; the scope of the solution is not limited in this regard.

[0014] In this manual and the sample solutions, any processing of personal information will be conducted only under legal grounds (such as obtaining the consent of the data subject or being necessary for the performance of a contract) and will only be carried out within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0015] Figure 1 A schematic diagram of example environment 100 is shown. Example environment 100 includes electronic device 110. Electronic device 110 runs an application 120 that supports voice input, such as an input method or other application with voice recognition capabilities. User 140 can interact with application 120 through electronic device 110. Electronic device 110 can present interface 150 through application 120. Electronic device 110 can communicate with server 130. In some cases, server 130 can provide background services for application 120 that supports voice input on electronic device 110, such as providing voice recognition services.

[0016] Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR / AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some cases, electronic device 110 may also support any type of user-facing interface (such as "wearable" circuitry).

[0017] Server 130 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. Server 130 may include, for example, computing systems / servers such as mainframes, edge computing nodes, computing devices in a cloud environment, etc. Server 130 can provide backend services for applications 120 in electronic devices 110 that support voice input.

[0018] A communication connection can be established between server 130 and electronic device 110. This communication connection can be established via wired or wireless means. The communication connection can include, but is not limited to, Bluetooth, mobile network, Universal Serial Bus (USB), and Wireless Fidelity (WiFi) connections. In some cases, server 130 and electronic device 110 can exchange signaling information through their communication connection.

[0019] The examples described herein can be implemented on electronic device 110. It should be noted that the operations described relative to electronic device 110 can be performed by a related application 120 running on electronic device 110. The interface described relative to electronic device 110 can be provided by application 120; for example, electronic device 110 presents the corresponding interface by running application 120. Furthermore, some operations described relative to electronic device 110 may require the assistance of server 130 to complete.

[0020] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of the scheme.

[0021] As mentioned above, with the application of speech recognition capabilities in input methods, users can use the speech function provided by the input method to convert spoken content into text and input the text into input objects such as text boxes. This improves input efficiency and convenience.

[0022] Typically, an input method's voice input function relies on the input method being the current input source or on the input method establishing an input context connection with the current foreground application. In some cases, if the current input source is a specific input method that does not have voice input functionality, and the user wants to use the voice input function of that specific input method, the user needs to manually switch to that specific input method before they can use the voice input function. This not only makes the operation cumbersome but also easily interrupts the current input flow.

[0023] In other scenarios, some standalone voice input applications can recognize text within speech and insert it at the current cursor position using operating system assistance or clipboard mechanisms. However, this approach doesn't essentially involve the input method taking over the input context, making it difficult to submit the streaming results of speech recognition to the original host input box in real time via the input method receiver. The final insertion position depends on the cursor and operating system assistance, and its stability is affected by system permissions, focus state, and application compatibility.

[0024] In view of this, this paper proposes a scheme for voice input. Figure 2 A flowchart of a process 200 for voice input is shown, depending on certain scenarios. Process 200 can be implemented at electronic device 110. See below for reference. Figure 2 An overview of schemes for voice input is provided.

[0025] like Figure 2 As shown in box 210, in response to a first trigger event, the electronic device 110 switches from a first input method to a second input method. The first trigger event is used to trigger the voice recognition function of the second input method, and may include, but is not limited to, triggering preset shortcut keys, triggering preset controls, preset gesture operations, etc. The first input method refers to the input method currently in use on the electronic device 110 when the first trigger event is detected, and may include, but is not limited to, input methods provided by the operating system or third-party input methods not provided by the operating system. The second input method refers to an input method with voice recognition input capability. It should be understood that the second input method can be any input method with voice input capability, and is not limited to a specific product.

[0026] As an example, when user 140 is editing a document using a first input method (such as the default input method provided by the operating system or a third-party input method), they can trigger a first trigger event by pressing a preset voice input shortcut key. If the electronic device 110 detects this first trigger event, it switches from the first input method to the second input method without requiring user 140 to manually perform the input method switching operation.

[0027] In box 220, electronic device 110 can verify whether a first connection has been established between the second input method and the first input object. If it is determined that the first connection between the second input method and the first input object has been established, electronic device 110 can activate the voice recognition function of the second input method.

[0028] The first input object refers to the input object used to receive input from the second input method. In some cases, the first input object may include, but is not limited to, an application, a window, a focused control, a text service session, or other objects capable of receiving input. For example, the first input object may be a document editor that the user is editing, a message input box in an instant messaging tool, or the code editing area of ​​a code editor.

[0029] The first connection refers to a receiving object, connection object, or controller object within the input method framework used to submit input content (such as text, symbols, emoticons, or images) generated by the second input method to the current input position of the first input object. In some cases, the first connection may include, but is not limited to, an input receiver, an input controller, a text service connection object, or other objects with equivalent text submission functionality.

[0030] Speech recognition refers to the technology of converting spoken content into input content such as text, symbols, emojis, or images. Speech recognition functionality provides the ability to recognize speech. In some cases, speech recognition can include, but is not limited to, automatic speech recognition, end-to-end speech recognition, and streaming speech recognition.

[0031] In this way, the electronic device 110 uses whether the second input method has established a first connection for submitting text to the first input object as the activation condition for speech recognition. The speech recognition function is activated only when the first connection is established. This avoids the situation where "the input source has been switched but the input context has not taken over," ensuring that the speech recognition result can be accurately submitted to the first input object.

[0032] In box 230, in response to receiving voice content, electronic device 110 inputs a voice recognition result into a first input object via a first connection. The voice recognition result may include input content obtained by performing voice recognition on the voice content using a voice recognition function, such as text content, image content, symbols, emoticons, etc.

[0033] In some cases, speech recognition results may include streaming intermediate recognition results and final recognition results. For example, during the reception of voice content, electronic device 110 can use its speech recognition function to perform real-time speech recognition on the voice content to obtain intermediate recognition results. Electronic device 110 can input the intermediate recognition results into a first input object in real time through a first connection, allowing user 140 to see the visual effect of the input content gradually appearing on the screen. After speech recognition is completed, electronic device 110 can input the final recognition result into the first input object through the first connection.

[0034] In this way, when using other input methods, if a user wishes to use an input method with speech recognition (i.e., a second input method) for voice input, the user only needs to trigger the speech recognition function of the second input method with a single operation, simplifying the process. Furthermore, the activation condition for speech recognition is whether the second input method has established a first connection for submitting text to the first input object. The speech recognition function is only activated when this first connection is established. This ensures that the speech recognition result is accurately submitted to the first input object.

[0035] Figure 3 A flowchart of an example process 300 for voice input is shown, based on several scenarios. Process 300 can be implemented at electronic device 110. Process 300 provides some exemplary implementations of process 200. Figure 4A and Figure 4B Example interfaces 400A and 400B for voice input are shown in some scenarios. See below for reference. Figure 3 and combined Figure 4A and Figure 4B To describe process 300.

[0036] like Figure 3 As shown in box 310, electronic device 110 responds to a first trigger event and obtains the input context.

[0037] In some cases, the first trigger event includes the detection of a preset trigger operation used to trigger the speech recognition function. If the preset trigger operation is detected, the electronic device 110 can determine that the first trigger event has been received, and thus obtain the input context. In some cases, the preset trigger operation may include the trigger operation of a second shortcut key. As an example, a user can pre-configure a shortcut key (e.g., a combination key) for the speech recognition function of a second input method. If the user wishes to trigger the speech recognition function of the second input method, they can do so by triggering, for example, the combination key on a keyboard. The electronic device 110 can determine that the first trigger event has been received in response to the triggering of the combination key.

[0038] Alternatively and / or additionally, the first triggering event may include a triggering operation on a preset control. For example, the electronic device 110 may display a preset control (e.g., a triggering control for the voice recognition function of a second input method) in the operating system's display interface. If a trigger to the preset control is received (e.g., a single click, double click, long press, or drag, etc.), the electronic device 110 can determine that the first triggering event has been received. Of course, the preset control may also be displayed in, for example, the interface of an input object, such as the menu bar of a text editor or other locations. This document does not limit the display location of the preset control.

[0039] Alternatively and / or additionally, the first triggering event may include a preset gesture operation. As an example, a user pre-configures a gesture (i.e., a preset gesture) to trigger the voice recognition function of a second input method. If the user wishes to trigger the voice recognition function of the second input method, the user can input the preset gesture (i.e., perform the preset gesture operation) via, for example, the touchscreen or touchpad of the electronic device 110. The electronic device 110 may determine that it has received the first triggering event in response to detecting the preset gesture operation.

[0040] In some cases, the input context can indicate the type of the first input method. This type can include, but is not limited to, a first type and a second type. The first type of input method can include the keyboard layout provided by the operating system (also known as the operating system's default input method or native input method). The second type of input method includes third-party input methods not provided by the operating system. For example, input methods developed independently by the input method service provider and installed independently by the user on the electronic device 110.

[0041] As an example, such as Figure 4A As shown, during the process of a user inputting text content into a text editor (i.e., the input object) using "Input Method A" (i.e., the first input method), the electronic device 110 can display the text editor's input window 402 and the toolbar 406 of "Input Method A" (i.e., the first input method). If a first trigger event is received, the electronic device 110 can obtain the type and identifier of "Input Method A" (e.g., program number, process number, name, etc.) as at least part of the input context.

[0042] Alternatively and / or additionally, the input context may also indicate the first input object. In some cases, the electronic device 110 may acquire information capable of indicating the first input object as at least part of the input context. For example, process number, window number, object name, etc. Alternatively and / or additionally, the input context may also indicate the input position within the first input object. As an example, such as Figure 4A As shown, the electronic device 110 can record the position of the cursor 404 in the input window 402 as at least part of the input context.

[0043] Alternatively and / or additionally, the input context may also indicate the triggering time of the first triggering event. For example, in response to receiving a first triggering event, electronic device 110 may acquire a timestamp of the first triggering event and include that timestamp as part of the input context. As an example, when user 140 presses a preset shortcut key for triggering the B input method's speech recognition function, electronic device 110 detects the first triggering event and acquires the input context, which records the first input object (i.e., the document), the input position indicated by cursor 404, the type of A input method, and the triggering time.

[0044] In box 320, electronic device 110 determines whether the preconditions are met. The preconditions may include at least one of the following: voice recognition function is enabled, first auxiliary function is enabled, and input method switching command is enabled. The first auxiliary function is used to provide input assistance for the input method, such as listening to key events on the user's keyboard.

[0045] Input method switching commands are used to switch input methods. In some cases, input method switching commands can be triggered by a shortcut key (e.g., a first shortcut key), which can be an input of software or hardware controls. For example, a user can trigger a shortcut key by pressing a combination of keys on a keyboard, or the electronic device 110 can trigger the input method switching command by simulating a combination of keys through an application or executable object. Alternatively, input method switching commands can be triggered by preset controls. For example, the electronic device 110 can display preset controls in the interface of an operating system or application.

[0046] If the prerequisites are not met, a prompt will appear in box 330 on electronic device 110, guiding the user to complete the corresponding prerequisite configuration. If the prerequisites are met, the subsequent steps will continue.

[0047] In box 340, electronic device 110 switches from the first input method to the second input method based on a switching method corresponding to the type of the first input method. In some cases, if the first input method belongs to the first type (e.g., a keyboard layout provided by the operating system), electronic device 110 can first switch from the first input method to the third input method, and then switch from the third input method to the second input method. If the first input method belongs to the second type (e.g., a third-party input method not provided by the operating system), electronic device 110 can switch directly from the first input method to the second input method. Differentiated switching methods will be discussed in detail below. Figure 7 and Figure 8 Provide a detailed description.

[0048] As an example, such as Figure 4A and Figure 4BAs shown, the electronic device 110 can switch from input method A to input method B (i.e., the second input method) in response to a first trigger event. The electronic device 110 can switch from the toolbar 406 displaying input method A to the toolbar 408 displaying input method B.

[0049] In box 350, electronic device 110 verifies whether a first connection has been established between the second input method and the first input object. If the first connection has not yet been established, in box 360, electronic device 110 waits for a first duration before re-verifying. This waiting mechanism is used to handle situations where there is a time lag in the establishment of the connection object during input method switching. If the first connection has been established, in box 370, electronic device 110 activates the voice recognition function of the second input method. Figure 4B Taking the example interface 400B as an example, when electronic device 110 switches from input method A to input method B, toolbar 408 indicates that input method B is active and the voice recognition function has been activated.

[0050] In some cases, electronic device 110 can determine whether the second input method is enabled to determine whether the first connection has been established. For example, as Figure 4B As shown, electronic device 110 can determine whether input method B is enabled to determine whether a first connection has been established. Alternatively and / or additionally, electronic device 110 can determine whether a connection object (e.g., an input receiver) is enabled to determine whether a first connection has been established. The connection object is used to input content to the first input object. Alternatively and / or additionally, electronic device 110 can determine whether the connection object was established after the triggering time of the first triggering event. If it is determined that the connection object was established after the triggering event of the first triggering event, electronic device 110 can determine that the first connection has been established. Alternatively and / or additionally, electronic device 110 can determine whether the connection object is associated with the input context corresponding to the first input method. If it is determined that the connection object is associated with the input context corresponding to the first input method, electronic device 110 can determine that the first connection has been established.

[0051] In box 380, electronic device 110 inputs the speech recognition result into the first input object via the first connection. Since the first connection has been verified and correctly bound to the first input object, the speech recognition result can be accurately submitted to the position indicated by cursor 404 in the input window 402 edited by the user at the trigger moment.

[0052] In this way, users can switch to the voice input state shown in example interface 400B with a single shortcut key trigger while in the editing state shown in example interface 400A, reducing the manual switching steps from input method A to input method B. Since the electronic device 110 does not activate the voice recognition function before the first connection is established, the number of invalid voice recognition sessions caused by the connection object not being ready during the switching process is reduced, and the invalid startup overhead of the voice recognition engine and the number of abnormal retries by the client are also reduced accordingly.

[0053] In some cases, the electronic device 110 can receive a second trigger event, which is used to disable the voice recognition function. In some cases, the second trigger event may include a shortcut key activation, a control selection, or a predetermined gesture, etc. The electronic device 110 can switch back from the second input method to the first input method in response to the second trigger event; for example, the electronic device 110 can switch from the second input method to the first input method. Figure 4B The input method shown has been switched back to B. Figure 4A The A input method is shown. Alternatively, the electronic device 110 may continue to receive input using the second input method in response to a second trigger event.

[0054] In some cases, electronic device 110 can establish a second connection between the second input method and the second input object in response to the second input object receiving text input instead of the first input object. Electronic device 110 can then input speech recognition results into the second input object via this second connection. The second input object can be different from the first input object; it can include any suitable input object capable of receiving input content, such as a text editor, an application's input box, or a browser's input box. In this way, when the input object changes, the second input method can automatically establish a connection with the switched input object.

[0055] Figure 5 A flowchart of an example process 500 for voice input under other scenarios is shown. Process 500 can be implemented at electronic device 110. Process 500 describes the process of guiding precondition configuration before enabling the voice recognition function. See below for reference. Figure 5 and combined Figures 6A to 6D To describe process 500.

[0056] like Figure 5 As shown, in box 510, electronic device 110 presents a first interface. As an example, such as... Figure 6A As shown, the electronic device 110 can present an example interface 600A, which may be, for example, the settings interface of the second input method, including a first configuration entry 602 and a control 604.

[0057] In box 520, electronic device 110 receives a first operation. In some cases, user 140 can perform a trigger operation on control 604 to request the voice recognition function to be enabled.

[0058] In some situations, if the primary accessibility function is disabled, the electronic device 110 needs to guide the user to enable the primary accessibility function before activating the voice recognition function. In box 530, the electronic device 110 presents a second interface. As an example, such as... Figure 6B As shown, the electronic device 110 can present an example interface 600B, which can display a configuration interface for, for example, the accessibility features of an operating system. The example interface 600B may include a control 606 (i.e., a second configuration entry point). The user can trigger the control 606 to control the enabling or disabling of accessibility permissions for the second input method. In other words, the second operation includes triggering the control 606.

[0059] In box 540, electronic device 110 responds to the second operation by switching the first accessibility function to an enabled state. As an example, when user 140 performs a trigger operation on control 606, electronic device 110 can switch the first accessibility function to an enabled state, enabling the second input method to obtain system-level detection and page positioning capabilities.

[0060] In this way, the electronic device 110 guides the user step-by-step to enable accessibility permissions, reducing the problem of function unavailability caused by missing permissions. Since accessibility permissions are permanently effective once enabled, there is no need to repeat the permission check and guidance process each time voice input is triggered, thereby reducing the permission verification calculation overhead of the electronic device 110 in subsequent voice input scenarios.

[0061] In some cases, if the input method switching command is disabled, the electronic device 110 may need to guide the user to enable the input method switching command. In box 550, the electronic device 110 presents a third interface. For example, as shown... Figure 6C As shown, the electronic device 110 can display an example interface 600C (i.e., the third interface), which displays the operating system's keyboard shortcut settings page. The example interface 600C includes option 608 (i.e., the third configuration entry). Option 608 is used to control whether the input method switching shortcut is enabled or disabled.

[0062] In box 560, electronic device 110 responds to a third operation (e.g., selecting option 608) by enabling the input method switching command. As an example, after selecting option 608, electronic device 110 can enable a first shortcut key, allowing it to trigger the input method switching operation by simulating this shortcut key when switching from a first type of input method (e.g., the system keyboard layout) to a second input method. Afterward, electronic device 110 can switch from example interface 600C to example interface 600D (i.e., the first interface) and enable the voice recognition function (i.e., enabling the voice recognition function). In this way, electronic device 110 ensures that the shortcut key is available when simulating the input method switching command is needed. Since the shortcut key configuration is persistent, subsequent triggering from the system keyboard layout scenario does not require repeated configuration guidance, reducing the number of times the guidance interface is displayed and the number of user interaction rounds.

[0063] In box 570, electronic device 110 switches the voice recognition function to the enabled state. As an example, such as... Figure 6C and Figure 6D As shown, in example interface 600D, control 604 switches to the second visual style (e.g., Figure 6D The switch shown in the diagram (indicates the voice recognition function is switched to the enabled state) indicates that the voice recognition function is switched to the enabled state. At this time, both the first auxiliary function and the input method switching command are enabled, the electronic device 110 has completed the configuration of the prerequisites, and the voice recognition function is successfully enabled.

[0064] In this way, the electronic device 110 sets explicit precondition verification logic during the function activation phase, avoiding direct entry into the voice takeover process when conditions are insufficient, thereby reducing the probability of takeover failure due to missing preconditions. Since the preconditions are permanently effective once configured, there is no need to repeat the condition detection and guidance process in subsequent voice input scenarios, thus reducing the computational resource consumption of the electronic device 110 for condition verification.

[0065] It should be noted that process 500 can occur before the voice recognition function is enabled. That is, the user can pre-enable the voice recognition function, the first auxiliary function, and the input method switching command. Process 500 can also occur after the first triggering event. For example, in process 300, if the detection does not meet the preconditions, the user can switch to the first interface shown in process 500 through the prompt presented in box 330.

[0066] Figure 7A flowchart of an example process 700 for voice input is shown, based on certain scenarios. Process 700 can be implemented at electronic device 110. Process 700 describes the process of switching from a first input method to a second input method and establishing a first connection when the first input method belongs to a first type (e.g., a keyboard layout provided by the operating system, such as the system English keyboard layout). See below for reference. Figure 7 To describe process 700.

[0067] like Figure 7 As shown, in box 702, electronic device 110 responds to a first trigger event and acquires an input context. The input context at least indicates that the type of the first input method is a first type. The first type of input method includes keyboard layouts provided by the operating system, such as the basic keyboard layouts of different languages, such as the system English keyboard layout, the system German keyboard layout, the system Japanese keyboard layout, etc.

[0068] In box 704, electronic device 110 verifies whether the input method switching command is enabled. In some cases, the input method switching command can be triggered by a shortcut key (e.g., a first shortcut key), which can be input from software or hardware controls. For example, a user can trigger a shortcut key by pressing a combination of keys on a keyboard, or electronic device 110 can trigger the input method switching command by simulating a combination of keys through an application or executable object. Alternatively, the input method switching command can be triggered by a preset control. For example, electronic device 110 can display a preset control in the interface of an operating system or application.

[0069] If the input method switching command is disabled, a prompt will be displayed in box 706 on the electronic device 110. For example, the prompt can guide the user to the system keyboard shortcuts page and select the relevant input method switching shortcuts, such as displaying... Figure 6C The example interface shown is 600C.

[0070] If the input method switching command is enabled, in box 708, electronic device 110 triggers an input method switching operation by simulating the input method switching command to switch from the first input method to the third input method. This action does not directly select the second input method, but rather causes the current first input object to move from the system keyboard layout state into a processing link where a connection can be established between the input method framework and the input method.

[0071] In some situations, electronic device 110 can simulate a user pressing the operating system's input method switching shortcut key to trigger an input source switching action. This simulated operation is imperceptible to the user. This bridging step is necessary because the first interface provided by the operating system only supports switching between third-party input methods and does not support direct switching from the system keyboard layout to a third-party input method. If the first interface is directly called to switch from the system keyboard layout to the second input method, the current input source displayed at the system level may have changed, but the connection object of the second input method has not yet established a binding relationship with the first input object, resulting in the input method menu being disabled.

[0072] In box 710, electronic device 110 determines whether it has switched to the third input method. If it has not switched to the third input method (i.e., the current input source has not left the initial system keyboard layout), then in box 712, electronic device 110 waits for a second period of time before re-determining.

[0073] If the third input method has already been switched to, then in box 714, the electronic device 110 switches from the third input method to the second input method by calling the first interface. The first interface is an interface provided by the operating system for selecting between input methods. In some cases, if the second input method has already been switched to directly after the simulated switching operation in box 708 (i.e., the third input method happens to be the second input method), the operation in box 714 can be skipped.

[0074] In box 716, electronic device 110 determines whether a first connection has been established between the second input method and the first input object. If the first connection has not yet been established, in box 718, electronic device 110 waits for a first duration and then re-determines. If the first connection has been established, in box 720, electronic device 110 determines that the first connection between the second input method and the first input object has been established, and then the voice recognition function of the second input method can be activated.

[0075] In this way, when the first input method belongs to the first type, the electronic device 110 adopts a two-stage strategy of "bridging switching plus interface selection" to complete the input method switching, avoiding the false success state that may occur when directly calling the first interface to switch from the system keyboard layout to the second input method. Since the electronic device 110 performs state verification in each stage before entering the next stage, the number of repeated switching and abnormal retries caused by unstable intermediate states is reduced, and the computational overhead and memory usage at the input method framework level for handling switching anomalies are also reduced accordingly.

[0076] Figure 8A flowchart of an example process 800 for voice input under certain circumstances is shown. Process 800 can be implemented at electronic device 110. Process 800 describes the process of switching from the first input method to the second input method and establishing a first connection when the first input method belongs to the second type (e.g., a third-party input method not provided by the operating system). See below for reference. Figure 8 To describe process 800.

[0077] like Figure 8 As shown in block 810, electronic device 110, in response to a first trigger event, acquires an input context. The input context at least indicates that the type of the first input method is a second type. The second type of input method includes third-party input methods not provided by the operating system.

[0078] In box 820, electronic device 110 switches from the first input method to the second input method by calling the first interface. Unlike process 700, since the first input method is already in the input method processing chain (i.e., it belongs to a third-party input method), the second input method can establish a connection object within this chain. Therefore, electronic device 110 does not need to perform the bridging action in the system keyboard layout scenario, but directly selects the second input method by calling the first interface.

[0079] In box 830, electronic device 110 determines whether a first connection has been established between the second input method and the first input object. If the first connection has not been established, in box 840, electronic device 110 waits for a first duration and then re-determines. If the first connection has been established, in box 850, electronic device 110 determines that the first connection between the second input method and the first input object has been established, and then the voice recognition function of the second input method can be activated.

[0080] In this way, when the first input method belongs to the second type, the electronic device 110 can directly switch by calling the first interface, reducing the number of switching steps. Since the stage of simulating shortcut key bridging is eliminated, the latency of input method switching is reduced, and the system resource usage for waiting for bridging to complete is also reduced accordingly.

[0081] A corresponding apparatus for implementing the above methods or processes is also provided.

[0082] Figure 9 Block diagrams of a device 900 for voice input in several scenarios are shown. Device 900 can be implemented as or included in electronic device 110. The various modules / components in device 900 can be implemented by hardware, software, firmware, or any combination thereof.

[0083] like Figure 9As shown, the device 900 includes a switching module 910 configured to switch from a first input method to a second input method in response to a first trigger event, the first trigger event being used to trigger the speech recognition function of the second input method; a verification module 920 configured to activate the speech recognition function of the second input method in response to the establishment of a first connection between the second input method and the first input object; and an input module 930 configured to input a speech recognition result into the first input object through the first connection in response to receiving speech content, the speech recognition result being obtained by the speech recognition function performing speech recognition on the speech content.

[0084] In some cases, the switching module 910 can also be configured to detect the detection of a preset trigger operation, which is used to trigger the speech recognition function.

[0085] In some cases, the switching module 910 can also be configured to acquire an input context, which at least indicates the type of the first input method; and to switch from the first input method to the second input method based on the switching method corresponding to that type.

[0086] In some cases, the switching module 910 can also be configured to switch from the first input method to the third input method and then from the third input method to the second input method based on the first input method belonging to the first type; or to switch directly from the first input method to the second input method based on the first input method belonging to the second type.

[0087] In some cases, the switching module 910 can also be configured to trigger an input method switching operation by simulating an input method switching command to switch from the first input method to the third input method; and to switch from the third input method to the second input method by calling a first interface, the first interface providing input method selection functionality.

[0088] In some cases, the switching module 910 can also be configured to switch from the first input method to the second input method by calling the first interface, which provides the input method selection function.

[0089] In some cases, the input context acquisition module 960 may also be configured to acquire an input context indicating at least one of the following: a first input object; an input position in the first input object; or the triggering time of a first triggering event.

[0090] In some cases, the precondition determination module 950 may also be configured to determine whether at least one of the following is satisfied: the speech recognition function is enabled; the first auxiliary function is enabled, which is used to provide input assistance for the input method; and the input method switching instruction is enabled, which is used to switch the input method.

[0091] In some cases, the switching module 910 can also be configured to present a first interface based on the voice recognition function being disabled, the first interface providing a first configuration entry for the voice recognition function; and in response to the first operation, switch the voice recognition function to the enabled state.

[0092] In some cases, the switching module 910 can also be configured to present a first interface based on the first assistive function being enabled; or to present a second interface based on the first assistive function being disabled, the second interface providing a second configuration entry for the first assistive function; in response to a second operation, switch the first assistive function to the enabled state; and switch the voice recognition function to the enabled state.

[0093] In some cases, the switching module 910 can also be configured to present a third interface based on the input method switching command being in an inactive state, the third interface providing a third configuration entry for the input method switching command; and in response to the third operation, to switch the input method switching command to an active state.

[0094] In some cases, the verification module 920 may also be configured to determine that a first connection between the second input method and the first input object has been established based on at least one of the following: the second input method is enabled; the connection object is enabled and is used to input content to the first input object; the connection object is established after the first triggering event; the connection object is associated with the input context corresponding to the first input method.

[0095] In some cases, the input module 930 may also be configured to determine the first input object as an input object for receiving input content before the first triggering event; or to determine the first input object as an input object for receiving input content after the first triggering event.

[0096] In some cases, the switching module 910 can also be configured to respond to a second trigger event, which is used to trigger the disabling of the voice recognition function and switch back from the second input method to the first input method; or to continue receiving input using the second input method.

[0097] In some cases, the input module 930 may also be configured to establish a second connection between the second input method and the second input object in response to the second input object receiving text input in place of the first input object; and to input the speech recognition result into the second input object via the second connection.

[0098] In some cases, the input module 930 may also be configured to input intermediate recognition results into the first input object via the first connection during the reception of voice content; and to input the final recognition result into the first input object via the first connection after the reception of voice content ends.

[0099] In some cases, the input module 930 can also be configured to update the intermediate recognition results already entered in the first input object as the voice content is received.

[0100] In some cases, the switching module 910 can also be configured to detect at least one of the following preset trigger operations: trigger operation of a second shortcut key; trigger operation of a preset control; preset gesture operation.

[0101] In some cases, the switching module 910 can also be configured to determine the first type of input method as the keyboard layout provided by the operating system, and the second type of input method as a third-party input method not provided by the operating system.

[0102] The modules included in device 900 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some cases, one or more modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units in device 900 can be implemented at least partially by one or more hardware logic components. By way of example, and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard parts (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0103] Figure 10 A block diagram of an electronic device 1000 in which one or more examples can be implemented is shown. It should be understood that... Figure 10 The electronic device 1000 shown is merely exemplary and should not be construed as limiting the functionality and scope of the examples described herein. Figure 10 The electronic device 1000 shown can be used to implement the device 900 discussed above.

[0104] like Figure 10As shown, electronic device 1000 is in the form of a general-purpose electronic device. Components of electronic device 1000 may include, but are not limited to, one or more processing units or processors 1010, memory 1020, storage device 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060. Processor 1010 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1020. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of electronic device 1000.

[0105] Electronic device 1000 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof). Storage device 1030 can be removable or non-removable media and may include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 1000.

[0106] Electronic device 1000 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 10 As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 1020 may include computer program product 1025 having one or more program modules configured to perform various methods or actions of various examples.

[0107] The communication unit 1040 enables communication with other electronic devices via a communication medium. Additionally, the functionality of the components of the electronic device 1000 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the electronic device 1000 can operate in a networked environment using logical connections to one or more other servers, networked personal computers, or another network node.

[0108] Input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1060 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 1000 can also communicate with one or more external devices (not shown) via communication unit 1040 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 1000, or with any device that enables electronic device 1000 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).

[0109] A computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. A computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.

[0110] The flowcharts and / or block diagrams of the methods, apparatus, devices, and computer program products referred to herein describe various aspects. It should be understood that each block of the flowcharts and / or block diagrams, as well as combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0111] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0112] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0113] The flowcharts and block diagrams in the accompanying figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products under various scenarios. In this respect, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0114] Various examples have been described above. The foregoing descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for voice input, comprising: In response to a first trigger event, the input method is switched from the first input method to the second input method, wherein the first trigger event is used to trigger the speech recognition function of the second input method; In response to the establishment of a first connection between the second input method and the first input object, the speech recognition function of the second input method is activated; as well as In response to receiving voice content, a voice recognition result is input into the first input object through the first connection. The voice recognition result is obtained by the voice recognition function performing voice recognition on the voice content.

2. The method according to claim 1, wherein the first triggering event includes the detection of a preset triggering operation, the preset triggering operation being used to trigger the speech recognition function.

3. The method according to claim 1, wherein switching from the first input method to the second input method comprises: Obtain the input context, which at least indicates the type of the first input method; as well as Based on the switching method corresponding to the type, switch from the first input method to the second input method.

4. The method according to claim 3, wherein switching from the first input method to the second input method based on the switching method comprises: Based on the fact that the first input method belongs to the first type. Switch from the first input method to the third input method; as well as Switch from the third input method to the second input method; or Based on the fact that the first input method belongs to the second type. Switch directly from the first input method to the second input method.

5. The method according to claim 4, wherein switching from the first input method to the third input method comprises: The input method switching operation is triggered by simulating an input method switching command, so as to switch from the first input method to the third input method; and Switching from the third input method to the second input method includes: By calling the first interface, the user can switch from the third input method to the second input method. The first interface provides an input method selection function.

6. The method according to claim 4, wherein directly switching from the first input method to the second input method comprises: By calling the first interface, the input method can be switched from the first input method to the second input method. The first interface provides the input method selection function.

7. The method of claim 3, wherein the input context further indicates at least one of the following: The first input object; The input position in the first input object; The triggering time of the first triggering event.

8. The method according to claim 1, wherein switching from the first input method to the second input method comprises: Switching from the first input method to the second input method is based on at least one of the following: The voice recognition function is enabled. The first auxiliary function is enabled, and the first auxiliary function is used to provide input assistance for the input method; The input method switching command is enabled and is used to switch input methods.

9. The method of claim 8, further comprising: Since the speech recognition function is disabled, a first interface is displayed, which provides a first configuration entry for the speech recognition function. as well as In response to the first operation, the voice recognition function is switched to the enabled state, whereby the first operation is an operation on the first configuration entry.

10. The method of claim 9, wherein switching the speech recognition function to the enabled state comprises: Since the first accessibility function is enabled, the first interface is displayed. or Since the first accessibility function is disabled. A second interface is presented, which provides a second configuration entry for the first auxiliary function; In response to the second operation, the first accessibility function is switched to the enabled state, wherein the second operation is an operation on the second configuration entry; as well as Switch the voice recognition function to enabled.

11. The method of claim 8, further comprising: Since the input method switching command is in an inactive state, a third interface is displayed, which provides a third configuration entry for the input method switching command. as well as In response to the third operation, the input method switching instruction is switched to the enabled state, wherein the third operation is an operation on the third configuration entry.

12. The method of claim 1, wherein the establishment of a first connection between the second input method and the first input object is determined based on at least one of the following: The second input method is enabled; The connection object is in an enabled state, and the connection object is used to input content to the first input object; The connection object is established after the first triggering event; The connection object is associated with the input context corresponding to the first input method.

13. The method of claim 1, wherein the first input object is an input object used to receive input content before the first triggering event; or The first input object is the input object used to receive input content after the first triggering event.

14. The method according to claim 1, further comprising: In response to a second triggering event, which is used to disable the speech recognition function, Switch back to the first input method from the second input method; or Continue receiving input using the second input method.

15. The method of claim 1, further comprising: In response to the second input object receiving text input in place of the first input object, a second connection is established between the second input method and the second input object; as well as The speech recognition result is input into the second input object via the second connection.

16. The method of claim 1, wherein the speech recognition result includes intermediate recognition results and a final recognition result, and wherein inputting the speech recognition result into the first input object includes: During the process of receiving the voice content, the intermediate recognition result is input into the first input object via the first connection; as well as After the reception of the voice content ends, the final recognition result is input into the first input object via the first connection.

17. The method of claim 16, wherein inputting the intermediate identification result comprises: As the voice content is received, the intermediate recognition results already entered in the first input object are updated.

18. An electronic device comprising: At least one processor; as well as At least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 17 when executed by the at least one processor.

19. A computer-readable storage medium having stored thereon computer-executable instructions that can be executed by a processor to implement the method according to any one of claims 1 to 17.

20. A computer program product tangibly stored in a computer storage medium and comprising computer-executable instructions that, when executed by a device, cause the device to perform the method according to any one of claims 1 to 17.