Input method, input device, electronic device, medium, and program product

By receiving and integrating text and voice input information on smart devices, and using artificial intelligence models to generate high-quality input information, the problem of inaccurate expression when users switch input modes is solved, achieving more natural and efficient information expression.

CN122195386APending Publication Date: 2026-06-12VIVO MOBILE COMM CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
VIVO MOBILE COMM CO LTD
Filing Date
2026-03-19
Publication Date
2026-06-12

AI Technical Summary

Technical Problem

In existing technologies, users often struggle to accurately express their input intentions when inputting information on smart devices, especially when switching between text and voice input modes, where inaccurate expression can occur.

Method used

An input method and apparatus are provided that receive text and voice input from a user and fuse the two information using an artificial intelligence model to generate high-quality first input information. The method supports simultaneous text and voice input and does not require manual operation during the input process, such as through a microphone icon.

Benefits of technology

It enables users to express their input intentions more naturally and accurately on electronic devices, improves the quality and efficiency of information input, and facilitates accurate expression during communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122195386A_ABST
    Figure CN122195386A_ABST
Patent Text Reader

Abstract

The application discloses an input method, an input device, an electronic device, a medium and a program product, and belongs to the technical field of electronic products. The input method comprises the following steps: receiving a character input and a voice input of a user; and displaying first input information based on character information corresponding to the character input and voice information corresponding to the voice input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic product technology, specifically to an input method, input device, electronic device, medium, and program product. Background Technology

[0002] In related technologies, on smart devices, users can input information using the system's built-in input method or a third-party input method. Current input methods typically support both text and voice input modes. Users can select between text and voice input modes using controls within the input box. However, based on these input methods, in some information input scenarios, there may be issues with accurately conveying the user's input intent. Summary of the Invention

[0003] This application provides an input method, input device, electronic device, medium, and program product that can solve the problem in related technologies that cannot accurately express the user's input intent.

[0004] In a first aspect, an input method is provided for use in an electronic device, the method comprising:

[0005] It accepts text and voice input from users;

[0006] Based on the text information corresponding to the text input and the voice information corresponding to the voice input, the first input information is displayed.

[0007] Secondly, an input device is provided for use in an electronic device, the device comprising:

[0008] The receiving module is used to receive the user's text input and voice input;

[0009] The display module is used to display the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input.

[0010] Thirdly, an electronic device is provided, comprising a processor and a memory, wherein the memory stores a program or instructions executable on the processor, the program or instructions, when executed by the processor, perform the steps of the method described in the first aspect.

[0011] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0012] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the steps of the method described in the first aspect.

[0013] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the steps of the method described in the first aspect.

[0014] In this embodiment, when a user inputs data on an electronic device, they can simultaneously perform text input and voice input. The electronic device displays first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input. That is, when a user inputs data, they can simultaneously use both voice and text to express information, which helps the user express information more naturally during communication using the electronic device, and thus facilitates the user's accurate expression of their input intentions. Attached Figure Description

[0015] Figure 1 This is one of the flowcharts illustrating an input method provided in some embodiments of this application;

[0016] Figure 2 This is a schematic diagram of the first input window in some embodiments of this application;

[0017] Figure 3 yes Figure 2 A diagram of the input box in the image;

[0018] Figure 4 This is a diagram illustrating the input box when voice and text input are performed simultaneously.

[0019] Figure 5(a) is a schematic diagram of a plain text input window in some embodiments of this application;

[0020] Figure 5(b) is a schematic diagram of a pure voice input window in some embodiments of this application;

[0021] Figure 5(c) is a schematic diagram of the first input window in some embodiments of this application;

[0022] Figure 5(d) is a schematic diagram of the first input information generated after mixed voice and text input in some embodiments of this application;

[0023] Figure 6(a) is a schematic diagram of a plain text input window in some embodiments of this application;

[0024] Figure 6(b) is a schematic diagram of the process of a user clicking the microphone icon in Figure 6(a) and making a voice input.

[0025] Figure 6(c) is a schematic diagram of the user continuing to input text and voice based on the click in Figure 6(b);

[0026] Figure 6(d) is a schematic diagram of the first input information generated based on the text input and voice input in Figures 6(a), 6(b) and 6(c);

[0027] Figure 7(a) is a schematic diagram after the first input information is generated in some embodiments of this application;

[0028] Figure 7(b) is a schematic diagram after the user clicks the return control in Figure 7(a);

[0029] Figure 7(c) is a schematic diagram of what happens after a user long-presses the voice input area in Figure 7(b);

[0030] Figure 8 This is a second schematic flowchart of an input method provided in some embodiments of this application;

[0031] Figure 9(a) is a schematic diagram of a plain text input window in some embodiments of this application;

[0032] Figure 9(b) is a schematic diagram of the process of a user pressing and holding the microphone icon in Figure 9(a) and making a voice input.

[0033] Figure 9(c) is a schematic diagram of the first input information generated based on the text input and voice input in Figures 9(a) and 9(b);

[0034] Figure 10(a) is a schematic diagram of the first input information generated when the input content of the voice input in Figure 9(b) is changed to "call folder";

[0035] Figure 10(b) is a schematic diagram of the first input information generated when the input content of the voice input in Figure 9(b) is modified to "call the chat history between AA and BB";

[0036] Figure 11 This is a schematic diagram of the structure of the input device in some embodiments of this application;

[0037] Figure 12 Schematic diagrams of the structure of electronic devices provided for some embodiments of this application;

[0038] Figure 13 A schematic diagram of the hardware structure of an electronic device provided for some embodiments of this application. Detailed Implementation

[0039] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0040] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same class, not limited in number; for example, the first object can be one or more. Furthermore, "or" in this application indicates at least one of the connected objects. For example, the scope of protection for "A or B" covers at least three scenarios: Scenario 1: including A but not B; Scenario 2: including B but not A; Scenario 3: including both A and B. In addition, the terms "A and / or B," "at least one of A and B," and "at least one of A or B" also cover at least the above three scenarios. The character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0041] The input methods, input devices, electronic devices, media, and program products provided in this application will be described in detail below with reference to the accompanying drawings and through some embodiments and application scenarios.

[0042] Please see Figure 1 , Figure 1 This is a flowchart illustrating an input method provided in an embodiment of this application. The input method is applied to an electronic device and includes the following steps:

[0043] Step 101: Receive text and voice input from the user.

[0044] The text input and voice input can be input by the user using an input method on an electronic device. The electronic device can receive the text input through an input window and the voice input through its microphone. The text input and voice input can be input by the user in various information input scenarios, such as input when chatting with other users through social applications, input when communicating with an intelligent agent, or input when searching for information using relevant query software.

[0045] Understandably, users can perform both text and voice input simultaneously. For example, if a user struggles to find the right words to describe an object, they can verbally express it, and the electronic device can refine their verbal expression during the subsequent fusion process to improve the quality of the initial input. Furthermore, users can perform text input first, followed by voice input; or vice versa; or they can alternate or interweave text and voice input. The specific timing of text and voice input can be chosen according to individual needs.

[0046] In some embodiments of this application, the text input may include the main text in the first input information, and the voice input may include at least one of the following: the main text in the first input information, an instruction for generating specific data, an instruction for calling specific data, a processing instruction for processing the input content of the text input, etc.

[0047] Step 102: Based on the text information corresponding to the text input and the voice information corresponding to the voice input, display the first input information.

[0048] The above-mentioned method of displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input may include: concatenating the input content of the text input with the input content of the voice input, performing a polishing process to obtain the first input information, and then displaying the first input information. Alternatively, the method of displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input may also include: concatenating the content generated based on the voice input with the input content of the text input, performing a polishing process to obtain the first input information, and then displaying the first input information. Alternatively, the method of displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input may further include: processing the input content of the text input according to the processing method indicated by the voice input to obtain the first input information, and then displaying the first input information, etc.

[0049] In some embodiments of this application, since users can perform voice input while inputting text, that is, while users are inputting voice, their hands may be editing text. Therefore, when performing voice input, users can directly speak the content they want to input. In other words, users do not need to coordinate with hand movements while performing voice input, such as pressing the microphone icon or other hand operations.

[0050] The first input information mentioned above can be presented in text form.

[0051] In this embodiment, when a user inputs data on an electronic device, they can simultaneously perform text input and voice input. The electronic device displays the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input. That is, when inputting data, the user can simultaneously use both voice and text to express information, which helps the user express information more naturally during communication using the electronic device, and thus facilitates the user's accurate expression of their input intentions.

[0052] Optionally, before receiving the user's text input and voice input, the method further includes:

[0053] Display a first input window, which includes a text input area and a voice input area.

[0054] In some embodiments of this application, receiving user text input and voice input may include: receiving the first text information input by the text input based on the text input area, and converting the received voice information into the second text information and displaying it on the voice input area.

[0055] The first input window can simultaneously receive mixed voice and text input from the user during the same input process.

[0056] It is understood that the text input area and the voice input area can be two different regions within the input box. For example, the voice input area and the text input area can be arranged along the height of the input box, or they can be distributed on the left and right sides of the input box, depending on the specific requirements. Furthermore, to facilitate user differentiation between the voice input area and the text input area, the text content in each area can be displayed in different formats. For example, please refer to [link to relevant documentation]. Figure 4 The text in the voice input area can be displayed in gray font, while the text in the text input area can be displayed in black font, so as to distinguish information entered through different input methods.

[0057] It should be noted that, as Figure 2 As shown, the first input window 210 may include not only an input box, but also a microphone icon and other related controls. Please refer to [link / reference]. Figure 3 The diagram below illustrates an input box in some embodiments of this application. The input box includes a text input area 212, a voice input area 211, a microphone icon 213, and a send control 214. When receiving voice input, the microphone icon 213 may be in a flashing state.

[0058] In some embodiments of this application, the input box can be divided into two areas: a text input area and a voice input area, which can simultaneously support keyboard typing and voice input. Because the written text entered by typing may differ from the text entered by voice in terms of linguistic logic, after input is complete, AI capabilities are used to combine the voice-input text with the typed text for unified processing to generate the final text content.

[0059] The input box allows users to input text via voice; alternatively, users can input information via voice while pausing during typing, or input information simultaneously. After simultaneous input, if no voice or text is entered for 2 seconds, the AI ​​will begin processing the results and generating the final result in the input box. If the result does not meet expectations, users can click the back control in the upper right corner of the input box to return to the input state for adjustments.

[0060] In this embodiment, a first input window is displayed before receiving text input and voice input from the user. The first input window includes a text input area and a voice input area, thus enabling the simultaneous reception of text input and voice input from the user.

[0061] Optionally, displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input includes:

[0062] The first input information is obtained by fusing the first text information in the text input area and the second text information in the voice input area based on an artificial intelligence (AI) model; wherein the second text information is obtained based on the voice information.

[0063] The first input information is displayed in the first input window.

[0064] Please refer to Figure 6, which is a schematic diagram of an input method provided in some embodiments of this application, including the following steps:

[0065] To activate: Click the microphone icon above the keyboard in Figure 6(a) to enter the first interaction mode shown in Figure 6(b) and start recording, thus initiating the process of AI model supplementing the dialogue.

[0066] As shown in Figures 6(b) and 6(c), during speech, a dual display is shown:

[0067] The small gray text above represents what the user says via voice, including phrases like "um," "that," and "just...".

[0068] The large black text below indicates what the user should type using the keyboard.

[0069] Simultaneous input: The input box can simultaneously accept user voice input and keyboard input.

[0070] As shown in Figure 6(d), the processing is complete: when speaking and typing stop, the AI ​​model begins to intervene in the processing, the small gray text disappears, and only the well-polished "beautiful words" that combine speech and text and are polished by the AI ​​model are displayed, that is, the first input information is displayed.

[0071] In this embodiment, the first input information is obtained by fusing the first text information in the text input area and the second text information in the voice input area based on an AI model. This helps to improve the quality of the generated first input information.

[0072] Optionally, receiving the user's text input and voice input includes:

[0073] The system receives the first text information input by the text input area and converts the received voice information into the second text information and displays it in the voice input area.

[0074] Displaying the first input information in the first input window includes:

[0075] Cancel the display of the first and second text information, and display the first input information and the return control.

[0076] In some embodiments of this application, the method further includes:

[0077] Upon receiving the user's first input to the return control, the first input information is de-displayed, and the first and second text information are re-displayed;

[0078] Upon receiving editing input from a user, the system edits the first text information and / or the second text information in response to the editing input.

[0079] Understandably, please see Figure 3The first input window may also include a send control 214. In some embodiments of this application, if the user is satisfied with the first input information, they can directly click the send control 214 to send the first input information. If the user is not satisfied with the first input information, they can click the return control 701 in Figure 7(a). At this time, the first input window returns to the state before the AI ​​model merges the text content in the text input area and the text content in the voice input area, that is, returns to the state shown in Figure 6(c), as shown in Figure 7(b). After returning to the state shown in Figure 6(c), the user can edit the text content in the text input area and the voice input area. For example, the user can modify the text in the voice input area and / or modify the text in the text input area.

[0080] It is understandable that after the user completes editing the text in the voice input area and / or the text in the text input area, the electronic device can fuse the text content in the text input area with the text content in the voice input area based on an artificial intelligence (AI) model to obtain the first input information.

[0081] In this implementation, upon receiving a first input from the user to the return control, the first input information is de-displayed, and the first and second text information are re-displayed. Upon receiving an edit input from the user, the first and / or second text information are edited in response to the edit input. This allows users to easily return to the interface before fusion if they are dissatisfied with the generated first input information, and to edit the text in the voice input area and / or the text in the text input area based on the edit input, thereby regenerating the first input information and further improving the quality of the generated first input information.

[0082] Optionally, after canceling the display of the first input information and redisplaying the first text information and the second text information, the method further includes:

[0083] Upon receiving a user's switching instruction, the editing area in the first input window is switched from the first editing area to the second editing area, wherein the first editing area and the second editing area are respectively the voice input area and the text input area.

[0084] In some embodiments of this application, the switching instruction can be various common touch inputs to the current editing area, such as long press input or double-click input to the current editing area.

[0085] It should be noted that since the text input area contains content manually entered by the user, while the text content in the voice input area is the result of the electronic device converting received speech into text, errors may occur during the speech-to-text conversion process. Therefore, the text content in the text input area is usually relatively accurate, while the text content in the voice input area may contain typos and other issues. This means that when a user clicks the back control, there is a high probability that they will need to manually modify the text in the voice input area. In some embodiments of this application, upon receiving the user's first input to the back control, the first input information is de-displayed, and the text input area and the voice input area are displayed. Simultaneously, the current editing area is designated as the voice input area; for example, the cursor can be displayed at the end of the voice input area to facilitate quick editing of the content in the voice input area.

[0086] Please refer to Figure 7(c). After the user finishes editing the text in the voice input area, they can enter a switching command in the voice input area. For example, the switching command can be a long press on the voice input area. At this time, a window as shown in Figure 7(c) can be displayed. This window can include two options: "Copy" and "Switch Editing Area". When the user selects the "Switch Editing Area" option, the cursor can be switched to the text input area, where the user can edit the text. When the user selects the "Copy" option, the text content in the voice input area can be copied.

[0087] In some embodiments of this application, upon receiving the user's first input to the return control, the first input information can be canceled, and the text input area and the voice input area can be displayed. At the same time, the current editing area can be determined as the text input area, which can be set as needed.

[0088] Understandably, once a user has finished editing the text input area, they can switch back to the voice input area by entering a switching command in the text input area.

[0089] In this embodiment, upon receiving a user's switching instruction, the editing area in the first input window is switched from the first editing area to the second editing area. The first editing area and the second editing area are respectively the voice input area and the text input area. This allows the user to easily switch between the editing areas so that the user can edit the text input area and the voice input area separately.

[0090] Optionally, displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input includes:

[0091] The speech information is subjected to intent recognition to obtain intent recognition information;

[0092] When the intent recognition information indicates that the text information in the text input area should be edited, the text in the text input area is edited according to the editing method indicated by the intent recognition information to obtain the first input information;

[0093] The first input information is displayed in the first input window.

[0094] The above editing methods can be of various types, such as "delete the second sentence", "modify this text", "check this text for typos", "generate detailed DocString comments for this code and complete the test cases for the two edge cases", etc.

[0095] In some embodiments of this application, the first input window may have a first interaction mode and a second interaction mode. When the first input window is in the first interaction mode, displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input includes: fusing the first text information in the text input area and the second text information in the voice input area based on an artificial intelligence (AI) model to obtain the first input information; wherein the second text information is obtained based on the voice information; and displaying the first input information in the first input window. Correspondingly, when the first input window is in the second interaction mode, displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input includes: performing intent recognition on the voice information to obtain intent recognition information; if the intent recognition information indicates that the text information in the text input area should be edited, editing the text in the text input area according to the editing method indicated by the intent recognition information to obtain the first input information; and displaying the first input information in the first input window. It is understood that the user can choose to enter the first interaction mode or the second interaction mode in the first input window.

[0096] Of course, in some embodiments of this application, the AI ​​model can also determine the current input mode based on the content of the voice input. When the voice input includes instructions, it automatically enters the second interaction mode; when the voice input is text content that connects with the text input, it automatically enters the first interaction mode.

[0097] In this embodiment, the voice information is subjected to intent recognition to obtain intent recognition information; when the intent recognition information indicates that the text information in the text input area is to be edited, the text in the text input area is edited according to the editing method indicated by the intent recognition information to obtain the first input information; the first input information is displayed in the first input window. In this way, the text content in the text input area can be automatically optimized with the assistance of an AI model through voice commands to improve the accuracy of the generated first input information.

[0098] Optionally, the editing method includes at least one of the following:

[0099] Delete the first sub-message in the text input area;

[0100] Identify typos in the text input area;

[0101] Replace the second sub-information in the text input area;

[0102] Add explanatory information to the third sub-information in the text input area;

[0103] Translate the fourth sub-information in the text input area into the target language;

[0104] Add a fifth piece of information to the text input area.

[0105] The first sub-information mentioned above can refer to at least a portion of the text information in the text input area, such as a specific character, word, or sentence within the text input area. Similarly, the second, third, and fourth sub-information mentioned above can also refer to at least a portion of the text information in the text input area, such as a specific character, word, or sentence. The fifth sub-information mentioned above can be various types of information, such as document information, image information, video information, audio information, or chat information.

[0106] It is understood that the above editing methods are only some examples in the embodiments of this application. In fact, the editing methods in the embodiments of this application can also be various other editing methods besides those listed above, and can be set as needed.

[0107] In this embodiment, by including at least one of the following editing methods: deleting the first sub-information in the text input area; identifying typos in the text input area; replacing the second sub-information in the text input area; adding explanatory information to the third sub-information in the text input area; translating the fourth sub-information in the text input area into the target language; and adding a fifth sub-information in the text input area, users can edit the text content in the text input area in various ways based on voice input, thereby facilitating users to edit the text content in the text input area based on voice input.

[0108] Optionally, displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input further includes:

[0109] When the intent recognition information indicates the retrieval of first data stored in the electronic device, the first data or its corresponding abbreviated form is displayed in the voice input area; or...

[0110] If the intent recognition information indicates that second data should be generated, the generated second data is displayed in the voice input area; or,

[0111] When the intent recognition information indicates the generation of third data and the text input area includes corresponding descriptive information, the third data is generated according to the intent recognition information and the descriptive information, and the third data is displayed.

[0112] The aforementioned first data can be various types of data stored in an electronic device, such as an emoji from a meme, image or video data from a photo album, or a file in a folder. The aforementioned thumbnail information can be simplified information about the first data. For example, when the first data is image data, the thumbnail information can be a thumbnail of the image data. As another example, when the first data is video data, the thumbnail information can be the video's cover image. As yet another example, when the first data is an animated emoji, the thumbnail information can be the cover frame image of the animated emoji.

[0113] The second and third data mentioned above can be data generated by AI, or data that can be downloaded from the network. For example, please refer to Figure 9(b) when a user is inputting text and their voice input is "Find the flowers I took pictures of yesterday," the first input information shown in Figure 9(c) can be obtained. As another example, please refer to Figure 10(a) for a scenario where a user accesses a file, displaying the corresponding file thumbnail. As yet another example, please refer to Figure 10(b) for a scenario where a user accesses chat history, displaying the chat history thumbnail information.

[0114] Please see Figure 8 The diagram below illustrates an input method provided in some embodiments of this application, including the following steps:

[0115] Step 1: The user opens the text keyboard.

[0116] Users open the text input keyboard to input text.

[0117] Step 2: Whether to enable the real-time text-to-speech mixing function. In some embodiments of this application, for better understanding of this function, it is named the "write-and-speak" function.

[0118] If this feature is not enabled, the process will end and the feature will not respond.

[0119] If you enable this feature, proceed to the next step.

[0120] How to activate: Long press.

[0121] Step 3: Perform speech recognition.

[0122] If the user has enabled this feature, voice recognition will proceed directly.

[0123] During the recognition process, the user's main interface input should not be affected. See the subsequent interface demonstration for specific results.

[0124] Commonly used keywords and phrases are shown in Table 1:

[0125] Table 1:

[0126]

[0127] Step 4: Analyze Intent

[0128] The intent of the speech recognition content is analyzed, and speech-to-text technology is used.

[0129] Analyze the user's voice input intent to obtain intent recognition information.

[0130] Based on the intent recognition information, specific operation processing is performed. In this embodiment, the tentative operation processing includes, but is not limited to: adding, deleting, modifying, and querying input text; other convenient operations can also be superimposed, such as inserting other text. Please refer to Table 2 for the intent recognition information obtained by performing intent recognition on the statements in Table 1:

[0131] Table 2:

[0132]

[0133] Step 5: Intent Processing

[0134] Follow up on the analyzed intent and process the input text directly.

[0135] For example, you can simply delete the content you want to delete, or modify the content you want to modify.

[0136] Step 6: Generate the final result and provide feedback.

[0137] The processing results are displayed directly on the interface.

[0138] The method provided in this application, within a multimodal input capability framework combining voice modality and typing input, can simultaneously handle secondary tasks in addition to the primary task. Combined with AI, it allows users to invoke AI capabilities at any time while typing to assist in completing the primary task, making the invocation of AI capabilities more natural and efficient. For ease of understanding, this application further explains the input method in conjunction with specific scenarios:

[0139] In some embodiments of this application, in an office setting, the input method may include the following process:

[0140] Users input document content via text;

[0141] At this point, you need to insert a recording. You can simply say, "Insert the recording of my previous conversation with Joe here."

[0142] After the AI ​​model recognizes the instruction, it directly inserts the corresponding voice recording at this point. This does not affect the content manually entered by the user.

[0143] In some embodiments of this application, in a code scenario, the input method may include the following process:

[0144] User action: Continuous keyboard typing.

[0145] Voice command: "Generate detailed DocString comments for this code and complete the test cases for the two edge cases."

[0146] AI Response: The AI ​​recognizes that "this code" refers to the function block where the cursor is currently located, automatically inserts a standard-format comment above the function, and automatically generates test file code in the split screen / bottom for the user to review later.

[0147] In some embodiments of this application, in a chat scenario, the input method may include the following process:

[0148] Main task (keyboard typing): The user is typing: "Let me show you the Corgi I photographed in Sanlitun yesterday..."

[0149] Auxiliary task (voice): The user doesn't need to stop, just say: "Post the picture of the corgi from yesterday here."

[0150] AI Response: Users simply click on the image, and the image is directly inserted as a thumbnail into the input box, while the text remains after it. This eliminates the need to open the photo album to find the image.

[0151] In some embodiments of this application, under on-the-fly translation, the input method may include the following process:

[0152] The keyboard inputs English text. If there's an English word you don't know how to write, you can use voice input to describe the relevant Chinese meaning. The AI ​​will then combine this information to output the final English text.

[0153] In some embodiments of this application, the above input method can also be implemented using instructions that combine fuzzy description with precise parameters:

[0154] User behavior:

[0155] Voiceover: "I'm looking for a keyboard that... looks retro, with a bit of a cyberpunk feel..."

[0156] Keyboard: Enter the precise parameters "price < 500" and "98-key layout".

[0157] AI fusion logic: AI extracts visual style descriptions (retro, cyberpunk) from speech as vector retrieval conditions, while keyboard input is used as a hard filter.

[0158] Final result: Instead of generating a text, directly display search results for "98-key layout keyboards" that match the "cyberpunk style" and are priced below 500 yuan.

[0159] In this embodiment, when the intent recognition information indicates the retrieval of first data stored in the electronic device, the first data or its corresponding abbreviation is displayed in the voice input area; or, when the intent recognition information indicates the generation of second data, the generated second data is displayed in the voice input area; or, when the intent recognition information indicates the generation of third data and the text input area includes corresponding descriptive information, the third data is generated according to the intent recognition information and the descriptive information, and the third data is displayed. In this way, various user intents can be processed to further improve the convenience of the user input process.

[0160] Optionally, the first data includes at least one of the following: image data in the album, file data in the folder, and chat history data;

[0161] The second data includes at least one of the following: dynamic image data or video data that matches the text information input by the text input;

[0162] The third data includes at least one of the following: product data, online image data, and online video data.

[0163] The image data can be various types of image data, such as picture data or video data.

[0164] In this embodiment, by including at least one of the following in the first data: image data in the album, file data in the folder, and chat history data; by including at least one of the following in the second data: dynamic image data and video data that match the text information entered by the text input; and by including at least one of the following in the third data: product data, online image data, and online video data, various user intentions can be processed to further improve the convenience of the user input process.

[0165] Optionally, displaying the first input window includes:

[0166] Display the initial input window;

[0167] Upon receiving a second input from the user in the initial input window, the first input window is displayed, and the first input window is controlled to be in a first interactive mode; or,

[0168] Upon receiving a third input from the user in the initial input window, the first input window is displayed, and the first input window is controlled to be in a second interaction mode.

[0169] In the case of receiving the same text input and voice input, the first input information corresponding to the first interaction mode is different from the first input information corresponding to the second interaction mode.

[0170] Understandably, in the first interaction mode, the first input information is a fused information obtained by directly combining the text input and the text information corresponding to the voice input. In the second interaction mode, the first input information is either edited based on the intention of the voice input, or information matching the voice input is generated based on the intention of the voice input. Therefore, when the same text input and voice input are received, the first input information corresponding to the first interaction mode is different from the first input information corresponding to the second interaction mode. For example, when the text input is "AAAB" and the voice input is "change B in the text input area to A", in the first interaction mode, the first input information might be: "AAAB change B in the text input area to A". However, in the second interaction mode, the first input information would be "AAAA".

[0171] In some embodiments of this application, the initial input window can be a plain text input window as shown in Figure 5(a) or a plain voice input window as shown in Figure 5(b). The second input can be a click on the microphone icon in Figure 5(a) or Figure 5(b), and the third input can be a long press on the microphone icon in Figure 5(a) or Figure 5(b). It is understood that the second and third inputs can also be other common touch inputs, such as a double-click for the second input and a single-click for the third input, etc., which can be set as needed. Please refer to Figure 5(c) for a schematic diagram of the first input window, and please refer to Figure 5(d) for a schematic diagram of the first input information generated after mixed voice and text input.

[0172] It is understandable that when users do not need to perform mixed input, they can also perform plain text input or plain voice input based on the initial input window.

[0173] The input method provided in this application redesigns the input interface of input methods in related technologies to meet users' demands for simultaneous text and voice input, thereby enabling natural communication between humans and intelligent agents.

[0174] In this embodiment, an initial input window is displayed; upon receiving a second input from the user in the initial input window, the first input window is displayed and controlled to be in a first interaction mode; or, upon receiving a third input from the user in the initial input window, the first input window is displayed and controlled to be in a second interaction mode. In this way, the user can select the interaction mode of the first input window as needed, that is, the user can set the interaction mode of the first input window according to their own interaction habits, so as to further improve the convenience of the information input process.

[0175] The input method provided in this application can be executed by an input device. This application uses an input device executing the input method as an example to illustrate the input device provided in this application.

[0176] Please see Figure 11 , Figure 11 This is a schematic diagram of the structure of an input device 1100 provided in an embodiment of this application. The input device 1100 is applied to an electronic device and includes:

[0177] The receiving module 1101 is used to receive text input and voice input from the user;

[0178] Display module 1102 is used to display first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input.

[0179] Optionally, the display module 1102 is further configured to display a first input window, wherein the first input window includes a text input area and a voice input area.

[0180] Optionally, the display module 1102 includes:

[0181] The fusion submodule is used to fuse the first text information in the text input area and the second text information in the voice input area based on an artificial intelligence (AI) model to obtain the first input information; wherein the second text information is obtained based on the voice information;

[0182] The display submodule is used to display the first input information in the first input window.

[0183] Optionally, the receiving module 1101 is specifically used to receive the first text information input by the text input area based on the text input area, and to convert the received voice information into the second text information and display it in the voice input area;

[0184] The display submodule is specifically used to cancel the display of the first text information and the second text information, and to display the first input information and the return control;

[0185] The display submodule is further configured to, upon receiving a first input from the user to the return control, cancel the display of the first input information and redisplay the first text information and the second text information;

[0186] The device further includes:

[0187] An editing module is used to edit the first text information and / or the second text information in response to receiving editing input from a user.

[0188] Optionally, the device further includes:

[0189] The switching module is used to switch the editing area in the first input window from the first editing area to the second editing area when a user switches the editing area. The first editing area and the second editing area are respectively the voice input area and the text input area.

[0190] Optionally, the device further includes:

[0191] The recognition submodule is used to perform intent recognition on the voice information to obtain intent recognition information;

[0192] The editing submodule is used to edit the text in the text input area according to the editing method indicated by the intent recognition information when the intent recognition information indicates that the text information in the text input area should be edited, so as to obtain the first input information;

[0193] The display submodule is also used to display the first input information in the first input window.

[0194] Optionally, the editing method includes at least one of the following:

[0195] Delete the first sub-message in the text input area;

[0196] Identify typos in the text input area;

[0197] Replace the second sub-information in the text input area;

[0198] Add explanatory information to the third sub-information in the text input area;

[0199] Translate the fourth sub-information in the text input area into the target language;

[0200] Add a fifth piece of information to the text input area.

[0201] Optionally, the display module 1102 is further configured to display the first data or its corresponding abbreviated information in the voice input area when the intent recognition information indicates the retrieval of first data stored in the electronic device; or,

[0202] The display module 1102 is further configured to display the generated second data in the voice input area when the intent recognition information indicates that second data should be generated; or,

[0203] The display module 1102 is further configured to generate the third data according to the intent recognition information and the description information, and display the third data, when the intent recognition information indicates the generation of third data and the text input area includes corresponding description information.

[0204] Optionally, the first data includes at least one of the following: image data in the album, file data in the folder, and chat history data;

[0205] The second data includes at least one of the following: dynamic image data or video data that matches the text information input by the text input;

[0206] The third data includes at least one of the following: product data, online image data, and online video data.

[0207] Optionally, the display module 1102 is specifically used to display the initial input window;

[0208] The display module 1102 is further configured to, upon receiving a second input from the user in the initial input window, display the first input window and control the first input window to be in a first interactive mode; or,

[0209] The display module 1102 is further configured to display the first input window and control the first input window to be in a second interactive mode when a third input from the user is received in the initial input window.

[0210] In the case of receiving the same text input and voice input, the first input information corresponding to the first interaction mode is different from the first input information corresponding to the second interaction mode.

[0211] In this embodiment, when a user inputs data on an electronic device, they can simultaneously perform text input and voice input. The electronic device displays the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input. That is, when inputting data, the user can simultaneously use both voice and text to express information, which helps the user express information more naturally during communication using the electronic device, and thus facilitates the user's accurate expression of their input intentions.

[0212] The input device 1100 in this embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This embodiment does not specifically limit the specific type of device.

[0213] The input device 1100 in this embodiment can be a device with an operating system. This operating system can be Android, iOS, or other possible operating systems; this embodiment does not specifically limit the specific operating system.

[0214] The input device 1100 provided in this application embodiment can achieve... Figure 1 The various processes implemented in the method embodiments achieve the same technical effect, and will not be described again here to avoid repetition.

[0215] In some embodiments, such as Figure 12 As shown, this application embodiment also provides an electronic device 1200, including a processor 1201, a memory 1202, and a program or instructions stored in the memory 1202 and executable on the processor 1201. When the program or instructions are executed by the processor 1201, they implement the various processes of the above-described input method embodiments and achieve the same technical effects. To avoid repetition, they will not be described again here.

[0216] Figure 13 A schematic diagram of the hardware structure of an electronic device according to an embodiment of this application.

[0217] The electronic device 1300 includes, but is not limited to, components such as: radio frequency unit 1301, network module 1302, audio output unit 1303, input unit 1304, sensor 1305, display unit 1306, user input unit 1307, interface unit 1308, memory 1309, and processor 1310.

[0218] The user input unit 1307 is used to receive text input and voice input from the user;

[0219] The processor 1310 is used to control the display unit 1306 to display the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input.

[0220] Optionally, the display unit 1306 is used to display a first input window, wherein the first input window includes a text input area and a voice input area.

[0221] Optionally, the processor 1310 is configured to fuse the first text information in the text input area and the second text information in the voice input area based on an artificial intelligence (AI) model to obtain the first input information; wherein the second text information is obtained based on the voice information;

[0222] The display unit 1306 is used to display the first input information in the first input window.

[0223] Optionally, the user input unit 1307 is configured to receive the first text information input by the text input area based on the text input area, and the processor 1310 is configured to convert the received voice information into the second text information and control the display unit 1306 to display it on the voice input area.

[0224] The display unit 1306 is used to cancel the display of the first text information and the second text information, and to display the first input information and the return control;

[0225] The display unit 1306 is configured to, upon receiving a first input from the user to the return control, cancel the display of the first input information and redisplay the first text information and the second text information;

[0226] The processor 1310 is configured to, upon receiving editing input from a user, edit the first text information and / or the second text information in response to the editing input.

[0227] Optionally, the processor 1310 is configured to switch the editing area in the first input window from a first editing area to a second editing area upon receiving a user's switching instruction, wherein one of the first editing area and the second editing area is the voice input area and the other is the text input area.

[0228] Optionally, the processor 1310 is configured to perform intent recognition on the voice information to obtain intent recognition information;

[0229] The processor 1310 is configured to, when the intent recognition information indicates that the text information in the text input area should be edited, edit the text in the text input area according to the editing method indicated by the intent recognition information to obtain the first input information;

[0230] The display unit 1306 is used to display the first input information in the first input window.

[0231] Optionally, the editing method includes at least one of the following:

[0232] Delete the first sub-message in the text input area;

[0233] Identify typos in the text input area;

[0234] Replace the second sub-information in the text input area;

[0235] Add explanatory information to the third sub-information in the text input area;

[0236] Translate the fourth sub-information in the text input area into the target language;

[0237] Add a fifth piece of information to the text input area.

[0238] Optionally, the display unit 1306 is configured to display the first data or its corresponding abbreviated information in the voice input area when the intent recognition information indicates the retrieval of first data stored in the electronic device; or,

[0239] The display unit 1306 is configured to display the generated second data in the voice input area when the intent recognition information indicates that second data should be generated; or...

[0240] The processor 1310 is configured to generate the third data according to the intent recognition information and the description information when the intent recognition information indicates that the third data is to be generated and the text input area includes corresponding description information, and to control the display unit 1306 to display the third data.

[0241] Optionally, the first data includes at least one of the following: image data in the album, file data in the folder, and chat history data;

[0242] The second data includes at least one of the following: dynamic image data or video data that matches the text information input by the text input;

[0243] The third data includes at least one of the following: product data, online image data, and online video data.

[0244] Optionally, the display unit 1306 is used to display the initial input window;

[0245] The display unit 1306 is configured to display the first input window upon receiving a second input from the user in the initial input window, and the processor 1310 is configured to control the first input window to be in a first interactive mode; or...

[0246] The display unit 1306 is used to display the first input window when a third input from the user is received in the initial input window, and the processor 1310 is used to control the first input window to be in a second interaction mode.

[0247] In the case of receiving the same text input and voice input, the first input information corresponding to the first interaction mode is different from the first input information corresponding to the second interaction mode.

[0248] Those skilled in the art will understand that the electronic device 1300 may also include a power supply (such as a battery) for supplying power to various components. The power supply may be logically connected to the processor 1310 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. Figure 13 The electronic device structure shown does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0249] It should be understood that, in this embodiment, the input unit 1304 may include a graphics processing unit (GPU) 13041 and a microphone 13042. The GPU 13041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1306 may include a display panel 13061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1307 includes a touch panel 13071 and other input devices 13072. The touch panel 13071 is also called a touch screen. The touch panel 13071 may include a touch detection device and a touch controller. Other input devices 13072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0250] The memory 1309 can be used to store software programs and various data. The memory 1309 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1309 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1309 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0251] Processor 1310 may include one or more processing units; optionally, processor 1310 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1310.

[0252] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described input method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0253] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0254] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described input method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0255] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0256] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0257] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0258] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An input method performed by an electronic device, characterized in that, The method includes: It accepts text and voice input from users; Based on the text information corresponding to the text input and the voice information corresponding to the voice input, the first input information is displayed.

2. The method according to claim 1, characterized in that, Before receiving the user's text input and voice input, the method further includes: Display a first input window, which includes a text input area and a voice input area.

3. The method according to claim 2, characterized in that, The step of displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input includes: The first text information in the text input area and the second text information in the voice input area are fused based on an artificial intelligence (AI) model to obtain the first input information; wherein the second text information is obtained based on the voice information. The first input information is displayed in the first input window.

4. The method according to claim 3, characterized in that, The receiving of user text input and voice input includes: The system receives the first text information input by the text input area and converts the received voice information into the second text information and displays it in the voice input area. Displaying the first input information in the first input window includes: Cancel the display of the first and second text information, and display the first input information and the return control; The method further includes: Upon receiving the user's first input to the return control, the first input information is de-displayed, and the first and second text information are re-displayed; Upon receiving editing input from a user, the system edits the first text information and / or the second text information in response to the editing input.

5. The method according to claim 4, characterized in that, After canceling the display of the first input information and redisplaying the first text information and the second text information, the method further includes: Upon receiving a user's switching instruction, the editing area in the first input window is switched from the first editing area to the second editing area, wherein the first editing area and the second editing area are respectively the voice input area and the text input area.

6. The method according to claim 2, characterized in that, The step of displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input includes: The speech information is subjected to intent recognition to obtain intent recognition information; When the intent recognition information indicates that the text information in the text input area should be edited, the text in the text input area is edited according to the editing method indicated by the intent recognition information to obtain the first input information; The first input information is displayed in the first input window.

7. The method according to claim 6, characterized in that, The editing method includes at least one of the following: Delete the first sub-message in the text input area; Identify typos in the text input area; Replace the second sub-information in the text input area; Add explanatory information to the third sub-information in the text input area; Translate the fourth sub-information in the text input area into the target language; Add a fifth piece of information to the text input area.

8. The method according to claim 6, characterized in that, The step of displaying the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input further includes: When the intent recognition information indicates the retrieval of first data stored in the electronic device, the first data or its corresponding abbreviated form is displayed in the voice input area; or... If the intent recognition information indicates that second data should be generated, the generated second data is displayed in the voice input area; or, When the intent recognition information indicates the generation of third data and the text input area includes corresponding descriptive information, the third data is generated according to the intent recognition information and the descriptive information, and the third data is displayed.

9. The method according to claim 8, characterized in that, The first data includes at least one of the following: image data in the album, file data in the folder, and chat history data; The second data includes at least one of the following: dynamic image data or video data that matches the text information input by the text input; The third data includes at least one of the following: product data, online image data, and online video data.

10. The method according to any one of claims 2 to 9, characterized in that, The display of the first input window includes: Display the initial input window; Upon receiving a second input from the user in the initial input window, the first input window is displayed, and the first input window is controlled to be in a first interactive mode; or, Upon receiving a third input from the user in the initial input window, the first input window is displayed, and the first input window is controlled to be in a second interaction mode. In the case of receiving the same text input and voice input, the first input information corresponding to the first interaction mode is different from the first input information corresponding to the second interaction mode.

11. An input device, applied to an electronic device, characterized in that, The device includes: The receiving module is used to receive the user's text input and voice input; The display module is used to display the first input information based on the text information corresponding to the text input and the voice information corresponding to the voice input.

12. The apparatus according to claim 11, characterized in that, The display module is also used to display a first input window, wherein the first input window includes a text input area and a voice input area.

13. The apparatus according to claim 12, characterized in that, The display module includes: The fusion submodule is used to fuse the first text information in the text input area and the second text information in the voice input area based on an artificial intelligence (AI) model to obtain the first input information; wherein the second text information is obtained based on the voice information; The display submodule is used to display the first input information in the first input window.

14. The apparatus according to claim 13, characterized in that, The receiving module is specifically configured to receive the first text information input by the text input area based on the text input area, and to convert the received voice information into the second text information and display it in the voice input area; The display submodule is specifically used to cancel the display of the first text information and the second text information, and to display the first input information and the return control; The display submodule is further configured to, upon receiving a first input from the user to the return control, cancel the display of the first input information and redisplay the first text information and the second text information; The device further includes: An editing module is used to edit the first text information and / or the second text information in response to receiving editing input from a user.

15. The apparatus according to claim 14, characterized in that, The device further includes: The switching module is used to switch the editing area in the first input window from the first editing area to the second editing area when a user switches the editing area. The first editing area and the second editing area are respectively the voice input area and the text input area.

16. The apparatus according to claim 12, characterized in that, The device further includes: The recognition submodule is used to perform intent recognition on the voice information to obtain intent recognition information; The editing submodule is used to edit the text in the text input area according to the editing method indicated by the intent recognition information when the intent recognition information indicates that the text information in the text input area should be edited, so as to obtain the first input information; The display submodule is used to display the first input information in the first input window.

17. The apparatus according to claim 16, characterized in that, The editing method includes at least one of the following: Delete the first sub-message in the text input area; Identify typos in the text input area; Replace the second sub-information in the text input area; Add explanatory information to the third sub-information in the text input area; Translate the fourth sub-information in the text input area into the target language; Add a fifth piece of information to the text input area.

18. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a program or instructions that can run on the processor, and the program or instructions, when executed by the processor, implement the steps of the input method as described in any one of claims 1-10.

19. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the input method as described in any one of claims 1-10.

20. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the steps of the input method as described in any one of claims 1-10.