Systems, methods, and apparatus for multi-modal text editing

By combining pressure sensors on electronic pens and touch screen devices to achieve seamless integration of handwriting and voice input, the problem of complex user interface switching in existing systems is solved, and the efficiency of text editing and user experience are improved.

CN120752638APending Publication Date: 2025-10-03HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380088329.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-14
Filing Date
2023-12-07
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing systems that support both voice input and pen input are rather confusing in user interface switching, resulting in low input efficiency and difficulty in meeting user needs in different text editing tasks.

Method used

By combining an electronic pen and a touch screen device, pressure sensors are used to enable seamless integration of handwriting and voice input, achieving seamless switching and editing between pen input and voice input, providing content insertion, correction, and formatting functions, and reducing dependence on user interface switching.

Benefits of technology

It improves the efficiency of text editing and user experience, achieves seamless integration of voice input and pen input by simplifying mode switching, and enhances the intuitiveness and accuracy of input.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120752638A_ABST
    Figure CN120752638A_ABST
Patent Text Reader

Abstract

The invention discloses a text editing system and method. A system may include a touch screen, a memory, and a processor. A voice input channel is activated, and one or more alternative word options are presented after text content is selected at a location of the touch screen through a preset shape. And determining an input type, and obtaining input information according to the input type. And replacing the selected text content by using the acquired input information.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] The present invention claims the benefit of priority to U.S. patent application Ser. No. 18 / 508,454, filed on Nov. 13, 2023, entitled “SYSTEM, METHOD AND DEVICE FOR MULTIMODAL TEXT EDITING,” and U.S. patent application Ser. No. 18 / 145,295, filed on Dec. 22, 2022, entitled “SYSTEM, METHOD AND DEVICE FOR MULTIMODAL TEXT EDITING,” the entireties of which are incorporated herein by reference. Technical Field

[0003] The present invention generally relates to text editing devices and methods, including enabling multi-modal text editing using an electronic pen. Background Art

[0004] In recent years, as natural language processing (NLP) technology has become increasingly reliable, voice input has become an increasingly popular method for text editing and word processing applications. While a skilled typist can type at 70 words per minute, voice input can reach over 400 words per minute. It is estimated that voice input is currently the third most popular text editing method and will rise to second place within the next five to ten years.

[0005] Despite its advantages, voice input can be inconvenient for some common text editing tasks. While a keyboard and / or mouse can be used in these situations, an electronic pen offers greater precision, simple and intuitive gesture control, direct pointing, and portability. Therefore, combining voice input with pen input can maximize efficiency in text content creation.

[0006] Current systems that support both voice and pen input can be confusing to use. For example, in many such systems, users inevitably switch between pen writing, pen gestures, voice input, voice commands, touch input, and soft keyboards primarily through the associated user interface (UI). This is an area where improvement is urgently needed. Summary of the Invention

[0007] The examples described in this article combine pen input with voice input to enhance the multimodal text editing experience while using touchscreen devices and electronic pens in a natural gesture. Pen input enables more intuitive post-editing and / or short input tasks based on voice-transcribed text, potentially benefiting from the advantages of precise pointing, handwriting, and other embedded sensing technologies. The examples described in this article provide for content insertion, content correction, and content formatting, potentially improving efficiency.

[0008] The examples described in this article can use the pressure sensor within the electronic pen to more conveniently enable and disable different modes. The examples described in this article provide contextual selection of homophones, which may help solve typical speech dictation problems. This can eliminate the need for complex mode switching routines that rely on the user interface (UI), thereby improving input efficiency.

[0009] The example provided in this article enables text editing of speech transcriptions while the touchscreen device continues to receive speech input, thus achieving seamless integration of speech and pen input.

[0010] According to one aspect of the present application, a computer system is provided. The computer system includes: a touch screen; a processor; and a memory coupled to the processor, wherein the memory stores instructions that, when executed by the processor, cause the system to perform the following operations when executing a text editing application: receiving a pressure signal indicating that a first pressure value is detected at a pen tip from an electronic pen in communication with the processor; enabling handwriting recognition in response to receiving the pressure signal indicating that the first pressure value is detected at the pen tip; receiving a touch input representing handwriting at a first position on the touch screen; and converting the touch input representing handwriting into rendered text content corresponding to the handwriting.

[0011] In some implementations, the system is further configured to perform the following operations: receive touch input at a second location on the touch screen; receive a request to enable voice recognition from the electronic pen; enable voice recognition; receive a signal representing voice input from a microphone in communication with the processor; and convert the signal representing voice input into rendered text content corresponding to the voice input.

[0012] In some implementations, the request to enable voice recognition is received via a pressure signal from the electronic pen indicating a second pressure value detected at the pen tip, the second pressure value being distinguishable from the first pressure value.

[0013] In some implementations, the electronic pen further includes a button, and the request to enable voice recognition is received via an input signal indicative of a button press on the button.

[0014] In some implementations, the system is further configured to perform the following operations: receive a touch input representing an ellipse at a third position of the touch screen; identify target text content, wherein the target text content is text content presented at the third position of the touch screen; determine one or more replacement text candidates corresponding to the target text content; display the one or more replacement text candidates as optional options for replacing the target text content; and in response to selecting one of the one or more replacement text candidates, replace the target text content with the selected one of the one or more replacement text candidates.

[0015] In some implementations, the system is further configured to perform the following operations: receive a touch input representing an ellipse at a fourth position on the touch screen; identify target text content, wherein the target text content is text content presented at the fourth position on the touch screen; receive an input representing replacement content; and replace the target text content with presented text content corresponding to the replacement content.

[0016] In some implementations, the received input is touch input representing handwriting.

[0017] In some implementations, the received input is speech input representing a spoken word.

[0018] In some implementations, the system is further configured to perform the following operations: receive a touch input representing a strikethrough at a fifth position on the touch screen; identify target text content, wherein the target text content is text content presented below the touch input representing the strikethrough; display one or more content format options near the target text content; receive a selection from one or more content format options; and modify the target text content based on the received selection from the one or more content format options.

[0019] In some implementations, the system is further configured to perform the following operations: receive a touch input representing a strikethrough at a sixth position on the touch screen; identify target text content, wherein the target text content is text content presented below the touch input representing the strikethrough; receive an instruction to enable voice recognition through the electronic pen; enable voice recognition; receive a signal representing voice input from a microphone communicating with the processor; recognize the voice input as a voice command corresponding to a content format option; and modify the target text content according to the content format option.

[0020] In some implementations, the system is further configured to perform the following operations: receive a touch input representing a strikethrough at a seventh position on the touch screen; identify target text content, wherein the target text content is text content presented below the strikethrough; receive an instruction to enable voice recognition through the electronic pen; enable voice recognition; receive voice input through a microphone communicating with the processor; recognize the voice input as a voice dictation; and replace the target text content with content corresponding to the voice dictation.

[0021] In some implementations, the system is further configured to perform the following operations: receive a touch input representing a strikethrough at an eighth position on the touch screen; identify target text content, wherein the target text content is text content presented below the strikethrough; receive an instruction to enable handwriting recognition through the electronic pen; enable handwriting recognition; remove the target text content from the eighth position on the touch screen; receive a touch input representing handwriting at the eighth position on the touch screen; and convert the touch input representing handwriting into presented text content corresponding to the handwriting.

[0022] In some implementations, the electronic pen also includes a button, and causes the system to perform the following operations: receive an input signal indicating a button press on the electronic pen from the electronic pen in communication with the processor; enable voice dictation in response to receiving the input signal indicating a button press on the electronic pen; receive touch input at a ninth position on the touch screen; receive voice input; and present text content corresponding to the voice dictation at the ninth position on the touch screen.

[0023] According to another aspect of the present application, a computer-implemented method for execution in a text editing application is provided. The method includes: receiving a pressure signal from an electronic pen indicating that a first pressure value is detected at a pen tip; in response to receiving the pressure signal indicating that the first pressure value is detected at the pen tip, enabling handwriting recognition; receiving a touch input representing handwriting at a first location on a touch screen; and converting the touch input representing handwriting into rendered text content corresponding to the handwriting.

[0024] In some implementations, the method further includes: receiving touch input at a second location on the touch screen; receiving a request to enable voice recognition from the electronic pen; enabling voice recognition; receiving a signal representing the voice input from a microphone; and converting the signal representing the voice input into rendered text content corresponding to the voice input.

[0025] In some implementations, the request to enable voice recognition is received via a pressure signal from the electronic pen indicating a second pressure value detected at the pen tip, the second pressure value being distinguishable from the first pressure value.

[0026] In some implementations, the electronic pen further includes a button, and the request to enable voice recognition is received via a pressure signal from the button. The method may further include: receiving an input signal from the electronic pen indicating a button press on the electronic pen; enabling voice dictation in response to receiving the input signal indicating the button press on the electronic pen; receiving a touch input at a ninth position on the touch screen, wherein the ninth position corresponds to a position before a series of presented content; receiving voice input; and presenting text content corresponding to the voice dictation at the ninth position on the touch screen.

[0027] In some implementations, the method further includes: receiving a touch input representing an ellipse at a third position of the touch screen; identifying target text content, wherein the target text content is text content presented at the third position of the touch screen; determining one or more replacement text candidates corresponding to the target text content; displaying the one or more replacement text candidates as optional options for replacing the target text content; and in response to selecting a selected one of the replacement text candidates, replacing the target text content with the selected one of the replacement text candidates.

[0028] In some implementations, the method further includes: receiving a touch input representing an ellipse at a fourth position of the touch screen; identifying target text content, wherein the target text content is text content presented at the fourth position of the touch screen; receiving an input representing replacement content; and replacing the target text content with presented text content corresponding to the replacement content.

[0029] According to another aspect of the present invention, a non-transitory computer-readable medium including instructions is provided. When executed by a processor, the instructions cause the processor to perform the following operations: receiving a pressure signal indicating that a first pressure value is detected at a tip of an electronic pen in communication with the processor; enabling handwriting recognition in response to receiving the pressure signal indicating that the first pressure value is detected at the tip of the pen; receiving a touch input representing handwriting at a first location on a touch screen; and converting the touch input representing handwriting into rendered text content corresponding to the handwriting.

[0030] According to another aspect of the present invention, a computer system is provided. The computer system includes: a touch screen; a processor; and a memory coupled to the processor, wherein the memory stores instructions that, when executed by the processor, cause the system to perform the following operations when executing a text editing application: receiving a request to enable voice input; in response to receiving the request to enable voice input, enabling voice recognition and displaying a voice cursor at a first position on the touch screen, wherein the voice cursor indicates a first position for presenting voice input; receiving a signal representing voice input from a microphone in communication with the processor; processing the voice input at the first position into presented text content corresponding to the voice input; receiving touch input at a second position on the touch screen; in response to receiving touch input at the second position on the touch screen, displaying an edit cursor at the second position on the touch screen, wherein the edit cursor indicates a second position different from the first position relative to the presented text content corresponding to the voice input for editing text content; the voice cursor is positioned independently relative to the voice input and the edit cursor is positioned independently relative to the touch input; and the voice input and the touch input are processed simultaneously.

[0031] In some implementations, the touch input is received by an electronic pen in communication with the processor contacting the touch screen.

[0032] In some implementations, the touch input is handwriting.

[0033] In some implementations, the touch input is a selection among the replacement text candidates.

[0034] According to another aspect of the present invention, a computer system is provided. The computer system includes: a touch screen; a processor; and a memory coupled to the processor, wherein the memory stores instructions that, when executed by the processor, cause the system to perform the following operations when executing a text editing application: activate a voice input channel, and after selecting text content at a location on the touch screen using a preset shape, present one or more alternative word options; determine an input type; obtain input information based on the input type; and replace the selected text content with the obtained input information.

[0035] In some implementations, when the instructions are executed by the processor, the system may also perform the following operations to obtain input information according to the input type: when the input type is voice input, the content of the voice input is identified; when the instructions are executed by the processor, the system may also perform the following operations to replace the selected text content with the obtained input information: when the content of the voice input is descriptive language, the selected text content is replaced with the obtained input information.

[0036] In some implementations, when the instructions are executed by the processor, the system may also obtain input information according to the input type through the following operations: when the input type is voice input, identify the content of the voice input; when the instructions are executed by the processor, the system may also use the obtained input information to replace the selected text content through the following operations: when the content of the voice input is a voice instruction, search for target text that meets the requirements of the content; and replace the selected text content with the target associative word selected by user input.

[0037] In some implementations, when the instruction is executed by the processor, it can also cause the system to obtain input information according to the input type through the following operations: when the input type is to select one of the one or more alternative word options, determine the target option selected from the one or more alternative word options through user input; when the instruction is executed by the processor, it can also cause the system to use the obtained information to replace the selected text content through the following operations: replace the selected text content with the target option.

[0038] In some implementations, the selected text content may be input through handwriting input, and the one or more alternative word options may include one or more words that look similar to the selected text content.

[0039] In some implementations, the selected text content may be input through voice input, and the one or more alternative word options may include one or more words that are homophones or homonyms of the selected text content.

[0040] In some implementations, one or more of the one or more alternative word options can be provided by an artificial intelligence model.

[0041] In some implementations, one or more of the one or more alternative word options can be provided by an input method tool.

[0042] In some implementations, when the instructions are executed by the processor, the system may also obtain input information according to the input type through the following operations: when the input type is written input, determine one or more written words; when the instructions are executed by the processor, the system may also use the obtained information to replace the selected text content through the following operations: replace the selected text content with the one or more written words.

[0043] According to another aspect of the present invention, a computer-implemented method for use in a text editing application is provided. The method comprises: enabling a voice input channel, and after selecting text content at a location on the touch screen using a preset shape, presenting one or more alternative word options; determining an input type; obtaining input information based on the input type; and replacing the selected text content with the obtained input information.

[0044] In some implementations, obtaining input information according to the input type may include: when the input type is voice input, identifying the content of the voice input; using the obtained input information to replace the selected text content may include: when the content of the voice input is descriptive language, using the obtained input information to replace the selected text content.

[0045] In some implementations, obtaining input information according to the input type may include: when the input type is voice input, identifying the content of the voice input; using the obtained input information to replace the selected text content may include: when the content of the voice input is a voice instruction, searching for a target text that meets the requirements of the content; and replacing the selected text content with a target associative word selected by user input.

[0046] In some implementations, obtaining input information according to the input type may include: when the input type is to select one of the one or more alternative word options, determining the target option selected from the one or more alternative word options through user input; using the obtained information to replace the selected text content may include: using the target option to replace the selected text content.

[0047] In some implementations, the selected text content may be input through handwriting input, and the one or more alternative word options may include one or more words that look similar to the selected text content.

[0048] In some implementations, the selected text content may be input through voice input, and the one or more alternative word options may include one or more words that are homophones or homonyms of the selected text content.

[0049] In some implementations, one or more of the one or more alternative word options can be provided by an artificial intelligence model.

[0050] In some implementations, one or more of the one or more alternative word options can be provided by an input method tool.

[0051] In some implementations, obtaining input information according to the input type may include: when the input type is writing input, determining one or more written words; and using the obtained information to replace the selected text content may include: replacing the selected text content with the one or more written words.

[0052] According to another aspect of the present invention, a non-transitory computer-readable medium storing instructions is provided. When executed by a processor of a computer system, the instructions cause the system to perform the following operations when executing a text editing application: activating a voice input channel, and after selecting text content at a location on the touch screen using a preset shape, presenting one or more alternative word options; determining an input type; obtaining input information based on the input type; and replacing the selected text content with the obtained input information.

[0053] In some implementations, when the instructions are executed by the processor, they may also cause the system to perform any one of the exemplary implementations of the above method. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Reference is now made, by way of example, to the accompanying drawings which illustrate exemplary embodiments of the present invention, in which:

[0055] Figure 1 is a schematic diagram of a first exemplary operating environment according to exemplary embodiments of the examples described herein;

[0056] Figure 2 is a high-level operational diagram of an exemplary computing device 200 according to examples described herein;

[0057] Figure 3 illustrates a simplified organization of software components that may be stored in the memory of an exemplary computing device according to examples described herein;

[0058] Figure 4is a simplified organization of components that may be connected to an input / output (I / O) interface in accordance with some embodiments (e.g., when the exemplary computing device operates as a touch screen device);

[0059] Figure 5 is a simplified organization of components that may be connected to an I / O interface according to some embodiments (e.g., when the exemplary computing device operates as an electronic pen);

[0060] Figure 6 is a flow chart of an exemplary text editing method 600 according to examples described herein;

[0061] Figure 7A 、 Figure 7B and Figure 7C shows an exemplary sentence presented on a touch screen according to examples described herein;

[0062] Figure 8 is a flow chart of an exemplary text editing method according to examples described herein;

[0063] Figure 9A 、 Figure 9B and Figure 9C shows an exemplary sentence presented on a touch screen according to examples described herein;

[0064] Figure 10 is a flow chart of an exemplary text editing method according to examples described herein;

[0065] Figure 11A and Figure 11B shows an exemplary sentence including textual content that may be presented on a touch screen according to examples described herein;

[0066] Figure 12A and Figure 12B shows an exemplary sentence including textual content that may be presented on a touch screen according to examples described herein;

[0067] Figure 13 is a flow chart of an exemplary text editing method according to examples described herein;

[0068] Figure 14A 、 Figure 14B 、 Figure 14C and Figure 14D shows an exemplary sentence including textual content that may be presented on a touch screen according to examples described herein;

[0069] Figure 15 is a flow chart of an exemplary text editing method according to examples described herein;

[0070] Figure 16A and Figure 16B shows an exemplary sentence including textual content that may be presented on a touch screen according to examples described herein;

[0071] Figure 17 is a flow chart of an exemplary text editing method according to examples described herein;

[0072] Figure 18A and Figure 18B shows an exemplary sentence including textual content that may be presented on a touch screen according to examples described herein;

[0073] Figure 19 is a flow chart of an exemplary text editing method according to examples described herein;

[0074] Figure 20A 、 Figure 20B and Figure 20C shows an exemplary sentence including textual content that may be presented on a touch screen according to examples described herein;

[0075] Figure 21 is a flow chart of an exemplary text editing method according to examples described herein;

[0076] Figure 22A 、 Figure 22B and Figure 22C shows an exemplary sentence including textual content that may be presented on a touch screen according to examples described herein;

[0077] Figure 23 is a flow chart of an exemplary text editing method according to examples described herein;

[0078] Figure 24A 、 Figure 24B 、 Figure 24C and Figure 24D shows an exemplary sentence including textual content that may be presented on a touch screen according to examples described herein;

[0079] Figure 25 An exemplary implementation of simultaneous voice input and text input according to the examples described herein is shown;

[0080] Figure 26A and Figure 26B A flowchart illustrating how various examples disclosed herein may be implemented together according to examples described herein;

[0081] Figure 27 An exemplary implementation of simultaneous voice input and text input according to the examples described herein is shown;

[0082] Like reference numerals may be used in different drawings to identify like components. DETAILED DESCRIPTION

[0083] The embodiments described herein can run on a variety of touchscreen devices, including dual-screen laptops, convertible laptops, standard laptops, tablets, smartphones, and digital whiteboards.

[0084] In the present invention, "screen" refers to the outer layer of the touch screen display facing the user.

[0085] As used herein, the terms "touch screen element" and "touch screen" refer to a combination of a display and a touch sensing system that is capable of functioning as an input device by receiving touch input. Non-limiting examples of touch screen displays include capacitive touch screens, resistive touch screens, infrared touch screens, and surface acoustic wave touch screens.

[0086] In this disclosure, the term "touch screen device" refers to a computing device having a touch screen element.

[0087] In the present invention, the term "application" refers to a software program including a set of instructions that can be executed by a processing device of an electronic device.

[0088] Figure 1 FIG. 1 is a schematic diagram of a first exemplary operating environment 100 of an exemplary embodiment. As shown in the figure, the first exemplary operating environment 100 includes an electronic pen 110 held by a user's hand 120 above a touch screen device 140 .

[0089] As shown, touch screen device 140 includes touch screen element 130. The touch screen element includes a touch panel (input device) and a display (output device). Therefore, touch screen element 130 can be used to present content and sense touches thereon. As described above, touch screen element 130 can also be described as touch screen 130. Touch screen 130 can implement one or more touch screen technologies. For example, touch screen 130 can be a resistive film touch screen, a surface capacitive touch screen, a projected capacitive touch screen, a surface acoustic wave (SAW) touch screen, an optical touch screen, an electromagnetic touch screen, etc.

[0090] The touch screen device 140 may include a touch screen device microphone 150. Figure 1 , the touch screen device microphone 150 is shown to be disposed at the lower left corner of the touch screen device 140 , but the touch screen device microphone 150 may be disposed at other locations of the touch screen device 140 , for example, at the back of the touch screen device 140 .

[0091] Although Figure 1A tablet computer is shown, but touch screen device 140 may be a smartphone, laptop computer, and / or other similar electronic devices that can be used to execute and display text editing applications thereon. Touch screen device 140 may be a computer system within the scope of the present invention.

[0092] The electronic pen 110 includes a pen tip 160. The touch screen device 140 can detect the position of the pen tip 160 of the electronic pen 110 on the touch screen 130. In this way, the electronic pen 110 can function as a stylus. In some examples described herein, the electronic pen can include one or more pressure sensors for detecting pressure at the pen tip 160. The pressure sensors can be used to determine multiple distinguishable pressure values ​​at the pen tip 160. The electronic pen 110 can be described as a digital pen and / or a smart pen. In some embodiments, the electronic pen can also include a button 170. In some embodiments, the electronic pen can also include an electronic pen microphone 180.

[0093] The touch screen device 140 can be communicatively coupled with the electronic pen 110. For example, the touch screen device 140 can communicate with the electronic pen 110 via Bluetooth. TM , near-field communication (NFC) or other forms of short-range wireless communication to communicate with the electronic pen 110.

[0094] Figure 2 is a high-level diagram of the operation of an exemplary computing device 200 according to an embodiment of the present invention. In at least some embodiments, the exemplary computing device 200 may be a touch screen device 140 ( Figure 1 ) and / or electronic pen 110 ( Figure 1 ) are examples and are not intended to be limiting.

[0095] The exemplary computing device 200 includes various components. For example, as shown, the exemplary computing device 200 may include a processor 202, an input / output (I / O) interface 204, a communication component 206, a memory 210, and / or a storage unit 208. As shown, the exemplary components of the exemplary computing device 200 communicate via a bus 212. The bus 212 shown in the figure provides communication between the components of the computing device 200. The bus 212 may be any suitable bus architecture, including a memory bus, a peripheral bus, or a video bus.

[0096] The processor 202 may include one or more processors, such as a microprocessor, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a dedicated logic circuit, or a combination thereof.

[0097] The communication component 206 enables the exemplary computing device 200 to communicate with other computers or computing devices and / or various communication networks. The communication component 206 may include one or more network interfaces for wired or wireless communication with a network (e.g., an intranet, the Internet, a peer-to-peer (P2P) network, a wide area network (WAN), and / or a local area network (LAN)) or other nodes. The one or more network interfaces may include wired links (e.g., Ethernet cables) and / or wireless links (e.g., one or more antennas) for communication within the network and / or communication outside the network.

[0098] The communication component 206 can enable the exemplary computing device 200 to send or receive communication signals. The communication signals can be sent or received according to one or more protocols or according to one or more standards. For example, the communication component 206 can enable the exemplary computing device 200 to communicate over a cellular data network, such as, for example, according to one or more standards, such as Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Evolution Data Optimized (EVDO), Long-term Evolution (LTE), etc. Additionally or alternatively, the communication component 206 can enable the exemplary computing device 200 to communicate over a cellular data network, such as, for example, according to one or more standards, such as, for example, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Evolution Data Optimized (EVDO), Long-term Evolution (LTE), etc. TM ,Bluetooth TM Or some combination of one or more networks or protocols to communicate. In some embodiments, all or part of the communication component 206 can be integrated into the components of the touch screen device 140. For example, the communication component 206 can be integrated into a communication chipset.

[0099] The exemplary computing device 200 may include one or more memories 210, which may include volatile (e.g., random access memory (RAM)) and non-volatile or non-transitory memory (e.g., flash memory, magnetic storage, and / or read-only memory (ROM)). The one or more non-transitory memories of the memory 210 store programs including software instructions that are executed by the processor 202 to perform the examples described herein, among other things. In the exemplary embodiment, these programs include software instructions for implementing an operating system (OS) and software applications.

[0100] In some examples, memory 210 may include software instructions for exemplary computing device 200, which are executed by processor 202 to perform the operations described herein. In some other examples, one or more data sets and / or modules may be provided by external memory (e.g., an external drive in wired or wireless communication with computing device 200) or by a transient or non-transitory computer-readable medium. Examples of non-transitory computer-readable media include RAM, ROM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, CD-ROM, or other portable memory.

[0101] Storage unit 208 may be one or more storage units and may include a mass storage unit, such as a solid state drive, a hard disk drive, a magnetic disk drive, and / or an optical disk drive. In some embodiments of exemplary computing device 200, storage unit 208 may be optional.

[0102] I / O interface 204 may be one or more I / O interfaces and may enable connection to one or more suitable input and / or output devices, such as a physical keyboard (not shown).

[0103] Figure 3 shows that can be stored in an exemplary computing device 200 ( Figure 2 ) in the memory 210. As shown, these software components include application software 310 and an operating system (OS) 320.

[0104] The application software 310 converts the exemplary computing device 200 ( Figure 2) is used in conjunction with OS 320 to enable it to operate as a device that performs specific functions. In some embodiments, application software 310 may include a virtual input device application.

[0105] The OS 320 is software. The OS 320 enables the application software 310 to access the processor 202, the I / O interface 204, the communication component 206, the memory 210, and / or the storage unit 208 ( Figure 2 ). OS 320 may be Apple TM iOS TM 、Android TM 、Microsoft TM Windows TM or Google TM ChromeOS TM wait.

[0106] OS 320 may include various modules, such as driver 330. Driver 330 provides a program interface to control and manage specific low-level interfaces that are typically connected to a specific type of hardware. For example, in the exemplary computing device 200 ( Figure 2 ) in an embodiment when operating as a touch screen device, the OS 320 may include a touch panel driver 330 and / or a display driver 330. In the exemplary computing device 200 ( Figure 2 ) as a touch screen device 140 communicating with the electronic pen 110 ( Figure 1 )In an embodiment of the runtime, the touch screen device 140 can be referred to as a computer system.

[0107] Reference below Figure 4 , which is a diagram according to some embodiments (e.g., when the exemplary computing device 200 ( Figure 2 ) is a simplified organization of components that can be connected to the I / O interface 204 when operating as a touch screen device 140. Figure 4 As shown, the I / O interface 204 can communicate with the touch panel 244 and the touch screen display 242. As described above, the touch panel 244 and the touch screen display 242 can constitute part of the touch screen element 130. The touch panel 244 can include various touch sensors for sensing touch input, which may depend on the touch sensing mode used by the touch screen element 130 (e.g., resistive sensor, capacitive sensor, SAW device, optical sensor, electromagnetic sensor, etc.).

[0108] In the exemplary computing device 200 ( Figure 2 ) in some embodiments where the touch screen device 140 is running, the application software 310 may be configured to display the touch screen device 140 by means of the display driver 330 ( Figure 3) presents content on the touch screen display 242, and touches on the touch panel 244 can be sensed by the touch panel driver 330.

[0109] As described above, the touch panel driver 330 ( Figure 3 ) can be coupled to the touch panel 244 to generate touch events. Display driver 330 ( Figure 3 ) can be coupled to the touch screen display 242 to present content on the touch screen display 242.

[0110] Reference below Figure 5 , which is a diagram according to some embodiments (eg, when the exemplary computing device is used as an electronic pen 110 ( Figure 1 ) is a simplified organization of components that can be connected to the I / O interface 204 (at runtime). Figure 5 As shown, the I / O interface 204 can communicate with one or more pressure sensors 510. In some embodiments, the pressure sensor 510 can be located at the electronic pen 110 ( Figure 1 ) and can be used to detect the tip 160 ( Figure 1 ) at multiple distinguishable pressure values.

[0111] Figure 6 6 is a flowchart of an exemplary text editing method 600 according to one embodiment of the present invention. The method 600 may be executed by one or more processors of a computing system (eg, a touch screen device 140 ( Figure 1 ) is executed by one or more processors).

[0112] At operation 602, the system receives an indication that the pen tip 160 ( Figure 1 ) detects a pressure signal of a first pressure value. These pressure signals can be generated by the electronic pen 110 ( Figure 1 ) of one or more pressure sensors 510 ( Figure 5 ) is generated. Electronic pen 110 ( Figure 1 ) can be used with the touch screen device 140 ( Figure 1 ) to communicate. Then, the touch screen device 140 ( Figure 1 ) can be obtained from the electronic pen 110 ( Figure 1 ) receives these pressure signals.

[0113] At operation 604 , in response to receiving a pressure signal indicating a first pressure value at the pen tip, the system enables handwriting recognition.

[0114] At operation 606, the system displays the touch screen 130 ( Figure 1 ) receives a touch input representing handwriting at a first position of the electronic pen 110 ( Figure 1 ) provides some examples of touch input, which can be referred to as pen input.

[0115] At operation 608, the system converts the touch input representing the handwriting into rendered text content corresponding to the handwriting. For example, in some implementations, the system can send the touch input corresponding to the handwriting to a stroke recognition engine to convert the handwriting into computer-renderable text. The stroke recognition engine can be located in a computing system, etc. Additionally or alternatively, the stroke recognition engine can be a cloud service or reside on a remote system. The stroke recognition engine can utilize optical character recognition (OCR) and / or natural language processing (NLP) neural networks to convert the handwriting into computer-renderable text. The system can receive the computer-renderable text from the stroke recognition engine and then display the computer-renderable text as rendered text content on the touch screen.

[0116] Reference below Figure 7A 、 Figure 7B and Figure 7C , these figures show an exemplary implementation of the method of the present invention after handwriting recognition is enabled. Figure 7A 、 Figure 7B and Figure 7C Both include example sentences that include text content that can be presented on a touch screen.

[0117] Figure 7A An exemplary sentence 702 is shown as being presented on touch screen 130: "We bought a pound of pears from the market and took it home." An editing cursor 710 is displayed after the word "pears." Editing cursor 710 may be a text editing cursor that updates its position and appearance based on a user's touch input (e.g., input using electronic pen 110) for selection and / or editing. Exemplary sentence 702 includes a highlighted area 708 between the words "pears" and "from." Figure 7A The electronic pen 110 is also shown near the highlighted area.

[0118] The highlighted area 708 may represent an area where a user may write using the electronic pen 110 , which, in some examples, the system may provide when receiving touch input at a first location on the touch screen 130 .

[0119] Figure 7B Shown with Figure 7A Example sentence 704 is similar to example sentence 702 in FIG. Example sentence 704 is displayed on touch screen 130. However, Figure 7BThe example sentence 704 in FIG. 7 also includes the handwritten phrase "and plums" on the highlighted area 708 between the words "pears" and "from." An editing cursor 710 is displayed after the word "plums," indicating the updated text editing position in response to the addition of the handwritten phrase. Figure 7B Also shown is the electronic pen 110 near the highlighted area 708. As shown, the user can edit a series of text contents on the touch screen 130 using the electronic pen 110.

[0120] Figure 7C An exemplary sentence 706 is shown displayed on the touch screen 130: “We bought a pound of pears and plums from the market and took it home.” The exemplary sentence 706 is rendered entirely in text content, illustrating that the system converts touch edits representing handwriting into rendered text content corresponding to the handwriting.

[0121] Figure 8 800 is a flowchart of an exemplary text editing method according to one embodiment of the present invention. The method 800 may be executed by one or more processors of a computing system (eg, a touch screen device 140 ( Figure 1 ) is executed by one or more processors).

[0122] At operation 802, the system displays a Figure 1 ) receives touch input at a second position.

[0123] At operation 804, the system receives a request to enable voice recognition. In some embodiments, the system may receive a request to enable voice recognition from the electronic pen 110 ( Figure 1 ) receives a request to enable voice recognition. For example, in some embodiments, the request to enable voice recognition can be indicated on the pen tip 160 ( Figure 1) is received by the user of the electronic pen upon detecting a second pressure value at the tip of the electronic pen. In some implementations, the second pressure value is distinguishable from the first pressure value. For example, a user of the electronic pen may apply a first pressure value at the tip of the pen. The first pressure value may indicate a request to enable handwriting recognition. The user of the electronic pen may then apply a second pressure value at the tip of the pen, which is distinguishable from the first pressure value. The second pressure value may indicate a request to enable voice recognition. In some implementations, for example, a lower pressure value at the tip of the pen may indicate a request to enable voice recognition, while a higher pressure value may indicate a request to enable handwriting recognition. In some implementations, "lower" and "higher" pressure values ​​may be distinguished by comparing the pressure measured by the pressure sensor with predefined thresholds. For example, a pressure greater than a first threshold but less than a second threshold may be identified as a "lower" pressure value for enabling voice input, while a pressure greater than the second threshold may be identified as a "higher" pressure value for enabling handwriting input. In another example, a pressure greater than the first threshold but less than the second threshold may be identified as a "higher" pressure value for enabling handwriting input, while a pressure greater than the second threshold may be identified as a "lower" pressure value for enabling voice input.

[0124] In some embodiments where the electronic pen includes a button, the request to enable voice recognition can be received as an input signal indicating a button press on the electronic pen. For example, a user of the electronic pen can apply a button press that can indicate a request to enable voice activation.

[0125] In some embodiments where the electronic pen does not include a button, the request to enable voice recognition may be received as an input signal indicating a "long press" of the electronic pen. A "long press" may be determined by the electronic pen and / or the touchscreen device. For example, a long press threshold representing a duration may be implemented. In some such examples, the request to enable voice recognition may be received when a first pressure value at the pen tip is detected for a duration equal to or greater than the long press threshold.

[0126] In some embodiments, a request to enable voice recognition can be received via a user interface (UI) contextual input button. For example, in some embodiments, a user of an electronic pen can apply a first pressure value at the tip of the pen. The first pressure value can indicate a request to enable voice recognition and / or display a UI. The UI can include one or more buttons and can include a UI contextual voice input button, etc., which can be clicked to enable voice recognition. The UI contextual input button can provide accessibility for users, for example, for users who may have difficulty applying different or greater pressure values ​​on the pen tip.

[0127] At operation 806 , the system enables speech recognition.

[0128] At operation 808 , the system receives a signal representing voice input. The signal may be received by the touch screen device microphone 150 . Additionally or alternatively, the signal may be received by the electronic pen microphone 180 and sent by the electronic pen 110 to the touch screen device 140 .

[0129] At operation 810, the system converts the signal representing the voice input into text content corresponding to the voice input. In some implementations, the text content can be presented at a second position on the touch screen. For example, in some implementations, the system can send a representation of the signal representing the voice input to a speech recognition engine to convert the representation of the signal representing the voice input into computer-renderable text. The speech recognition engine can be located in a computing system or the like. Alternatively or in addition, the speech recognition engine can be a service (e.g., a cloud service) or can reside on a remote system. The speech recognition engine can employ an NLP neural network. The system can receive computer-renderable text from the speech recognition engine and can display the computer-renderable text as text content on the touch screen.

[0130] Reference below Figure 9A 、 Figure 9B and Figure 9C , these figures show an exemplary implementation of the method of the present invention after enabling speech recognition. Figure 9A 、 Figure 9B and Figure 9C Both include example sentences that include text content that can be presented on a touch screen.

[0131] Figure 9A An exemplary sentence 902 is shown as being presented on the touch screen 130: "We bought a pound of pears from the market and took it home." A microphone icon 908 and the electronic pen 110 are displayed near a second location on the touch screen 130, between the words "pear" and "from." In this exemplary illustration, the microphone icon 908 indicates that voice recognition is enabled. The microphone icon 908 may also represent a voice cursor, indicating a location on the touch screen 130 where voice dictation can be presented.

[0132] Figure 9B Shown with Figure 9A Example sentence 904 is similar to example sentence 902 in FIG. However, example sentence 904 includes a highlighted area 708 between the words "pear" and "from." A microphone icon 908 is displayed near a second location on the touch screen, similar to example sentence 902 in FIG. Figure 9A Same as in. Figure 9BAlso included is the phrase "and plums" typed on highlighted area 708, which the system has received via a signal representing voice input.

[0133] Figure 9C An exemplary sentence 706 is shown: “We bought a pound of pears and plums from the market and took it home.” This exemplary sentence is rendered entirely in text content, illustrating that the system converts a signal representing a speech input into rendered text content corresponding to the speech input.

[0134] Figure 10 1 is a flowchart of an exemplary text editing method 1000 according to one embodiment of the present invention. The method 1000 may be executed by one or more processors of a computing system (eg, a touch screen device 140 ( Figure 1 ) is executed by one or more processors).

[0135] At operation 1002, the system displays a Figure 1 ) receives a touch input representing an ellipse at a third position.

[0136] At operation 1004, the system identifies target text content. The target text content may be content presented at a third position on the touch screen. The target text content may represent text content circled by the touch input representing the ellipse.

[0137] Reference below Figure 11A and Figure 11B . Figure 11A and Figure 11B Exemplary sentences 1102 , 1104 are shown that include textual content presented on the touch screen 130 . Figure 11A An exemplary sentence 1102 is listed:

[0138] “We bought a pound of pears from the marcket and took it home,”

[0139] Figure 11B An exemplary sentence 1104 is listed:

[0140] "We bought a pound of pears from the market and took it home."

[0141] Figure 11A and Figure 11B The pen input representing the ellipse 1106 is shown to be displayed around the target text content. Figure 11A In the example of FIG, the pen input representing the ellipse 1106 is displayed around the text content "marcket", while Figure 11B In the example of , pen input representing an ellipse 1106 is displayed around the text content "pears".

[0142] Return to Figure 10 In method 1000 , after operation 1004 , operation 1006 is performed next.

[0143] At operation 1006, the system determines one or more replacement candidates corresponding to the target text content. For example, in some implementations, the target text content may correspond to a misspelled word, such as Figure 11A . In such implementations, one or more replacement candidates may correspond to the correct spelling of the misspelled word. For example, in some implementations, the system may send a representation of the target text content to a spell-checking application. The spell-checking application may be an application local to the system, or the spell-checking application may be a cloud-based application, or may reside on a remote system, etc. The spell-checking application may determine that the target text content may represent a misspelled word and, if so, may provide one or more correctly spelled words as one or more replacement candidates to the system.

[0144] As another example, the target text content may correspond to a word with one or more homophones, e.g. Figure 11B As shown in the example in . In such implementations, the one or more replacement candidates can correspond to one or more homophones of the target text content. For example, in some implementations, the system can send a representation of the target text content to a homophone application. The homophone application can be an application local to the system, or the homophone application can be a cloud-based application, or can reside on a remote system, etc. The homophone application can determine that the target text content can correspond to one or more homophones, and if so, can provide the one or more homophones to the system as one or more replacement candidates.

[0145] In some embodiments, the target text content may correspond to a word that lacks capital letters, a word that contains grammatical errors, a word that has synonyms, and / or punctuation marks, etc. In such an implementation, one or more replacement text candidates may be associated with replacement candidates corresponding to a corrected version of the target text content and / or one or more synonyms of the target text content. In such an implementation, the system may use a corresponding application, such as a grammar application and / or a thesaurus application, to determine the one or more replacement candidates.

[0146] At operation 1008 , the system displays one or more replacement text candidates as selectable options to replace the target text content.

[0147] Reference below Figure 12A and Figure 12B . Figure 12A and Figure 12B Exemplary sentences 1102, 1104 are shown, which include text content presented on the touch screen 130. Figure 11A and Figure 11B As shown, the pen input of the ellipse 1106 has been enclosed Figure 12A and Figure 12B The text content in is displayed. Figure 12A and Figure 12B One or more replacement text candidates 1202 , 1204 , and 1206 are displayed as selectable options for replacing the target text content. Figure 12A and Figure 12B Also shown is an electronic pen 110 adjacent to the rendered text content.

[0148] exist Figure 12A In the example of , the replacement text candidate "market" 1106 is displayed as a selectable option near the circled misspelled word "marcket". Figure 12B In the example of , the replacement text candidates "pairs" 1204 and "pares" 1206 are displayed as selectable options near the circled word "pears". In some implementations, Figure 12A and Figure 12B In the examples shown, the user can select one of the displayed replacement text candidates 1202, 1204, and 1206. The user can use the electronic pen 110 or the like to make the selection.

[0149] In some examples where the target text content corresponds to punctuation marks, replacement text candidates representing alternative punctuation marks may be displayed as selectable options.

[0150] Return to Figure 10 In method 1000, after operation 1008, operation 1010 is performed next.

[0151] At operation 1010, in response to selecting one of the one or more replacement text candidates, the system replaces the target text content with the selected one of the one or more replacement text candidates. Figure 12A , the system can replace the misspelled word "marcket" with the replacement text candidate "market" 1202. For another example, combined with Figure 12B, the system may replace the word “pears” with the text candidate “pairs” 1204 or “pares” 1206 , depending on the selection.

[0152] Figure 13 1 is a flowchart of an exemplary text editing method 1300 according to one embodiment of the present invention. The method 1300 may be executed by one or more processors of a computing system (e.g., a touch screen device 140 ( Figure 1 ) is executed by one or more processors).

[0153] At operation 1302, the system displays the touch screen 130 ( Figure 1 ) receives a touch input representing an ellipse at a fourth position of the .

[0154] At operation 1304, the system identifies target text content. The target text content may be content presented at a fourth position on the touch screen. The target text content may represent text content circled by the touch input representing the ellipse.

[0155] Reference below Figure 14A . Figure 14A An exemplary sentence 1102 is shown that includes textual content presented on the touch screen 130 . Figure 14A The exemplary sentence 1102 in cites:

[0156] "We bought a pound of pears from the marcket and took it home."

[0157] Pen input representing an ellipse 1106 has been displayed around the misspelled word "marcket."

[0158] Return to Figure 13 , after operation 1304, operation 1306 is performed next.

[0159] At operation 1306 , the system receives input representing replacement content.

[0160] In some embodiments, the received input may be a touch input, for example, the received input may be a pen input representing handwriting.

[0161] Reference below Figure 14B , the figure shows that with Figure 14A Example sentence 1402 is similar to example sentence 1102 in . Figure 14B The example sentence 1402 in includes rendering the touch input representing handwriting at the highlighted area 708. Figure 14BIn the example of , the presentation includes the word “market.” An edit cursor 710 is displayed after the word “market,” indicating a location for subsequent text editing. Figure 14B The electronic pen 110 is also shown near the highlighted area 708 .

[0162] Return to Figure 13 In operation 1306, in some embodiments, the received input representing the replacement content may be a voice input representing a spoken word. The voice input may be received by the touch screen device microphone 150 and / or the electronic pen microphone 180 and may be sent by the electronic pen 110 to the touch screen device 140. In some embodiments, the voice input and the touch input may be performed simultaneously, such as in combination with Figure 24A 、 Figure 24B 、 Figure 24C and Figure 24D Details.

[0163] Reference below Figure 14C , which shows an example sentence 1404. The example sentence 1402 recites: "Webought a pound of pears from the and took it home," and includes a highlighted region 708 between the words "the" and "market." Figure 14C Also included is a microphone icon 908, indicating that the system is receiving voice input. As described above, the microphone icon 908 can also represent a voice cursor, indicating the location on the touch screen 130 where the voice dictation can be presented. Figure 14C In the example of , a user may speak the word "market" into a microphone associated with the system (eg, touch screen device microphone 150 and / or electronic pen microphone 180).

[0164] In some examples, the highlighted area 708 indicates an area where the user can write using the electronic pen 110. Figure 14B As shown. In some examples, highlighted area 708 represents an area where rendered text content representing voice input can be displayed. In some examples, highlighted area 708 may appear in response to touch input received at a location on touch screen 130 corresponding to the indicated insertion point. In some examples, highlighted area 708 may appear in response to an indication to use a voice input mode (e.g., by pressing button 170 on electronic pen 110) or in response to an indication to use handwriting. As described above, a user's selection of a particular input mode (e.g., handwriting input or voice input) can be indicated by applying different pressure values ​​at pen tip 160.

[0165] The system may display one or more replacement text candidates 1202, 1204, 1206 as selectable options for replacing the target text content (e.g. Figure 12A and Figure 12B In some embodiments (shown in FIG. 14 ), the user may forgo selecting one or more replacement text candidates 1202, 1204, 1206. In some such embodiments, a highlighted region 708 may appear at a location on the touch screen 130 in response to an indication of using handwriting input and / or voice input. As described above, the indication of using handwriting input and / or voice input may be received as a pressure signal indicating a pressure value detected at the pen tip 160 and / or as an input signal indicating a button press at the button 170 ( FIG. 14 ) of the electronic pen 110, etc.

[0166] Return to Figure 13 In method 1300, after operation 1306, operation 1308 is performed next.

[0167] At operation 1308 , the system replaces the target content with the rendered text content corresponding to the replacement content.

[0168] In some examples where the received input is a touch input representing handwriting, the system can send a representation of the touch input representing the handwriting to a stroke recognition engine to convert the handwriting into computer-renderable text content. The stroke recognition engine can be located in a computing system, etc. Additionally or alternatively, the stroke recognition engine can be a cloud service or reside on a remote system. The stroke recognition engine can utilize OCR and / or NLP neural networks to convert the handwriting into computer-renderable text content. In some implementations, the system can receive the computer-renderable text content from the stroke recognition engine and then display the computer-renderable text content as text content on the touch screen.

[0169] In some examples where the received input is a voice input representing a spoken word, the system can send the representation of the voice input to a speech recognition engine to convert the representation of the voice input into computer-renderable text content. In some embodiments, the speech recognition engine can be an engine local to the computing system. In other embodiments, the speech recognition engine can be a service, such as a cloud service, or it can reside on a remote system. The speech recognition engine can employ an NLP neural network. In some implementations, the system can receive computer-renderable text content from the service and can display the computer-renderable text content as text content on the touch screen.

[0170] Reference below Figure 14D , which shows an exemplary sentence 1406 displayed on the touch screen 130. The exemplary sentence 1402 lists:

[0171] "We bought a pound of pears from the market and took it home."

[0172] Example sentence 1406 shows that when the system receives voice input (such as Figure 14C ) or after receiving a touch input representing handwriting (as shown in Figure 14B The text content that appears after the .

[0173] Figure 15 1 is a flowchart of an exemplary text editing method 1500 according to one embodiment of the present invention. The method 1500 may be executed by one or more processors of a computing system (e.g., a touch screen device 140 ( Figure 1 ) is executed by one or more processors).

[0174] At operation 1502, the system displays a Figure 1 ) receives a touch input representing a deletion line at a fifth position.

[0175] At operation 1504, the system identifies target text content. The target text content may be content presented at the fifth position of the touch screen. The target text content may represent text content presented below the touch input representing the strikethrough.

[0176] At operation 1506 , the system displays one or more content formatting options adjacent to the target text content.

[0177] Reference below Figure 16A , which illustrates an exemplary sentence 1602 that includes textual content that may be presented on the touch screen 130. The exemplary sentence 1602 lists:

[0178] "We bought a pound of pears from the market and took it home."

[0179] In this example, the target text content is the word "market" which is displayed with a strikethrough 1606. A selection 1608 of a plurality of content formatting options is displayed adjacent to the target text content. Figure 16A In the example of , the display selections 1608 of the content format options include "copy", "paste", (indicates bold, italic, and underline options), Search, and Translate.

[0180] Reference again Figure 15 At operation 1508, the system receives a selection of one or more content options. For example, the system may receive a selection of a content format option. (Indicates options highlighted in bold).

[0181] At operation 1510 , the system modifies the target text content based on the received selection of one or more content options.

[0182] For example, the following reference Figure 16B , which shows an exemplary sentence 1604 displayed on the touch screen 130. Figure 16B As shown, the selection in the received content format option When the system can highlight the word "market" in bold, the generated example sentence 1604 (where the word "market" is highlighted in bold) lists the following:

[0183] “We bought a pound of pears from the and took it home.”

[0184] Figure 17 1700 is a flowchart of an exemplary text editing method 1700 according to one embodiment of the present invention. The method 1700 may be executed by one or more processors of a computing system (e.g., a touch screen device 140 ( Figure 1 ) is executed by one or more processors).

[0185] At operation 1702, the system displays the touch screen 130 ( Figure 1 ) receives a touch input representing a deletion line at the seventh position.

[0186] At operation 1704, the system identifies target text content. The target text content may be the content presented at the seventh position of the touch screen. The target text content may represent the text content presented below the touch input representing the strikethrough.

[0187] At operation 1706, the system receives an instruction to enable voice recognition via the electronic pen. In some embodiments, the instruction may be received via a pressure signal representing a pressure value at the pen tip. In some embodiments, the instruction may be received as an input signal indicating a button press at the electronic pen.

[0188] In some embodiments where the electronic pen does not include a button, the instruction to enable voice recognition may be received as an input signal instructing the electronic pen to "long press." A "long press" may be determined by the electronic pen and / or the touchscreen device. For example, a long press threshold representing a duration may be implemented. In some such examples, the instruction to enable voice recognition may be received when a first pressure value at the pen tip is detected for a duration equal to or greater than the long press threshold.

[0189] In some examples, the indication of using a pen input can be received as a pressure signal indicating a first pressure value at the pen tip (e.g., at a certain location on a touch screen). In some examples, the indication of using a voice input can be received as an input signal indicating a button press on a button or as a pressure signal indicating the detection of a second pressure value at the pen tip (e.g., at a certain location on a touch screen). In some embodiments, voice input and touch input can be performed simultaneously, such as in combination with Figure 24A 、 Figure 24B 、 Figure 24C and Figure 24D Details.

[0190] At operation 1708 , the system enables speech recognition.

[0191] At operation 1710, the system receives voice input. The voice input can be received via a microphone in communication with a processor. In some embodiments, the microphone can be a component of a computer system. In some embodiments, the microphone can be a component of an electronic pen.

[0192] Reference below Figure 18A , which illustrates an exemplary sentence 1602 that includes textual content that may be presented on the touch screen 130. The exemplary sentence 1602 lists:

[0193] “We bought a pound of pears from the and took it home.”

[0194] In this example, the target text content is the word "market" which is displayed with a strikethrough 1606. A selection 1608 of a plurality of content formatting options is displayed adjacent to the target text content. Figure 18A In the example of FIG, the selection 1608 of the plurality of content format options includes “copy”, “paste”, (indicating options highlighted in bold, italics, and underline), "Search," and "Translate." A microphone icon 908 is displayed near the word "market," indicating that voice recognition has been enabled. As described above, the microphone icon 908 may also represent a voice cursor, indicating a location on the touch screen 130 where voice dictation may be rendered.

[0195] Return to Figure 17At operation 1712, the system recognizes the voice input as a voice command corresponding to a content formatting option. In some implementations, examples of content formatting options may include "bold," "highlight," "italic," "underline," or "delete," among others. In some examples, the voice command may correspond to one of the displayed content formatting options for selection 1608 of the plurality of content formatting options. In some examples, the voice command may not correspond to one of the displayed content formatting options for selection 1608 of the plurality of content formatting options.

[0196] For example, in some implementations, the system can send a signal representing the voice input to a voice recognition engine, which processes the voice input using a voice recognition algorithm or the like. The voice recognition engine can employ an NLP neural network. In some examples, the voice recognition engine can be an engine local to the computing system. In other examples, the voice recognition engine can be a service, such as a cloud service, or it can reside on a remote system. In some examples, the voice recognition engine can determine that the voice input corresponds to a voice command. In some examples, the voice recognition engine or the system can then compare the voice command with a set of predefined voice commands to determine the operation requested by the voice command. The system can then perform the operation.

[0197] At operation 1714, the system modifies the target text content according to the content formatting options. For example, in response to recognizing the voice command "bold", the system may highlight the target text content in bold, such as Figure 18B As shown in the exemplary sentence 1604. Figure 18B The following exemplary sentence 1604 is shown displayed on touch screen 130:

[0198] “We bought a pound of pears from the and took it home.”

[0199] The word "market" is highlighted in bold in the sentence. It should be noted that in some examples, multiple different voice commands can be used to modify the same target text content without the user having to repeatedly select the same target text content. For example, after formatting the word "market" according to the voice instruction "bold", the word "market" may still be the target text content (e.g., indicated by keeping it highlighted and / or keeping the displayed content formatting options). Other voice input and / or touch input can also be provided to further modify the same target text content (e.g., add additional formatting). The target text content may remain unchanged until the user provides other input (e.g., touch input) at a different location.

[0200] Figure 19 1900 is a flowchart of an exemplary text editing method 1900 according to one embodiment of the present invention. The method 1900 may be executed by one or more processors of a computing system (e.g., a touch screen device 140 ( Figure 1 ) is executed by one or more processors).

[0201] At operation 1902, the system displays the touch screen 130 ( Figure 1 ) receives a touch input representing a deletion line at an eighth position.

[0202] At operation 1904, the system identifies target text content. The target text content may be content presented at the eighth position of the touch screen. The target text content may represent text content presented below the touch input representing the strikethrough.

[0203] At operation 1906, the system receives an instruction to enable voice recognition via the electronic pen. In some embodiments, the instruction may be received via a pressure signal representing a pressure value at the pen tip. In some embodiments, the instruction may be received as an input signal indicating a button press at the electronic pen.

[0204] The system may display one or more content format options (such as Figure 16A and Figure 18A In some embodiments (as shown), the user may forgo selecting one or more content format options. Instead, the system may receive an indication of using handwriting input and / or voice input, and thus, a highlighted area may appear at a certain location on the touch screen. As described above, the indication of using handwriting input and / or voice input may include receiving a pressure signal indicating a pressure value at a pen tip and / or receiving an input signal indicating a button press on a button, etc.

[0205] At operation 1908, the system enables speech recognition.

[0206] At operation 1910, the system receives voice input. The voice input can be received via a microphone in communication with a processor. In some embodiments, the microphone can be a component of a computer system. In some embodiments, the microphone can be a component of an electronic pen.

[0207] Reference below Figure 20A , which illustrates an exemplary sentence 1602 that includes textual content that may be presented on the touch screen 130. The exemplary sentence 1602 lists:

[0208] “We bought a pound of pears from the and took it home.”

[0209] In this example, the target text content is the word "market," which is displayed with a strikethrough 1606. A microphone icon 908 is displayed near the word "market," indicating that voice recognition has been enabled. As described above, the microphone icon 908 can also represent a voice cursor, indicating a location on the touch screen 130 where voice dictation can be presented.

[0210] Return to Figure 19 At operation 1912, the system recognizes the speech input as a speech dictation.

[0211] For example, in some implementations, the system can send a signal representing the voice input to a speech recognition engine, which processes the voice input using a speech recognition algorithm or the like. The speech recognition engine can employ an NLP neural network. In some examples, the speech recognition engine can be an engine local to the computing system. In other examples, the speech recognition engine can be a service, such as a cloud service, or it can reside on a remote system. In some examples, the speech recognition engine can determine that the voice input corresponds to a voice dictation, and can then convert a representation of the signal representing the voice input into computer-renderable text content. In some examples, the system can receive computer-renderable text content from the speech recognition engine, and can display the computer-renderable text content as text content on a touch screen.

[0212] At operation 1914 , the system replaces the target text content with content corresponding to the voice dictation.

[0213] Reference below Figure 20B , which illustrates an exemplary sentence 2002 that includes textual content that may be presented on the touch screen 130. The exemplary sentence 2002 lists:

[0214] "We bought a pound of pears from the grocery store and took it home."

[0215] In this example, highlighted area 708 is displayed below the phrase "grocery store." A microphone icon 908 is displayed near the phrase "grocery store," indicating that voice recognition has been enabled. As described above, microphone icon 908 may also represent a voice cursor, indicating a location on touch screen 130 where the voice dictation may be rendered. The phrase "grocery store" represents the rendered content corresponding to the voice dictation, which replaces the target text content "market."

[0216] After the speech input stops, the sentence can be Figure 20C appears as shown. Figure 20C The following exemplary sentences 2002 are listed:

[0217] "We bought a pound of pears from the market and took it home."

[0218] exist Figure 20C In the example of , exemplary sentence 2002 is displayed on touch screen 130 . Figure 20C The microphone icon 908 and the highlighted area 708 are not shown.

[0219] In some embodiments, the system may have multiple ways to determine that voice input has ceased. For example, a voice input timeout threshold may be implemented. In such an example, once voice input begins, the system may determine that voice input has ceased after receiving a period of voice silence that is equal to or greater than the voice input timeout threshold. In some embodiments, the voice input timeout threshold is user-adjustable.

[0220] Additionally or alternatively, the user may instruct the voice input to cease by performing certain actions. For example, the user may use a relevant UI element displayed on the touch screen 130 to "turn off" the voice input. In another example, the user may use an electronic pen to perform an action, such as applying a certain amount of pressure to the touch screen with the electronic pen, double-clicking with the electronic pen, briefly pressing an electronic pen button, and / or pressing a button on a relevant physical keyboard.

[0221] Figure 21 2 is a flowchart of an exemplary text editing method 2100 according to one embodiment of the present invention. The method 2100 may be executed by one or more processors of a computing system (e.g., a touch screen device 140 ( Figure 1 ) is executed by one or more processors).

[0222] At operation 2102, the system displays the touch screen 130 ( Figure 1 ) receives a touch input representing a deletion line at an eighth position.

[0223] At operation 2104, the system identifies target text content. The target text content may be the content presented at the eighth position of the touch screen. The target text content may represent the text content presented below the touch input representing the strikethrough.

[0224] Reference below Figure 22A , which illustrates an exemplary sentence 1602 that includes textual content that may be presented on the touch screen 130. The exemplary sentence 1602 lists:

[0225] “We bought a pound of pears from the and took it home.”

[0226] In this example, the target text content is the word "market," which is displayed with a strikethrough.

[0227] Return to Figure 21 At operation 2106, the system receives an instruction to enable handwriting recognition via the electronic pen. In some embodiments, the instruction may be received via a pressure signal representing a pressure value at the pen tip. The pressure value received at operation 2106 may be distinguished from the pressure value received at operation 1906.

[0228] At operation 2108, the system enables handwriting recognition.

[0229] At operation 2110, the system removes the target text content from the eighth position of the touch screen. Optionally, a highlighted area can be displayed at the eighth position, where a touch input can be received.

[0230] The system may display one or more content format options (such as Figure 16A and Figure 18A In some embodiments (as shown), the user may forgo selecting one or more content format options. Instead, the system may receive an indication of using handwriting input and / or voice input, and thus, a highlighted area may appear at a certain location on the touch screen. As described above, the indication of using handwriting input and / or voice input may include receiving a pressure signal indicating a pressure value at a pen tip and / or receiving an input signal indicating a button press on a button, etc.

[0231] At operation 2112, the system receives touch input representing handwriting at an eighth location on the touch screen.

[0232] Reference below Figure 22B , which illustrates an exemplary sentence 2202 that includes textual content that may be presented on the touch screen 130. The exemplary sentence 2202 lists:

[0233] "We bought a pound of pears from the grocery store and took it home."

[0234] In this example, the phrase "grocery store" is displayed in handwriting on the highlighted area 708 between the words "the" and "and." An editing cursor 710 is displayed after the phrase "grocery store," indicating a location for subsequent text editing. An electronic pen 110 with a pen tip 160 is displayed near the phrase "grocery store."

[0235] Return to Figure 21 At operation 2114 , the system converts the touch input representing the handwriting into rendered text content corresponding to the handwriting.

[0236] For example, in some implementations, the system can send a representation of a touch input representing handwriting to a stroke recognition engine to convert the handwriting into computer-renderable text. The stroke recognition engine can be located in a computing system, etc. Additionally or alternatively, the stroke recognition engine can be a cloud service or reside on a remote system. The stroke recognition engine can utilize OCR and / or NLP neural networks to convert the handwriting into computer-renderable text. The system can receive the computer-renderable text from the stroke recognition engine and then display the computer-renderable text as text content on the touch screen.

[0237] Reference below Figure 22C , which shows an exemplary sentence 2002 on the touch screen 130. The exemplary sentence 2002 lists:

[0238] "We bought a pound of pears from the grocery store and took it home."

[0239] Figure 22C An example of text content presented after converting a touch input representing handwriting into presented text content corresponding to the handwriting is shown.

[0240] Figure 23 2300 is a flowchart of an exemplary text editing method 2300 according to one embodiment of the present invention. The method 2300 may be executed by one or more processors of a computing system (e.g., a touch screen device 140 ( Figure 1 In exemplary method 2300, before touch input is received at a certain location on the touch screen, voice dictation can be enabled by the electronic pen.

[0241] At operation 2302, the system receives an input signal indicating a button press on an electronic pen in communication with a processor. In some implementations, voice dictation can be enabled through other suitable input devices such as a keyboard, a mouse, or a headset with input buttons, for example, by button presses on a keyboard, a mouse, or a headset.

[0242] At operation 2304 , in response to receiving an input signal indicating a button press on the electronic pen, the system enables voice recognition.

[0243] At operation 2306, the system receives a touch input at a ninth position of the touch screen. In some embodiments, the ninth position may correspond to a position before a series of presented content.

[0244] At operation 2308, the system receives voice input. The voice input can be received via a microphone in communication with the processor. In some embodiments, the microphone can be a component of the computer system. In some embodiments, the microphone can be a component of the electronic pen.

[0245] At operation 2310 , the system presents text content corresponding to the voice dictation at a ninth position on the touch screen.

[0246] As mentioned above, this enables voice dictation using the electronic pen without having to use the touch screen beforehand.

[0247] Reference below Figure 24A 、 Figure 24B 、 Figure 24C and Figure 24D , these figures relate to specific embodiments of the present application. Figure 24A 、 Figure 24B 、 Figure 24C and Figure 24D An exemplary text stream 2402 is shown displayed on touch screen 130 according to one embodiment of the present invention.

[0248] exist Figure 24A 、 Figure 24B 、 Figure 24C and Figure 24D In each of the figures in FIG, exemplary text flow 2402 lists:

[0249] "To do so, I will give you a complete account of the system, and expound the actual teachings of".

[0250] exist Figure 24A 、 Figure 24B 、 Figure 24C and Figure 24DIn each of the figures in FIG, an exemplary text flow 2402 includes a first phrase 2404, which includes a first portion 2406 and a second portion 2408. The first phrase 2404 is identified with a dark background and lists:

[0251] "I will give you a complete account of the system, and expound the actual teachings of".

[0252] The microphone icon 908 is displayed near the first phrase 2404, which, together with the dark background, indicates that the first phrase is the rendered text content corresponding to the voice input. As described above, the microphone icon 908 can also represent a voice cursor, indicating the location on the touch screen 130 where the voice dictation can be rendered.

[0253] A first portion 2406 of the first phrase 2404 is identified using the label "Confirmation Text," and a second portion 2408 of the first phrase 2404 is identified using the label "Assumption." The first portion 2406 includes the following words:

[0254] "I will give you a complete account of the system,and".

[0255] The second portion 2408 includes the remaining words of the first phrase 2404, namely:

[0256] "expound the actual teachings of".

[0257] The second portion 2408 is also highlighted using a light font.

[0258] In some examples, upon receiving voice input, the system may send a signal representing the voice input to a speech recognition engine, which processes the voice input using a speech recognition algorithm or the like. The speech recognition engine may employ a natural language processing (NLP) neural network. In some examples, the speech recognition engine may be a local engine on the computing system. In other examples, the speech recognition engine may be a service, such as a cloud service, or may reside on a remote system. In some examples, the speech recognition engine may determine that the voice input corresponds to a spoken dictation and may subsequently convert a representation of the signal representing the voice input into computer-renderable text content. The system may then receive the computer-renderable text from the speech recognition engine and display the computer-renderable text as text content on the touch screen. The computer-renderable text may be received in one of multiple states. For example, the system may receive the computer-renderable text in a confirmed state, indicating that the service has determined that the relevant computer-readable text is accurate. Additionally or alternatively, the system may receive the computer-readable text in a hypothetical state, indicating that the computer-renderable text is still being processed by the service.

[0259] As mentioned above, in Figure 24A 、 Figure 24B 、 Figure 24C and Figure 24D In some examples of specific embodiments shown, text in a "confirmed" state can indicate text that has been confirmed by a speech recognition engine. In some examples, text in a confirmed state can remain active for a preset time to receive pen gesture input for punctuation, writing, and editing, such as deleting or adding new lines and / or adding and removing regions.

[0260] As described above, in some examples, text in a "hypothetical" state can represent text that has not yet been confirmed by the speech recognition engine and can therefore be dynamic because the text may still be processed by the service. In some examples, text in a "hypothetical" state may disable pen functionality.

[0261] Figure 24B Shown Figure 24A Other elements other than the elements in . Specifically, Figure 24B The electronic pen 110 is shown near the word "complete" in the first portion 2406. A strikethrough 1606 is displayed on the word "complete," indicating that the text in the "confirmed" state has received a pen gesture edit while continuing to receive voice input. The system can recognize the word "complete" as the target text content.

[0262] Figure 24C Shown Figure 24A Other features other than the features in . Specifically, Figure 24CA selection 1608 among a plurality of content formatting options adjacent to the target text content is shown. Figure 24C In the example, selection 1608 includes "copy", "paste", (indicates bold, italic, and underline options), Search, and Translate.

[0263] exist Figure 24C In the example shown in Figure 2, the voice command provided is "underline".

[0264] Figure 24D Shown Figure 24C Other features than those shown. Specifically, Figure 24D Shows the word " complete ”, indicating that the system has modified the target text content according to the received selections in one or more content formatting options.

[0265] It should be noted that text in the "confirmed" state can be edited in other ways, including any of the above-mentioned pen editing functions.

[0266] As mentioned above 24A to 24D As described above, in some embodiments, voice dictation and handwriting recognition can be enabled at the same time. In some specific embodiments, the "hypothetical" state may only last for a few seconds. In some examples, all presented text can be in a "confirmed" state, i.e., the "hypothetical" text is not displayed. In some such examples, the presented text can be regarded as ordinary editable text. Therefore, the user can use an electronic pen and / or keyboard and / or mouse (or other suitable input device) to edit the presented text without ending the current dictation (e.g., without switching between voice input and touch input modes).

[0267] Reference below Figure 25 , which shows an exemplary implementation of simultaneous speech recognition and text editing input according to some examples. Figure 25 First, second, third, and fourth examples 2510 , 2520 , 2530 , 2540 of rendered text displayed on the touch screen 130 are shown, each including a microphone icon 908 and / or an edit cursor 710 .

[0268] The first example 2510 displays the following text:

[0269] "The quick brown fox jumps over the lazy dog. Frequently this is the sentence used to test out new typewriters,presumably because it includes every letter of the alphabet."

[0270] An editing cursor 710 is displayed over the presentation of the word "typewriters," indicating a location on the touch screen 130 where subsequent text editing can occur. The editing cursor 710 can be a text editing cursor that updates its position and appearance based on a user's touch input selections and / or text editing operations. For example, the position of the editing cursor 710 can reflect the location of a detected touch input (e.g., touching the touch screen 130 using an electronic pen 110, a finger, etc.), or can reflect the location of the most recent text input using a keyboard or electronic pen 110, etc. A microphone icon 908 is displayed after the word "alphabet," indicating a location on the touch screen 130 where subsequent voice dictation can occur. The microphone icon 908 can represent the location of the voice cursor, and therefore can reflect the location of the most recent input of voice transcription data, and can appear when voice dictation is enabled.

[0271] The second example 2520 displays the same text as the first example 2510. The second example 2520 shows that the word "typewriters" is highlighted, indicating that the user has selected the word (e.g., using the electronic pen 110, using the example described above) for text editing. It should be noted that the microphone icon 908 remains in the position after the word "alphabet", as in the first example 2510. Therefore, touch input (e.g., selecting text content to be edited) does not affect the location in the text where voice input can be provided. In this way, the user can continue to provide voice input at the location where the voice transcription was last entered, while also editing the text content elsewhere in the document (using a different input method), which is very convenient.

[0272] The third example 2530 displays the following text:

[0273] "The quick brown fox jumps over the lazy dog. Frequently this is the sentence used to test out new keyboards,presumably because it includes everyletter of the alphabet."

[0274] As described above, the text of the third example 2530 differs from the text of the first and second examples 2510 and 2520 in that the word "typewriters" is replaced with the word "keyboards," which is the result of a user performing a text edit (e.g., the user provides touch input to select a synonym for the target word "typewriters," the user types a replacement word using a keyboard, or the user provides handwriting input to replace the target word, e.g., using the various examples described previously). The editing cursor 710 is displayed after the word "keyboards," indicating a location on the touch screen 130 where subsequent text editing can be performed. Notably, the microphone icon 908 is displayed in the same location after the word "alphabet" as in the first example 2510.

[0275] The fourth example 2540 displays the following text:

[0276] "The quick brown fox jumps over the lazy dog. Frequently this is the sentence used to test out new keyboards,presumably because it includes everyletter of the alphabet. This is known as a pangram."

[0277] As described above, the text of fourth example 2540 differs from the text of third example 2530 in that an additional sentence, "This is known as a pangram," has been added to the end of the text. In this example, the additional sentence is the result of transcribing the voice input. The editing cursor 710 remains at its previous position, namely after the word "keyboards," as in third example 2530. However, the microphone icon 908 is now displayed after the word "pangram," reflecting the last position of the text presented by voice dictation.

[0278] like Figure 25As shown in the various examples of the present invention, the present invention provides various examples of how different input modes can be used to simultaneously input and / or modify text content at different locations in the text. It is worth noting that input using a first input mode (e.g., touch input) can be located at a first position indicated by a first cursor (e.g., edit cursor 710), while input using a second input mode (e.g., voice input) can be located at a different second position indicated by a second cursor (e.g., microphone icon 908). The corresponding positions of the first and second cursors can change according to the input of the corresponding first and second modes. In other words, the position of the first cursor may not be affected by the input using the second mode, and conversely, the position of the second cursor may not be affected by the input using the first mode. This helps to improve the overall efficiency of the system because it can process two parallel input streams using two different modes at the same time without switching between the two input modes. For example, a user can provide continuous voice input, enter new text at one location, and simultaneously manually edit previously presented text using touch input (e.g., using electronic pen 110) at a different location. It should be noted that although touch input is described as an example of an input mode for text editing, in some examples, other input modes (eg, keyboard input, mouse input, etc.) may also be used for text editing simultaneously with voice input.

[0279] Using speech recognition and text editing simultaneously can have a variety of uses. For example, in a meeting and / or lecture, a first party may be speaking, and the speech may be transcribed using speech recognition. At the same time, a second party can use touch input to edit the transcribed text (e.g., formatting, marginalia, incorrectly recognized words) without interrupting the speaker.

[0280] For example, users can use touch input to edit the transcription of their own voice dictation (e.g., formatting, annotations, incorrectly recognized words), and regardless of the text editing position, the voice transcription position can reflect the latest position of the voice dictation, which is convenient to use.

[0281] Reference below Figure 26A and Figure 26B , these figures show a flowchart 2600 of one embodiment of the present invention according to some examples. It should be understood that the flowchart 2600 shows an exemplary combination of various examples disclosed herein. That is, the present invention is described in detail in FIG. Figure 6 、 Figure 8 、 Figure 10 、 Figure 13 、 Figure 15 、 Figure 17 、 Figure 19 、 Figure 21 and Figure 23The examples disclosed in the flowcharts and the like can be performed in combination. According to the input sequence provided by the user, the system can determine which methods shown in the above flowcharts can be executed.

[0282] At operation 2602, a user may enable a text box using a text editing cursor.

[0283] Operations 2604 and 2606 represent different methods by which a user can enable voice input / voice recognition. As shown in flow chart 2600, different methods can produce different results.

[0284] For example, at operation 2604 , the user may enable voice input / voice recognition using hardware, such as using pressure at the tip of a pen, a button on an electronic pen, or other buttons on a different input device, etc. Thus, operation 2608 may then be performed.

[0285] At operation 2608 , transcription of the voice dictation may begin at the location of the edit cursor and the voice cursor (microphone icon) may be enabled. After operation 2608 , transcription of the voice dictation may continue, with operation 2614 being performed.

[0286] At operation 2614, dictation continues and the voice cursor (microphone icon) position may change depending on the location of the transcribed text.

[0287] As another example, at operation 2606, the user can enable voice input / voice recognition using a contextual voice button in the UI or by mouse / touch selection. The contextual voice button can be enabled using a pen gesture or by mouse / touch selection (e.g., drawing an ellipse on the touch screen, applying pressure to a certain location on the touch screen, or drawing a strikethrough at a certain location on the touch screen). Accordingly, the user can proceed to operation 2610.

[0288] At operation 2610 , a voice cursor (microphone icon) may be enabled at the gesture-pointed location, and transcription of the voice dictation may begin at the gesture-pointed location, regardless of the location of the editing cursor.

[0289] After operation 2610, operation 2612 is performed next.

[0290] At operation 2612, the system can present the text as "hypothetical" text on the touch screen. As described above, the speech engine and / or NLP can still process the "hypothetical" text based on the context.

[0291] At operation 2614, the system may update the voice cursor (microphone icon) position based on the location of the transcribed text.

[0292] As described above, when the text changes from the "hypothetical" state to the "confirmed" state, operation 2616 may be continued.

[0293] At operation 2616, the speech engine and / or NLP can confirm the text. After confirming the text, the user can choose to perform operation 2618 and continue dictating, or continue to perform operation 2620 and edit.

[0294] At operation 2620, the user can edit the content. For example, the user can edit the content by moving the editing cursor using various input devices and then performing normal text editing. For example, as shown, the user can edit the content using pen gestures, handwriting, a mouse and / or keyboard, and / or touching a touch screen. Depending on the actions performed by the user, different results may occur.

[0295] For example, an edit made by a user may result in a continue execution operation 2624 .

[0296] At operation 2624, voice input may be enabled. The user may enable voice input by selecting or drawing an ellipse via a contextual voice button on the UI, a pen gesture, or a mouse / touch gesture. After operation 2624, operation 2610 may continue as described above.

[0297] As another example, the edits made by the user at operation 2620 may result in the execution of operation 2622 .

[0298] At operation 2622, the text may be updated. The updated content may be based on user input.

[0299] After operations 2622 and 2624, operation 2626 may be performed.

[0300] At operation 2626, voice dictation can be disabled. For example, the user can disable dictation by clicking the voice cursor (microphone icon), applying pressure to the pen tip, pressing a button on the electronic pen, or pressing another button. As another example, dictation may be disabled due to a timeout. In some examples, the timeout can be set and / or adjusted by the user. The timeout can specify a period of time from the time voice dictation is enabled, e.g., after which voice dictation is automatically disabled. Alternatively, the timeout can specify a period of silence for voice input, e.g., after which voice dictation is automatically disabled.

[0301] After operation 2626, operation 2628 is performed next.

[0302] At operation 2628, the voice cursor is canceled. The voice cursor may be automatically canceled due to deactivation of voice dictation.

[0303] Reference below Figure 27, which shows a flowchart 2700 of one embodiment of the present invention according to some examples. It should be understood that flowchart 2700 can be implemented by a combination of the above examples. For example, based on the input sequence provided by the user, the system can determine which methods shown in the above flowcharts can be executed. Figure 27 The method can be performed by a computer system (e.g., Figure 2 The computing device 200, which may be the touch screen device 140, is executed when executing a text editing application.

[0304] At operation 2702, a gesture or a preset shape (which can be detected as a touch input, such as a touch input by the electronic pen 110) is detected to select text content. The gesture or shape can be detected at the location where the selected text content is displayed on the touch screen. For example, as described above, a pen gesture (e.g., a strikethrough gesture) or a preset shape (e.g., an ellipse) can be detected as input for selecting text content in a text editing application. Other gestures (e.g., a vertical line) or shapes (e.g., a rectangle) can be used.

[0305] At operation 2704 , in response to detecting a gesture or shape that selects text content, a voice input channel is initiated. This may include, for example, enabling a microphone of the computer system, such as microphone 150 of touch screen device 140 or microphone 180 of electronic pen 110 .

[0306] At operation 2706, in response to detecting a gesture or shape for selecting text content, one or more alternative word options (also referred to as recommended word options) are presented. It should be understood that operations 2704 and 2706 can be performed in any order and can be performed in parallel. The one or more alternative word options may include one or more alternative words, one or more alternative phrases, one or more alternative sentences, one or more alternative paragraphs and / or a combination thereof, which may depend on the length of the selected text content (for example, if the selected text content is a word, the one or more alternative word options may be one or more alternative words; if the selected text content is a sentence, the one or more alternative word options may be one or more alternative sentences). One or more alternative word options may be displayed.

[0307] Presenting the one or more alternative word options may include presenting the one or more alternative word options at a location on the touch screen relative to the location of the detected gesture or shape, etc. For example, the one or more alternative word options may be presented near (e.g., above, next to, etc.) the selected text content, as described above. A target option may be selected from the one or more alternative word options by, for example, using touch input on the target option.

[0308] The one or more alternative word options may include one or more options that are similar to the selected text content (e.g., similar in appearance) and / or one or more options that are homophones of the selected text content. The one or more alternative word options may depend on the method by which the selected text content is entered. For example, if the selected text content is entered by handwriting input, the one or more alternative word options may include one or more similar-looking options (e.g., the one or more alternative word options may include one or more words that differ by one letter, or common misspellings); if the selected text content is entered by voice input, the one or more alternative word options may include one or more homophones of the selected text content.

[0309] In some examples, one or more of the one or more alternative word options can be provided by an artificial intelligence model. For example, a prompt can be provided to a trained large language model (LLM) or other machine learning model to generate one or more alternative word options based on the context of the selected text content. For example, a computer system can prompt the LLM by initiating an API call to a remote system hosting the LLM and receiving one or more generated alternative word options in response.

[0310] In some examples, one or more of the one or more alternative word options may be provided by an input method tool. An input method tool may be software in a computer system that is designed to facilitate text input and suggest modifications or improvements, such as a predictive text algorithm or a spell-checking algorithm. For example, the computer system may provide selected text content to a predictive text algorithm or a spell-checking algorithm (which may be software executed on the computer system), and display the text recommended by the predictive text algorithm or the spell-checking algorithm as one or more alternative word options.

[0311] At operation 2708, the input type is determined. The computer system can determine the type (nature) of the input received from the user. For example, the input type can be determined to be voice input, handwriting input, or selection of one of one or more alternative word options.

[0312] At operation 2710, input information is obtained. The computer system may obtain input information based on the determined input type. For example, if the input type is voice input, the input information may be voice data; if the input type is handwriting input, the input information may be handwriting text; if the input type is selecting an alternative word option, the input information may be a selected target option.

[0313] For example, if the input type is voice input, the input information is the content of the voice input, which can be recognized (for example, using a suitable voice recognition algorithm) as descriptive or descriptive language (rather than command or prompt language). Such descriptive or descriptive language can be recognized as replacement text.

[0314] For another example, if the input type is voice input, the content of the voice input may be a voice instruction that is recognized as a command or prompt (e.g., voice recognition can detect a voice command such as "search"). The command or prompt may include requirements or criteria, such as a search query (e.g., "search for types of dogs"). The computer system may execute the command or prompt (e.g., by initiating an API call to a remote search engine) to search for target text that meets the requirements or criteria contained in the voice input. In some examples, one or more results of the search may be used to populate one or more alternative word options.

[0315] For another example, if the input type is to select one of one or more alternative word options, the acquired input information may be the target option selected by the user from the one or more alternative word options.

[0316] For another example, if the input type is handwriting input, the acquired input information may be one or more handwritten words recognized from the handwriting input (eg, using a handwriting recognition algorithm).

[0317] At operation 2712, the selected text content is replaced according to the input information obtained. The selected text content can be replaced with the input information itself, or the selected text content can be replaced with new text content according to the input information (for example, the input information can be a command or prompt for new text content). For example, if the input information is descriptive or descriptive language, the selected text content can be replaced with text content in the descriptive or descriptive language (for example, using a speech recognition speech-to-text algorithm). For another example, if the input information is a voice indication of a command or prompt, the selected text content can be replaced with one or more results obtained from the command or prompt (for example, one of the one or more search results selected by the user). For another example, if the input information is a selection of a target option from one or more alternative word options, the selected text content can be replaced with the target option. For another example, if the input information is one or more written words, the selected text content can be replaced with one or more written words. In various examples, the present invention provides methods, systems, and devices for text editing applications capable of executing advanced text replacement functions with integrated voice and touch input.

[0318] The various examples of the present invention provide a sophisticated text editing platform that goes beyond traditional text selection and replacement methods by integrating speech recognition and artificial intelligence to provide context-appropriate word options. Various input types can be supported, including voice input, written input, and selection input (e.g., selecting options generated by an artificial intelligence model or other algorithm). The various examples of the present invention can enhance the user experience by providing a fast and intuitive text editing method on a device with touch screen functionality.

[0319] Although the present invention describes methods and processes by steps performed in a certain order, one or more steps in the methods and processes may be omitted or changed as appropriate. In appropriate cases, one or more steps may be performed in an order other than the order described.

[0320] Although the present invention has been described at least in part in terms of methods, it will be understood by those skilled in the art that the present invention is also directed to various components for performing at least some aspects and features of the methods, whether through hardware components, software, or any combination thereof. Accordingly, the technical solutions of the present invention may be embodied in the form of software products. Suitable software products may be stored in pre-recorded storage devices or other similar non-volatile or non-transitory computer-readable media, including DVDs, CD-ROMs, USB flash drives, removable hard drives, or other storage media. The software product includes instructions tangibly stored thereon that enable a processing device (e.g., a personal computer, a server, or a network device) to perform examples of the methods disclosed herein.

[0321] The present invention may be embodied in other specific forms without departing from the subject matter of the claims. The exemplary embodiments described are intended to be illustrative in all respects and not restrictive. Selected features from one or more of the above-described embodiments may be combined to create alternative embodiments not explicitly described, and features suitable for such combinations are understood to be within the scope of the present invention.

[0322] All values ​​and subranges within the disclosed ranges are also disclosed. Furthermore, while the systems, devices, and processes disclosed and illustrated herein may include a specific number of elements / components, the systems, devices, and components may be modified to include more or fewer such elements / components. For example, while any disclosed element / component may be referenced as a single quantity, the embodiments disclosed herein may be modified to include multiple such elements / components. The subject matter described herein is intended to cover and encompass all suitable technical variations.

Claims

1. A computer system, characterized in that: include: touchscreen; processor; A memory coupled to the processor, wherein the memory stores instructions that, when executed by the processor, cause the system to perform the following operations when executing a text editing application: activating a voice input channel and presenting one or more alternative word options after selecting text content at a location on the touch screen using a preset shape; Determine the input type; Acquire input information according to the input type; The selected text content is replaced with the acquired input information.

2. The system according to claim 1, wherein: When executed by the processor, the instructions further cause the system to obtain input information according to the input type by performing the following operations: When the input type is voice input, identifying the content of the voice input; when the processor executes the instruction, further causing the system to replace the selected text content with the acquired input information by performing the following operations: When the content of the voice input is descriptive language, the selected text content is replaced with the acquired input information.

3. The system according to claim 1, wherein: When executed by the processor, the instructions further cause the system to obtain input information according to the input type by performing the following operations: When the input type is voice input, identifying the content of the voice input; When executed by the processor, the instructions further cause the system to replace the selected text content with the acquired input information by performing the following operations: When the content of the voice input is a voice instruction, searching for a target text that meets the requirements of the content; The selected text content is replaced with the target associated word selected by the user input.

4. The system according to claim 1, wherein: When executed by the processor, the instructions further cause the system to obtain input information according to the input type by performing the following operations: When the input type is to select one of the one or more alternative word options, determining a target option selected from the one or more alternative word options through user input; When executed by the processor, the instructions further cause the system to replace the selected text content with the acquired information by performing the following operations: Replaces the selected text content with the target option.

5. The system according to claim 1, wherein: The selected text content is input by handwriting input, and the one or more alternative word options include one or more words that look similar to the selected text content.

6. The system according to claim 1, wherein: The selected text content is input through voice input, and the one or more alternative word options include one or more words that are homophones or homonyms of the selected text content.

7. The system according to claim 1, wherein: One or more of the one or more alternative word options are provided by an artificial intelligence model.

8. The system according to claim 1, wherein: One or more of the one or more alternative word options are provided by an input method tool.

9. The system according to claim 1, wherein: When executed by the processor, the instructions further cause the system to obtain input information according to the input type by performing the following operations: When the input type is handwritten input, determining one or more handwritten words; When executed by the processor, the instructions further cause the system to replace the selected text content with the acquired information by performing the following operations: The selected text content is replaced with the one or more written words.

10. A computer-implemented method for execution in a text editing application, characterized in that The method comprises: activating a voice input channel and presenting one or more alternative word options after selecting text content at a location on the touch screen using a preset shape; Determine the input type; Acquire input information according to the input type; The selected text content is replaced with the acquired input information.

11. The method according to claim 10, characterized in that The acquiring of input information according to the input type includes: When the input type is voice input, identifying the content of the voice input; The replacing the selected text content with the acquired input information includes: When the content of the voice input is descriptive language, the selected text content is replaced with the acquired input information.

12. The method according to claim 10, characterized in that The acquiring of input information according to the input type includes: When the input type is voice input, identifying the content of the voice input; The replacing the selected text content with the acquired input information includes: When the content of the voice input is a voice instruction, searching for a target text that meets the requirements of the content; The selected text content is replaced with the target associated word selected by the user input.

13. The method according to claim 10, characterized in that The acquiring of input information according to the input type includes: When the input type is to select one of the one or more alternative word options, determining a target option selected from the one or more alternative word options through user input; The replacing the selected text content with the acquired information includes: Replaces the selected text content with the target option.

14. The method according to claim 10, characterized in that The selected text content is input by handwriting input, and the one or more alternative word options include one or more words that look similar to the selected text content.

15. The method according to claim 10, characterized in that The selected text content is input through voice input, and the one or more alternative word options include one or more words that are homophones or homonyms of the selected text content.

16. The method according to claim 10, characterized in that One or more of the one or more alternative word options are provided by an artificial intelligence model.

17. The method according to claim 10, wherein: One or more of the one or more alternative word options are provided by an input method tool.

18. The method according to claim 10, wherein: The acquiring of input information according to the input type includes: When the input type is handwritten input, determining one or more handwritten words; Replacing the selected text content with the acquired information includes: The selected text content is replaced with the one or more written words.

19. A non-transitory computer-readable medium storing instructions, characterized in that: When the instructions are executed by a processor of a computer system, the system performs the following operations when executing a text editing application: activating a voice input channel and presenting one or more alternative word options after selecting text content at a location on the touch screen using a preset shape; Determine the input type; Acquire input information according to the input type; The selected text content is replaced with the acquired input information.