Method for voice typing and apparatus

The method uses multiple AI models to accurately interpret voice inputs, addressing ambiguity in voice typing by distinguishing commands and content, enhancing user experience through precise text and system operations.

WO2026113254A1PCT designated stage Publication Date: 2026-06-04HUAWEI TECH CO LTD

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-04-30
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing voice typing methods often fail to accurately interpret user needs, particularly in cases where input content and commands are interleaved, leading to ambiguity and inefficiency.

Method used

A method involving multiple AI models to process voice inputs, distinguishing between text operation commands and system commands, and providing user options for ambiguous cases, ensuring accurate intent prediction and user interaction.

Benefits of technology

Enhances user experience by improving the accuracy of intent prediction and enabling efficient handling of mixed commands and content, allowing for precise text and system operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025092474_04062026_PF_FP_ABST
    Figure CN2025092474_04062026_PF_FP_ABST
Patent Text Reader

Abstract

A method and an apparatus for voice typing. The method includes: obtaining digital text based on a voice input; processing the digital text via a first model to obtain an initial intent prediction result, where the initial intent prediction result includes at least one set of tokens from the digital text and predicted intent(s) corresponding to each of the at least one set of tokens; and, generating a first output, where the first output is generated based on processing the digital text as at least one command and / or input content based on the initial intent prediction result. According to the technical solution, more accurate interpretation of user needs is expected.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR VOICE TYPING AND APPARATUSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] The present application claims priority to PCT patent application No. PCT / CN2024 / 134471, filed on November 26, 2024 and entitled “A method for voice typing and apparatus” , which is incorporated herein by reference in its entirety.TECHNICAL FIELD

[0002] Embodiments of the present application relate to the field of artificial intelligence, and more specifically, to a method for voice typing and apparatus.BACKGROUND

[0003] Voice typing is widely used with the development of electronic technologies. The voice typing allows users to convert spoken words into written text.

[0004] However, present voice typing methods often fall short in accurate interpretation of user needs.SUMMARY

[0005] Embodiments of the present application provide a method for voice typing and an apparatus, which can support a more accurate interpretation of user needs.

[0006] According to a first aspect, an embodiment of the present application provides a method for voice typing, including: obtaining digital text based on a voice input; processing the digital text via a first model to obtain an initial intent prediction result, where the initial intent prediction result includes at least one set of tokens from the digital text and predicted intent (s) corresponding to each of the at least one set of tokens; and, generating a first output, where the first output is generated based on processing the digital text as at least one command and / or input content based on the initial intent prediction result.

[0007] According to the above technical solution, the digital text can be converted from the voice input using techniques such as automatic speech recognition (ASR) . The digital text could be processed via the first model to obtain the initial intent prediction result which may include predicted intents corresponding to each set of tokens from the digital text, so that the digital text can be processed as at least one command and / or input content based on the initial intent prediction result, which allows intent prediction in cases of input content and commands are present in the digital text in an interleaved manner. A better user experience is expected.

[0008] With reference to the first aspect, in some embodiments, the at least one command includes: at least one text operation command and / or at least one system command.

[0009] According to the above technical solution, the text operation command can be used to operate on text (e.g., on the input content) , for example, the text operation commands include text edit commands and text analysis commands. The system command can be used to trigger system operations. The technical solutions provided in this application supports distinguishing between different kinds of commands and corresponding operations (e.g., text edit, text analysis and system operation) .

[0010] With reference to the first aspect, in some embodiments, the at least one set of tokens includes a first set of tokens and / or a second set of tokens from the digital text, the predicted intent corresponding to the first set of tokens is a command and the predicted intent corresponding to the second set of tokens is input content.

[0011] According to the above technical solution, the first model may be an artificial intelligence model which can be installed in mobile devices. The initial intent prediction result may be used as a reference for processing the first set of tokens and the second set of tokens. It is to be noted that the initial intent prediction result may not correspond with exact intent of a user. For example, the first set of tokens is predicted to be a command by the first model but it can in fact be input content. The digital text may include input content and / or command. For example, for digital text “I am going to eat at 5p.m. and sleep after that” , input content is included in the digital text. For another example, for digital text “I am going to eat at 5p.m. and sleep after that, change it to 6p.m. ” , input content and command coexist. For another example, for digital text “send a message to A” , a command is included in the digital text.

[0012] With reference to the first aspect, in some embodiments, the at least one set of tokens includes the first set of tokens and a first confidence score of the first set of tokens is equal to or higher than a first threshold, where the first confidence score of the first set of tokens indicates a probability of the first set of tokens being a command.

[0013] According to the above technical solution, the first model may also output intent prediction scores which indicate a probability of each of the set of tokens being arguments of a command. Based on the intent predication scores, the first confidence score of the first set of tokens is considered to be higher than the first threshold. Therefore, the first set of tokens are predicated to be a command as at least part of the initial intent prediction result.

[0014] In some embodiments, based on the intent prediction scores, the probability of a set of tokens being a command and a probability of the set of tokens being input content can be determined. If the probability of the set of tokens being a command is higher than the probability of the set of tokens being input content, the set of tokens is predicated to be a command.

[0015] With reference to the first aspect, in some embodiments, the method further includes: obtaining command arguments corresponding to the first set of tokens from the digital text via a second model, and the first output is generated based on the command arguments obtained via the second model.

[0016] According to the above technical solution, the second model and the first model may be a same model. Optionally, it could be a different model from the first model. There are different cases related to command arguments obtained via the second model, for example, there may be a lack of arguments corresponding to the first set of tokens, a confidence in the command arguments may be high or low. For different cases, the first set of tokens may be processed differently and the first output may be different.

[0017] With reference to the first aspect, in some embodiments, the method further includes: if there is a lack of command arguments corresponding to the first set of tokens, elongating an utterance duration for obtaining longer digital text.

[0018] According to the above technical solution, the digital text processed using technical solutions provided in this application can be separated from large piece of text based on pauses of user speaking. If there is a lack of command arguments corresponding to the first set of tokens, the positions for separation of the digital text from the large piece of text may be inappropriate, resulting in an incomplete statement for the digital text. It is possible that the remaining command arguments are present beside the digital text, e.g., before the starting position of the digital text (past utterance relative to the digital text) or after the ending position of the digital text (future utterance relative to the digital text) in the large piece of text. The electronic device may elongate the utterance duration for the digital text to obtain longer digital text, i.e., to obtain past utterance or future utterance relative to the present digital text. For example, for the large piece of text “I am going to eat at 5p.m. and sleep after that, change it to 6p.m. ” , “and sleep after that, change it to 6p.m” may be obtained as the original digital text, the electronic device may elongate the utterance duration to obtain a longer digital text including the past utterance “I am going to eat at 5p.m. ” .

[0019] With reference to the first aspect, in some embodiments, if a second confidence score of the first set of tokens is higher than a second threshold, the method further includes: processing the first set of tokens as a command, where the first output is generated based on a result of processing the first set of tokens as a command, where the second confidence score of the first set of tokens indicates a confidence in the command arguments obtained via the second model for the first set of tokens.

[0020] According to the above technical solution, the first set of tokens is highly possible to be a command if the second confidence score in the command arguments is high. Therefore the first set of tokens can be processed as a command and the first output may be based on processing the first set of tokens as a command. It is to be noted that although the first set of tokens are processed as a command, there may be tokens that are essential for processing the first set of tokens as a command in other parts of the digital text, e.g., the input content.

[0021] With reference to the first aspect, in some embodiments, if a second confidence score of the first set of tokens is equal to or lower than a second threshold, the method further includes: extracting command arguments corresponding to the first set of tokens from the digital text via a third model, where the third model has a processing capability better than a processing capability of the second model, and the first output is generated based on an extraction result via the third model, where the second confidence score of the first set of tokens indicates a confidence in the command arguments obtained via the second model for the first set of tokens.

[0022] Optionally, if the second confidence score is equal to or lower than a second threshold, the first set of tokens can also be processed as a command as a trial and the third model is not needed in this scenario.

[0023] According to the above technical solution, if a confidence in the command arguments is low, the second model may have inefficient capability to predict intent for the first set of tokens, a model with a better capability may be employed to extract the command arguments.

[0024] With reference to the first aspect, in some embodiments, the at least one set of tokens includes a third set of tokens from the digital text, a first confidence score of the third set of tokens is higher than a third threshold and lower than a first threshold, and the method further includes: extracting command arguments corresponding to the third set of tokens from the digital text via a third model, where the third model has a processing capability better than a processing capability of the first model, and the first output is generated based on an extraction result via the third model, where the first confidence score of the third set of tokens indicates a probability of the third set of tokens being a command; or providing a user with at least two options, where for the at least two options, the third set of tokens are processed differently, and the first output is generated based on the at least two options.

[0025] According to the above technical solution, it is hard to predict intent for the third set of tokens by the first model. In some embodiments, a model with better capability can be employed to extract commands for the third set of tokens. In other embodiments, at least two options could be provided for the user to select to resolve ambiguity.

[0026] With reference to the first aspect, in some embodiments, if the extraction result via the third model is successful, the method further includes: processing the first set of tokens as a command, where the first output is generated based on a result of processing the first set of tokens as a command; and / or processing the third set of tokens as a command, where the first output is generated based on a result of processing the third set of tokens as a command.

[0027] According to the above technical solution, as the command extraction via the third model is successful, the first set of tokens and / or the third set of tokens is highly possible to be commands and can be processed as commands to generate the first output.

[0028] With reference to the first aspect, in some embodiments, if the result of processing the first set of tokens as a command is unsuccessful, the method further includes: processing the first set of tokens as input content, where the first output is generated based on processing the first set of tokens as input content; and / or if the result of processing the third set of tokens as a command is unsuccessful, the method further includes: processing the third set of tokens as input content, where the first output is generated based on processing the third set of tokens as input content.

[0029] According to the above technical solution, as processing the first set of tokens and / or the third set of tokens as commands is unsuccessful, the first set of tokens and / or the third set of tokens can be input content rather than commands. The first set of tokens and / or the third set of tokens can be processed as input content and used to generate the first output. For example, text including the first set of tokens and / or the third set of tokens can be displayed on screen.

[0030] With reference to the first aspect, in some embodiments, if the extraction result via the third model is unsuccessful, the method further includes: processing the first set of tokens as input content, where the first output is generated based on processing the first set of tokens as input content; and / or processing the third set of tokens as input content where the first output is generated based on processing the third set of tokens as input content.

[0031] According to the above technical solution, as command extraction via the third model is unsuccessful, the first set of tokens and / or the third set of tokens can be input content rather than commands. The electronic device may generate the first output based on processing the first set of tokens and / or the third set of tokens as input content.

[0032] With reference to the first aspect, in some embodiments, if the extraction result via the third model is unsuccessful, the method further includes: providing a user with at least two options, where the first output is generated based on the at least two options and the first set of tokens are processed differently for the at least two options; and / or providing a user with at least two options, where the first output is generated based on the at least two options and the third set of tokens are processed differently for the at least two options.

[0033] According to the above technical solution, as command extraction via the third model is unsuccessful, there is ambiguity for intent prediction on the first set of tokens and / or the third set of tokens. The electronic device may provide user with options, where the options may correspond to processing the first set of tokens differently (e.g., one option correspond to processing the first set of tokens as input content and another option correspond to processing the first set of tokens as a command) and / or processing the third set of tokens differently.

[0034] With reference to the first aspect, in some embodiments, the at least one set of tokens includes the second set of tokens and a first confidence score of the second set of tokens is equal to or lower than a third threshold, and the method further includes: processing the second set of tokens as input content, where the first output is generated based on processing the second set of tokens as input content, where the first confidence score of the second set of tokens indicates a probability of the second set of tokens being a command.

[0035] According to the above technical solution, the second set of tokens is highly possible to be input content and can be processed as input content in priority.

[0036] With reference to the first aspect, in some embodiments, the method further includes: receiving a first user input, where the first user input is used to cancel the first output; and generating a second output, where the first set of tokens and / or the second set of tokens and / or the third set of tokens are processed differently corresponding to the second output from that corresponding to the first output.

[0037] According to the above technical solution, the first output may be displayed on screen or delivered to the user and the user may find that first output is undesired. In this scenario, the user may use commands like undo / cut to cancel present first output. Then the user may re-specks the same utterance, and the electronic device may process the digital text differently from the processing manner corresponding to the first output.

[0038] With reference to the first aspect, in some embodiments, the method further includes: providing at least one alternative option, where the first set of tokens and / or the second set of tokens and / or the third set of tokens are processed differently corresponding to the at least one alternative option from that corresponding to the first output.

[0039] According to the above technical solution, the digital text is processed in a specific way to generate the first output. If there is ambiguity related with the digital text, the electronic device may provide alternative options for user to choose, which may improve user experience.

[0040] With reference to the first aspect, in some embodiments, the method further includes: receiving a second user input, where the second user input is a selection of one of the least one option; generating a second output corresponding to the second user input.

[0041] According to the above technical solution, the second output is generated based on user selecting one of the at least one option.

[0042] With reference to the first aspect, in some embodiments, the at least one text operation command includes at least one text edit command and / or at least one text analysis command.

[0043] With reference to the first aspect, in some embodiments, the at least one text edit command includes one or more of:at least one replace command, at least one add input command, at least one delete command, at least one add reference command.

[0044] According to a second aspect, an embodiment of the present application provides an apparatus, wherein the apparatus is configured to perform the method in the first aspect or any optional implementation of the first aspect.

[0045] According to a third aspect, an embodiment of this application provides a computer-readable storage medium, where the computer-readable storage medium stores instructions, and when the instructions run on a device, the device is enabled to perform the method according to the method in the first aspect or any optional implementation of the first aspect.

[0046] According to a fourth aspect, an embodiment of this application provides a computer program product, when the computer program product runs on a device, the device is enabled to perform the method according to the method in the first aspect or any optional implementation of the first aspect.

[0047] According to a fifth aspect, an embodiment of this application provides an chip system, including a memory and a processor, where the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that a device on which the chip system is disposed performs the method according to the method in the first aspect or any optional implementation of the first aspect.DESCRIPTION OF DRAWINGS

[0048] FIG. 1 is a schematic diagram of a hardware structure of an electronic device 100 according to an embodiment of this application;

[0049] FIG. 2 is a block diagram of a software structure of the electronic device 100 according to an embodiment of this application;

[0050] FIG. 3 is a schematic diagram of a method for voice typing according to an embodiment of this application;

[0051] FIG. 4 is a schematic diagram showing an initial intent prediction result according to an embodiment of this application;

[0052] FIG. 5 is a schematic diagram of an example architecture for labeling tokens with different class labels;

[0053] FIG. 6 is a schematic diagram showing handling case when missing one or more edit arguments;

[0054] FIG. 7 is a schematic diagram showing handling case of low confidence in classified arguments;

[0055] FIG. 8 is a schematic diagram of showing handling case when language model is less confident in classification between input content and edit command;

[0056] FIG. 9 is a schematic diagram showing user decision-making support for ambiguous options;

[0057] FIG. 10 is a schematic diagram showing complete pipeline for intent segmentation, processing edit instruction, ambiguity detection and user decision-making support for ambiguous options;

[0058] FIG. 11 is a schematic diagram of a hardware structure of an apparatus according to an embodiment of this application.DESCRIPTION OF EMBODIMENTS

[0059] The following describes the technical solutions in this application with reference to the accompanying drawings.

[0060] Terms used in the following embodiments of this application are merely intended to describe specific embodiments, but are not intended to limit this application. Terms “one” , “a” , “the” , “the foregoing” , “this” , and “the one” of singular forms used in this specification and the appended claims of this application are also intended to include plural forms like “one or more” , unless otherwise specified in the context clearly.

[0061] Reference to “an embodiment” , “some embodiments” , or the like described in this specification indicates that one or more embodiments of this application include a specific feature, structure, or characteristic described with reference to the embodiments. Therefore, in this specification, statements, such as “in an embodiment” , “in some embodiments” , “in some other embodiments” , and “in other embodiments” , that appear at different places do not necessarily mean referring to a same embodiment, instead, but mean “one or more but not all of the embodiments” , unless otherwise specified. The terms “include” , “comprise” , “have” , and their variants all mean “include but are not limited to” , unless otherwise specified.

[0062] In order to describe the method for voice typing provided by the embodiment of the present application more clearly, an electronic device for executing the method is first introduced below.

[0063] In some embodiments, the electronic device may be a portable electronic device that further includes other functions such as a personal digital assistant function and / or a music player function, for example, a mobile phone, a tablet computer, or a wearable electronic device having a wireless communication function (for example, a smartwatch) . An example embodiment of the portable electronic device includes but is not limited to a portable electronic device using  or another operating system. The portable electronic device may alternatively be another portable electronic device, for example, a laptop computer (Laptop) . It should be further understood that, in some other embodiments, the electronic device may alternatively be a desktop computer, but not a portable electronic device.

[0064] FIG. 1 illustrates a schematic diagram of a hardware structure of an electronic device 100 according to an embodiment of this application.

[0065] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (universal serial bus, USB) port 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headset jack 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display 194, and a subscriber identity module (subscriber identity module, SIM) card interface 195. The sensor module 180 may include a pressure sensor 180A, a gyro sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, an optical proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, and the like.

[0066] It may be understood that the structure shown in this embodiment of this application does not constitute a specific limitation on the electronic device 100. In some other embodiments of this application, the electronic device 100 may include more or fewer components than those shown in the figure, or combine some components, or split some components, or have different component arrangements. The components shown in the figure may be implemented by hardware, software, or a combination of software and hardware.

[0067] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (application processor, AP) , a modem processor, a graphics processing unit (graphics processing unit, GPU) , an image signal processor (image signal processor, ISP) , a controller, a memory, a video codec, a digital signal processor (digital signal processor, DSP) , a baseband processor, and / or a neural-network processing unit (neural-network processing unit, NPU) . Different processing units may be independent components, or may be integrated into one or more processors.

[0068] The controller may be a nerve center and a command center of the electronic device 100. The controller may generate an operation control signal based on an instruction operation code and a time sequence signal, to complete control of instruction reading and instruction execution.

[0069] A memory may be further disposed in the processor 110, and is configured to store instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. The memory may store instructions or data that has been used or cyclically used by the processor 110. If the processor 110 needs to use the instructions or the data again, the processor may directly invoke the instructions or the data from the memory. This avoids repeated access, reduces waiting time of the processor 110, and improves system efficiency.

[0070] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an inter-integrated circuit (inter-integrated circuit, I2C) interface, an inter-integrated circuit sound (inter-integrated circuit sound, I2S) interface, a pulse code modulation (pulse code modulation, PCM) interface, a universal asynchronous receiver / transmitter (universal asynchronous receiver / transmitter, UART) interface, a mobile industry processor interface (mobile industry processor interface, MIPI) , a general-purpose input / output (general-purpose input / output, GPIO) interface, a subscriber identity module (subscriber identity module, SIM) interface, a universal serial bus (universal serial bus, USB) port, and / or the like.

[0071] The I2C interface is a two-way synchronization serial bus, and includes one serial data line (serial data line, SDA) and one serial clock line (serial clock line, SCL) . In some embodiments, the processor 110 may include a plurality of groups of I2C buses. The processor 110 may be separately coupled to the touch sensor 180K, a charger, a flash, the camera 193, and the like through different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K through the I2C interface, so that the processor 110 communicates with the touch sensor 180K through the I2C bus interface, to implement a touch function of the electronic device 100.

[0072] The I2S interface may be used for audio communication. In some embodiments, the processor 110 may include a plurality of groups of I2S buses. The processor 110 may be coupled to the audio module 170 through the I2S bus, to implement communication between the processor 110 and the audio module 170. In some embodiments, the audio module 170 may transmit an audio signal to the wireless communication module 160 through the I2S interface, to implement a function of answering a call through a Bluetooth headset.

[0073] The PCM interface may also be used for audio communication, and sample, quantize, and code an analog signal. In some embodiments, the audio module 170 may be coupled to the wireless communication module 160 through a PCM bus interface. In some embodiments, the audio module 170 may also transmit an audio signal to the wireless communication module 160 through the PCM interface, to implement a function of answering a call through a Bluetooth headset. Both the I2S interface and the PCM interface may be used for audio communication.

[0074] The UART interface is a universal serial data bus, and is used for asynchronous communication. The bus may be a two-way communication bus, and converts to-be-transmitted data between serial communication and parallel communication. In some embodiments, the UART interface is usually configured to connect the processor 110 to the wireless communication module 160. For example, the processor 110 communicates with a Bluetooth module in the wireless communication module 160 through the UART interface, to implement a Bluetooth function. In some embodiments, the audio module 170 may transmit an audio signal to the wireless communication module 160 through the UART interface, to implement a function of playing music through a Bluetooth headset.

[0075] The MIPI may be configured to connect the processor 110 to a peripheral component such as the display 194 or the camera 193. The MIPI includes a camera serial interface (camera serial interface, CSI) , a display serial interface (display serial interface, DSI) , and the like. In some embodiments, the processor 110 communicates with the camera 193 through the CSI, to implement a photographing function of the electronic device 100. The processor 110 communicates with the display 194 through the DSI, to implement a display function of the electronic device 100.

[0076] The GPIO interface may be configured by software. The GPIO interface may be configured as a control signal or a data signal. In some embodiments, the GPIO interface may be configured to connect the processor 110 to the camera 193, the display 194, the wireless communication module 160, the audio module 170, the sensor module 180, or the like. The GPIO interface may alternatively be configured as an I2C interface, an I2S interface, a UART interface, an MIPI, or the like.

[0077] The USB port 130 is a port that conforms to a USB standard specification, and may be specifically a mini USB port, a micro USB port, a USB Type-C port, or the like. The USB port 130 may be configured to connect to a charger to charge the electronic device 100, or may be configured to transmit data between the electronic device 100 and a peripheral device, or may be configured to connect to a headset for playing audio through the headset. The port may be further configured to connect to another electronic device such as an AR device.

[0078] It may be understood that an interface connection relationship between the modules illustrated in this embodiment of this application is merely an example for description, and constitutes no limitation on the structure of the electronic device 100. In some other embodiments of this application, the electronic device 100 may alternatively use an interface connection manner different from that in the foregoing embodiment, or use a combination of a plurality of interface connection manners.

[0079] The charging management module 140 is configured to receive a charging input from a charger. The charger may be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 140 may receive a charging input of a wired charger through the USB port 130. In some embodiments of wireless charging, the charging management module 140 may receive a wireless charging input through a wireless charging coil of the electronic device 100. The charging management module 140 supplies power to the electronic device through the power management module 141 while charging the battery 142.

[0080] The power management module 141 is configured to connect to the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives an input of the battery 142 and / or the charging management module 140, to supply power to the processor 110, the internal memory 121, an external memory, the display 194, the camera 193, the wireless communication module 160, and the like. The power management module 141 may be further configured to monitor parameters such as a battery capacity, a battery cycle count, and a battery health status (electric leakage or impedance) . In some other embodiments, the power management module 141 may alternatively be disposed in the processor 110. In some other embodiments, the power management module 141 and the charging management module 140 may alternatively be disposed in a same device.

[0081] A wireless communication function of the electronic device 100 may be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor, the baseband processor, and the like.

[0082] The antenna 1 and the antenna 2 are configured to: transmit and receive an electromagnetic wave signal. Each antenna in the electronic device 100 may be configured to cover one or more communication frequency bands. Different antennas may be further multiplexed, to improve antenna utilization. For example, the antenna 1 may be multiplexed as a diversity antenna in a wireless local area network. In some other embodiments, the antenna may be used in combination with a tuning switch.

[0083] The mobile communication module 150 may provide a wireless communication solution that is applied to the electronic device 100 and that includes 2G / 3G / 4G / 5G or the like. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (low noise amplifier, LNA) , and the like. The mobile communication module 150 may receive an electromagnetic wave through the antenna 1, perform processing such as filtering or amplification on the received electromagnetic wave, and transmit the electromagnetic wave to the modem processor for demodulation. The mobile communication module 150 may further amplify a signal modulated by the modem processor, and convert the signal into an electromagnetic wave for radiation through the antenna 1. In some embodiments, at least some functional modules in the mobile communication module 150 may be disposed in the processor 110. In some embodiments, at least some functional modules in the mobile communication module 150 may be disposed in a same device as at least some modules of the processor 110.

[0084] The modem processor may include a modulator and a demodulator. The modulator is configured to modulate a to-be-sent low-frequency baseband signal into a medium-high frequency signal. The demodulator is configured to demodulate a received electromagnetic wave signal into a low-frequency baseband signal. Then, the demodulator transmits the low-frequency baseband signal obtained through demodulation to the baseband processor for processing. The low-frequency baseband signal is processed by the baseband processor and then transmitted to the application processor. The application processor outputs a sound signal through an audio device (which is not limited to the speaker 170A, the receiver 170B, or the like) , or displays an image or a video through the display 194. In some embodiments, the modem processor may be an independent component. In some other embodiments, the modem processor may be independent of the processor 110, and is disposed in a same device as the mobile communication module 150 or another functional module.

[0085] The wireless communication module 160 may provide a wireless communication solution that is applied to the electronic device 100 and that includes a wireless local area network (wireless local area network, WLAN) (for example, a wireless fidelity (wireless fidelity, Wi-Fi) network) , Bluetooth (Bluetooth, BT) , a global navigation satellite system (global navigation satellite system, GNSS) , frequency modulation (frequency modulation, FM) , a near field communication (near field communication, NFC) technology, an infrared (infrared, IR) technology, or the like. The wireless communication module 160 may be one or more components integrating at least one communication processing module. The wireless communication module 160 receives an electromagnetic wave through the antenna 2, performs frequency modulation and filtering processing on an electromagnetic wave signal, and sends a processed signal to the processor 110. The wireless communication module 160 may further receive a to-be-sent signal from the processor 110, perform frequency modulation and amplification on the signal, and convert the signal into an electromagnetic wave for radiation through the antenna 2.

[0086] In some embodiments, the antenna 1 and the mobile communication module 150 in the electronic device 100 are coupled, and the antenna 2 and the wireless communication module 160 in the electronic device 100 are coupled, so that the electronic device 100 can communicate with a network and another device by using a wireless communication technology. The wireless communication technology may include a global system for mobile communications (global system for mobile communications, GSM) , a general packet radio service (general packet radio service, GPRS) , code division multiple access (code division multiple access, CDMA) , wideband code division multiple access (wideband code division multiple access, WCDMA) , time-division code division multiple access (time-division code division multiple access, TD-CDMA) , long term evolution (long term evolution, LTE) , BT, a GNSS, a WLAN, NFC, FM, an IR technology, and / or the like. The GNSS may include a global positioning system (global positioning system, GPS) , a global navigation satellite system (global navigation satellite system, GLONASS) , a BeiDou navigation satellite system (BeiDou navigation satellite system, BDS) , a quasi-zenith satellite system (quasi-zenith satellite system, QZSS) , and / or a satellite based augmentation system (satellite based augmentation system, SBAS) .

[0087] The electronic device 100 may implement a display function through the GPU, the display 194, the application processor, and the like. The GPU is a microprocessor for image processing, and is connected to the display 194 and the application processor. The GPU is configured to: perform mathematical and geometric computation, and render an image. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0088] The display 194 is configured to display an image, a video, and the like. The display 194 includes a display panel. The display panel may be a liquid crystal display (liquid crystal display, LCD) , an organic light-emitting diode (organic light-emitting diode, OLED) , an active-matrix organic light emitting diode (active-matrix organic light emitting diode, AMOLED) , a flexible light-emitting diode (flexible light-emitting diode, FLED) , a mini-LED, a micro-LED, a micro-OLED, a quantum dot light emitting diode (quantum dot light emitting diode, QLED) , or the like. In some embodiments, the electronic device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0089] The electronic device 100 may implement a photographing function through the ISP, the camera 193, the video codec, the GPU, the display 194, the application processor, and the like.

[0090] The ISP is configured to process data fed back by the camera 193. For example, during photographing, a shutter is pressed, and light is transmitted to a photosensitive element of the camera through a lens. An optical signal is converted into an electrical signal, and the photosensitive element of the camera transmits the electrical signal to the ISP for processing, to convert the electrical signal into a visible image. The ISP may further perform algorithm optimization on noise, brightness, and complexion of the image. The ISP may further optimize parameters such as exposure and a color temperature of a photographing scenario. In some embodiments, the ISP may be disposed in the camera 193.

[0091] The camera 193 is configured to capture a static image or a video. An optical image of an object is generated through the lens, and is projected onto the photosensitive element. The photosensitive element may be a charge coupled device (charge coupled device, CCD) or a complementary metal-oxide-semiconductor (complementary metal-oxide-semiconductor, CMOS) phototransistor. The photosensitive element converts an optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert the electrical signal into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0092] The digital signal processor is configured to process a digital signal, and may process another digital signal in addition to the digital image signal. For example, when the electronic device 100 selects a frequency, the digital signal processor is configured to perform Fourier transform on frequency energy.

[0093] The video codec is configured to: compress or decompress a digital video. The electronic device 100 may support one or more video codecs. In this way, the electronic device 100 may play or record videos in a plurality of coding formats, for example, moving picture experts group (moving picture experts group, MPEG) -1, MPEG-2, MPEG-3, and MPEG-4.

[0094] The NPU is a neural-network (neural-network, NN) computing processor, quickly processes input information by referring to a structure of a biological neural network, for example, by referring to a mode of transmission between human brain neurons, and may further continuously perform self-learning. Applications such as intelligent cognition of the electronic device 100 may be implemented through the NPU, for example, image recognition, facial recognition, speech recognition, and text understanding.

[0095] The external memory interface 120 may be used to connect to an external storage card, for example, a micro SD card, to extend a storage capability of the electronic device 100. The external storage card communicates with the processor 110 through the external memory interface 120, to implement a data storage function. For example, files such as music and videos are stored in the external storage card.

[0096] The internal memory 121 may be configured to store computer-executable program code. The executable program code includes instructions. The processor 110 runs the instructions stored in the internal memory 121, to perform various function applications of the electronic device 100 and data processing. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store an operating system, an application required by at least one function (for example, a voice playing function or an image playing function) , and the like. The data storage area may store data (such as audio data and an address book) created during use of the electronic device 100, and the like. In addition, the internal memory 121 may include a high-speed random access memory, or may include a nonvolatile memory, for example, at least one magnetic disk storage device, a flash memory, or a universal flash storage (universal flash storage, UFS) .

[0097] The electronic device 100 may implement an audio function, for example, music playing and recording, through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headset jack 170D, the application processor, and the like.

[0098] The audio module 170 is configured to convert digital audio information into an analog audio signal for output, and is also configured to convert analog audio input into a digital audio signal. The audio module 170 may be further configured to:code and decode an audio signal. In some embodiments, the audio module 170 may be disposed in the processor 110, or some functional modules in the audio module 170 are disposed in the processor 110.

[0099] The speaker 170A, also referred to as a "horn" , is configured to convert an audio electrical signal into a sound signal. The electronic device 100 may be used to listen to music or answer a call in a hands-free mode over the speaker 170A.

[0100] The receiver 170B, also referred to as an "earpiece" , is configured to convert an electrical audio signal into a sound signal. When a call is answered or speech information is received through the electronic device 100, the receiver 170B may be put close to a human ear to listen to a voice.

[0101] The microphone 170C, also referred to as a "mike" or a "mic" , is configured to convert a sound signal into an electrical signal. When making a call or sending a voice message, a user may make a sound near the microphone 170C through the mouth of the user, to input a sound signal to the microphone 170C. At least one microphone 170C may be disposed in the electronic device 100. In some other embodiments, two microphones 170C may be disposed in the electronic device 100, to collect a sound signal and implement a noise reduction function. In some other embodiments, three, four, or more microphones 170C may alternatively be disposed in the electronic device 100, to collect a sound signal, implement noise reduction, and identify a sound source, so as to implement a directional recording function and the like.

[0102] The headset jack 170D is configured to connect to a wired headset. The headset jack 170D may be a USB port 130, or may be a 3.5 mm open mobile terminal platform (open mobile terminal platform, OMTP) standard interface or cellular telecommunications industry association of the USA (cellular telecommunications industry association of the USA, CTIA) standard interface.

[0103] The pressure sensor 180A is configured to sense a pressure signal, and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A may be disposed on the display 194. There are a plurality of types of pressure sensors 180A, such as a resistive pressure sensor, an inductive pressure sensor, and a capacitive pressure sensor. The capacitive pressure sensor may include at least two parallel plates made of conductive materials. When a force is applied to the pressure sensor 180A, capacitance between electrodes changes. The electronic device 100 determines pressure intensity based on the change in the capacitance. When a touch operation is performed on the display 194, the electronic device 100 detects intensity of the touch operation through the pressure sensor 180A. The electronic device 100 may also calculate a touch location based on a detection signal of the pressure sensor 180A. In some embodiments, touch operations that are performed at a same touch location but have different touch operation intensity may correspond to different operation instructions. For example, when a touch operation whose touch operation intensity is less than a first pressure threshold is performed on a Messages application icon, an instruction for viewing an SMS message is performed. When a touch operation whose touch operation intensity is greater than or equal to the first pressure threshold is performed on the Messages application icon, an instruction for creating a new SMS message is performed.

[0104] The gyro sensor 180B may be configured to determine a moving posture of the electronic device 100. In some embodiments, an angular velocity of the electronic device 100 around three axes (namely, axes x, y, and z) may be determined through the gyro sensor 180B. The gyro sensor 180B may be configured to implement image stabilization during photographing. For example, when the shutter is pressed, the gyro sensor 180B detects an angle at which the electronic device 100 jitters, calculates, based on the angle, a distance for which a lens module needs to compensate, and allows the lens to cancel the jitter of the electronic device 100 through reverse motion, to implement image stabilization. The gyro sensor 180B may also be used in a navigation scenario and a somatic game scenario.

[0105] The barometric pressure sensor 180C is configured to measure barometric pressure. In some embodiments, the electronic device 100 calculates an altitude through the barometric pressure measured by the barometric pressure sensor 180C, to assist in positioning and navigation.

[0106] The magnetic sensor 180D includes a Hall sensor. The electronic device 100 may detect opening and closing of a flip cover by using the magnetic sensor 180D. In some embodiments, when the electronic device 100 is a clamshell phone, the electronic device 100 may detect opening and closing of a flip cover based on the magnetic sensor 180D. Further, a feature such as automatic unlocking of the flip cover is set based on a detected opening or closing state of the flip cover.

[0107] The acceleration sensor 180E may detect accelerations of the electronic device 100 in various directions (usually on three axes) . When the electronic device 100 is still, a magnitude and a direction of gravity may be detected. The acceleration sensor 180E may be further configured to identify a posture of the electronic device, and is used in an application such as switching between a landscape mode and a portrait mode or a pedometer.

[0108] The distance sensor 180F is configured to measure a distance. The electronic device 100 may measure the distance in an infrared manner or a laser manner. In some embodiments, in a photographing scenario, the electronic device 100 may measure a distance through the distance sensor 180F to implement quick focusing.

[0109] The optical proximity sensor 180G may include, for example, a light-emitting diode (LED) and an optical detector, for example, a photodiode. The light-emitting diode may be an infrared light-emitting diode. The electronic device 100 emits infrared light by using the light-emitting diode. The electronic device 100 detects infrared reflected light from a nearby object through the photodiode. When sufficient reflected light is detected, the electronic device 100 may determine that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 may determine that there is no object near the electronic device 100. The electronic device 100 may detect, by using the optical proximity sensor 180G, that the user holds the electronic device 100 close to an ear for a call, to automatically turn off a screen for power saving. The optical proximity sensor 180G may also be used in a smart cover mode or a pocket mode to automatically perform screen unlocking or locking.

[0110] The ambient light sensor 180L is configured to sense ambient light brightness. The electronic device 100 may adaptively adjust brightness of the display 194 based on the sensed ambient light brightness. The ambient light sensor 180L may also be configured to automatically adjust white balance during photographing. The ambient light sensor 180L may also cooperate with the optical proximity sensor 180G to detect whether the electronic device 100 is in a pocket, to avoid an accidental touch.

[0111] The fingerprint sensor 180H is configured to collect a fingerprint. The electronic device 100 may use a feature of the collected fingerprint to implement fingerprint-based unlocking, application lock access, fingerprint-based photographing, fingerprint-based call answering, and the like.

[0112] The temperature sensor 180J is configured to detect a temperature. In some embodiments, the electronic device 100 executes a temperature processing policy through the temperature detected by the temperature sensor 180J. For example, when the temperature reported by the temperature sensor 180J exceeds a threshold, the electronic device 100 lowers performance of a processor near the temperature sensor 180J, to reduce power consumption for thermal protection. In some other embodiments, when the temperature is less than another threshold, the electronic device 100 heats the battery 142 to prevent the electronic device 100 from being shut down abnormally due to a low temperature. In some other embodiments, when the temperature is less than still another threshold, the electronic device 100 boosts an output voltage of the battery 142 to avoid abnormal shutdown caused by a low temperature.

[0113] The touch sensor 180K is also referred to as a "touch panel" . The touch sensor 180K may be disposed on the display 194, and the touch sensor 180K and the display 194 constitute a touchscreen. The touch sensor 180K is configured to detect a touch operation performed on or near the touch sensor. The touch sensor may transfer the detected touch operation to the application processor to determine a type of the touch event. A visual output related to the touch operation may be provided through the display 194. In some other embodiments, the touch sensor 180K may also be disposed on a surface of the electronic device 100 at a location different from that of the display 194.

[0114] The bone conduction sensor 180M may obtain a vibration signal. In some embodiments, the bone conduction sensor 180M may obtain a vibration signal of a vibration bone of a human vocal-cord part. The bone conduction sensor 180M may also be in contact with a body pulse to receive a blood pressure beating signal. In some embodiments, the bone conduction sensor 180M may also be disposed in the headset, to obtain a bone conduction headset. The audio module 170 may obtain a speech signal through parsing based on the vibration signal that is of the vibration bone of the vocal-cord part and that is obtained by the bone conduction sensor 180M, to implement a speech function. The application processor may parse heart rate information based on the blood pressure beating signal obtained by the bone conduction sensor 180M, to implement a heart rate detection function.

[0115] The button 190 includes a power button, a volume button, and the like. The button 190 may be a mechanical button, or may be a touch button. The electronic device 100 may receive a key input, and generate a key signal input related to a user setting and function control of the electronic device 100.

[0116] The motor 191 may generate a vibration prompt. The motor 191 may be configured to provide an incoming call vibration prompt and a touch vibration feedback. For example, touch operations performed on different applications (for example, photographing and audio playback) may correspond to different vibration feedback effects. The motor 191 may also correspond to different vibration feedback effects for touch operations performed on different areas of the display 194. Different application scenarios (for example, a time reminder, information receiving, an alarm clock, and a game) may also correspond to different vibration feedback effects. A touch vibration feedback effect may be further customized.

[0117] The indicator 192 may be an indicator light, and may be configured to indicate a charging status and a power change, or may be configured to indicate a message, a missed call, a notification, and the like.

[0118] The SIM card interface 195 is configured to connect to a SIM card. The SIM card may be inserted into the SIM card interface 195 or removed from the SIM card interface 195, to implement contact with or separation from the electronic device 100. The electronic device 100 may support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 may support a nano-SIM card, a micro-SIM card, a SIM card, and the like. A plurality of cards may be simultaneously inserted into a same SIM card interface 195. The plurality of cards may be of a same type or different types. The SIM card interface 195 is compatible with different types of SIM cards. The SIM card interface 195 is also compatible with an external storage card. The electronic device 100 interacts with a network through the SIM card, to implement functions such as conversation and data communication. In some embodiments, the electronic device 100 uses an embedded SIM (embedded SIM, eSIM) card. The eSIM card may be embedded into the electronic device 100, and cannot be separated from the electronic device 100.

[0119] It should be understood that a calling card in embodiments of this application includes but is not limited to a SIM card, an eSIM card, a universal subscriber identity module (universal subscriber identity module, USIM) , a universal integrated circuit card (universal integrated circuit card, UICC) , and the like.

[0120] A software system of the electronic device 100 may use a layered architecture, an event-driven architecture, a microkernel architecture, a micro service architecture, or a cloud architecture. In an embodiment of this application, an Android system with a layered architecture is used as an example to describe a software structure of the electronic device 100.

[0121] FIG. 2 is a block diagram of a software structure of the electronic device 100 according to an embodiment of this application. In a layered architecture, software is divided into several layers, and each layer has a clear role and task. The layers communicate with each other through a software interface. In some embodiments, the Android system is divided into four layers: an application layer, an application framework layer, an Android runtime (Android runtime) and system library, and a kernel layer from top to bottom. The application layer may include a series of application packages.

[0122] As shown in FIG. 2, the application packages may include applications such as Camera, Gallery, Calendar, Phone, Map, Navigation, WLAN, Bluetooth, Music, Videos, and Messages.

[0123] The application framework layer provides an application programming interface (application programming interface, API) and a programming framework for an application at the application layer. The application framework layer includes some predefined functions.

[0124] As shown in FIG. 2, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.

[0125] The window manager is configured to manage a window program. The window manager may obtain a size of the display, determine whether there is a status bar, perform screen locking, take a screenshot, and the like.

[0126] The content provider is configured to: store and obtain data, and enable the data to be accessed by an application. The data may include a video, an image, an audio, calls that are made and answered, a browsing history and bookmarks, an address book, and the like.

[0127] The view system includes visual controls such as a control for displaying a text and a control for displaying an image. The view system may be configured to construct an application. A display interface may include one or more views. For example, a display interface including an SMS message notification icon may include a text display view and an image display view.

[0128] The phone manager is configured to provide a communication function for the electronic device 100, for example, management of a call status (including answering, declining, or the like) .

[0129] The resource manager provides various resources such as a localized character string, an icon, an image, a layout file, and a video file for an application.

[0130] The notification manager enables an application to display notification information in a status bar, and may be configured to convey a notification message. The notification manager may automatically disappear after a short pause without requiring user interaction. For example, the notification manager is configured to: notify download completion, give a message notification, and the like. The notification manager may alternatively be a notification that appears in a top status bar of the system in a form of a graph or a scroll bar text, for example, a notification of an application that is run on the background, or may be a notification that appears on the screen in a form of a dialog window. For example, text information is displayed in the status bar, an announcement is given, the electronic device vibrates, or the indicator light blinks.

[0131] The Android runtime includes a kernel library and a virtual machine. The Android runtime is responsible for scheduling and management of the Android system.

[0132] The kernel library includes two parts: a function that needs to be invoked in Java language and a kernel library of Android.

[0133] The application layer and the application framework layer run on the virtual machine. The virtual machine executes Java files of the application layer and the application framework layer as binary files. The virtual machine is configured to implement functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0134] The system library may include a plurality of functional modules, for example, a surface manager (surface manager) , a media library (media library) , a three-dimensional graphics processing library (for example, OpenGL ES) , and a 2D graphics engine (for example, SGL) .

[0135] The surface manager is configured to: manage a display subsystem and provide fusion of 2D and 3D layers for a plurality of applications.

[0136] The media library supports playback and recording in a plurality of commonly used audio and video formats, and static image files. The media library may support a plurality of audio and video coding formats, for example, MPEG-4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0137] The three-dimensional graphics processing library is configured to implement three-dimensional graphics drawing, image rendering, composition, layer processing, and the like.

[0138] The 2D graphics engine is a drawing engine for 2D drawing.

[0139] The kernel layer is a layer between hardware and software. The kernel layer includes at least a display driver, a camera driver, an audio driver, and a sensor driver.

[0140] It should be understood that the technical solutions in embodiments of this application may be applied to systems such as Android, iOS, and Harmony.

[0141] Intent prediction in language input tasks has historically been a challenging endeavor, often necessitating explicit gestures, such as a separate button for input versus command to edit. To date, the most successful implementation of automatic intent prediction in text input and editing has been spelling auto-correction. However, with the significant advancements in language understanding enabled by large language models (LLMs) , this area of research has experienced a resurgence of interest and promise. Notably, there is now a possibility to somewhat automatically discern user intent without relying on explicit gestures in most cases, effectively mimicking human-like understanding.

[0142] Determining user intent can be inherently challenging due to the inherent ambiguity that often accompanies human communication. In fact, even human interlocutors may struggle to clearly discern the intent behind another user's statement, necessitating further clarification to resolve any uncertainty.

[0143] Voice typing is a technology for converting voice signals into digital text, which offers a faster and more convenient alternative to traditional typing.

[0144] Present voice typing solutions are limited in their ability to accurately understand human needs. For example, they are limited in their ability to facilitate voice-based editing of text, instead relying on manual typing and touch-and-type actions to correct errors. While some solutions do offer voice-based text correction capabilities, they often fall short in providing a seamless and intuitive way to handle ambiguous situations, where multiple correct interpretations of a statement are possible.

[0145] Voice typing solutions often fall short in providing seamless, interleaved voice command editing capabilities. While a limited number of solutions do allow for the editing of textual content using voice commands in a continuous, interleaved manner, these systems typically lack the ability to detect ambiguity or discern user intent, resulting in a rigid interpretation of commands that may not always align with the user's intention to use a different interpretation of a statement. This lack of flexibility and nuance can lead to frustration and inefficiency for users, highlighting the need for more advanced and sophisticated voice-based editing solutions.

[0146] The intent prediction can be combined with the voice typing to enhance, which allows accurate interpretation of user needs.

[0147] The proposed disclosure offers a solution for parsing and segmenting between input content and command (s) (e.g., editing command (s) , also referred to as text edit command (s) or edit command (s) in the following) within a single, uninterrupted statement that combines both input content and command in natural language, especially when made through voice-based input. In addition, this disclosure features a pipeline that facilitates the identification and resolution of ambiguous inputs, incorporating user feedback as needed to ensure accurate intent interpretation, thereby enabling a mechanism to deal with ambiguity in user intent without requiring additional clarification or rephrasing.

[0148] Following aspects are provided in this application:

[0149] segmentation of input content and command (s) (e.g., edit command (s) ) in a continuous interleaved user input statement;

[0150] allowing user to seamlessly perform command (s) (e.g., editing command (s) ) without having to indicate which part of spoken statement is meant as command (s) . Allows for simpler interaction that is easy for the user to learn;

[0151] method of detecting ambiguity in user statements i.e. for statements that can have multiple possible correct interpretations and executions;

[0152] allowing to take appropriate actions in case of multiple possible correct interpretations of statements;

[0153] resolving ambiguities in text editing through user interaction when provided with multiple choices;

[0154] allowing user to choose from multiple possible correct answers without having to rephrase or provide additional clarification on input statement;

[0155] The technical solutions provided in this application is applicable to devices that use microphone sensor for speech input and voice typing. These include smartphones, tablets, smartwatch, laptops, VR / AR headsets etc.

[0156] FIG. 3 is a schematic diagram of a method for voice typing provided according to an embodiment of this application. The method could be performed by the electronic device 100 as aforementioned.

[0157] The voice typing allows users to interact with applications running on the electronic device through spoken language. The spoken language can be converted into digital texts and input into the applications. The applications may support virtual assistants, smart home devices, car infotainment services, navigation services, and the like. For example, the user can use the voice typing to search routes from a starting point and an ending point for a navigation service.

[0158] At 310, the electronic device obtains digital text based on a voice input.

[0159] The voice input can be at least part of a contextual session between a user and the electronic device or can be at least part of an audio file. For example, the electronic device may include a microphone and the voice input can be directly input to the electronic device from the microphone by the user. For another example, the voice input can be an audio file and a user may use the audio file as the voice input to perform the method for voice typing. The audio file can be transmitted from other devices to the electronic device, downloaded from networks or pre-stored in storage units of the electronic device.

[0160] The voice input may be separated from context based on pauses of user speaking. The voice input may correspond to a period of time referred to as an utterance durance.

[0161] The voice input can be processed into the digital text using automatic speech recognition (ASR) , which is a technology that enables the conversion of spoken language into written text.

[0162] The digital text can be also referred to as a prompt. In some implementations, cursor (or placeholder token) can be incorporated into the digital text. For example, the user can add new content (new text) at the cursor location by manual typing, voice typing or other means such as copy and paste operation.

[0163] Additionally, the cursor (or its placeholder token) can be incorporated into the prompt, enabling the seamless addition of new text at the cursor location.

[0164] At 320, the electronic device processes the digital text via a first model to obtain an initial intent prediction result.

[0165] The initial intent prediction result indicates at least one set of tokens from the digital text and predicted intent (s) corresponding to each of the at least one set of tokens. The predicted intent (s) corresponding to each of the at least set of tokens may include command and input content.

[0166] At 330, the electronic device generates a first output, where the first output is generated based on processing the digital text as at least one command and / or input content based on the initial intent prediction result.

[0167] Based on the initial intent prediction result, the voice input can be processed as at least one command and / or input content.

[0168] Input content and commands (e.g., text edit commands) may co-exist in user statement (voice input) and could be interleaved.

[0169] The input content is information or data conveyed by the user via spoken words. It typically represents the main message or details that the user wants the system to record, process or respond to. The input content can be updated or modified based on one or more commands (e.g., via a text edit command) . The input content (or modified input content based on one or more commands) can be basis for further processing, it can be displayed on a screen of the electronic device and used as a search query or input of models, applications and the like. Optionally, the input content or the modified input content may not be displayed on screen of the electronic device (the electronic device may have a screen or does not have a screen) . In the following, the input content or the modified input content is displayed on the screen is shown as an example to introduce details of the embodiments of this application.

[0170] It is to be noted that the prompt is the digital text directly converted from the voice input, and the input content may be part of the prompt and can be processed based on part of the prompt (e.g., text editing commands) .

[0171] The commands may indicate operations that the user want the system to perform. The commands may be classified according to their functions, for example, the commands can be classified into: text operation commands (e.g., text edit command or text analysis command) and system commands.

[0172] The text edit command may be used to edit a prompt (from the voice input) , e.g., to edit the input content or to edit other commands in the prompt. There may be different kinds of text edit commands including but not limited to: replace commands, delete commands, insert commands, and so on.

[0173] The replace command can be used to replace argument (s) in a prompt, e.g., to replace arguments in the input content or other commands in the prompt. For example, for a voice input / digital text “I am going to eat at 5p.m. and sleep after that, change it to 6p.m. ” , “change it to 6p.m. ” corresponds to a replace command and is used to replace the argument “5p.m. ” in the input content “I am going to eat at 5p.m. and sleep after that” . The modified input content that can be displayed on the screen is “I am going to eat at 6p.m. and sleep after that” .

[0174] The delete command can be used to delete argument (s) in a prompt. For example, for digital text “I am going to have eggs, milk and bread for breakfast, no eggs” , “no eggs” corresponds to a delete command and is used to delete the argument “eggs” in the input content “I am going to have eggs, milk and bread for breakfast” . The modified input content which can be displayed on the screen is “I am going to have milk and bread for breakfast” .

[0175] The insert command can be used to insert arguments or reference information in a prompt. For example, for digital text “The result is obtained based on related reports of 2024, I mean the yearly economic report” , “I mean the yearly economic report” is used to insert or add information “the yearly economic report” to the input content “The result is obtained based on related reports of 2024” . The modified input content which can be displayed on the screen is “The result is obtained based on related reports (the yearly economic report) of 2024” .

[0176] Example (voice input) : “I am going to celebrate Halloween today I mean tomorrow and eat a lot of candies” .

[0177] Here the user is inputting content (the text) while also providing an edit command (correcting "today" to "tomorrow" ) .

[0178] It is to be noted that in the previous description, although the commands are shown to be used to edit the input content, some commands can be used to edit other commands. For example, for digital text “Set an alarm at 6 p.m., change it to 7 p.m. ” , “Set an alarm at 6 p.m. ” is a system command and “change it to 7p.m. ” is a replace command. The replace command is used to replace the argument “6p.m. ” in the system command to “7p.m. ” . The electronic device may set an alarm at 6 p.m. according to the digital text.

[0179] The text analysis command may be used to analyze a prompt, e.g., the input content in the prompt. There may be different kinds of text analysis commands including but not limited to: word count commands, grammar check commands, find commands and so on.

[0180] The word count command, the grammar check command and the find command can be used to perform word count, grammar check for text and find specified content in a prompt.

[0181] The system command may be used to control the electronic device or applications to perform certain operations. For example, the system command may include “send a message to A” or “start the camera app” and so on.

[0182] Some examples of such system commands used with input methods include “copy this text to a new document in docs app” or “send this as a message to <contact name> on app 1” etc.

[0183] User-uttered commands may be intended for text editing or manipulation, such as editing existing text, word count, or text reformatting. Alternatively, commands may be general system instructions, like "send" or "help" , which are often used in conjunction with text operations.

[0184] Therefore, the digital text can be processed as input content and / or at least one text operation command and / or at least one system command.

[0185] It is to be noted that the commands may also be classified differently and there may be other types of commands, e.g., there may be other subtypes of commands for the text edit commands, text analysis commands or system commands. In the following, the text edit commands are used as an example to introduce details of the embodiments provided in this application.

[0186] The proposed methods are primarily described within the context of text editing and analysis, but the underlying approach is equally applicable to system commands. Specifically, the techniques for segmenting and classifying commands, detecting uncertainty and resolving uncertainty can be seamlessly extended to system commands used during voice-based text input.

[0187] In this application, the first output in 330 may be an interface displayed on a screen of the electronic device. For example, the interface could include output text based on the digital text (voice input) , which may be processed via one or more text editing commands indicated by the digital text. For another example, if the voice input / digital text is used to trigger system operation, e.g., send a message to someone, the first output may be an indication notifying that corresponding operations will be performed.

[0188] In this application, an initial intent prediction process could be performed to obtain the initial intent prediction result in 320.

[0189] In some implementations, a first model or a set of templates could be employed to perform the initial intent prediction process. The first model can be a lightweight classification model.

[0190] The set of templates may correspond to different type of commands (also referred to as command templates in the following) .

[0191] The command templates can be predefined rules or patterns. If a specific set of tokens of the ASR results matches a specific command template, the set of tokens can be determined as a specific type of command; otherwise, the set of tokens may be determined as input content.

[0192] For example, a command template for a system command can be “play [artist name] ’s [song name] ” . For a set of tokens “play XX’s YY” , where XX is an artist name and YY is a song name, as the set of tokens match the command template, the set of tokens can be determined as a system command. For another example, a command template for a replace command can be “change [A] to [B] ” , and if a set of tokens match the command template, the set of tokens can be determined as a replace command.

[0193] The lightweight classification model is a kind of artificial intelligence model which can process human language. Similar to popular large language models (LLM) , the lightweight classification model can be based on deep learning architectures, it can be pretrained on large datasets to learn general language patterns or be fine-tuned for specific tasks. For example, the lightweight classification model can be fine-tuned for intent prediction tasks as described in this application. However, the lightweight classification model is relatively smaller in size (with lower processing capabilities) and can be installed on mobile devices.

[0194] The initial intent prediction result from the initial intent prediction process of the lightweight classification model may include sets of tokens corresponding to input content and / or commands. As aforementioned in previous description about different types of the commands, the command can be classified into text operation commands and system commands. The initial intent prediction result from the lightweight classification model can be shown as FIG. 4.

[0195] FIG. 4 is a schematic diagram showing an initial intent prediction result according to an embodiment of this application. As shown in FIG. 4, the ASR result (prompt based on the voice input) can be used as an input to the lightweight classification model, and input content and / or text operation commands (including text editing commands and / or text analysis commands) and / or system command (also referred to as the system action commands) can be obtained based on the lightweight classification model.

[0196] For obtaining the initial intent prediction result, the prompt needs to be divided into smaller units, e.g., tokens.

[0197] Most language processing models operate on smaller, more manageable units of text. Tokenization is the process of dividing text into distinct, individual elements called tokens, which can be words, phrases, or even smaller units. This process involves applying a set of predefined rules to break down text into a fixed number of tokens, each of which is represented by a unique numerical vector. By converting text into tokens, tokenization enables the transformation of text into a format that machine learning models can understand, allowing them to process and analyze language effectively. Tokens can be thought of as words in a language but instead of being unlimited like possible words, the number of tokens is fixed.

[0198] Intent segmentation (the initial intent prediction process) involves labeling each token with its purpose, distinguishing between input content and commands.

[0199] In some implementations, a first set of tokens and / or a second set of tokens may be obtained from the prompt based on the initial intent prediction process (i.e., the at least one set of tokens in 320 includes the first set of tokens and / or the second set of tokens) . The predicted intent corresponding to the first set of tokens is command and / or the predicted intent corresponding to the second set of tokens is input content. The digital text can be processed as at least one command and / or input content based on the initial intent prediction result.

[0200] In the following, how the digital text can be processed as at least one command and / or input content based on the initial intent prediction result will be detailed.

[0201] In some implementations, by the initial intent prediction process, the lightweight classification model may also output intent prediction scores for each token. The initial intent prediction result is obtained based on the intent prediction scores.

[0202] For example, the intent prediction scores can be used to indicate a probability of a token corresponding to a command. In this scenario, a first confidence score of a set of tokens can be calculated based on the intent prediction scores to indicate a probability of the set of tokens being a command, e.g., by calculate an average / weighted average, sum / weighted sum, maximum, minimum of the intent prediction scores of the set of tokens. The intent prediction score or the first confidence score can be in forms of percent, level, etc. In this application, a confidence score of the first set of tokens may indicate a probability of the first set of tokens being a command is high, e.g., equal to or higher than a first threshold; a first confidence score of the second set of tokens may indicate a probability of the second set of tokens being a command is low (i.e., high possibility to be input content) , e.g., equal to or lower than a third threshold. The third threshold is lower than the first threshold.

[0203] In some embodiments, a third set of tokens may also be obtained by the lightweight classification model, a first confidence score of the third set of tokens may indicate a probability of the third set of tokens corresponding to a command is higher than the third threshold and lower than the first threshold.

[0204] Optionally, the intent prediction scores can be used to indicate a probability of a token corresponding to input content, a first confidence score of a set of tokens can also be calculated to indicate a probability of the set of tokens being / corresponding to input content. To introduce details about the method for voice typing, the intent prediction scores indicating a probability of a token being a command is described, details of the intent prediction scores indicating a probability of a token corresponding input content can be deduced from the present description.

[0205] In some implementations, a first confidence score of a set of tokens may also include all intent prediction scores for the tokens in the set of tokens, for example, all intent prediction scores for a set of tokens can be compared with a threshold and predicted intent of the set of tokens can be determined.

[0206] In some implementations, based on the intent predication scores, tokens are grouped into different sets and can be predicted as input content or commands. In some implementations, consecutive tokens with intent prediction scores equal to or higher than a threshold (e.g., the first threshold) may be grouped into a set and predicted as a command and consecutive tokens with intent prediction scores equal to or lower than a threshold (e.g., the third threshold) may be grouped into a set and predicted as input content. For example, for a phrase with 6 tokens: ABCDEF with intent prediction scores being 10%, 30%, 70%, 80%, 80%, 75%, respectively, as the intent prediction scores for the tokens CDEF are high (e.g., equal to or higher than 60%) , CDEF can be grouped into a set and can be predicted to be a command and the intent prediction scores for the tokens AB are low (e.g., equal to or lower than 30%) , AB can be grouped into a set and predicted to be input content.

[0207] In this scenario, if intent prediction scores corresponding to all tokens in the set of tokens are equal to or higher than the first threshold (or equal to or lower than the third threshold) , the first confidence score of the set of tokens can be considered as equal to or higher than the first threshold (or equal to or lower than the third threshold) .

[0208] Optionally, for the grouping or determination of command, other parameters may be considered (e.g., language characteristics, sentence structure or grammatical analysis) and not all intent prediction scores of tokens that are grouped to be a set may be equal to or higher than the first threshold (e.g., few tokens may have relatively low intent prediction scores) .

[0209] Optionally, the lightweight classification model may directly decide that a specific set of tokens being a command or input content without the need of outputting the intent prediction scores or the first confidence score.

[0210] For different sets of tokens with different intent prediction scores, processes on corresponding sets of tokens are different. It is to be noted that apart from the first set of tokens, the second set of tokens and the third set of tokens, other sets of tokens may be obtained based on the voice input. In the following, the processes on the first set of tokens, the second set of tokens and the third set of tokens will be detailed and other repetition of different sets of tokens can be deduced from the present description.

[0211] Based on the first confidence score of a set of tokens as aforementioned and a second confidence score of the set of tokens (which will be detailed in the following) , there may be ambiguity in user intent. The method provided in this application can be used to solve ambiguity involved in the voice input.

[0212] The problem formulation involves assigning a classification score (the first confidence score or the second confidence scores corresponding to each label which will be introduced in the following) to each token or word prediction, which serves as a measure of the model's confidence or certainty in its predictions. This score is utilized to detect perceived ambiguity or uncertainty in user statements, which can arise from various sources, including inherent ambiguity in the user's input or errors in speech recognition that result in a less meaningful statement. In cases where the model encounters ambiguity, it assigns lower scores to one or both of intent classification and argument classification.

[0213] The statement might fall under any of the following categories.

[0214] Case 1: low confidence in input (input content) vs command (e.g., edit intent) . Case 1 will be introduced by the example of the third set of tokens.

[0215] Case 2: high confidence edit intent, high confidence in predicted function arguments but one or more arguments are missing.

[0216] case 3: high confidence in edit intent but low confidence in classified function arguments.

[0217] In this application, as the probability for the first set of tokens being a command is high (e.g., the first confidence score is high) , the first set of tokens will be processed as a command in priority.

[0218] Command arguments corresponding to the first set of tokens can be obtained via a second model via an argument extraction process. The voice input can be processed as at least one command and / or input content based on the command arguments, that is, the first output is generated based on the command arguments.

[0219] The second model may also be installed on mobile devices like the first model and have lower processing capabilities than cloud models such as LLM.

[0220] The second model can be a same model as that used to perform the initial intent prediction process, e.g., the lightweight classification model. The argument extraction process and the initial intent prediction process may be performed simultaneously, for example, in order to improve the accuracy of the initial intent prediction result, the argument extraction process can also be performed.

[0221] Optionally, the second model can be a different model (e.g., referred to as a lightweight token classification model) from the model used to perform the initial intent prediction process. The two models may have a similar structure and can be pre-trained and / or fine-tuned differently for performing the initial intent prediction process and the argument extraction process.

[0222] In some embodiments, the arguments extraction process can be performed based on the set of templates. For example, for digital text “play XX’s YY” , the “XX” is labeled as an artist name and the “YY” is labeled as a song name. For another example, for digital text “change A to B” , the “A” is labeled as replace_original, and the “B” is labeled as replace_modified (the two labels will be introduced in the following) .

[0223] During the argument extraction process, tokens in the prompt can be classified into different argument classes / sub classes, and may correspond to different command classes / sub-classes, which may be represented by different class labels.

[0224] FIG. 5 is a schematic diagram of an example architecture for labeling tokens with different class labels.

[0225] As shown in FIG. 5, for labeling tokens with different class labels (by the argument extraction process) , a second model could be used. In some implementations, the second model can be a neural network with three layers: a bottom layer, a middle layer and a top layer. The bottom layer is used to input token embeddings. The token embeddings are then passed through the middle layer which may include one or more transformer blocks, the output of the last transformer block (final hidden state) is passed through the top layer, which may be a classification head.

[0226] The class labels output by the top layer may include one or more of: replace_original, replace_modified, delete, add_input, add_reference_before, add_reference_after, command, text. The class labels: replace_original, replace_modified, delete, add_input, add_reference_before, add_reference_after may be labels for command arguments (referred to as argument labels) . Apart from the argument labels, labels for intent (referred to as intent labels) may also be output by the top layer, for example, the intent labels may include command and text (input content as aforementioned) .

[0227] The class labels “replace_original” and the “replace_modified” are labels for arguments of a replace command. For example, for a voice input “I am going to eat at 5p.m. and sleep after that, change it to 6p.m. ” , the tokens “5p.m. ” can be labeled as “replace_original” and the tokens “6p.m. ” can be labeled as “replace_modified” .

[0228] The class label “delete” is a label for a delete command and can refer to tokens that need to be deleted from the input content. For example, for a voice input “I am going to have eggs, milk and bread for breakfast, no eggs” , the “eggs” in the “I am going to have eggs, milk and bread for breakfast” can be labeled as “delete” .

[0229] The class labels “add_input, add_reference_before and add_reference_after” may be labels for an insert command, the label “add_input” can be used to refer to tokens that need to be inserted to the input content, the labels “add_reference_before and the add_reference_after” can be used to refer to tokens that need to be inserted to the input content at a specific position, e.g., before or after a specific word.

[0230] The model assigns labels corresponding to each possible argument required for performing a successful edit.

[0231] Furthermore, command tokens can be categorized into more specific subtypes, such as text editing commands (which will be the primary focus of subsequent sections) , text analysis commands, or system commands that perform actions related to text, allowing for a more nuanced understanding of the intent behind each token.

[0232] Each token from the existing text (input content) and the given command can be categorized into specific classes or sub-classes that correspond to the arguments required to execute the command as intended by the user. For instance, in the case of a text editing command (e.g., a replace command) , such as “change A to B in the text” where A and B are strings, the tokens from substring A are classified as the 'reference string' argument (replace_original) , while the tokens from string B are classified as the 'replacement string' argument (replace modified) .

[0233] It is to be noted there may be other labels for other types of commands, which will not be detailed here.

[0234] For each type of command, there may be essential arguments for the command to be operated. For example, for the replace command, the argument labeled as “replace_original” and “replace_modified” may be essential for performing the replace command.

[0235] These arguments are essential for successfully executing the edit operation. Similarly, for system commands like sending a message to a contact, the command is not only classified as the correct command type, but the tokens for the contact's name and the app name are also identified and classified as required arguments, enabling the command to be executed accurately.

[0236] For example, for a replacement command, for the text editing to be successful, following assumptions need to be met:

[0237] the replacement text (desired output text) is present in the prompt;

[0238] the reference text is present in the prompt.

[0239] Notably, the input prompt (also referred to as a prompt input to the model) is constructed by combining the text and command, in any order. Under these assumptions, editing operations can be effectively formulated as a multi-class / multi-label token classification problem, allowing for successful completion. This approach imposes no constraints on the command's syntax or structure, nor does it require the reference text to be explicitly stated in the command. For instance, implicit reference to the text can be inferred from the command, as illustrated in the example below.

[0240] Example (voice input) : i am going to eat at 5 p.m. and sleep after that. change it to 6p.m.

[0241] Here, the reference text (also referred to as replace_original, 5 p.m. ) and replacement text (also referred to as replace_modified, 6 p.m. ) are both present in prompt. This is allowed even though command does not clearly specify the reference.

[0242] Furthermore, this method can be readily extended to accommodate re-dictation scenarios, offering flexibility and versatility in editing tasks.

[0243] After the argument extracting process, the electronic device may decide if there is a lack of command arguments corresponding to the first set of tokens.

[0244] If there is a lack of command arguments corresponding to the first set of tokens, the electronic device may elongate an utterance durance for obtaining longer digital text.

[0245] As aforementioned, the utterance durance for the voice input (also the digital text) can be based on pauses of user speaking. In some cases, electronic device may separate the voice input (digital text) from the context inappropriately, resulting in an incomplete statement in the digital text. If there is a lack of command arguments corresponding to a set of tokens, it is possible that the remaining / missed command arguments are present beside / neighboring to the digital text, e.g., before the starting position of the digital text (past utterance relative to the digital text) or after the ending position of the digital text (future utterance relative to the digital text) in the large piece of text. The electronic device may elongate the utterance duration for the digital text to obtain longer digital text, i.e., to obtain past utterance or future utterance relative to the present digital text. For example, for the large piece of text “I am going to eat at 5p.m. and sleep after that, change it to 6p.m. ” , “and sleep after that, change it to 6p.m” may be obtained as the original digital text, the electronic device may elongate the utterance duration to obtain a longer digital text including the past utterance “I am going to eat at 5p.m. ” .

[0246] Optionally, if there is still lack of command arguments after elongating utterance durance, the electronic device may further elongate the utterance durance. If certain condition is fulfilled (e.g., times of elongating the utterance durance is over a threshold or the time utterance is over a time threshold) and there is still a lack of command arguments, the electronic device may invoke other complex models such as the LLM to perform the argument extraction process.

[0247] In some implementations, the electronic device may remind the user to finish the command if there is a lack of arguments corresponding to the first set of tokens. For example, for a voice input “please transmit this message to …” , the electronic device may display a reminder on the screen or remind the user via speaker, e.g., a reminder “please confirm the recipient” could be displayed on the screen or delivered via the speaker.

[0248] In addition to the labels, for each label, there may be a corresponding confidence score which is higher than threshold, where the confidence score indicates a confidence in the label. The confidence scores can be output by the top layer as in FIG. 5. In this scenario, the confidence scores for the labels “command” and “text” may be the intent prediction scores as aforementioned. For a specific token in a command, it may correspond to two labels and two corresponding confidence scores, where one is the label “command” and the other is the detailed command argument, and for a specific token in input content, there may be one label “text” .

[0249] Beyond predicting classes (intent classes and argument classes) , the model (the second model) can also be designed to output a confidence score that indicates the model's certainty in its predicted classification for each token. In many cases, this is a default output of classification models in current machine learning frameworks, where a trained multi-label classification model produces a probability value between 0 and 1, typically generated by a sigmoid function. This value can be interpreted as a measure of the model's confidence in its prediction. Note that this approach is just one method for obtaining a classification score (confidence score) , and alternative approaches are also possible.

[0250] If there is not a lack of command arguments corresponding to the first set of tokens, the electronic device may decide if a second confidence score of the first set of tokens is higher than a second threshold, where the second confidence score indicate a confidence in the command arguments obtained via the second model for the first set of tokens.

[0251] For each set of tokens which is predicted to be a command based on the initial intent prediction result, the second confidence score of the set of tokens can be obtained based on the confidence scores, which indicates a confidence in the command arguments for the set of tokens. The second confidence score can be represented by an average (or weighted average) , sum (or weighted sum) , minimum, or maximum of all the confidence scores.

[0252] In some implementations, the operation of elongating the utterance durance can be performed if the second confidence score for the extracted command arguments is higher than a second threshold (case 2 as aforementioned) .

[0253] Case 2: high confidence in command (e.g., edit intent) , high confidence in predicted function arguments but one or more arguments are missing.

[0254] Solution: use longer context. in this case, it is likely that some arguments are missing from the input prompt. The context length is increased for previously present input text and waiting for user to finish command or speak the remaining arguments.

[0255] Solution: prompt user to complete the command.

[0256] FIG. 6 is a schematic diagram showing handling case when missing one or more edit arguments.

[0257] As shown in FIG. 6, at step 1, a voice input is input to an electronic device to obtain ASR result. The ASR result can be obtained based on ASR technology. The ASR result of the voice input is then input to a first model (the lightweight classification model) at step 11. If the first confidence score of the first set of tokens is equal to or higher than the first threshold (β) , a second model (which may be a same model as the first model) is used to extract command arguments corresponding to the first set of tokens from the digital text of the ASR result at step 13. If there is a lack of arguments for the first set of tokens, longer context is used to obtain the mission arguments at step 5. Other steps in FIG. 6 will be detailed with reference to the following figures.

[0258] As aforementioned, the second confidence score can be obtained for each set of tokens. If the second confidence score of the first set of tokens is higher than the second threshold, the electronic device can process the first set of tokens as a command and the first output can be generated based on a result of processing the first set of tokens as a command.

[0259] Optionally, the second confidence score of a set of tokens can be all the confidence scores of the tokens in the set of tokens. For example, the electronic can use each confidence score in the first set of tokens and determine if all the confidence scores are higher than the second threshold.

[0260] In the scenario where the second confidence score of the first set of tokens is higher than the second threshold, the first set of tokens is highly possible to be a command, the electronic device may perform text operation or system operation as the command indicated.

[0261] If the second confidence score of the first set of tokens is equal to or lower than the second threshold, the first set of tokens may not be a command, the electronic device may extract command arguments corresponding to the first set of tokens via a third model. The third model has a processing capability better than a processing capability of the second model and the voice input is processed based on an extraction result via the third model.

[0262] In this scenario, the second model may have inefficient capability to decide the first set of tokens as a command. The third model has higher processing capability than the second model.

[0263] The third model may be a cloud model such as the LLM. The electronic may employ the model to do argument extraction.

[0264] The argument extraction by the third model may be successful or unsuccessful. For example, if there is a lack of command arguments based on a result of the extraction of the third model, the command extraction by the third model may be considered as unsuccessful.

[0265] If the argument extraction by the third model is successful, the electronic device may process the first set of tokens as a command.

[0266] If the argument extraction performed by the third model is unsuccessful, there is ambiguity related to the first set of tokens. The electronic device may provide a user with at least two options and the first set of tokens are processed differently for the at least two options.

[0267] In some implementations, the at least two options may include options for processing the first set of tokens as input content and a command. In some other implementations, the at least two options may include options for processing the first set of tokens as different commands.

[0268] For example, for a voice input “I am going to eat at 5p.m. and sleep at 9p.m., change it to 8p.m. ” , it may be difficult to decide whether to output “I am going to eat at 5p.m. and sleep at 8p.m. ” or “I am going to eat at 8p.m. and sleep at 9p.m. ” . The electronic device may display the two options on the screen for the user to select. The user may select one option from the at least two option and the selected option may be displayed in a more conspicuous way. For example, the selected option may be zoomed in to be displayed on the screen. In this scenario, the first output can be in the form of an interface and the interface is based on the provided options and user selection.

[0269] In some implementations, there may be default options for the at least two options.

[0270] For example, for a voice input “I am going to eat at 5p.m. and sleep at 9p.m., change it to 8p.m. ” , the electronic device may display a default option which processes the voice input as input content, e.g., the electronic device may display “I am going to eat at 8p.m. and sleep at 9p.m” on the screen. The electronic device may also display an alternative option which process the voice input as a command, e.g., the electronic device may display “I am going to eat at 5p.m. and sleep at 8p.m” .

[0271] For another example, for a voice input “what is the weather like today” , the electronic device may display a default option which processes the voice input as input content, e.g., the electronic device may display “what is the weather like today” on the screen. The electronic device may also display an alternative option which process the voice input as a command of weather inquiry, e.g. the electronic device may display “sunny” . The default option may be displayed in a more conspicuous way than the alternative option, e.g., with a bigger size.

[0272] If the argument extraction performed by the third model is unsuccessful, the electronic device may also process the first set of tokens as input content, then the first output may be generated based on processing the first set of tokens as input content.

[0273] Optionally, if the second confidence score of the first set of tokens is equal to or lower than the second threshold, the electronic device may also provide at least one option for user to choose and display as the user’s selection from the at least one option.

[0274] Case 3: High confidence in Edit intent but low confidence in classified function arguments.

[0275] Solution: invoke LLM

[0276] Solution: provide multiple edit options to user to make a choice.

[0277] FIG. 7 is a schematic diagram showing handling case of low confidence in classified arguments.

[0278] As shown in FIG. 7, if a second confidence score of a set of tokens (e.g., the first set of tokens) is equal to or lower than a second threshold (e.g., α2) , the electronic device may invoke a third model at step 14, which has a better processing capability than the second model. Optionally, the electronic device may also provide the user with at least one option and process as user’s selection. If the second confidence score of the set of tokens is higher than the second threshold, the electronic device may process the set of tokens as a command, for example, perform string edit processing, string analysis operation, or string reformat operation.

[0279] The details of FIG. 7 have already been covered in previous description and will not be repeated for brevity.

[0280] As aforementioned, the first set of tokens are processed as a command if the second confidence score of the first set of tokens is higher than the second threshold or the argument extraction by the third model (e.g., LLM) is successful.

[0281] The electronic device may perform operations as indicated by the command corresponding to the first set of tokens. For example, if the first set of tokens correspond to a text operation command, the electronic device may perform string edit processing, string analysis, or string reformat operation according to the text operation command.

[0282] In some implementations, the processing of the first set of tokens as a command may be unsuccessful, the first set of tokens will be reprocessed as input content, the processing result can be used to generate the first output. For example, text including the first set of tokens (a first output) can be displayed on the screen.

[0283] However, the first set of tokens may be incorrectly predicted.

[0284] In some implementations, a first user input from the user may be input to the electronic device and the electronic device receives the first user input, where the first user input is used to cancel the first output. The electronic device may reprocess the first set of tokens differently and display a second output on the screen, the first set of tokens are processed in a different way corresponding to the second output from that corresponding to the first output. In this scenario, the user may re-speak the voice input again and the electronic device may process the first set of tokens differently.

[0285] Optionally, the electronic device may also provide at least one option corresponding to at least one alternative output and the first set of tokens are processed in a different way corresponding to the at least one alternative output from that corresponding to the first output. For example, the first set of tokens is processed as a command corresponding to the first output, the at least one option may correspond to processing the first set of tokens as input content or other types of commands.

[0286] The electronic device may receive a second user input from the user, which is a selection operation of one of the at least one choice. The electronic device can generate a second output corresponding to the second user input.

[0287] In this application, a probability of the second set of tokens being a command is low (e.g., the first confidence score of the second set of tokens is low) and can be processed as input content in priority.

[0288] The electronic device may display text including the second set of tokens. For example, the electronic device may generate a first output and display it on the screen based on processing the second set of tokens as input content.

[0289] For the second set of tokens, a first user input may also be provided by the user to cancel the first output, which may trigger the electronic device to process the second set of tokens differently and generate a second output.

[0290] Optionally, the electronic device may also provide at least one alternative option for user to choose when generating the first output and generate a second output based on the user’s choice (via a second user input) . Details can refer to introduction of the first set of tokens and will not be repeated for brevity.

[0291] In this application, the first confidence score of the third set of tokens being a command is higher than the third threshold and lower than the first threshold. In other words, either a possibility of the third set of tokens being a command or input content is low.

[0292] In some implementations, the electronic device may then use a third model (e.g., LLM) to extract command arguments corresponding to the third set of tokens.

[0293] In some other implementations, the electronic device may provide at least two options and the third set of tokens are processed differently for the at least two options. For example, the at least two options may include an option corresponding to processing the third set of tokens as a command and another option corresponding to processing the third set of tokens as input content.

[0294] Case 1: low confidence in input vs edit intent.

[0295] Solution 1: generate outputs with both options (processing as input and command) and let the user choose.

[0296] Solution 2: invoke LLM

[0297] The argument extraction by the third model may be successful or unsuccessful. For example, if there is a lack of command arguments based on the result of the extraction of the third model, the command extraction by the third model may be considered as unsuccessful.

[0298] If the argument extraction by the third model is successful, the electronic device may process the third set of tokens as a command, then the first output is generated based on a result of processing the third set of tokens as a command.

[0299] In some implementations, the result of processing of the third set of tokens as a command is unsuccessful, the third set of tokens will be reprocessed as input content. The electronic may display a first output based on processing the third set of tokens as input content.

[0300] In other implementations, the processing of the third set of tokens as a command is successful, the electronic may display a first output based on processing the third set of tokens as a command.

[0301] For the third set of tokens, a first user input may also be provided by the user to cancel the first output, which may trigger the electronic device to process the third set of tokens differently and display a second output.

[0302] Optionally, the electronic device may also provide at least one alternative output for user to choose when displaying the first output and display a second output based on the user’s choice (via a second user input) . Details can refer to introduction of the first set of tokens and will not be repeated here for brevity.

[0303] If the argument extraction performed by the third model is unsuccessful, the electronic device may provide a user with at least two options and the third set of tokens are processed differently for the at least two options or the electronic device may directly process the third set of tokens as input content, the first output is generated based on the at least two options or processing the third set of tokens as input content.

[0304] FIG. 8 is a schematic diagram of showing handling case when language model is less confident in classification between input content and edit command.

[0305] As shown in FIG. 8, if a first confidence score of the third set of tokens is higher than the third threshold (α) and lower than the first threshold (β) (i.e., a probability of the third set of tokens being a command or input content is low) , the electronic device may either invoke the third model (LLM) or provide the user with options. The processing of the third set of tokens can be determined based on extraction result of the third model or the user selection.

[0306] As aforementioned by processing of the first set of tokens, the second set of tokens and the third set of tokens, user feedback may be demanded in cases of model uncertainty as mentioned above, and is recorded in following forms:

[0307] (1) choosing one of provided options if the electronic device provides at least one option and / or alternatives for user selection.

[0308] (2) manually correcting an error

[0309] For example, user can add new content or revise present content by manual typing at the cursor location.

[0310] (3) undo / cut followed by reiteration

[0311] For example, the user may cancel present processing result and re-speak same voice input and the electronic device may process the voice input in a different way.

[0312] For each of the feedback scenarios, following actions can be taken:

[0313] Alternative output is chosen.

[0314] The interaction is recorded and used to update / finetune weights for user-calibrated intent recognition.

[0315] Addition of new patterns for matching for future interactions.

[0316] FIG. 9 is a schematic diagram showing user decision-making support for ambiguous options.

[0317] FIG. 9 is one case as aforementioned, where the first threshold and the third threshold can be considered as the same. As aforementioned, a probability of a set of tokens being input content can also be used to perform the initial intent prediction process. In fact, for a specific set of tokens, a sum of a probability of the set of tokens being a command (denoted as Pcommand) and a probability of the set of tokens being input content (denoted as Pinput content) is 1 (100%) . For example, if Pcommand is 30%, Pinput content is 70%. If we set the first threshold and the third threshold as 50%, and process a set of tokens as a command when Pinput content = Pcommand, steps in FIG. 9 can be obtained.

[0318] As shown in FIG. 9, at 901, classification scores can be obtained, the classification scores can be the intent predication scores as aforementioned or other kinds of scores that could be used to derive Pcommand and Pcommand.

[0319] For a specific set of tokens, if a possibility of the set of tokens being a command is lower than a possibility of the set of tokens being input content, the set of tokens can be processed as input content in priority and can be shown on screen at 902. The electronic device may also provide user with alternative options, which corresponds to process the set of tokens differently at 903. At 904, the user may select one alternative option (make a choice) . The processing of the set of tokens can be used as a sample to update model weights at 905. For example, the user may have a specific way of speaking which corresponds to a command, the processing for the set of tokens can be used to update model weights (as an example, it can be used to update the model that is used to output the classification score) .

[0320] If a possibility of the set of tokens being a command is equal to or higher than a possibility of the set of tokens being input content, the set of tokens can be processed as a command in priority at 906. Then the electronic device may display the processing result on screen. The user may use undo / cut command to cancel the existing output at 907 and may re-speak same input action (voice input) at 908 if the user finds the displayed processing result is undesirable. The set of tokens may be processed as input content at 909.

[0321] It is to be noted that although the previous description mainly describes commands are processed in one way, if there are more than one commands in the digital text, execution order of the commands may influence the result to be displayed on the screen or the result of execution. For example, for digital text “I am going to eat at 5p.m. and sleep after that, change it to 6p.m., change the 6 p.m. to 7p.m. ” , “change the 6 p.m. to 7p.m. ” and “change it to 6p.m. ” corresponds to replace commands. The replace command 1 “change the 6 p.m. to 7p.m. ” is used to replace the argument “6p.m. ” to “7p.m. ” in the replace command 2 “change it to 6p.m. ” . The replace command 2 is changed to “change it to 7p.m. ” based on the replace command 1. Then the revised replace command 2 can be used to change the “5p.m. ” in the input content “I am going to eat at 5p.m. and sleep after that” . The modified input content that can be displayed on the screen is “I am going to eat at 7p.m. and sleep after that” . However, it is possible that replace command 1 is used to replace the argument “5p.m. ” to “6p.m. ” . The replace command may be determined to be a replace command at first but after the execution of the replace command 1, the execution of the replace command 2 will be failed and the replace command 2 will be determined as input command in the end. In this scenario, the modified input content that can be displayed on the screen is “I am going to eat at 6p.m. and sleep after that, change the 6 p.m. to 7p.m. ” .

[0322] FIG. 10 is a schematic diagram showing complete pipeline for intent segmentation, processing edit instruction, ambiguity detection and user decision-making support for ambiguous options.

[0323] At step 1, voice input is input to an electronic device and is transcribed from ASR to obtain ASR results.

[0324] At step 2, the ASR results is matched with a set of command templates.

[0325] The command templates can be predefined rules or patterns. If a specific set of tokens of the ASR results matches a specific command template, corresponding operations or instruction could be performed; otherwise, the set of tokens may be processed as input content or handled in other ways.

[0326] For example, a command template for a system command can be “play [artist name] ’s [song name] . For a voice input “play XX’s YY” , where XX is an artist name and YY is a song name. As the voice input matches the command template, the voice input can be processed, i.e., the YY from XX can be played.

[0327] At step 3, if the matching between the ASR results and a command template is successful, command arguments is extracted based on the command template.

[0328] At step 11, if the matching between the ASR results and the set of command templates is failed, the voice input is input into a lightweight classification model (first model as aforementioned) for an initial intent prediction process.

[0329] At step 12, a first confidence score which indicates a probability of the voice input being a command is output by the lightweight classification model, the first confidence score is compared with several thresholds (α, β) to obtain an initial intent prediction result.

[0330] At step 13, if the probability of the voice input being a command is equal to or higher than β, the voice input could be processed as a command in priority. Command arguments corresponding to the voice input can be extracted by a lightweight token classification model (the second model as aforementioned) .

[0331] At step 4, after the extraction by the lightweight token classification model (as in step 13) or based on the command template (as in step 3) , whether there is a lack of command arguments for a specific type of command or not is determined.

[0332] At step 5, if there is a lack of command arguments, longer context is used to obtain the remaining arguments.

[0333] At step 6, if there is not a lack of command arguments, a confidence in the arguments is assessed (based on the second confidence score as aforementioned) .

[0334] At steps 14 and 16, if the confidence in the arguments is equal to or lower than α2, a model with higher capability (e.g., LLM) than the lightweight token classification model is invoked and used to extract arguments from the voice input.

[0335] At step 7, if the confidence in the arguments is higher than α2 or extraction by the LLM (step 16) is successful, corresponding operation as a command (e.g., string edit processing, string analysis operation or string reformat operation) can be performed.

[0336] If the probability of the voice input being a command is higher than α and lower than β, LLM could be used to extract arguments as in steps 14 and 16.

[0337] At step 15, optionally, if the probability of the voice input being a command is higher than α and lower than β, options could be provided to users, for different options, the voice input may be processed differently.

[0338] At step 17, user makes a choice.

[0339] At step 21, whether the operations as a command of step 7 are successful is judged.

[0340] At step 22, if operations of step 7 are failed, or extraction by the LLM is failed, or if the possibility of the voice input being a command is lower than α, the voice input is processed as input content.

[0341] In some embodiments, if the user does not make a choice as in step 17, the voice input is processed as input content in default after step 15. Optionally, the voice input is processed as input content in default if the probability of the voice input being a command is higher than α and lower than β.

[0342] At step 23, although the voice input is processed as input content as step 22, if alternative options are available, options could be provided for user selection.

[0343] At step 24, the user may make choice from the alternative options.

[0344] At steps 32, 33 and 34, if the processing as a command as step 7 is successful, although the confidence in arguments judged by step 6 is higher than α2, there is possibility that the voice input should not be processed as a command. In this scenario, the user may use undo / cut command to cancel existing output. The user then re-speaks same input action (voice input) and the voice input can be processed as input content.

[0345] At steps 42 and 43, if the processing as a command as step 7 is successful, but in a case where the confidence in the arguments judged by step 6 is equal to or lower than α2, there is possibility that the voice input should not be processed as a command. In this scenario, the electronic device may provide options for user selection and the options may corresponding to processing the voice input in a different way from step 7. The user may make a choice from the options.

[0346] The pipeline presented above performs intent segmentation and processes the text for argument extraction after being transcribed from ASR (automatic speech recognition) systems. However, a similar approach can be applied to a multi-modal end-to-end model that performs speech transcription and token classification. Such a model also uses existing text as its input as it is required for the contextual information.

[0347] The proposed pipeline is not limited to English and Chinese and can be extended to other languages as well.

[0348] In this disclosure, following aspect are provided:

[0349] 1. Supporting segmentation between input content and commands (e.g., editing commands) within a single, uninterrupted statement that combines both content and command in natural language.

[0350] 2. Pipeline to detect ambiguous statements i.e. the statements that can have multiple possible correct interpretations either between input content and command (e.g., edit command) or multiple possible commands (e.g., edit commands) .

[0351] 3. Pipeline to resolve ambiguous statements without need of further clarification or rephrasing.

[0352] FIG. 11 is a schematic diagram of a hardware structure of an apparatus 10 according to an embodiment of this application. An apparatus 10 shown in FIG. 11 includes a memory 11, a processor 12, a communications interface 13, and a bus 14. The memory 11, the processor 12, and the communications interface 13 implement communication connection to each other through the bus 14.

[0353] The memory 11 may store a program. When the program stored in the memory 11 is executed by the processor 12, the processor 12 is configured to perform steps of the foregoing method embodiments.

[0354] The processor 12 may use a general-purpose CPU, a microprocessor, an ASIC, a GPU, or one or more integrated circuits, and is configured to execute a related program, to perform the foregoing method embodiments.

[0355] The processor 12 may alternatively be an integrated circuit chip and has a signal processing capability. In an implementation process, steps of the foregoing method embodiments may be accomplished by using an integrated logic circuit of hardware in the processor 12 or instructions in a form of software.

[0356] It should be noted that although the apparatus 10 show only the memory, the processor, and the communications interface, in a specific implementation process, a person skilled in the art should understand that the apparatuses 10 may further include another component necessary for normal operation. In addition, based on a specific requirement, a. person skilled in the art should understand that the apparatuses 10 may further include hardware components for implementing other additional functions. In addition, a person skilled in the art should understand that the apparatuses 10 may include only components required for implementing the embodiments of this application, and there is need to include all components shown in FIG. 11.

[0357] It will be understood that a person skilled in the art may make various modifications and variations to this application without departing from the scope of this application. This application is intended to cover these modifications and variations of this application provided that they fall within the scope of protection defined by the following claims and their equivalent technologies.

[0358] A person skilled in the art will understand that embodiments of this application may be provided as a method, an apparatus (or system) , a computer-readable storage medium, or a computer program product. Therefore, this application may use a form of a hardware-only embodiment, a software-only embodiment, or an embodiment with a combination of software and hardware. Moreover, this application may use a form of a computer program product that is implemented on one or more computer-usable storage media (including but not limited to a disk memory, an optical memory, and the like) that include computer-usable program code.

[0359] This application is described with reference to the flowcharts and / or block diagrams of the method, the device (system) , and the computer program product according to this application. It should be understood that computer program instructions may be used to implement each process and / or each block in the flowcharts and / or the block diagrams and a combination of a process and / or a block in the flowcharts and / or the block diagrams. The computer program instructions may be provided for a general-purpose computer, a dedicated computer, an embedded processor, or a processor of another programmable data processing device to generate a machine, so that the instructions executed by the computer or the processor of the another programmable data processing device generate an apparatus for implementing a specific function in one or more procedures in the flowcharts and / or in one or more blocks in the block diagrams.

[0360] The computer program instructions may alternatively be stored in a computer-readable memory that can indicate a computer or another programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate an artifact that includes an instruction apparatus. The instruction apparatus implements a specific function in one or more procedures in the flowcharts and / or in one or more blocks in the block diagrams.

[0361] The computer program instructions may alternatively be loaded onto a computer or another programmable data processing device, so that a series of operations and steps are performed on the computer or the another programmable device, so that computer-implemented processing is generated. Therefore, the instructions executed on the computer or the another programmable device provide steps for implementing a specific function in one or more procedures in the flowcharts and / or in one or more blocks in the block diagrams.

[0362] Throughout the present disclosure, a processor, a processor system, an application processor, a baseband processor, a processor circuit, or a processor core may be collectively referred to as a processor. A processor may include one or more of a central processing unit (CPU) , a digital signal processor (DSP) , a microprocessor unit (MPU) , a microcontroller unit, (MCU) , a graphics processing unit (GPU) , a field programmable gate array (FPGA) , an artificial intelligence (AI) processor, or a neural network processing unit (NPU) , or a combination of at least two of these integrated circuit forms.

[0363] Throughout the present disclosure, a memory may include one or more of the following storage media: a RAM, a static random access memory (SRAM) , a dynamic random access memory (DRAM) , a phase-change memory (PCM) , a resistive random access memory (ReRAM) , a magneto-resistive random access memory (MRAM) , a ferroelectric random access memory (FRAM) , a cache, a register, a read-only memory (ROM) , a flash memory, an erasable programmable read-only memory (EPROM) , a hard disk, and / or the like. In an example, the computer program instructions used to execute embodiments contained herein may be stored in a non-volatile memory. When a terminal runs, part or all of corresponding computer program instructions may be loaded into a memory that has a higher transmission speed with a corresponding processor, for example, the instructions may be loaded into at least a part of a memory such that the processor executes the computer program instructions to perform the steps in of embodiments described herein.

Claims

1.A method for voice typing, comprising:obtaining digital text based on a voice input;processing the digital text via a first model to obtain an initial intent prediction result, wherein the initial intent prediction result comprises at least one set of tokens from the digital text and predicted intent (s) corresponding to each of the at least one set of tokens; and,generating a first output, wherein the first output is generated based on processing the digital text as at least one command and / or input content based on the initial intent prediction result.2.The method according to claim 1, wherein the at least one command comprises: at least one text operation command and / or at least one system command.3.The method according to claim 1 or 2, wherein the at least one set of tokens comprises a first set of tokens and / or a second set of tokens, the predicted intent corresponding to the first set of tokens is a command and the predicted intent corresponding to the second set of tokens is input content.4.The method according to claim 3, wherein the at least one set of tokens comprises the first set of tokens and a first confidence score of the first set of tokens is equal to or higher than a first threshold, wherein the first confidence score of the first set of tokens indicates a probability of the first set of tokens being a command.5.The method according to claim 4, wherein the method further comprises:obtaining command arguments corresponding to the first set of tokens from the digital text via a second model, and the first output is generated based on the command arguments obtained via the second model.6.The method according to claim 5, wherein the method further comprises:if there is a lack of command arguments corresponding to the first set of tokens, elongating an utterance duration for obtaining longer digital text.7.The method according to claim 5 or 6, wherein if a second confidence score of the first set of tokens is higher than a second threshold, the method further comprises:processing the first set of tokens as a command, wherein the first output is generated based on a result of processing the first set of tokens as a command, wherein the second confidence score of the first set of tokens indicates a confidence in the command arguments obtained via the second model for the first set of tokens.8.The method according to claim 5 or 6, wherein if a second confidence score of the first set of tokens is equal to or lower than a second threshold, the method further comprises:extracting command arguments corresponding to the first set of tokens from the digital text via a third model, wherein the third model has a processing capability better than a processing capability of the second model, and the first output is generated based on an extraction result via the third model, wherein the second confidence score of the first set of tokens indicates a confidence in the command arguments obtained via the second model for the first set of tokens.9.The method according to any one of claims 3 to 8, wherein the at least one set of tokens comprises a third set of tokens from the digital text, a first confidence score of the third set of tokens is higher than a third threshold and lower than a first threshold, and the method further comprises:extracting command arguments corresponding to the third set of tokens from the digital text via a third model, wherein the third model has a processing capability better than a processing capability of the first model, and the first output is generated based on an extraction result via the third model, wherein the first confidence score of the third set of tokens indicates a probability of the third set of tokens being a command;orproviding a user with at least two options, wherein for the at least two options, the third set of tokens are processed differently, and the first output is generated based on the at least two options.10.The method according to claim 8 or 9, wherein if the extraction result via the third model is successful, the method further comprises:processing the first set of tokens as a command, wherein the first output is generated based on a result of processing the first set of tokens as a command; and / orprocessing the third set of tokens as a command, wherein the first output is generated based on a result of processing the third set of tokens as a command.11.The method according to claim 7 or 10, whereinif the result of processing the first set of tokens as a command is unsuccessful, the method further comprises: processing the first set of tokens as input content, wherein the first output is generated based on processing the first set of tokens as input content; and / orif the result of processing the third set of tokens as a command is unsuccessful, the method further comprises: processing the third set of tokens as input content, wherein the first output is generated based on processing the third set of tokens as input content.12.The method according to claim 8 or 9, wherein if the extraction result via the third model is unsuccessful, the method further comprises:processing the first set of tokens as input content, wherein the first output is generated based on processing the first set of tokens as input content; and / orprocessing the third set of tokens as input content wherein the first output is generated based on processing the third set of tokens as input content.13.The method according to claim 8 or 9, wherein if the extraction result via the third model is unsuccessful, the method further comprises:providing a user with at least two options, wherein the first output is generated based on the at least two options and the first set of tokens are processed differently for the at least two options; and / orproviding a user with at least two options, wherein the first output is generated based on the at least two options and the third set of tokens are processed differently for the at least two options.14.The method according to any one of claims 3 to 13, wherein the at least one set of tokens comprises the second set of tokens and a first confidence score of the second set of tokens is equal to or lower than a third threshold, and the method further comprises:processing the second set of tokens as input content, wherein the first output is generated based on processing the second set of tokens as input content, wherein the first confidence score of the second set of tokens indicates a probability of the second set of tokens being a command.15.The method according to any one of claims 7, 10, 11, 12, 14, wherein the method further comprises:receiving a first user input, wherein the first user input is used to cancel the first output; and,generating a second output, wherein the first set of tokens and / or the second set of tokens and / or the third set of tokens are processed differently corresponding to the second output from that corresponding to the first output.16.The method according to any one of claims 7, 10, 11, 12, 14, wherein the method further comprises:providing at least one alternative option, wherein the first set of tokens and / or the second set of tokens and / or the third set of tokens are processed differently corresponding to the at least one alternative option from that corresponding to the first output.17.The method according to claim 16, wherein the method further comprises:receiving a second user input, wherein the second user input is a selection of one of the least one option; and,generating a second output corresponding to the second user input.18.The method according to any one of claims 2 to 17, wherein the at least one text operation command comprises at least one text edit command and / or at least one text analysis command.19.The method according to claim 18, wherein the at least one text edit command comprises one or more of: at least one replace command, at least one add input command, at least one delete command, at least one add reference command.20.An apparatus, wherein the apparatus is configured to perform the method according to any one of claims 1 to 19.21.A computer-readable storage medium, wherein the computer-readable storage medium stores instructions, and when the instructions run on a device, the device is enabled to perform the method according to any one of claims 1 to 19.22.A computer program product, wherein when the computer program product runs on a device, the device is enabled to perform the method according to any one of claims 1 to 19.23.A chip system, comprising a memory and a processor, wherein the memory is configured to store a computer program, and the processor is configured to invoke the computer program from the memory and run the computer program, so that a device on which the chip system is disposed performs the method according to any one of claims 1 to 19.