Writing implement for voice assistant activation
The writing implement with integrated sensors and microphone allows for intuitive voice assistant activation by detecting pen movements and proximity, enhancing user experience and privacy through seamless integration with existing user behaviors.
Patent Information
- Application Number
- PCT/CN2024/123541
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-11
- Filing Date
- 2024-10-09
- Publication Date
- 2026-01-15
AI Technical Summary
Current methods for activating voice assistants with smart pens lack intuitiveness and seamlessness, often requiring specific keywords, button presses, or complex gestures, which can be cumbersome and interrupt user workflow, and do not leverage the pen's microphone for natural activation.
A writing implement equipped with a microphone and sensors to detect trigger events, such as lifting the pen to the mouth and speaking, combined with proximity detection, to activate voice assistants without keywords, providing visual feedback and context-aware actions.
Enhances user experience by offering a natural, intuitive, and privacy-protected voice assistant activation method that minimizes false activations and integrates seamlessly with existing user behaviors.
Smart Images

Figure CN2024123541_15012026_PF_FP_ABST
Abstract
Description
WRITING IMPLEMENT FOR VOICE ASSISTANT ACTIVATION
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to International Patent Application No. PCT / CN2024 / 104899, filed on July 11, 2024, the contents of which are incorporated herein by reference.TECHNICAL FIELD
[0003] The technical field relates to voice assistant activation, and more specifically to writing implements, and systems and methods for the keywordless activation of a voice assistant associated with a computing device via the writing implement.BACKGROUND
[0004] In recent years, there has been a significant surge in the adoption and utilization of voice assistant technology across various consumer devices. Voice assistants, for instance, powered by artificial intelligence (AI) and natural language processing (NLP) algorithms, offer users a convenient and hands-free way to interact with digital devices and access a wide range of services and information. The emergence of Large Language Models (LLM) has revolutionized the landscape of conversational interaction. LLMs, such as OpenAI’s GPT (Generative Pre-trained Transformer) models, are built upon vast amounts of text data and trained to generate human-like responses to user input, enabling more fluid and engaging interactions with digital systems.
[0005] The future of interfaces is shifting from traditional graphical user interfaces (GUI) and touch-based interfaces to conversational user interfaces (CUI) and voice interfaces, driven by their naturalness, accessibility, ubiquity, and efficiency. The activation of voice interaction is a critical aspect of leveraging voice interfaces effectively. The activation process serves as the gateway for users to engage with voice-enabled systems, triggering the system to listen for commands, queries, or prompts, while preserving privacy, enhancing user control, and supporting hands-free operation. By designing intuitive and contextually relevant activation mechanisms, developers can create voice interfaces that are both efficient and user-friendly, enhancing overall user experience and satisfaction.
[0006] The current methods for waking up voice assistants include saying a keyword (or wake word) , long-pressing a power button, touch an icon on a graphical user interface and tapping the back of a smartphone. As an example, there are multiple ways to activate AppleTM’s voice assistant SiriTM. Users can say “SiriTM” or “Hey SiriTM, ” then ask a question or make a request. Users can also activate SiriTM with a button. On an iPhoneTM with Face IDTM, this can include pressing and holding the side button. On an iPhone with a Home button, this can include pressing and holding the Home button. With EarPodsTM, this can include pressing and holding the centre or call button. With CarPlayTM, this can include pressing and holding the voice command button on the steering wheel, or touching and holding the Home button on the CarPlayTM Home Screen. With SiriTM Eyes Free, this can include pressing and holding the voice command button on the steering wheel. Additionally, back tapping is a possible interaction on iPhonesTM. Users can map the action to activating SiriTM. As further examples, to activate GoogleTM Assistant, users can open the GoogleTM Assistant app, or say “Hey GoogleTM” . On HuaweiTM phones, users can say the keyword “Xiaoyi Xiaoyi” or long press the power button. On HONORTM phones, users can say keywords or long press power button. Additionally, there is a feature called “awakening by breath” , in which a user can lift the phone to their mouth within a certain distance and start speaking to activate the voice assistant. These methods can be used for instance to activate voice assistant on smartphones. In typical academic research about pen–voice multimodal interaction, voice input is from devices other than pen, such as from a headphone that users wear.
[0007] Some commercial devices that put a microphone on a “recording pen” . On some such devices, users can press a button on the pen to activate the transcription function. Smart pens such as Microsoft SurfaceTM Pen, Apple PencilTM, SamsungTM S-PenTM are designed to be used on tablets. Some of them have buttons on them. Some uses double tapping on the pen shaft to enable mode switching. Some academic research studied activation of voice assistants and pen–voice interaction. Some authors proposed activating speech input by bringing the smartphone to the mouth, but this method requires numerous pieces of hardware. The user’s movement of bringing the phone close to their mouth can be monitored using an inertial measurement unit (IMU) and a proximity sensor. When the likelihood of the phone being moved close to the mouth exceeds a certain threshold, the camera is triggered to capture an image, which is then further analyzed by algorithms to confirm that the user is activating the voice assistant using this method. Upon activation, the user is provided feedback of successful activation through vibration.
[0008] The current methods for waking up voice assistants often lack the intuitiveness and seamlessness desired for a natural user experience. Keyword activation requires users to speak a specific wake word or keyword to activate the voice assistant. This can feel unnatural or forced. Users may forget the wake word or find it difficult to remember and pronounce correctly, especially if it’s not a commonly used term. Some devices offer the option to activate the voice assistant by long-pressing the power button. While this method provides a physical shortcut, it may not be intuitive for users who are accustomed to using the power button for other functions such as locking the device or taking screenshots. Tapping an icon on the device’s user interface (UI) to activate the voice assistant is a common method but requires users to locate and interact with a specific on-screen element. This may interrupt users’ workflow. Tapping the back of the phone to is not widely known or accessible across all devices, limiting its utility and adoption. The “awakening by breath” feature by HONORTM includes too many rules the user needs to learn in order to perform a successful activation. The standardized actions include ensuring the phone is lifted with an upward motion, at a normal speed, in less than 1 second, with the distance between the mouth and the bottom microphone being ≤ 7 cm, maintaining a 60° angle, with a gap of ≤2.5 seconds between speech and action, then speaking out the corresponding command word. The feature is designed for a phone, not a smart pen.
[0009] Smart pens are typical input devices on tablets. However, smart pens such as Microsoft SurfaceTM Pen, Apple PencilTM, SamsungTM S-PenTM do not have microphones on them. Although some recording pens have microphones on them, they either do not work on tablets, or they simply use a button to activate transcription function. The functions of recording pens are limited, and they do not use the microphone or user gesture to activate a voice assistant.
[0010] There is a need for voice assistant activation to be more natural and intuitive, enhancing the overall usability and effectiveness of voice-enabled technology. Smart pens are common input devices for tablets and PCs, yet voice assistant and automatic speech recognition (ASR) activation with pens is a current gap in industry and academia. An intuitive gesture is needed for activating voice assistant and ASR with smart pens, in order to protect users’ privacy, minimize false activation, and achieve better recording quality.SUMMARY
[0011] The present disclosure provides a writing implement that allows for a method for activating an application by using trigger events and detecting voice input from the writing implement’s microphone. Users can receive visual feedback on the implement or the on-screen UI of an associated computing device when activation is successful. Example mappings to perform specific actions based on the context on the computing device and the users’ utterances. Additionally, an activation method that combines the writing implement’s position sensor (s) and microphone (s) is provided. In some embodiments, users can lift the writing implement close to their mouth and start speaking to activate ASR and the application. In some embodiments, incorporating proximity detection ensures that the ASR system is only activated when the pen is within a certain proximity to the user’s mouth.
[0012] In accordance with an aspect, a method for keywordless activation of a voice assistant associated with a computing device via a writing implement is provided. The writing implements detects a trigger event indicative that a user intends to activate the voice assistant, records by a microphone an utterance of the user, and transmits, by the writing implement, the recording to the computing device.
[0013] In some embodiments, detecting the trigger event comprises measuring a movement of the writing implement, wherein the trigger event is detected if the movement measurement complies with a movement condition.
[0014] In some embodiments, the movement condition comprises a condition corresponding to the writing implement being lifted.
[0015] In some embodiments, measuring the movement of the writing implement comprises using a position sensor comprising an accelerometer.
[0016] In some embodiments, measuring the movement of the writing implement comprises using a position sensor comprising an inertial measurement unit.
[0017] In some embodiments, measuring a movement of the writing implement comprises measuring at least one of a yaw angle, a pitch angle, a roll angle and a velocity of the writing implement.
[0018] In some embodiments, detecting the trigger event comprises using a model trained to accept the movement measurement as input and to provide a prediction of whether the trigger event has occurred as output.
[0019] In some embodiments, detecting the trigger event comprises measuring a contact of the user with the writing implement, wherein the trigger event is detected if the contact measurement complies with a contact condition.
[0020] In some embodiments, the contact condition comprises a condition corresponding to a double tap being applied to the writing implement.
[0021] In some embodiments, detecting the trigger event comprises measuring a push of a button of the writing implement, wherein the trigger event is detected if the button push measurement complies with a button condition.
[0022] In some embodiments, the button condition comprises a condition corresponding to the button being continually pushed over a duration greater than a configurable duration threshold.
[0023] In some embodiments, detecting the trigger event comprises measuring, by the microphone, a distance between the mouth of the user and the writing implement, wherein the trigger event is detected if the distance measurement complies with a distance condition.
[0024] In some embodiments, measuring the distance between the mouth of the user and the writing implement is performed by the microphone and at least one additional microphone of the writing implement and / or of the computing device.
[0025] In some embodiments, the distance condition comprises a condition corresponding to the distance being less than a configurable distance threshold.
[0026] In some embodiments, the recording is transmitted only if the distance measurement complies with the distance condition.
[0027] In some embodiments, the writing implement provides a visual feedback while the microphone is recording the utterance.
[0028] In some embodiments, the writing implement provides a visual feedback in response to the trigger event detection.
[0029] In some embodiments, the computing device displays a graphical user interface element in response to the trigger event detection.
[0030] In some embodiments, the computing device receives the recording from the writing implement, converts the recording to a text, determines a context from a set of possible contexts, determines an action based on the text and / or the context, and performs the action.
[0031] In some embodiments, the set of possible contexts comprises a home screen context and, in response to determining that the context is the home screen context, the computing device determines whether the text corresponds to a query or to a command, in response to determining that the text corresponds to a query, performs the action of transmitting the recording and / or the text to the voice assistant for further processing, and in response to determining that the text corresponds to a command, performs the action of performing the command.
[0032] In some embodiments, the set of possible contexts comprises a text entry box context and, in response to determining that the context is the text entry box context, the computing device determines whether the text corresponds to content or to a generation prompt, in response to determining that the text corresponds to content performs the action of filling the text entry box with the text, and in response to determining that the text corresponds to a generation prompt performs the action of generating a generated text from the generation prompt and to file the text entry box with the generated text.
[0033] In some embodiments, the set of possible contexts comprises an app context and, in response to determining that the context is the app context, the computing device determines whether an app content is selected, in response to determining that no app content is selected performs the action of generating a generated content from the text, and in response to determining that an app content is selected performs the action of performing a command corresponding to the text.
[0034] In accordance with another aspect, a writing implement for keywordless activation of a voice assistant associated with a computing device is provided. The writing implement has an activation means to detect a trigger event indicative that a user intends to activate the voice assistant, a microphone to record an utterance of the user when the trigger event is detected, and a communication means to transmit the recording to the computing device.
[0035] In some embodiments, the activation means comprises at least one position sensor configured to measure a movement of the writing implement, wherein the trigger event is detected if the movement measurement complies with a movement condition.
[0036] In some embodiments, the movement condition comprises a condition corresponding to the writing implement being lifted.
[0037] In some embodiments, the position sensor comprises an accelerometer.
[0038] In some embodiments, the position sensor comprises an inertial measurement unit.
[0039] In some embodiments, the movement measurement comprises at least one of a yaw angle, a pitch angle, a roll angle and a velocity of the writing implement.
[0040] In some embodiments, the writing implement has a memory having stored thereon a model trained to accept the movement measurement as input and to provide a prediction of whether the trigger event has occurred as output.
[0041] In some embodiments, the activation means comprises at least one tactile sensor configured to measure a contact of the user with the writing implement, wherein the trigger event is detected if the contact measurement complies with a contact condition.
[0042] In some embodiments, the contact condition comprises a condition corresponding to a double tap being applied to the writing implement.
[0043] In some embodiments, the activation means comprises at least one button configured to measure a button push, wherein the trigger event is detected if the button push measurement complies with a button condition.
[0044] In some embodiments, the button condition comprises a condition corresponding to the button being continually pushed over a duration greater than a configurable duration threshold.
[0045] In some embodiments, the activation means comprises the microphone, the microphone being configured to measure a distance between the mouth of the user and the writing implement, wherein the trigger event is detected if the distance measurement complies with a distance condition.
[0046] In some embodiments, the activation means further comprises at least one additional microphone of the writing implement and / or of the computing device, the microphone and the additional microphone configured to measure the distance between the mouth of the user and the writing implement.
[0047] In some embodiments, the distance condition comprises a condition corresponding to the distance being less than a configurable distance threshold.
[0048] In some embodiments, the communication means is configured to transmit the recording only if the distance measurement complies with the distance condition.
[0049] In some embodiments, the writing implement has a light-emitting device configured to emit light while the microphone is recording the utterance.
[0050] In some embodiments, the writing implement has a light-emitting device configured to emit light in response to the trigger event detection.
[0051] In some embodiments, the writing implement is configured to cause the computing device to display a graphical user interface element in response to the trigger event detection.
[0052] In accordance with a further aspect, a system comprising the writing implement and the computing device associated with a voice assistant described above is provided. The computing device has a communication means to receive the recording from the writing implement, and a processor to convert the recording to a text, determine a context of the computing device from a set of possible contexts, determine an action based on the text and / or the context, and perform the action.
[0053] In some embodiments, the set of possible contexts comprises a home screen context, and wherein, in response to determining that the context is the home screen context, the processor in configured to determine whether the text corresponds to a query or to a command, in response to determining that the text corresponds to a query, perform the action of transmitting the recording and / or the text to the voice assistant for further processing, and in response to determining that the text corresponds to a command, perform the action of performing the command.
[0054] In some embodiments, the set of possible contexts comprises a text entry box context, and wherein, in response to determining that the context is the text entry box context, the processor in configured to determine whether the text corresponds to content or to a generation prompt, in response to determining that the text corresponds to content, perform the action of filling the text entry box with the text, and in response to determining that the text corresponds to a generation prompt, perform the action of generating a generated text from the generation prompt and to file the text entry box with the generated text.
[0055] In some embodiments, the set of possible contexts comprises an app context, and wherein, in response to determining that the context is the app context, the processor in configured to determine whether an app content is selected, in response to determining that no app content is selected, perform the action of generating a generated content from the text, and in response to determining that an app content is selected, perform the action of performing a command corresponding to the text.
[0056] In another aspect, embodiments of this disclosure provide a computer readable storage medium, comprising one or more instructions, wherein when the one or more instructions are run on a computer, the computer performs any of the methods disclosed herein.
[0057] In another aspect, embodiments of this disclosure provide a non-transitory computer-readable medium storing instruction the instructions causing a processor in a device to implement any of the methods disclosed herein.
[0058] In another aspect, embodiments of this disclosure provide a device configured to perform any of the methods disclosed herein.
[0059] In another aspect, embodiments of this disclosure provide a processor, configured to execute instructions to cause a device to perform any of the methods disclosed herein.
[0060] In another aspect, embodiments of this disclosure provide an integrated circuit configure to perform any of the methods disclosed herein.
[0061] According to one aspect of this disclosure, there is provided a module comprising: one or more circuits for performing any of the methods disclosed herein.
[0062] According to one aspect of this disclosure, there is provided an apparatus comprising: one or more processors functionally connected to one or more memories for performing any of the methods disclosed herein.
[0063] According to one aspect of this disclosure, there is provided an apparatus configured to perform any of the methods disclosed herein.
[0064] In some embodiments the apparatus comprises one or more units configured to perform the above-described method.
[0065] According to one aspect of this disclosure, there is provided one or more non-transitory, computer-readable storage media comprising computer-executable instructions, wherein the instructions, when executed, cause at least one processing unit, at least one processor, or at least one circuits to perform any of the methods disclosed herein.
[0066] According to one aspect of this disclosure, there is provided one or more computer-readable storage media storing a computer program, wherein, when the computer program is executed by an apparatus, the apparatus is enabled to implement any of the methods disclosed herein.
[0067] According to one aspect of this disclosure, there is provided a computer program product including one or more instructions, wherein, when the instructions are executed by an apparatus, the apparatus is enabled to implement any of the methods disclosed herein.
[0068] According to one aspect of this disclosure, there is provided a computer program, wherein, when the computer program is executed by a computer, an apparatus is enabled to implement any of the methods disclosed herein.
[0069] According to one aspect of this disclosure, there is provided a system comprising a node for performing any of the methods disclosed herein.BRIEF DESCRIPTION OF THE DRAWINGS
[0070] For a better understanding of the embodiments described herein and to show more clearly how they may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings which show at least one exemplary embodiment.
[0071] Figure 1 is a schematic of a system including a writing implement and a computing device, in accordance with an embodiment.
[0072] Figure 2 is a schematic of a writing implement, in accordance with an embodiment.
[0073] Figure 3 is a schematic of a computing device, in accordance with an embodiment.
[0074] Figure 4A is a flowchart of a method for keywordless activation of a voice assistant associated with a computing device via a writing implement, in accordance with an embodiment.
[0075] Figure 4B is a flowchart of one possible embodiment of a step of the method illustrated in figure 4A.
[0076] Figure 5A to 5H are illustrations of possible movement triggers that can serve to activate a voice assistant via a writing implement, in accordance with eight possible embodiments.DETAILED DESCRIPTION
[0077] It will be appreciated that, for simplicity and clarity of illustration, where considered appropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements or steps. In addition, numerous specific details are set forth in order to provide a thorough understanding of the exemplary embodiments described herein. However, it will be understood by those of ordinary skill in the art that the embodiments described herein may be practised without these specific details. In other instances, well-known methods, procedures and components have not been described in detail so as not to obscure the embodiments described herein. Furthermore, this description is not to be considered as limiting the scope of the embodiments described herein in any way but rather as merely describing the implementation of the various embodiments described herein.
[0078] In the present disclosure, the word “pen” or the expression “writing implement” will be used to describe indistinctly any traditional or digital writing instrument that can be equipped with a microphone and a means to detect a voice assistant activation trigger event, including for instance digital or smart pens, traditional ink pens such as dip pens, fountain pens and disposable pens, pencils, mechanical pencils, brushes, etc.
[0079] In the present disclosure, the expressions “activation trigger” or “trigger event” will be used to describe any event caused by a user to let a system knows that they intend to activate a function and provide, e.g., a voice command and / or query. Activation triggers can be said to be “performed” by a user and detected by a system, e.g., a computing device such as a smartphone or a tablet, or a pen. Activation triggers traditionally include uttering a keyword, or wake word. Activation triggers that do not use such an utterance, for instance relying on touch, proximity and / or movement detection, can be said to allow for keywordless activation. The triggered application may be, for example, a voice assistant, a dictation application or an editing application or the like.
[0080] With reference to figure 1, an exemplary system 100 is depicted, including a device 200, such as a writing implement, configured to be used the methods described herein. The device 200 may be in communication with a second computing device 300 via one or more communication links 230, 330, including wired and / or wireless communication links, combinations thereof, links which traverse one or more networks, including local area networks, wide-area networks, the internet, and the like.
[0081] The devices 200 and 300 may be any suitable computing devices, such as, but not limited to, mobile computing devices such as phones, smart phones, tablets, laptop computers, handheld and / or wearable devices, and the like, or fixed computing devices, such as desktop computers, servers, kiosks, and the like.
[0082] The devices 200 and 300 each includes at least one processor 210, 310, memory 220, 320 and a communications interface 230, 330.
[0083] The processors 210 and 310 may include a central processing unit (CPU) , a microcontroller, a microprocessor, a processing core, a field-programmable gate array (FPGA) , a graphics processing unit (GPU) , or similar. The processors 210 and 310 may include multiple cooperating processors. The processors 210 and 310 may cooperate with the memory 220 and 320 to realize the functionality described herein.
[0084] The memory 220 and 320 may include a combination of volatile (e.g., Random Access Memory or RAM) and non-volatile memory (e.g., read-only memory or ROM, Electrically Erasable Programmable Read Only Memory or EEPROM, flash memory) . All or some of the memory 220 and 320 may be integrated with the processors 210 and 310. The memory can store data 222 and 322 as well as applications 224 and 334, each including a plurality of computer-readable instructions executable by the processor 210 and 310. The execution of the instructions by the processors 210 and 310 configures the devices 200 and 300 to perform the actions discussed herein. In particular, the applications 224 and 324 stored in the memory 220 and 320 can include applications to perform methods for language processing. When executed by the processors 210 and 310, the applications 224 and 324 configure the processors 210 and 310 and / or the device 200 and 300 to perform various functions discussed in greater detail below.
[0085] With reference to figure 2, an exemplary writing implement 200 is shown in accordance with an embodiment. Broadly described, the writing implement includes at least one transmitter 230, at least one sound sensor 240, at least one activation means such as a position sensor 250 or a button 260.
[0086] The writing implement 200 includes at least one sound sensor such as a microphone, used to record an utterance from the user. Although the exemplary embodiment of figure 1 includes two microphones located at different points along the length of the pen’s body, there is no specific requirement for the number or position of mics.
[0087] The writing implement 200 includes at least one communication means 230 to remotely connect the pen 200 to a computing device 300 such as a tablet or a PC. The communication means 230 corresponds at least to a transmitter configured to transmit a signal such as a radio signal to a corresponding receiver installed in the computing device 300. In some embodiments, the communication means 230 is a radio transceiver. In some embodiments, a Bluetooth adaptor is used, although other communication devices and / or protocols can be used, for instance NearLinkTM by HuaweiTM.
[0088] The writing implement 200 includes at least one activation means configured to detect a trigger event, for instance, a trigger event caused by a user to indicate that the activation of a voice assistant is intended. The user can cause the detection of a trigger event for instance by moving or touching the writing implement in specific, preconfigured ways, or by speaking with their mouth in close proximity with the microphone 240. Trigger events can be simple, i.e., caused by one condition such as a movement, touch or sound condition being met, or can be complex, i.e., causes by more than one condition being met, for instance all at once, or in a predetermined specific order.
[0089] In some embodiments, the activation means includes at least one position sensor configured to measure movements of the writing implement 200, for instance in order to detect gestures performed by the user holding the writing implement. The position sensors can for instance include one or more accelerometers and / or one or more gyroscopes. The accelerometer is energy efficient, which makes it possible for it to be operational often and / or for prolonged period while preserving battery capacity. A gyroscope can be provided but is not essential. In some embodiments, a position sensor corresponds to an inertial measurement unit (IMU) 250. The position sensors are configured to measure movement, yielding a movement measurement including a number of parameters such as a yaw angle, a pitch angle and / or a roll angle of the writing implement, and / or a velocity of the writing implement. In some embodiments, the movement measurement can be aggregated as an acceleration in the direction of the gravity-time curve. The movement measurement can be compared to predefined and / or configurable movement conditions that are indicative of certain specific movements. As an example, a movement condition can correspond to parameter values that are indicative of the position sensors measuring that the user is actively lifting the writing implement. Figures 5A to 5H show exemplary positions resulting from hand gestures of the user while holding the writing implement which can be used to define conditions characterizing eight possible types of trigger events that can be detected by a position sensor. In some embodiments, some or all of the features included in the movement measurement can be input in a machine learning model trained to output a prediction of whether the features comply with a movement condition, i.e., whether the features are indicative of a trigger event having occurred.
[0090] In some embodiments, the activation means includes at least one tactile sensor (not shown) , configured to measure a contact of the user with the writing implement. Any suitable tactile sensor using any transduction mechanism or any combination thereof can be used. For instance, the tactile sensors can include one or more capacitive tactile sensors, one or more piezoresistive tactile sensors, one or more magnetic tactile sensors, one or more piezoelectric tactile sensors, and / or one or more optical tactile sensors. The contact measurement obtained by the tacile sensor (s) can for instance represent a signal-time curve. The contact measurement can be compared to predefined and / or configurable contact conditions that are indicative of certain specific manipulations of the writing implement. As an example, a contact condition can correspond to a series of signal values that are indicative of the writing implement having a predefined and / or configurable of tap (s) applied, e.g., by the user, for instance two taps, corresponding to a double tap manipulation.
[0091] In some embodiments, the activation means includes at least one button 260, configured to measure a button push, i.e., to transmit a signal of whether the button is being pushed. Any suitable type of button can be used. For instance, button 260 can correspond to a single-pole push button. The button push measurement obtained by the button 260 can for instance represent a binary state timeline. The push-button measurement can be compared to predefined and / or configurable button conditions that are indicative of certain specific patterns of pushing the button 260. As an example, a button push condition can correspond to a binary state timeline indicative of the button being pushed a predetermined and / or configurable number of times in rapid succession, or being continually pushed for a duration greater than a predetermined and / or configurable duration threshold.
[0092] In some embodiments, the activation means includes the microphone 240. The microphone can be configured to capture an acoustic signal of the user speaking and estimate therefrom a measure of the distance between the user’s mouth and the microphone. In some embodiments, the microphone 240 is configured to transmit the acoustic signal to the processor 210 of the writing implement 200 for processing to determine the distance measurement. In some embodiments, the processor 210 of the writing implement 200 is further configured to transmit the acoustic signal to the computing device 300 for processing to determine the distance measurement. In some embodiments, characteristics such as sound amplitude and / or spectrogram characteristics, and / or features such as the volume and / or other features of the user’s voice are extracted from the signal of a single microphone 240 in order to estimate the distance measure, for instance, by using a machine learning model such as a convolutional neural network model trained to classify audio segments based on an estimated distance. In some embodiments, more than one microphone is advantageously leveraged. In some embodiments, at least one additional microphone 245 is provided on the writing implement 200. The benefit of using two or more mics on the pen is that it makes it possible to estimate voice direction based on the differences in sound recorded by different mics on the pen, giving us a more accurate estimate of the pen’s posture relative to users’ mouth. In some embodiments, at least one additional microphone 340 is provided in the computing device 300. The benefit of using microphone (s) on pen and microphone (s) on the computing device, e.g., on a tablet / PC, is that it enhances recording quality and reduces environmental noise, e.g., the sound recorded only by tablet / PC’s mic but not the pen’s mic can be considered as environmental noise and filtered out. The distance measure can be compared to predefined and / or configurable distance conditions. As an example, a distance condition can correspond to a predetermined and / or configurable distance threshold for the measurement of the mouth-microphone distance, below which a trigger event is deemed to occur.
[0093] In some embodiments, the behaviour of the writing implement 200 can be controlled by events more complex than what is expressed by movement, contact, button and / or distance conditions. In some embodiments, the writing implement 200 is configured, upon detecting a first trigger event such as a trigger event based on a movement measurement, to initiate recording by the microphone 240, and, based on the recorded acoustic signal, to further estimate a distance measurement and make a determination of whether a distance condition is met. In some embodiments, the writing implement 200 is configured to start transmitting the audio recording to the computing device 300 only upon determining that the distance condition is met.
[0094] It can be appreciated that other activation means are possible. For instance, there are other ways of detecting the “lift pen to mouth” gesture, such as using a proximity sensor or a camera. However, the means taught herein, such as using an accelerometer, are advantageously among the most economic and energy-efficient ways.
[0095] In some embodiments, a light-emitting device 270, i.e., a status light, can be provided to provide users with visual feedback of the status of the pen, for instance, to convey that activation is successful. As an example, a light-emitting diode can be activated to indicate that a trigger event has been detected, that the microphone 240 is recording an utterance, and / or that the utterance is being transmitted to the computing device 300. In some embodiments, a plurality of status lights corresponding to a plurality of colours, or a multicolour light, can be configured to indicate different statuses based on colour. As an example, the status light 270 can be activated with a first colour to indicate that a movement condition has been met and therefore that the microphone 240 is recording, and with a second colour to indicate that a distance condition has further been met and that the transmitter 230 is further transmitting the recording.
[0096] With reference to figure 3, an exemplary computing device 300 is shown in accordance with an embodiment. Broadly described, the computing device includes at least one receiver 230 and a monitor 380.
[0097] The computing device 300 includes at least one communication means 330 to remotely connect with the writing implement 200. The communication means 330 corresponds at least to a receiver configured to receive a signal such as a radio signal transmitted by a corresponding transmitter 230 installed in the writing implement 200. In some embodiments, the communication means 330 is a radio transceiver. The communication means 330 of the computing device is compatible at least with one of the technologies, standards and protocols used by the communication means 230 of the writing implement 200.
[0098] In some embodiments, the computing device 300 includes at least one sound sensor 340 such as microphone configured to record the voice of the user. In some embodiments, the acoustic signal acquired by microphone 340 is processed alongside the signal acquired by microphone (s) 240 and / or 245 of the writing implement 200 to determine the distance measurement, as detailed above.
[0099] The computing device 300 includes a monitor 380 configured to display a graphical user interface (GUI) . In some embodiments, the GUI is configured to display an element conveying an indication of the status of the writing implement 200, such as an icon or a dialogue window, in addition or as an alternative to a status light 270 of the writing implement. In some embodiments, the GUI defines a context 385 taken from a set of possible contexts. For instance, the set of possible contexts can include a “home screen context” corresponding to the monitor 380 displaying the home screen of the device, i.e., with no application having the focus, a “text entry box context” corresponding to the focus being on an input element of the GUI, and an “app context” corresponding to the monitor 380 displaying the content of an application, i.e., with the application having the focus. The context 385 can be used along with the recording of the user utterance, converted into text, to determine an appropriate action in association with various commands or queries included in the user utterance. In some embodiments, subcontexts are possible. As an example, when in “app context” , different actions can be taken depending on whether content is selected or not.
[0100] The table below provides exemplary mappings from different context and user speech to actions performed.
[0101] As an example, in the home screen context, the user could give a query such as “how’s the weather like” and the voice assistant in the system can be activated and respond to the user’s query by providing weather information. Similarly, a voice command such as “open drawing app” can trigger the system to perform the voice command accordingly.
[0102] As another example, in a text entry box context, if the user speaks out the desired content of the entry box, automatic speech recognition (ASR) can be used to transcribe the content, i.e. convert the utterance to text, and fill in the text box therewith. In the same context, if the user provides a prompt such as “write a comment for this product” , the computing device 300 can generate the text accordingly, for instance using a trained generative neural network such as a large language model and providing the prompt as input, and fill in the text box.
[0103] As a further example, in an app context, if the user provides a prompt to generate content like “write a paragraph about sea” or “draw a duck wearing pants” , the computing device 300 can generate the content based on the prompt, for instance using a trained generative neural network such as a large language model of a multimodal generative model and providing the prompt as input. If content is selected or circled in some applications, the user can provide a voice command to perform certain functions such as delete, translate, highlight, or rewrite the selected content.
[0104] It can be appreciated that the above-listed mappings are provided as examples only, and do not include every possible mapping that can be achieved by a voice assistant activated via a writing implement 200.
[0105] With reference to figure 4A, the flowchart of an exemplary method 400 for keywordless activation using a writing implement is shown. Broadly described, method 400 starts with the computing device entering a specific context in step 410 before an event trigger is detected in step 420. Detecting the event trigger starts the recording 430 and feedback 435 steps. Speech recognition is performed on the recording in step 440 so that the appropriate action can be determined and taken on the computing device in step 450.
[0106] In first step 410, a specific context is entered by the computing device. The context scenarios include but are not limited to the computer being on a particular screen such as a home screen or a specific application, a text entry box appearing or being active on the screen, and some content having been selected or circled in a specific application via an input device.
[0107] In subsequent step 420, the user can perform a trigger event that is detected by the system, i.e., perform some actions to activate microphone on the pen. The possible activation methods include but are not limited to double tap on pen, pressing a physical button on pen, pressing a haptic button on pen, performing a pen gesture, long pressing a button on pen, etc. In some embodiments, the trigger events may include but are not limited to the lift pen gesture.
[0108] With reference to figure 4B, in embodiments including detecting gestures such as “lift pen” as a means of detecting trigger events, step 420 can include performing steps 422 to 426. In step 422, a position sensor such as an accelerometer on the pen is activated and a gesture detection module, e.g., a “lift pen to speak” module is turned on. When the user lifts the pen, the system detects the pen lifting gesture based on position sensor data. The gesture detection using sensor, e.g., accelerometer, data can be performed in a number of ways. In various embodiments, rule-based method and / or machine learning methods can be used for gesture detection. Rule-based detection can include storing the data from a suitable duration, e.g., 2 seconds. Different features, such as the yaw, pitch, and roll angles can be calculated based on the sensor data, as well as tridimensional velocity. Filters such as Kalman filter, One Euro filter, etc. can be applied on sensor data to smooth out any noise.
[0109] In some embodiments, the acceleration in the direction of the gravity can be calculated as acc_g. The peaks and / or wave crests of the acc_g can be calculated and, based on the values and time differences of the peaks, events, or conditions, such as “acc_g increasing to a positive value, then decreasing to a negative value, then approaching zero” can be detected. If those events exist, it can be confirmed that a pen lifting gesture has been performed.
[0110] In some embodiments, the features can be used as input to one or more classifier, including for instance decision trees, logistic regression, k-Nearest Neighbour (kNN) , Support Vector Machine (SVM) , random forest, and deep leaning methods such as Convolutional Neural Network (CNN) .
[0111] In some embodiments, if an activating gesture, e.g., the pen lifting gesture is detected, the microphone starts to record in step 424 to allow for proximity detection in subsequent step 426.
[0112] In some embodiments, when users speak, the system can perform voice activity detection (VAD) and proximity detection based on the recording in subsequent step 426, for instance in order to detect if a user’s mouth is in close proximity to the pen. Performing VAD can include first, applying noise reduction (e.g., spectral subtraction) on a period of microphone data. The features can be extracted from this period of signal, including for instance signal energy, signal correlation, fundamental frequency, Mel-frequency cepstrum, etc. The features can be provided as input to one or more classifier, including for instance decision trees, logistic regression, kNN, SVM, random forest, and deep leaning methods such as CNN.
[0113] If VAD and proximity detections identify that users are speaking close to the pen, noticeable feedback can be given to the user. Types of user feedback include but are not limited to status light on pen, GUI on PC or tablet, etc. Automatic speech recognition can be performed as the user speaks, and actions can be executed based on different context and user input.
[0114] In some embodiments, proximity can be defined as “any part of the pen being within X cm of the mouth” . The threshold X can be predefined and / or customized by users.
[0115] It can be appreciated that the triggering gesture if not limited to lifting the pen to the mouth. Possible gestures include but are not limited to horizon, landscape, portrait, touch mouth, touch nose, side of mouth, etc. Figures 5A to 5H show some possible, but not all gestures.
[0116] Step 420 does not require in this embodiment, in particular if users are provided with alternative triggers to lifting the pen to the mouth to speak. This also makes it possible to provide earlier visual feedback of a trigger event being detected. Preferably, the pen should be able to record clearly what users are saying.
[0117] It can be appreciated that different activation methods can be used, and that more than one activation methods can be provided in a single embodiment. As an example, the position sensor can be combined with one or more other proposed activation means.
[0118] Once the trigger event is detected, in step 430, the microphone on the pen can be activated, if it has not been activated in step 424, and can start recording and / or transmitting sound.
[0119] In step 435, the user can be given noticeable feedback once their speech is detected, being recorded, transmitted and / or processed. Types of user feedback include but are not limited to a status light on pen and / or the GUI on the computing device.
[0120] In step 440, as a user starts to speak, the system detects user intention based on users’ voice input. As the user speaks, automatic speech recognition can be performed to convert the recording of the user utterance into text.
[0121] Finally, in step 450, the computing device can analyze the recording or the text corresponding to the user utterance along with the current context to determine which actions are appropriate given the utterance and the context, and execute said actions.
[0122] The present disclosure offers a number of benefits. Some are detailed below.
[0123] Proposing a “lift pen to speak” gesture to activate voice interaction and ASR provides a privacy protection benefit. By lifting the pen to mouth, users can speak directly to the pen’s microphone. This will protect users’ privacy significantly.
[0124] Using proximity detection to only activate ASR when the pen is within certain proximity to the user’s mouth provides the benefit of minimizing false activation. With the combination of the lift pen gesture and voice activity and proximity detection, the system can distinguish between intentional gestures and incidental movements, minimizing false activations and increasing robustness.
[0125] Using a smart pen’s microphone for ASR activation offers a simpler and more intuitive user experience. Using the pen position sensor (s) and microphone (s) to trigger for activating a voice assistant capitalizes on existing user behaviour and eliminates the need for additional training.
[0126] Proposing mappings of context to user speech and perform actions accordingly provides the benefit of context-adaptive response.
[0127] Giving the user visual feedback when activation is successful advantageously provides the user with a sense of control. When users receive visual cues that their activations were successful, they feel more in control.
[0128] In the present disclosure, the terms “a” or “an” are defined to mean “at least one” , that is, these terms do not exclude a plural number of items, unless stated otherwise.
[0129] In the present disclosure, terms such as “substantially” , “generally” and “about” , which modify a value, condition or characteristic of a feature of an example embodiment, should be understood to mean that the value, condition or characteristic is defined within tolerances that are acceptable for the proper operation of the example embodiment for its intended application.
[0130] In the present disclosure, unless stated otherwise, the terms “connected” and “coupled” , and derivatives and variants thereof, refer herein to any structural or functional connection or coupling, either direct or indirect, between two or more elements. For example, the connection or coupling between the elements can be acoustical, mechanical, optical, electrical, thermal, logical, or any combinations thereof.
[0131] In the present disclosure, expressions such as “match” , “matching” and “matched” , including variants and derivatives thereof, are intended to refer herein to a condition in which two or more elements are either the same or within some predetermined tolerance of each other. That is, these terms are meant to encompass not only “exactly” or “identically” matching the two elements but also “substantially” , “approximately” or “subjectively” matching the two or more elements, as well as providing a higher or best match among a plurality of matching possibilities.
[0132] In the present disclosure, the expression “based on” is intended to mean “based at least partly on” , that is, this expression can mean “based solely on” or “based partially on” , and so should not be interpreted in a limited manner. More particularly, the expression “based on” could also be understood as meaning “depending on” , “representative of” , “indicative of” , “associated with” or similar expressions.
[0133] In the present disclosure, the terms “system” and “network” may be used interchangeably in different embodiments of this application. “At least one” means one or more, and “aplurality of” means two or more. The term “and / or” describes an association relationship of associated objects, and indicates that three relationships may exist. For example, A and / or B may indicate the following three cases: Only A exists, both A and B exist, and only B exists, where A and B may be singular or plural. The character “ / ” indicates an “or” relationship between associated objects. “At least one of the following items (pieces) ” or a similar expression thereof indicates any combination of these items, including a single item (piece) or any combination of a plurality of items (pieces) . For example, “at least one of A, B, or C” includes: only A; only B; only C; A and B; A and C; B and C; or A, B, and C, and “at least one of A, B, and C” may also be understood as including: only A; only B; only C; A and B; A and C; B and C; or A, B, and C. In addition, unless otherwise specified, ordinal numbers such as “first” and “second” in embodiments of this application are used to distinguish between a plurality of objects, and are not used to limit a sequence, a time sequence, priorities, or importance of the plurality of objects.
[0134] A person skilled in the art should understand that embodiments of this application may be provided as a method, an apparatus (or system) , computer-readable storage medium, or a computer program product. Therefore, this application may use a form of a hardware-only embodiment, a software-only embodiment, or an embodiment with a combination of software and hardware. Moreover, this application may use a form of a computer program product that is implemented on one or more computer-usable storage media (including but not limited to a disk memory, an optical memory, and the like) that include computer-usable program code.
[0135] This application is described with reference to the flowcharts and / or block diagrams of the method, the device (system) , and the computer program product according to this application. It should be understood that computer program instructions may be used to implement each process and / or each block in the flowcharts and / or the block diagrams and a combination of a process and / or a block in the flowcharts and / or the block diagrams. The computer program instructions may be provided for a general-purpose computer, a dedicated computer, an embedded processor, or a processor of another programmable data processing device and enable a machine to execute the instructions. When executed by any computer or the processor of a programmable data processing device, the instructions cause the apparatus to implement specific functions as described in one or more procedures in the flowcharts and / or one or more blocks in the block diagrams. The computer program instructions may alternatively be stored in a computer-readable memory that can indicate a computer or another programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate an artifact that includes an instruction apparatus. The instruction apparatus implements a specific function in one or more procedures in the flowcharts and / or one or more blocks in the block diagrams.
[0136] The computer program instructions may alternatively be loaded onto a computer or another programmable data processing device, so that a series of operations and steps are performed on the computer or another programmable device, so that computer-implemented processing is generated. Therefore, the instructions executed on the computer or on another programmable device provide steps for implementing specific functions as described in one or more procedures in the flowcharts and / or one or more blocks in the block diagrams.
[0137] It is clear that a person skilled in the art can make various modifications and variations to this application without departing from the scope of this disclosure. This disclosure is intended to cover these modifications and variations of this application provided that they fall within the scope of protection defined by the following claims and their equivalent technologies.
Claims
1.A method for keywordless activation of a voice assistant associated with a computing device via a writing implement, the method comprising:detecting, by the writing implement, a trigger event indicative that a user intends to activate the voice assistant;recording, by a microphone of the writing implement, an utterance of the user; andtransmitting, by the writing implement, the recording to the computing device.2.The method of claim 1, wherein detecting the trigger event comprises measuring a movement of the writing implement, wherein the trigger event is detected if the movement measurement complies with a movement condition.3.The method of claim 2, wherein the movement condition comprises a condition corresponding to the writing implement being lifted.4.The method of claim 2 or 3, wherein measuring the movement of the writing implement comprises using a position sensor comprising an accelerometer.5.The method of claim 4, wherein measuring the movement of the writing implement comprises using a position sensor comprising an inertial measurement unit.6.The method of any one of claims 2 to 5, wherein measuring a movement of the writing implement comprises measuring at least one of a yaw angle, a pitch angle, a roll angle and a velocity of the writing implement.7.The method of any one of claims 2 to 6, wherein detecting the trigger event comprises using a model trained to accept the movement measurement as input and to provide a prediction of whether the trigger event has occurred as output.8.The method of any one of claims 1 to 7, wherein detecting the trigger event comprises measuring a contact of the user with the writing implement, wherein the trigger event is detected if the contact measurement complies with a contact condition.9.The method of claim 8, wherein the contact condition comprises a condition corresponding to a double tap being applied to the writing implement.10.The method of any one of claims 1 to 9, wherein detecting the trigger event comprises measuring a push of a button of the writing implement, wherein the trigger event is detected if the button push measurement complies with a button condition.11.The method of claim 10, wherein the button condition comprises a condition corresponding to the button being continually pushed over a duration greater than a configurable duration threshold.12.The method of any one of claims 1 to 11, wherein detecting the trigger event comprises measuring, by the microphone, a distance between the mouth of the user and the writing implement, wherein the trigger event is detected if the distance measurement complies with a distance condition.13.The method of claim 12, wherein measuring the distance between the mouth of the user and the writing implement is performed by the microphone and at least one additional microphone of the writing implement and / or of the computing device.14.The method of claim 12 or 13, wherein the distance condition comprises a condition corresponding to the distance being less than a configurable distance threshold.15.The method of any one of claims 12 to 14, wherein the recording is transmitted only if the distance measurement complies with the distance condition.16.The method of any one of claims 1 to 15, further comprising providing, by the writing implement, a visual feedback while the microphone is recording the utterance.17.The method of any one of claims 1 to 15, further comprising providing, by the writing implement, a visual feedback in response to the trigger event detection.18.The method of any one of claims 1 to 17, further comprising displaying, by the computing device, a graphical user interface element in response to the trigger event detection.19.The method of any one of claims 1 to 18, further comprising:receiving, by the computing device, the recording from the writing implement;converting, by the computing device, the recording to a text;determining, by the computing device, a context from a set of possible contexts;determining, by the computing device, an action based on the text and / or the context; andperforming, by the computing device, the action.20.The method of claim 19, wherein the set of possible contexts comprises a home screen context, and further comprising, in response to determining that the context is the home screen context:determining, by the computing device, whether the text corresponds to a query or to a command;in response to determining that the text corresponds to a query, performing, by the computing device, the action of transmitting the recording and / or the text to the voice assistant for further processing; andin response to determining that the text corresponds to a command, performing, by the computing device, the action of performing the command.21.The method of claim 19 or 20, wherein the set of possible contexts comprises a text entry box context, and further comprising, in response to determining that the context is the text entry box context:determining, by the computing device, whether the text corresponds to content or to a generation prompt;in response to determining that the text corresponds to content, performing, by the computing device, the action of filling the text entry box with the text; andin response to determining that the text corresponds to a generation prompt, performing, by the computing device, the action of generating a generated text from the generation prompt and to file the text entry box with the generated text.22.The method of any one of claims 19 to 21, wherein the set of possible contexts comprises an app context, and further comprising, in response to determining that the context is the app context:determining, by the computing device, whether an app content is selected;in response to determining that no app content is selected, performing, by the computing device, the action of generating a generated content from the text; andin response to determining that an app content is selected, performing, by the computing device, the action of performing a command corresponding to the text.23.A writing implement for keywordless activation of a voice assistant associated with a computing device, the writing implement comprising:an activation means configured to detect a trigger event indicative that a user intends to activate the voice assistant;a microphone configured, in response to the trigger event, to record an utterance of the user; anda communication means configured to transmit the recording to the computing device.24.The writing implement of claim 23, wherein the activation means comprises at least one position sensor configured to measure a movement of the writing implement, wherein the trigger event is detected if the movement measurement complies with a movement condition.25.The writing implement of claim 24, wherein the movement condition comprises a condition corresponding to the writing implement being lifted.26.The writing implement of claim 24 or 25, wherein the position sensor comprises an accelerometer.27.The writing implement of claim 26, wherein the position sensor comprises an inertial measurement unit.28.The writing implement of any one of claims 24 to 27, wherein the movement measurement comprises at least one of a yaw angle, a pitch angle, a roll angle and a velocity of the writing implement.29.The writing implement of any one of claims 24 to 28, further comprising a memory having stored thereon a model trained to accept the movement measurement as input and to provide a prediction of whether the trigger event has occurred as output.30.The writing implement of any one of claims 23 to 29, wherein the activation means comprises at least one tactile sensor configured to measure a contact of the user with the writing implement, wherein the trigger event is detected if the contact measurement complies with a contact condition.31.The writing implement of claim 30, wherein the contact condition comprises a condition corresponding to a double tap being applied to the writing implement.32.The writing implement of any one of claims 23 to 31, wherein the activation means comprises at least one button configured to measure a button push, wherein the trigger event is detected if the button push measurement complies with a button condition.33.The writing implement of claim 32, wherein the button condition comprises a condition corresponding to the button being continually pushed over a duration greater than a configurable duration threshold.34.The writing implement of any one of claims 23 to 33, wherein the activation means comprises the microphone, the microphone being configured to measure a distance between the mouth of the user and the writing implement, wherein the trigger event is detected if the distance measurement complies with a distance condition.35.The writing implement of claim 34, wherein the activation means further comprises at least one additional microphone of the writing implement and / or of the computing device, the microphone and the additional microphone configured to measure the distance between the mouth of the user and the writing implement.36.The writing implement of claim 34 or 35, wherein the distance condition comprises a condition corresponding to the distance being less than a configurable distance threshold.37.The writing implement of any one of claims 34 to 36, wherein the communication means is configured to transmit the recording only if the distance measurement complies with the distance condition.38.The writing implement of any one of claims 23 to 37, further comprising a light-emitting device configured to emit light while the microphone is recording the utterance.39.The writing implement of any one of claims 23 to 37, further comprising a light-emitting device configured to emit light in response to the trigger event detection.40.The writing implement of any one of claims 23 to 39, configured to cause the computing device to display a graphical user interface element in response to the trigger event detection.41.A system comprising the writing implement and the computing device associated with a voice assistant of any one of claims 23 to 40, wherein the computing device comprises:a communication means configured to receive the recording from the writing implement; anda processor configured to:convert the recording to a text,determine a context of the computing device from a set of possible contexts,determine an action based on the text and / or the context, andperform the action.42.The system of claim 41, wherein the set of possible contexts comprises a home screen context, and wherein, in response to determining that the context is the home screen context, the processor in configured to:determine whether the text corresponds to a query or to a command;in response to determining that the text corresponds to a query, perform the action of transmitting the recording and / or the text to the voice assistant for further processing; andin response to determining that the text corresponds to a command, perform the action of performing the command.43.The system of claim 41 or 42, wherein the set of possible contexts comprises a text entry box context, and wherein, in response to determining that the context is the text entry box context, the processor in configured to:determine whether the text corresponds to content or to a generation prompt;in response to determining that the text corresponds to content, perform the action of filling the text entry box with the text; andin response to determining that the text corresponds to a generation prompt, perform the action of generating a generated text from the generation prompt and to file the text entry box with the generated text.44.The system of any one of claims 41 to 43, wherein the set of possible contexts comprises an app context, and wherein, in response to determining that the context is the app context, the processor in configured to:determine whether an app content is selected;in response to determining that no app content is selected, perform the action of generating a generated content from the text; andin response to determining that an app content is selected, perform the action of performing a command corresponding to the text.
Citation Information
Patent Citations
Voice assistant-enabled client application with user view context and multi-modal input support
CN117099077A
Activating voice command functionality from a stylus
US20140362024A1
Stylus pen, terminal device set and controlling method thereof
US20210240286A1
Pointing device with integrated audio input
US7233321B1
Methods and systems for transforming speech into visual text
WO2023160994A1