Voice input method and device based on Linux system, terminal and medium
Patent Information
- Application Number
- CN202611207899.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-11
- Publication Date
- 2026-09-08
AI Technical Summary
[0004]本发明实施例提供了一种基于Linux系统的语音输入方法、装置、终端及介质,以解决现有技术中基于Linux语音输入系统无法进行光标映射的技术问题
[0009] The speech input method, device, terminal, and medium based on the Linux system provided in this invention listen to cursor position events from the Wayland synthesizer through an input method framework to obtain the relative coordinates of the cursor within the current client window. In response to the triggering of a speech input interaction command, it calls the window management private protocol interface of the UKUI desktop environment to query the window status bits maintained in the Wayland synthesizer and identify the currently active window. It then uses the window management private protocol interface to obtain the window reference coordinates of the currently active window in the global screen coordinate system. Based on the window reference coordinates and the relative coordinates, it performs spatial mapping calculations to generate the absolute position information of the cursor in the global screen coordinate system. It encapsulates the input context using the absolute position information and injects the input context into the speech recognition service interface for use by the speech recognition model. In response to the recognition result output by the speech recognition model, it submits the recognition result to the window corresponding to the current input focus. By querying the window status bits through a private protocol, identifying the currently active window, and finally submitting the recognition result to the window corresponding to the current input focus, it ensures accurate input focus recognition and result submission in a multi-window environment. By utilizing the window management private protocol interface of the UKUI desktop environment, and under the security constraint that the Wayland protocol strictly prohibits ordinary clients from obtaining global information, the window reference coordinates of the currently active window in the global screen coordinate system were successfully obtained. Further, combined with the relative coordinates of the cursor within the current client window detected by the input method framework, the absolute position information of the cursor in the global screen coordinate system was generated through spatial mapping calculation. The generated absolute position information was used to encapsulate the input context and injected into the speech recognition service interface, constructing a spatially aware speech input context and improving the accuracy of system-level interaction. This enables the speech input system in the Linux system to accurately and effectively map local cursor offsets to the global window reference.
Smart Images

Figure CN122711031A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice input technology, and in particular to a voice input method, device, terminal and medium based on the Linux system. Background Technology
[0002] With the evolution of the Linux desktop operating system, next-generation display server protocols, represented by Wayland, are gradually replacing the traditional X11 architecture. The Wayland protocol effectively solves the security and performance issues in the traditional graphics stack by introducing strict centralized control of synthesizers and a sandbox mechanism, achieving a smoother and more efficient graphics rendering experience. Against this backdrop, deeply integrating voice interaction capabilities into the desktop environment, giving it the same system-level responsiveness as keyboard input, has become an important trend in improving human-computer interaction efficiency.
[0003] However, this architectural evolution has also brought new challenges to input assistance technologies that rely on global information awareness. Existing voice input systems typically require a floating interactive interface to be presented in real time, and this interface must accurately follow the user's current input cursor movement to provide "what you see is what you get" visual feedback. In traditional graphics system architectures, input method components usually have access to global window hierarchy information and can easily calculate the absolute position of the cursor in the screen's physical coordinate system, thus achieving UI following. However, under the Wayland architecture, due to the design principle of security isolation, the system enforces a separation mechanism between "input localization" and "display globalization." Specifically, the input context information obtained by the input method framework through the Wayland protocol interface only includes the local coordinate offset of the cursor relative to the current client window surface; and the window manager, in order to prevent malicious applications from stealing screen information, no longer exposes the absolute position of the window in the global screen space and the window hierarchy state to ordinary clients. This separation of coordinate space leads to a structural problem for voice input systems: they cannot effectively map local cursor offsets to the global window reference. Summary of the Invention
[0004] This invention provides a voice input method, device, terminal, and medium based on a Linux system to solve the technical problem that existing Linux-based voice input systems cannot perform cursor mapping.
[0005] In a first aspect, embodiments of the present invention provide a voice input method based on a Linux system, comprising: By listening to cursor position events from the Wayland synthesizer through the input method framework, the relative coordinates of the cursor within the current client window can be obtained. In response to the triggering of voice input interaction commands, the window management private protocol interface of the UKUI desktop environment is invoked to query the window status bits maintained in the Wayland synthesizer and identify the currently active window; The window reference coordinates of the currently active window in the global screen coordinate system are obtained using the window management private protocol interface. Based on the window reference coordinates and the relative coordinates, spatial mapping calculations are performed to generate the absolute position information of the cursor in the global screen coordinate system; The absolute location information is used to encapsulate the input context, and the input context is injected into the speech recognition service interface for the speech recognition model to call. In response to the recognition result output by the speech recognition model, the recognition result is submitted to the window corresponding to the current input focus.
[0006] Secondly, embodiments of the present invention also provide a voice input device based on a Linux system, comprising: The listening module is used to listen for cursor position events from the Wayland synthesizer through the input method framework and obtain the relative coordinates of the cursor within the current client window. The calling module is used to respond to the triggering of voice input interaction commands, call the window management private protocol interface of the UKUI desktop environment, query the window status bits maintained in the Wayland synthesizer, and identify the currently active window; The acquisition module is used to obtain the window reference coordinates of the currently active window in the global screen coordinate system using the window management private protocol interface; The calculation module is used to perform spatial mapping calculations based on the window reference coordinates and the relative coordinates to generate the absolute position information of the cursor in the global screen coordinate system; The injection module is used to encapsulate the input context using the absolute position information and inject the input context into the speech recognition service interface for the speech recognition model to call. The submission module is used to submit the recognition result output by the speech recognition model to the window corresponding to the current input focus in response to the recognition result.
[0007] Thirdly, embodiments of the present invention also provide a terminal, including: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the Linux-based voice input method as described in any of the above embodiments.
[0008] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the voice input method based on the Linux system provided in the above embodiments.
[0009] The speech input method, device, terminal, and medium based on the Linux system provided in this invention listen to cursor position events from the Wayland synthesizer through an input method framework to obtain the relative coordinates of the cursor within the current client window. In response to the triggering of a speech input interaction command, it calls the window management private protocol interface of the UKUI desktop environment to query the window status bits maintained in the Wayland synthesizer and identify the currently active window. It then uses the window management private protocol interface to obtain the window reference coordinates of the currently active window in the global screen coordinate system. Based on the window reference coordinates and the relative coordinates, it performs spatial mapping calculations to generate the absolute position information of the cursor in the global screen coordinate system. It encapsulates the input context using the absolute position information and injects the input context into the speech recognition service interface for use by the speech recognition model. In response to the recognition result output by the speech recognition model, it submits the recognition result to the window corresponding to the current input focus. By querying the window status bits through a private protocol, identifying the currently active window, and finally submitting the recognition result to the window corresponding to the current input focus, it ensures accurate input focus recognition and result submission in a multi-window environment. By utilizing the window management private protocol interface of the UKUI desktop environment, and under the security constraint that the Wayland protocol strictly prohibits ordinary clients from obtaining global information, the window reference coordinates of the currently active window in the global screen coordinate system were successfully obtained. Further, combined with the relative coordinates of the cursor within the current client window detected by the input method framework, the absolute position information of the cursor in the global screen coordinate system was generated through spatial mapping calculation. The generated absolute position information was used to encapsulate the input context and injected into the speech recognition service interface, constructing a spatially aware speech input context and improving the accuracy of system-level interaction. This enables the speech input system in the Linux system to accurately and effectively map local cursor offsets to the global window reference. Attached Figure Description
[0010] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a flowchart illustrating the voice input method based on the Linux system provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart illustrating the voice input method based on the Linux system provided in Embodiment 2 of the present invention; Figure 3 This is a schematic diagram of the structure of the voice input device based on the Linux system provided in Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the terminal provided in Embodiment 4 of the present invention. Detailed Implementation
[0011] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0012] Example 1 Figure 1 This is a flowchart of a voice input method based on a Linux system provided in Embodiment 1 of the present invention. This embodiment is applicable to situations involving voice input processing in a Linux operating system. The method can be executed by a voice input device based on a Linux system, and specifically includes the following steps: Step 110: Listen for cursor position events from the Wayland synthesizer through the input method framework to obtain the relative coordinates of the cursor within the current client window.
[0013] In this embodiment, an input method framework based on the Fcitx5 architecture can be deployed in a Linux desktop environment. The voice input function is integrated into the Fcitx5 framework as a plug-in module. During system initialization, the voice input plug-in is loaded and registered to the Fcitx5 instance. The voice input plug-in subscribes to input context state change events from the input method framework by calling the application programming interface provided by Fcitx5. These state change events may include cursor position update events and input focus change events, etc.
[0014] When a user inputs data or moves the cursor using a mouse / touchpad within a client application window, such as a text editor, terminal, or browser, the client application can send a request to the Wayland compositor to update the cursor position via the Wayland client library. The Wayland compositor, acting as a display server, manages the surfaces of all windows. Upon receiving the request, it encapsulates the geometric information of the cursor within the corresponding window surface into a protocol event, according to the standard specifications of the Wayland text input protocol.
[0015] Since the Fcitx5 framework acts as the text input manager in the Wayland architecture, the Wayland synthesizer forwards the protocol event containing cursor position information to the Fcitx5 framework. Upon receiving this event, the Fcitx5 framework triggers its internal event callback mechanism. The corresponding callback signal can be captured using a registered, listening voice input plugin. The plugin parses the parameters carried in the callback event, extracting the cursor's coordinates relative to the top-left corner of the current client window. In the Wayland protocol definition, this coordinate data belongs to surface local coordinates, meaning its origin (0,0) is the top-left corner of the visible area of the current client window, with the X-axis positive to the right and the Y-axis positive downwards. The voice input plugin determines the extracted values as the cursor's relative coordinates within the current client window and stores these relative coordinates in the plugin's internal context cache structure for later use.
[0016] Step 120: In response to the triggering of the voice input interaction command, call the window management private protocol interface of the UKUI desktop environment to query the window status bits maintained in the Wayland synthesizer and identify the currently active window.
[0017] For example, querying the window state bits maintained in the Wayland synthesizer to identify the currently active window may include: listening to the state_changed event of the ukui_window_management interface in the window management private protocol; parsing the state flags carried by the event; determining whether the state flags contain an active flag with a value of 0x1; if so, determining that the window corresponding to the ukui_window_management interface is the currently active window.
[0018] In this embodiment, the ukui_window_management protocol is a proprietary extension protocol designed by the UKUI desktop environment to compensate for the shortcomings of the Wayland standard protocol in global window management capabilities. Although the standard Wayland protocol, due to its security sandbox mechanism, restricts the client's perception of the global window state, this proprietary protocol exposes necessary window management information to authorized system-level components, such as the input method plugin in this embodiment, by registering the extended global interface ukui_window_management_manager within the Wayland synthesizer. The protocol defines a specific window object (ukui_window_management) and uses global shortcut key listening modules to capture user voice input interaction commands to describe the window's geometric and state attributes.
[0019] When a user presses a preset shortcut key, the voice input plugin receives the corresponding trigger signal and initiates the voice input process. In response to the trigger signal, the voice input plugin needs to determine the window with the current input focus, i.e., identify the currently active window. Because the standard Wayland protocol, for security reasons, does not provide an interface for directly querying the global window state to ordinary clients, this embodiment utilizes the UKUI desktop environment's extended window management private protocol interface, namely the ukui_window_management interface, to implement the global window state query function. The voice input plugin, acting as a client of this private protocol, initiates a query request to the Wayland synthesizer or reads its maintained window state cache. Under the ukui_window_management protocol framework, the Wayland synthesizer maintains a set of state attributes for each top-level window object. By calling this protocol interface, the state information of each window object in the current system can be obtained.
[0020] To accurately identify the currently active window from multiple windows, the plugin parses the acquired window state information. For example, according to the definition of the ukui_window_management protocol, the window state is represented by a state flag, which can be a bitmask value used to reflect whether the window is in different states such as active, maximized, or full-screen. The plugin parses the state flag and performs bitwise operations to detect whether it contains a preset active identifier. When the binary value of the state flag contains an "active" flag with a value of 0x1, the window is determined to be in the active state of the user's current operation.
[0021] Using the above method, the plugin can filter out a unique currently active window from the list of windows maintained by the Wayland synthesizer, thus determining the correct target object for subsequently obtaining the global reference coordinates of that window and finally submitting the recognition results.
[0022] Step 130: Use the window management private protocol interface to obtain the window reference coordinates of the currently active window in the global screen coordinate system.
[0023] For example, obtaining the window reference coordinates of the currently active window in the global screen coordinate system using the window management private protocol interface may include: subscribing to the geometry event of the ukui_window_management interface in the window management private protocol; in response to receiving the geometry event, parsing the x and y fields in the event parameters; and using the values corresponding to the x and y fields as the absolute position of the top-left corner of the currently active window in the global screen coordinate system, as the window reference coordinates.
[0024] Furthermore, before responding to the triggering of the voice input interaction command, the method may also include the following: the input method framework continuously listens for focus state change events; when a focus state change or cursor position change is detected, it actively calls the window management private protocol interface to obtain the corresponding window reference coordinates and relative coordinates, and updates them to the pre-cache area in local memory.
[0025] The plugin can send a subscription request to the ukui_window_management interface corresponding to the currently active window to listen for geometry events published by that interface. The geometry event is triggered by the Wayland compositor when it detects a change in the window's position or size, and can notify subscribers of the window's latest geometry parameters.
[0026] The plugin enters event listening mode. When the Wayland compositor detects a change in the position or state of the currently active window, or in response to the plugin's initial query request, the compositor sends a geometry event to the plugin callback. Upon receiving the geometry event, the plugin parses the parameters carried in the event. According to the definition of the ukui_window_management interface, these parameters can include a set of data describing the rectangular area of the window, where the x-field represents the horizontal offset of the top-left corner of the window on the screen, and the y-field represents the vertical offset of the top-left corner of the window on the screen. The values corresponding to the parsed x and y fields are used to determine the absolute position of the top-left corner of the currently active window in the global screen coordinate system. To handle situations where the window may be dragged by the user or automatically adjusted by the system, the plugin continuously listens for geometry events. Once a new event is captured, the locally cached window reference coordinates are immediately updated, ensuring the real-time performance and accuracy of the coordinate data.
[0027] And the response to the triggering of voice input interaction commands can include: In response to a global shortcut key event, the latest window reference coordinates and relative coordinates are read from the pre-cached area to facilitate the spatial mapping calculation.
[0028] Step 140: Perform spatial mapping calculation based on the window reference coordinates and the relative coordinates to generate the absolute position information of the cursor in the global screen coordinate system.
[0029] The voice input plugin maintains two key coordinate data in memory: one is the reference coordinates (xwindow, ywindow) of the top-left corner of the currently active window in the global screen coordinate system, obtained through the aforementioned ukui_window_management private protocol; the other is the relative coordinates (xrel, yrel) of the cursor relative to the top-left corner of the window, obtained through the input method framework's monitoring.
[0030] Since the two coordinate systems mentioned above are defined in different reference frames—(xwindow, ywindow) with the top-left corner of the screen as the origin, and (xrel, yrel) with the top-left corner of the window as the origin—a spatial mapping calculation is required to convert the cursor's local position into a global position. For example, the plugin treats the window's reference coordinates as the basic offset and the cursor's relative coordinates as the internal increment. The plugin performs addition operations on the X-axis and Y-axis components: xscreen = xwindow + xrel, yscreen = ywindow + yrel, where xscreen represents the absolute pixel position of the cursor in the horizontal direction of the screen, and yscreen represents the absolute pixel position of the cursor in the vertical direction of the screen. After completing the above calculations, the plugin encapsulates the generated (xscreen, yscreen) value pairs into structured absolute position information. This absolute position information accurately reflects the cursor's true geometric position on the physical display device, solving the problem of coordinate system spatial fragmentation and providing accurate coordinate parameters for subsequent precise tracking and positioning of the voice interaction interface on the screen.
[0031] Furthermore, the spatial mapping calculation based on the window reference coordinates and the relative coordinates may include: reading the latest window reference coordinates and relative coordinates from the pre-cached area, and performing the spatial mapping calculation using the read data. When the voice input plugin receives a voice input interaction command and prepares to perform cursor positioning calculation, it will not immediately initiate a time-consuming and potentially delayed protocol interface query request. Optionally, the plugin may preferentially access a pre-cached area that it has pre-allocated and maintained in its local memory. The pre-cached area stores the latest coordinate data that the plugin updates and writes in real time during the background continuous monitoring process. The plugin reads the window reference coordinates and cursor relative coordinates stored in the pre-cached area at the current moment. Since the data is updated in real time in the background in response to changes in window focus or cursor movement events, the data in the pre-cached area always remains synchronized with the current desktop environment, accurately reflecting the instantaneous window state at the moment the command is triggered. The two sets of coordinate values read from the pre-cached area are used as calculation parameters and substituted into the spatial mapping algorithm. Using the above method, the plugin directly completes the calculation using the latest pre-stored data, avoiding the waiting time for cross-process communication queries at the trigger moment. This ensures the accuracy of the calculation results while efficiently generating the absolute position information of the cursor in the global screen coordinate system.
[0032] Step 150: Encapsulate the input context using the absolute position information and inject the input context into the speech recognition service interface for the speech recognition model to call.
[0033] After calculating the absolute position of the cursor in the global screen coordinate system, the voice input plugin initiates the input context construction process. The plugin creates an input context object to carry the environment information of the current session. This object contains not only the audio stream data to be input but also environmental data used to assist decoding. The plugin uses the calculated absolute position information—the precise coordinates of the cursor on the screen (xscreen, yscreen)—as key spatial context parameters to fill designated fields in the object. Furthermore, to enhance the richness of the context, the plugin can also encapsulate surrounding text fragments and the type identifier of the current application, such as a document editor or instant messaging software, along with the coordinate information.
[0034] The plugin invokes the speech recognition service interface through a standardized inter-process communication mechanism, injecting the encapsulated input context object into the speech recognition service. Upon receiving the request, the speech recognition service loads the input context object into the internal state machine of the recognition engine. During subsequent speech recognition, the backend speech recognition model, while processing the speech signal, calls and parses the absolute position information in this context.
[0035] For example, a speech recognition model can infer the current user's interaction scenario based on absolute coordinates, thus adaptively loading a language model or hot word library matching that scenario. Simultaneously, by combining the contextual range provided by the coordinates, the model can more accurately utilize existing text content around the cursor for semantic completion and homophone disambiguation. Using this method, the speech recognition model no longer processes audio signals in isolation, but makes comprehensive decisions based on the specific input location and editing environment, thereby significantly reducing the recognition error rate and outputting text results that better match the current context and user intent.
[0036] To further improve the accuracy and intelligence of speech recognition, the speech recognition service can employ a speech recognition system built on the Sherpa inference framework. For example, the voice input plugin calls the client API provided by the Sherpa inference framework to inject the input context, encapsulated with absolute positional information, into the framework. The Sherpa inference framework, as the backend inference engine, loads a pre-trained large-model speech recognition engine. Unlike traditional models, this large model possesses powerful semantic context awareness capabilities. During inference, the model not only parses the acoustic features of the speech signal but also deeply analyzes the absolute positional information in the injected input context. For example, using absolute coordinates, the large model can accurately determine the current input scenario, such as coding, document editing, or instant messaging, thereby adaptively activating specific language model branches or hotword libraries matching that scenario. Simultaneously, combined with the text context around the cursor position, the large model can perform high-precision homophone disambiguation and grammatical error correction. Through the streaming scheduling mechanism of the Sherpa inference framework, the large model can output high-quality, semantically optimized recognition results while ensuring low latency, thus significantly improving the user's voice input experience in the Linux desktop environment.
[0037] Step 160: In response to the recognition result output by the speech recognition model, submit the recognition result to the window corresponding to the current input focus.
[0038] The response to the recognition result output by the speech recognition model may include: receiving the recognized text generated by the speech recognition model through the input method framework; simulating a text input event using the input method framework and encapsulating the recognized text into a submission string that conforms to the input method framework protocol; and injecting the submission string into the input buffer of the window corresponding to the current input focus through the window management private protocol interface or the D-Bus interface, so that the window inserts the recognized text at the current cursor position.
[0039] The voice input plugin obtains the processing status of the speech recognition model in real time through an event listening channel established by the input method framework. Once the speech recognition model completes the decoding of the audio signal and generates the final recognized text, the plugin receives the recognized text string through this listening channel.
[0040] The plugin utilizes the text submission mechanism provided by the input method framework to simulate regular keyboard input. Optionally, the plugin encapsulates the received recognized text as content into a submission string that conforms to the input method framework protocol standard, such as Fcitx or IBus. Using this method, semantic-level text data can be converted into a sequence of text input events that the window system can recognize.
[0041] To accurately deliver the converted text to the target location, the plugin selects an appropriate transmission path based on the current system's operating environment and injects the encapsulated submission string into the input buffer of the window corresponding to the current input focus. Optionally, under the Wayland composition architecture based on the UKUI desktop environment, the plugin prioritizes calling the aforementioned ukui_window_management window management private protocol interface, using its provided window communication channel to directly write the submission string to the target window; in certain specific scenarios or compatibility modes, the plugin can also send a text insertion signal to the target application process via the D-Bus (Desktop Bus) system bus interface.
[0042] After receiving the submitted string, the target window writes the string content into its own text editing buffer. Based on the previously obtained cursor position information, the window application automatically inserts the recognized text at the current editing cursor position, achieving accurate execution of the user's voice commands.
[0043] This embodiment listens for cursor position events from the Wayland synthesizer through the input method framework to obtain the relative coordinates of the cursor within the current client window. In response to a voice input interaction command, it calls the window management private protocol interface of the UKUI desktop environment to query the window status bits maintained in the Wayland synthesizer and identify the currently active window. It then uses the window management private protocol interface to obtain the window reference coordinates of the currently active window in the global screen coordinate system. Based on the window reference coordinates and the relative coordinates, it performs spatial mapping calculations to generate the absolute position information of the cursor in the global screen coordinate system. The absolute position information is used to encapsulate the input context, which is then injected into the speech recognition service interface for use by the speech recognition model. In response to the recognition result output by the speech recognition model, the recognition result is submitted to the window corresponding to the current input focus. By querying the window status bits through the private protocol, identifying the currently active window, and finally submitting the recognition result to the window corresponding to the current input focus, accurate input focus recognition and result submission are ensured in a multi-window environment. By utilizing the window management private protocol interface of the UKUI desktop environment, and under the security constraint that the Wayland protocol strictly prohibits ordinary clients from obtaining global information, the window reference coordinates of the currently active window in the global screen coordinate system were successfully obtained. Further, combined with the relative coordinates of the cursor within the current client window detected by the input method framework, the absolute position information of the cursor in the global screen coordinate system was generated through spatial mapping calculation. The generated absolute position information was used to encapsulate the input context and injected into the speech recognition service interface, constructing a spatially aware speech input context and improving the accuracy of system-level interaction. This enables the speech input system in the Linux system to accurately and effectively map local cursor offsets to the global window reference.
[0044] In a preferred embodiment of this example, before responding to the triggering of a voice input interaction command, the method may further include the following steps: obtaining the internal unique identifier of the currently active window; binding the current input context with the currently active window using the internal unique identifier; updating the corresponding window reference coordinates and relative coordinates to the pre-cache area of local memory according to the binding relationship in response to a change in focus state or cursor position; and optimizing the spatial mapping calculation based on the window reference coordinates and the relative coordinates as follows: reading the latest coordinate data associated with the internal unique identifier from the pre-cache area, and performing the spatial mapping calculation using the read data. The internal unique identifier of the currently active window, such as a window ID or UUID, can be queried and obtained from the system through the window management private protocol interface of the UKUI desktop environment. The plugin uses the internal unique identifier to establish a strict binding relationship between the current input context and the currently active window, thereby establishing the index basis for subsequent data association.
[0045] When the plugin enters background continuous monitoring mode, once it detects a switch of focus from one window to another, or a movement of the cursor within the current window, it triggers a data update process based on the previously established binding relationship. The plugin calls the window management private protocol interface to re-acquire the window's baseline and relative coordinates corresponding to the current focused window or cursor position, and writes or overwrites this latest coordinate data, along with its internal unique identifier, to the pre-cache area maintained in local memory. Using this method, the pre-cache area always stores the latest window position state. Based on this pre-caching mechanism, when calculations are required in response to voice input interaction commands, the plugin no longer initiates time-consuming real-time interface queries, but directly accesses the pre-cache area in local memory. The plugin uses the internal unique identifier of the currently active window as a search key to quickly locate and read the latest coordinate data associated with that identifier in the pre-cache area. This optimization avoids communication delays at the moment of command triggering, further improving the efficiency of the calculation process and the real-time nature of the data.
[0046] Example 2 Figure 2This is a flowchart illustrating a voice input method based on a Linux system provided in Embodiment 2 of the present invention. This embodiment is an optimization based on the above embodiment. After encapsulating the input context using the absolute position information, the method may further include the following steps: validating the encapsulated input context; if the absolute position information in the input context is valid, then determining to use the cursor following mode, calculating the initial display position of the voice interaction interface directly below the cursor, and adjusting the display position of the voice interaction interface above the cursor when the initial display position exceeds the screen display boundary; if the absolute position information in the input context is invalid, then determining to use the desktop preset mode, calling the preset desktop tray area coordinates as the display position; and drawing the voice interaction interface according to the determined display position.
[0047] See Figure 2 The Linux-based voice input method includes: Step 210: Listen for cursor position events from the Wayland synthesizer through the input method framework to obtain the relative coordinates of the cursor within the current client window.
[0048] Step 220: In response to the triggering of the voice input interaction command, the window management private protocol interface of the UKUI desktop environment is called to query the window status bits maintained in the Wayland synthesizer and identify the currently active window.
[0049] Step 230: Use the window management private protocol interface to obtain the window reference coordinates of the currently active window in the global screen coordinate system.
[0050] Step 240: Perform spatial mapping calculation based on the window reference coordinates and the relative coordinates to generate the absolute position information of the cursor in the global screen coordinate system.
[0051] Step 250: Encapsulate the input context using the absolute position information and perform validity verification on the encapsulated input context.
[0052] To ensure the integrity and accuracy of the input context passed to the speech recognition service and prevent recognition errors or service crashes due to data anomalies, the encapsulated input context needs to be validated. For example, this could include field integrity checks and numerical logic checks. The plugin scans the data structure of the input context object to confirm whether it fully contains the preset key parameter fields. These key fields include at least: the absolute position information of the cursor in the global screen coordinate system (xscreen and yscreen), the internal unique identifier of the currently active window, and the timestamp at the time of encapsulation. If any key field is missing or the data type is mismatched, the input context is directly deemed invalid. After ensuring field integrity, further logical validity checks are performed on the specific numerical values within the fields. For example, for absolute position information, the plugin compares it with the system's current screen resolution parameters to verify whether xscreen and yscreen are within the valid screen display range, i.e., the coordinate values are not negative and are less than the maximum width and height of the screen, in order to exclude out-of-bounds data caused by coordinate calculation errors; for timestamps, the plugin verifies whether the difference between it and the current system time is within a preset threshold range to prevent expired data caused by system delays or clock synchronization problems from being used.
[0053] The plugin will only mark an input context object as valid and allow it to proceed to the next stage of the process, i.e., injection into the speech recognition service interface, if both the field integrity check and the numerical logic check mentioned above are passed simultaneously. Conversely, if any anomalies such as missing data, out-of-bounds values, or logical errors are found during the validation process, the plugin will intercept the submission of the input context and trigger an exception handling mechanism. This mechanism may include logging errors for later investigation or attempting to reacquire the context data, thereby effectively shielding the speech recognition model from interference by dirty data and ensuring the robustness and stability of the entire voice interaction system.
[0054] Step 260: If the absolute position information in the input context is verified to be valid, then the cursor following mode is determined to be used. The initial display position of the voice interaction interface is calculated directly below the cursor. When the initial display position exceeds the screen display boundary, the display position of the voice interaction interface is adjusted to above the cursor.
[0055] If the verification result shows that the absolute position information in the input context is valid, the voice input plugin determines the current display strategy as cursor following mode. In this mode, the plugin calculates the initial display position of the voice interaction interface, such as the voice input floating box, based on the valid absolute position information. Using the current screen coordinates (xcursor, ycursor) of the cursor as a reference point, and combining the width W and height H of the voice interaction interface itself, the plugin calculates a coordinate point that centers the interface horizontally and is directly below the cursor as the initial display position. For example, the plugin can set the initial coordinates (xinit, yinit) of the upper left corner of the interface to (xcursor-W / 2, ycursor+δ), where δ is the preset vertical safety distance between the cursor and the interface. After calculating the initial display position, screen boundary detection logic is further executed to prevent incomplete display. The plugin compares the bottom vertical coordinate yinit+H of the interface corresponding to the initial display position with the lower boundary of the current screen display area. If it detects that the bottom area of the interface exceeds the lower boundary of the screen, or that most of the interface is obscured by the screen edge, it indicates that the vertical space below the current cursor position is insufficient to fully accommodate the voice interaction interface. In this case, the display position of the voice interaction interface can be adjusted from below the cursor to above the cursor, and the display coordinates (xadj, yadj) of the upper left corner of the interface can be recalculated. For example, it can be set to (xcursor-W / 2, ycursor-H-δ). This ensures that the voice interaction interface can be displayed completely and clearly in the effective visible area of the screen, avoiding being truncated by the bottom edge of the screen and optimizing the user's visual interaction experience.
[0056] Step 270: If the absolute position information in the input context is invalid, then the desktop preset mode is determined to be used, the preset desktop tray area coordinates are called as the display position, and the voice interaction interface is drawn according to the determined display position.
[0057] If the absolute position information in the input context is determined to be invalid, the voice input plugin will abandon the cursor following mode and instead adopt the desktop preset mode as the current display strategy.
[0058] The plugin retrieves preset coordinates for the desktop tray area from the system configuration. These preset coordinates are typically set in the taskbar at the bottom of the screen, near the system tray area, or in the lower right corner of the screen, where system notifications are frequently displayed. These areas are usually not completely obscured by full-screen applications and conform to users' visual habits of searching for system status information. After determining the preset coordinates as the display position, the plugin initializes the drawing parameters of the voice interaction interface based on the coordinates, controlling the graphics rendering engine to draw and display the voice interaction interface at the specified position on the screen. Using this mechanism, even in scenarios where the accurate cursor position cannot be obtained, the voice interaction interface can still appear within the visible range of the screen in a default manner, thus ensuring the usability of the voice interaction function and the accessibility of the interface.
[0059] Step 280: Inject the input context into the speech recognition service interface, and in response to the recognition result output by the speech recognition model, submit the recognition result to the window corresponding to the current input focus.
[0060] This embodiment, after encapsulating the input context using the absolute position information, may further include the following steps: validating the encapsulated input context; if the absolute position information in the input context is valid, determining to use cursor following mode, calculating the initial display position of the voice interaction interface directly below the cursor, and adjusting the display position of the voice interaction interface above the cursor when the initial display position exceeds the screen display boundary; if the absolute position information in the input context is invalid, determining to use desktop preset mode, calling the preset desktop tray area coordinates as the display position; and drawing the voice interaction interface according to the determined display position. Validity verification actively identifies and filters abnormal coordinate data, preventing the interface rendering logic from crashing or misaligning due to incorrect input, ensuring stable system operation in complex environments. When the data is valid, cursor following combined with boundary detection is used for adaptive adjustment, ensuring the interface is always in the optimal viewing area, reducing user eye movement, and providing a smooth and intuitive interactive experience. When the data is invalid, automatically switching to desktop preset mode is used as a fault-tolerant solution, ensuring the interactive interface always appears in a fixed accessible position, preventing functional failure.
[0061] Example 3 Figure 3 This is a schematic diagram of the structure of the voice input device based on the Linux system provided in Embodiment 3 of the present invention. See also... Figure 3 The Linux-based voice input device includes: The listening module 310 is used to listen for cursor position events from the Wayland synthesizer through the input method framework and obtain the relative coordinates of the cursor within the current client window. Module 320 is invoked in response to the triggering of voice input interaction commands to invoke the window management private protocol interface of the UKUI desktop environment, query the window status bits maintained in the Wayland synthesizer, and identify the currently active window; The acquisition module 330 is used to obtain the window reference coordinates of the currently active window in the global screen coordinate system using the window management private protocol interface; The calculation module 340 is used to perform spatial mapping calculations based on the window reference coordinates and the relative coordinates to generate the absolute position information of the cursor in the global screen coordinate system. The injection module 350 is used to encapsulate the input context using the absolute position information and inject the input context into the speech recognition service interface for the speech recognition model to call. The submission module 360 is used to submit the recognition result to the window corresponding to the current input focus in response to the recognition result output by the speech recognition model.
[0062] The Linux-based voice input device provided in this embodiment listens to cursor position events from the Wayland synthesizer through the input method framework to obtain the relative coordinates of the cursor within the current client window. In response to the triggering of a voice input interaction command, it calls the window management private protocol interface of the UKUI desktop environment to query the window status bits maintained in the Wayland synthesizer and identify the currently active window. It then uses the window management private protocol interface to obtain the window reference coordinates of the currently active window in the global screen coordinate system. Based on the window reference coordinates and the relative coordinates, it performs spatial mapping calculations to generate the absolute position information of the cursor in the global screen coordinate system. It encapsulates the input context using the absolute position information and injects the input context into the speech recognition service interface for use by the speech recognition model. In response to the recognition result output by the speech recognition model, it submits the recognition result to the window corresponding to the current input focus. By querying the window status bits through the private protocol, identifying the currently active window, and finally submitting the recognition result to the window corresponding to the current input focus, it ensures accurate input focus recognition and result submission in a multi-window environment. By utilizing the window management private protocol interface of the UKUI desktop environment, and under the security constraint that the Wayland protocol strictly prohibits ordinary clients from obtaining global information, the window reference coordinates of the currently active window in the global screen coordinate system were successfully obtained. Further, combined with the relative coordinates of the cursor within the current client window detected by the input method framework, the absolute position information of the cursor in the global screen coordinate system was generated through spatial mapping calculation. The generated absolute position information was used to encapsulate the input context and injected into the speech recognition service interface, constructing a spatially aware speech input context and improving the accuracy of system-level interaction. This enables the speech input system in the Linux system to accurately and effectively map local cursor offsets to the global window reference.
[0063] Based on the above embodiments, the device further includes: A continuous monitoring module is used to continuously monitor focus state change events using the input method framework. The update module is used to actively call the window management private protocol interface to obtain the corresponding window reference coordinates and relative coordinates when a change in focus state or cursor position is detected, and update them to the pre-cache area in local memory. Accordingly, the computing module includes: The first calculation unit is used to read the latest window reference coordinates and relative coordinates from the pre-cached area, and to perform the spatial mapping calculation using the read data.
[0064] Based on the above embodiments, the device further includes: The validation module is used to validate the encapsulated input context. The display adjustment module is used to determine the cursor following mode if the absolute position information in the input context is valid, calculate the initial display position of the voice interaction interface directly below the cursor, and adjust the display position of the voice interaction interface above the cursor when the initial display position exceeds the screen display boundary. The mode determination module is used to determine that if the absolute position information in the input context is invalid, the desktop preset mode is adopted and the preset desktop tray area coordinates are used as the display position. The drawing module is used to draw the voice interaction interface based on the determined display position.
[0065] Based on the above embodiments, the submission module includes: A text receiving unit is used to receive the recognized text generated by the speech recognition model through the input method framework; The encapsulation unit is used to simulate text input events using the input method framework and encapsulate the recognized text into a submission string that conforms to the input method framework protocol. An injection unit is used to inject the submitted string into the input buffer of the window corresponding to the current input focus through the window management private protocol interface or the D-Bus interface, so that the window inserts the recognized text at the current cursor position.
[0066] Based on the above embodiments, the acquisition module includes: The subscription unit is used to subscribe to the geometry event of the ukui_window_management interface in the window management private protocol; The corresponding unit is used to parse the x and y fields in the event parameters in response to receiving the geometry event; As a unit, it is used to take the values corresponding to the x and y fields as the absolute position of the upper left corner of the currently active window in the global screen coordinate system, and as the window reference coordinates.
[0067] Based on the above embodiments, the calling module includes: The listening unit is used to listen for the state_changed event of the ukui_window_management interface in the window management private protocol; The parsing unit is used to parse the status flags carried by the event; The judgment unit is used to determine whether the status flags contain an active flag with a value of 0x1; if it does, the window corresponding to the ukui_window_management interface is determined to be the currently active window.
[0068] Based on the above embodiments, the device further includes: The identifier acquisition module is used to acquire the internal unique identifier of the currently active window; A binding module is used to bind the current input context to the currently active window using the internal unique identifier; The update module is used to update the corresponding window base coordinates and relative coordinates to the pre-cache area of local memory in response to changes in focus state or cursor position, according to the binding relationship. The computing module includes: The second calculation unit is used to read the latest coordinate data associated with the internal unique identifier from the pre-cached area and to perform the spatial mapping calculation using the read data.
[0069] The Linux-based voice input device provided in this embodiment of the invention can execute the Linux-based voice input method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0070] Example 4 Figure 4 This is a schematic diagram of the structure of a terminal provided in Embodiment 4 of the present invention. Figure 4 A block diagram is shown of an exemplary terminal 12 suitable for implementing embodiments of the present invention. Figure 4 The terminal 12 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of the present invention.
[0071] like Figure 4 As shown, terminal 12 is presented in the form of a general-purpose computing terminal. The components of terminal 12 may include, but are not limited to: one or more processors or processing units 16, system memory 28, and bus 18 connecting different system components (including system memory 28 and processing unit 16).
[0072] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.
[0073] Terminal 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by terminal 12, including volatile and non-volatile media, removable and non-removable media.
[0074] System memory 28 may include computer system readable media in the form of volatile memory, such as RAM 30 and / or cache 32. Terminal 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 4 Not shown; usually referred to as a "hard drive"). Although Figure 4 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0075] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of the present invention.
[0076] Terminal 12 can also communicate with one or more external devices 14 (e.g., keyboard, pointing terminal, display 24, etc.), and with one or more terminals that enable a user to interact with terminal 12, and / or with any terminal (e.g., network card, modem, etc.) that enables terminal 12 to communicate with one or more other computing terminals. This communication can be performed via I / O interface 22. Furthermore, terminal 12 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 20. As shown, network adapter 20 communicates with other modules of terminal 12 via bus 18. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with terminal 12, including but not limited to: microcode, terminal drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0077] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28, such as implementing the voice input method based on the Linux system provided in the embodiments of the present invention.
[0078] Example 5 Embodiment 5 of the present invention also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform any of the Linux-based voice input methods provided in the above embodiments.
[0079] The computer storage medium of this invention can be any combination of one or more computer-readable media. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0080] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0081] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including, but not limited to, wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0082] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or terminal. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0083] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A voice input method based on a Linux system, characterized in that, include: By listening to cursor position events from the Wayland synthesizer through the input method framework, the relative coordinates of the cursor within the current client window can be obtained. In response to the triggering of voice input interaction commands, the window management private protocol interface of the UKUI desktop environment is invoked to query the window status bits maintained in the Wayland synthesizer and identify the currently active window; The window reference coordinates of the currently active window in the global screen coordinate system are obtained using the window management private protocol interface. Based on the window reference coordinates and the relative coordinates, spatial mapping calculations are performed to generate the absolute position information of the cursor in the global screen coordinate system; The absolute location information is used to encapsulate the input context, and the input context is injected into the speech recognition service interface for the speech recognition model to call. In response to the recognition result output by the speech recognition model, the recognition result is submitted to the window corresponding to the current input focus.
2. The method according to claim 1, characterized in that, Prior to the triggering of the voice input interaction command, the method further includes: The input method framework continuously monitors focus state change events; When a change in focus state or cursor position is detected, the window management private protocol interface is actively invoked to obtain the corresponding window reference coordinates and relative coordinates, and updated to the pre-cache area in local memory. The spatial mapping calculation based on the window reference coordinates and the relative coordinates includes: The latest window reference coordinates and relative coordinates are read from the pre-cached area, and the spatial mapping calculation is performed using the read data.
3. The method according to claim 2, characterized in that, After encapsulating the input context using the absolute position information, the method further includes: Perform validity validation on the encapsulated input context; If the absolute position information in the input context is verified to be valid, then the cursor following mode is determined to be used, the initial display position of the voice interaction interface is calculated directly below the cursor, and when the initial display position exceeds the screen display boundary, the display position of the voice interaction interface is adjusted to above the cursor. If the absolute position information in the input context is invalid, then the desktop preset mode is used, and the preset desktop tray area coordinates are used as the display position. Draw the voice interaction interface based on the determined display location.
4. The method according to claim 1, characterized in that, The recognition result in response to the output of the speech recognition model includes: The input method framework receives the recognized text generated by the speech recognition model. The input method framework is used to simulate text input events, and the recognized text is encapsulated into a submission string that conforms to the input method framework protocol; The submitted string is injected into the input buffer of the window corresponding to the current input focus through the window management private protocol interface or the D-Bus interface, so that the window inserts the recognized text at the current cursor position.
5. The method according to claim 1, characterized in that, The step of obtaining the window reference coordinates of the currently active window in the global screen coordinate system using the window management private protocol interface includes: Subscribe to the geometry event of the ukui_window_management interface in the aforementioned window management private protocol; In response to receiving the geometry event, parse the x and y fields in the event parameters; The values corresponding to the x and y fields are used as the absolute position of the top-left corner of the currently active window in the global screen coordinate system, and are used as the window's reference coordinates.
6. The method according to claim 1, characterized in that, The step of querying the window state bits maintained in the Wayland synthesizer to identify the currently active window includes: Listen for the state_changed event of the ukui_window_management interface in the window management private protocol; Parse the status flags carried by the event; Determine whether the status flags contain an active flag with a value of 0x1; If it is included, then the window corresponding to the ukui_window_management interface is determined to be the currently active window.
7. The method according to claim 2, characterized in that, Prior to the triggering of a voice input interaction command, the method further includes: Obtain the internal unique identifier of the currently active window; The current input context is bound to the currently active window using the internal unique identifier; In response to changes in focus state or cursor position, the corresponding window base coordinates and relative coordinates are updated to the pre-cache area of local memory according to the binding relationship. The spatial mapping calculation based on the window reference coordinates and the relative coordinates includes: The latest coordinate data associated with the internal unique identifier is read from the pre-cached area, and the spatial mapping calculation is performed using the read data.
8. A voice input device based on a Linux system, characterized in that, include: The listening module is used to listen for cursor position events from the Wayland synthesizer through the input method framework and obtain the relative coordinates of the cursor within the current client window. The calling module is used to respond to the triggering of voice input interaction commands, call the window management private protocol interface of the UKUI desktop environment, query the window status bits maintained in the Wayland synthesizer, and identify the currently active window; The acquisition module is used to obtain the window reference coordinates of the currently active window in the global screen coordinate system using the window management private protocol interface; The calculation module is used to perform spatial mapping calculations based on the window reference coordinates and the relative coordinates to generate the absolute position information of the cursor in the global screen coordinate system; The injection module is used to encapsulate the input context using the absolute position information and inject the input context into the speech recognition service interface for the speech recognition model to call. The submission module is used to submit the recognition result output by the speech recognition model to the window corresponding to the current input focus in response to the recognition result.
9. A terminal, characterized in that, include: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the Linux-based voice input method as described in any one of claims 1-7.
10. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the voice input method based on a Linux system as described in any one of claims 1-7.