A method for controlling the screen and display using a language

CN122575348APending Publication Date: 2026-08-14朱勤勤
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-01
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0002]基于目前电子终端设备上要进行语音操作时,需要用语言具体说出屏幕上按钮上的文字或标记,很麻烦,而且对于一些下载的第三方软件如果要进行具体的精细的语音操作很不方便甚至有时无法完成,需要手动进行操作,给使用者带来不便

Benefits of technology

[0006]启动语音操作时,显示的简单的标记m可以是数字如1、2、3…或字母如A、B、C…或简单的文字等也可以是原屏幕相应位置显示内容的前1-3个文字或字母或数字或符号等多种内容,也可以与原屏幕交互元素的位置稍有不同,也可以一闪一闪的,也可以是在一定范围内短距离移动的,也可以有线段指引所指标记,也可以是任意颜色的,也可以是半透明的,也可以是笔画纤细的等多种形式,只要方便使用者看清楚原屏幕的显示内容,方便选择相应的简单的标记m即可。当然简单的标记m的内容、形式也可以自由组合,也可以根据用户不同的需求,让用户进行选择和自由组合。不同使用者的选择也可以存储在电子设备终端或云端上,当语音控制启动时,通过识别不同的使用者的声纹,来启用相应使用者所选择的不同的简单的标记m的显示形式和内容。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575348A_ABST
    Figure CN122575348A_ABST
Patent Text Reader

Abstract

This invention discloses a method for controlling a screen and display using voice. The technical problem this design aims to solve is that voice operation on electronic terminal devices requires specifically speaking the text or markings on buttons, which is cumbersome. Sometimes, precise voice commands are not possible for certain third-party software. This design addresses this by having the electronic terminal scan the current screen, map its position coordinates, display a simple marker 'm', and then receive the user's voice commands. This generates touch click events at the corresponding position coordinates to operate the current screen. This allows users to control electronic terminal devices with simple language while protecting user privacy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This design relates to the field of speech signal processing technology, specifically to a method for controlling a screen and displaying speech signals. Background Technology

[0002] Currently, when performing voice operations on electronic terminal devices, users need to specifically speak the text or markings on the buttons on the screen, which is cumbersome. Moreover, for some downloaded third-party software, it is inconvenient or sometimes impossible to perform specific and precise voice operations, requiring manual operation, which causes inconvenience to users. Summary of the Invention

[0003] The technical problem this design aims to solve is that voice operation on electronic terminal devices requires specifically speaking the text or markings on the buttons, which is cumbersome. Furthermore, it can be difficult to perform precise voice operations on certain third-party software.

[0004] The technical solution adopted by this design to solve its technical problem is: When voice operation is initiated, software or programs on the electronic terminal device automatically scan the current screen and display simple markers 'm' at or near the corresponding positions of the interactive elements on the screen. This allows the user to issue simple voice commands containing the marker 'm' to operate the electronic terminal device. The electronic terminal device then generates a touch click event at the corresponding coordinates based on the user's voice command. For example, when a user wants to select the "TV series" button on the current screen, after initiating voice operation, the electronic terminal device immediately scans the current screen and displays simple markers 'm' at or near the positions of various interactive elements. In this case, the simple marker 'm' corresponding to the "TV series" button is "9". The user only needs to say a command containing "9", i.e., "click 9". The electronic terminal device receives the command through the microphone, analyzes it, and performs the corresponding operation on the interactive element "TV series" on the original screen corresponding to the simple marker 'm' "9", similar to using a remote control, mouse, keyboard, or other methods. This generates a touch click event at the coordinates of "TV series", thus completing the voice control.

[0005] This design expands the applicability of voice control by scanning the current screen, mapping position coordinates, displaying a simple marker 'm', receiving voice commands from the user, and generating touch click events at the corresponding position coordinates to operate the current screen. This minimizes manual operation. Displaying the simple marker 'm' reduces the difficulty for users to issue voice commands, sometimes saving time, and also better protects user privacy when using voice control of the screen.

[0006] When voice operation is initiated, the displayed simple marker 'm' can be numbers such as 1, 2, 3… or letters such as A, B, C… or simple text. It can also be the first 1-3 characters, letters, numbers, or symbols displayed at the corresponding position on the original screen. Its position can differ slightly from the interactive elements on the original screen; it can flash, move short distances within a certain range, be indicated by a line, and be of any color, semi-transparent, or have fine strokes—all in the form of facilitating the user's clear view of the original screen's content and easy selection of the corresponding simple marker 'm'. Of course, the content and form of the simple marker 'm' can be freely combined, allowing users to choose and combine them according to their different needs. Different users' selections can also be stored on the electronic device terminal or in the cloud. When voice control is initiated, the display form and content of the corresponding simple marker 'm' are activated by recognizing the user's voiceprint.

[0007] This design can also be implemented as follows: a simple marker 'm' can be displayed, or it could be a coordinate axis. For example, the X-axis could be at the bottom of the screen with numbers, and the Y-axis on the left with numbers. The user can issue a command based on the coordinates of the button they want to click. For instance, if the user wants to click the "Movie" button, and the approximate coordinates of this button are 2 on the X-axis and 7 on the Y-axis, the user can issue the command "Click 2-7". The electronic device can then determine the target location, calculate, and generate a touch click event at that location. When using this coordinate positioning for voice operation, it's not necessary to scan the original screen; the coordinate axis can be displayed directly, and the user can issue a voice command for the corresponding location. The electronic device can then determine the target location, calculate, and generate a touch click event at that location. Alternatively, it can be directly... By marking the coordinate axes outside the screen of the electronic device, such as on the frame outside the screen, the user can directly perform voice operations based on the external markings. This eliminates the need for any simple markers to be displayed on the screen; the user can determine the coordinate position and then issue commands containing the position coordinates. However, since the distance between the coordinate axes and interactive elements on the screen may be relatively large, errors can easily occur. Therefore, auxiliary lines can be displayed. Alternatively, after the user issues an operation command, the electronic device can display the corresponding position on the screen, which could be a small dot or other shape, for the user to confirm. After confirmation, the electronic device can then proceed to the next step.

[0008] This design can also be implemented as follows: When voice operation is initiated, after scanning the screen, simple markers 'm' are displayed. These markers are interconnected, driving the electronic terminal device to interact with the interactive elements on the current screen. For example, when voice operation is initiated, after scanning the screen, simple markers 'm' are displayed. When the user issues a command, the electronic terminal device analyzes and locates the target position based on the relative positions or relationships between the interactive elements. Following the user's command, the electronic terminal device generates a touch click event for the target position coordinates based on the relative positions or relationships of the interactive elements. This method can also be used when the user uses coordinate positioning.

[0009] This design can also be implemented as follows: after voice operation is initiated, the electronic terminal device will rescan the changed screen after completing one voice command, displaying a simple marker 'm' to await the user's next voice command. Alternatively, after initiating voice operation, the screen can be rescanned conditionally or periodically to keep up with screen changes. The simple marker 'm' program can also automatically close if the user does not issue a voice command for a period of time. Whether to scan the screen, when to scan, whether to display, when to display, and for how long to display the simple marker 'm' can also be enabled or set by the user via voice or manual operation. This design can be a complete voice control system or a series of programs integrated into a voice control system. Manual operation functionality can also be available on this interface after voice operation is initiated.

[0010] This design can also be implemented as follows: when voice operation is initiated, the software or program can automatically activate or the user can issue a command to activate the hidden interactive elements on the current screen, and then scan the screen to display a simple marker 'm'. For example, when an electronic device is playing a video, some interactive element buttons are hidden. In this case, when voice operation is initiated, the software or program can automatically activate the screen or the user can say "tap the screen" to make the hidden interactive elements appear on the screen, and then scan the current screen to display a simple marker 'm'. At this point, the user can issue a voice command to complete the operation.

[0011] This design can also be implemented as follows: the sliding interactive element can be displayed using multiple simple markers m, such as "1, 2, 3, 4, 5...". If the user wants to slide to position 3, they can simply say "slide to 3" during voice operation. For more precise adjustments, such as sliding slightly above 3 but not quite to position 4, the user can say "slide to 3.1". This allows for fine-tuning.

[0012] This design can also be implemented in this way: to facilitate voice control, interactive elements with three or more characters or words displayed on the screen will have a simple marker "m" or the first 1-3 characters, words, numbers, or letters of the original interactive element displayed before them. When voice operation is initiated, the user can directly read out the simple marker "m" or the first 1-3 characters, words, numbers, or letters on the interactive element, without having to read out all the information on the interactive element, to operate that element. For some simple interactive elements, such as the "back" button, the user does not need to read out the corresponding simple marker "m" to complete the "back" operation.

[0013] This design can also be implemented in this way: when using voice-controlled electronic terminal devices, the software or program can combine AI technology to handle more complex commands and provide a personalized experience. For example, by combining AI technology for learning, the user issues a command to allow the software or program to learn. Through a complete voice operation, the software or program remembers the operation and names it. This learning can be voice-controlled or manual. The next time, the user can simply speak a command containing the name given during the initial command to allow the device to complete the final target command. The interface at the start of this operation can be the first interface recorded during AI learning, or one of the interfaces recorded during AI learning. Interactive elements for setting simple markers like "m" can be added to the interface displaying the simple marker "m"—for example, a button called "m settings" can be added to configure the software or program of this design.

[0014] This design can also be implemented in this way: all clickable interactive elements and page elements on the screen can be scanned and marked with a simple "m" to facilitate voice operation of any operable point displayed on the screen. To prevent the reception of erroneous or irrelevant signals during voice control, the electronic device terminal can be configured to display a confirmation method at an appropriate time to verify whether the operation conforms to the user's wishes.

[0015] This design can also be like this: the above-mentioned contents can be freely combined under appropriate conditions, and two or more contents can exist at the same time or overlap in a certain form. When two or more contents exist at the same time, some of the contents can be added, deleted or adjusted.

[0016] The outstanding design features of this invention are: This design utilizes an electronic terminal device that, when controlled by voice, scans the current screen, maps the location coordinates, displays a simple marker 'm', receives the user's voice commands, and generates touch click events at the corresponding location coordinates to operate the current screen. This allows users to control the electronic terminal device with simple language while protecting user privacy. Attached Figure Description

[0017] Figure 1 For this design, the simple markings m (1-5, 6-14, A, B) are displayed on the original screen; Figure 2 The process of voice-controlled screen for this design. Detailed Implementation

[0018] The specific method of using this invention is as follows: After the user turns on the electronic terminal device, the user activates the voice control screen software or program by speaking a certain language. This software or program scans the screen and displays a simple mark 'm' at or next to the location of the interactive element displayed on the original screen. The user sees the simple mark 'm' at the corresponding position of the button to be operated on the screen, such as "10". The user says "click 10", and the microphone of the electronic terminal device receives the voice signal. The software or program analyzes it and, in a manner similar to using a remote control, mouse, or keyboard, or other methods, generates a touch click event operation on the corresponding interactive element on the original screen corresponding to the simple mark 'm' "10", thus completing the task given by the user.

[0019] This design is primarily used for controlling electronic terminal devices via voice.

[0020] The above description is merely a specific implementation of this embodiment, but the protection scope of this embodiment is not limited thereto. Any changes or substitutions within the technical scope disclosed in this embodiment should be covered within the protection scope of this embodiment. Therefore, the protection scope of this embodiment should be determined by the protection scope of the claims.

Claims

1. A method for controlling screen display using voice commands, characterized in that: When voice operation is initiated, software or programs on the electronic terminal device automatically scan the current screen and display simple markers 'm' at or near the corresponding positions of the interactive elements on the screen. This allows the user to issue simple voice commands containing the marker 'm' to operate the electronic terminal device. The electronic terminal device then generates a touch event corresponding to the coordinates of the given location based on the user's voice command. For example, if a user wants to select the "TV series" button on the current screen, after initiating voice operation, the electronic terminal device immediately scans the screen and displays simple markers 'm' at or near the locations of various interactive elements. In this case, the simple marker 'm' corresponding to the "TV series" button is "9". The user simply needs to say a command containing "9", i.e., "click 9". The electronic terminal device receives the command through the microphone, analyzes it, and, similar to using a remote control, mouse, keyboard, or other methods, performs a corresponding operation on the interactive element "TV series" corresponding to the simple marker 'm' "9" on the original screen, generating a touch event at the coordinates of "TV series", thus completing the voice control.

2. The method for controlling screen display with a voice as described in claim 1, characterized in that: When voice operation is initiated, the displayed simple marker m can be a number such as 1, 2, 3... or a letter such as A, B, C... or simple text, or it can be the first 1-3 characters, letters, numbers, or symbols displayed at the corresponding position on the original screen. It can also be slightly different in position from the interactive elements on the original screen, flashing, moving short distances within a certain range, or indicated by a line segment. It can be any color, semi-transparent, or have fine strokes, etc., as long as it allows the user to clearly see the content displayed on the original screen and easily select the corresponding simple marker m. Of course, the content and form of the simple marker m can be freely combined, allowing users to choose and combine them according to their different needs. Different users' selections can also be stored on the electronic device terminal or in the cloud. When voice control is initiated, the display form and content of the different simple marker m selected by the user are activated by recognizing the user's voiceprint.

3. The method for controlling screen display with a voice as described in claim 1, characterized in that: This design can also be implemented as follows: a simple marker 'm' can be displayed, or it could be a coordinate axis. For example, the X-axis could be at the bottom of the screen with numbers, and the Y-axis on the left with numbers. The user can issue a command based on the coordinates of the button they want to click. For instance, if the user wants to click the "Movie" button, and the button's approximate coordinates are 2 on the X-axis and 7 on the Y-axis, the user can issue the command "click 2-7". The electronic device can then determine the target location, calculate, and generate a touch event for that location. When using this coordinate positioning for voice operation, it's not necessary to scan the original screen; the coordinate axis can be displayed directly, and the user can issue a voice command for the corresponding location. The electronic device can then determine the target location, calculate, and generate a touch event for that location. Touch click events at the target location; alternatively, coordinate axes can be directly marked outside the screen of the electronic device terminal, such as on the frame outside the screen. Users can directly perform voice operations based on the marks outside the screen, so the screen of the electronic device terminal does not need to display any simple markers. Users can determine the coordinate position and then issue commands containing the position coordinates. Of course, since the distance between the coordinate axis and the interactive elements on the screen may be relatively large, errors are prone to occur. Therefore, some auxiliary lines can be displayed. Alternatively, after the user issues an operation command, the electronic device terminal can display the corresponding position to be operated on the screen. The position can be a small dot or other shape to allow the user to confirm again. After confirmation, the electronic device terminal will proceed to the next step.

4. The method for controlling screen display with language as described in claim 1, characterized in that: This design can also be implemented as follows: when voice operation is initiated, after scanning the screen, simple markers 'm' are displayed, which are related to each other to drive the electronic terminal device to operate on the interactive elements on the current screen. For example, when voice operation is initiated, after scanning the screen, simple markers 'm' are displayed. When the user issues a command, the electronic terminal device analyzes and locates the target position based on the relative position or relationship between the interactive elements. According to the user's command, the electronic terminal device generates a touch click event on the target position coordinates based on the relative position or relationship between the interactive elements. This method can also be used when the user uses coordinate positioning.

5. The method for controlling screen display by language as described in claim 1, characterized in that: This design can also be implemented as follows: after voice operation is initiated, the electronic terminal device will scan the changed screen again after completing one voice command, displaying a simple marker 'm' to await the user's next voice command; alternatively, after initiating voice operation, the screen can be rescanned conditionally or at regular intervals to keep up with screen changes; or the simple marker 'm' program can automatically close if the user does not issue a voice command for a period of time. Whether to scan the screen, when to scan the screen, whether to display the simple marker 'm', when to display it, and for how long can all be enabled or set by the user via voice or manual operation. This design can be a complete voice control system or a series of programs integrated into a voice control system; manual operation functions can also be available on this interface after voice operation is initiated.

6. The method for controlling screen display by language as described in claim 1, characterized in that: This design can also be implemented as follows: when voice operation is initiated, the software or program can automatically activate or the user can issue a command to activate the hidden interactive elements on the current screen, and then scan the screen to display a simple marker 'm'. For example, when an electronic device is playing a video, some interactive element buttons are hidden. In this case, when voice operation is initiated, the software or program can automatically activate the screen or the user can say "tap the screen" to make the hidden interactive elements appear on the screen, and then scan the current screen to display a simple marker 'm'. At this point, the user can issue a voice command to complete the operation.

7. The method for controlling screen display by language as described in claim 1, characterized in that: This design can also be like this: the sliding interactive element can be displayed using multiple simple markers m, such as simple markers m "1, 2, 3, 4, 5..." If the user wants to slide to position 3, then when performing voice operation, the user can say "slide to 3". If more precision is needed, such as sliding to a position slightly above 3 but not quite to position 4, the user can say "slide to 3.1", and so on for fine adjustment.

8. The method for controlling screen display by language as described in claim 1, characterized in that: This design can also be implemented in this way: to facilitate voice control, interactive elements with three or more characters or words displayed on the screen will have a simple marker 'm' or the first 1-3 characters, words, numbers, or letters of the original interactive element displayed in front of them. When voice operation is initiated, the user can directly read out the simple marker 'm' or the first 1-3 characters, words, numbers, or letters on the interactive element without having to read out all the information on the corresponding interactive element to operate it. For some simple interactive elements, such as the "back" button, the user does not need to read out the corresponding simple marker 'm' to complete the "back" operation.

9. The method for controlling screen display by language as described in claim 1, characterized in that: This design can also be implemented as follows: when using voice-controlled electronic terminal devices, this software or program can combine AI technology to process more complex commands and provide a personalized experience. For example, it can learn using AI technology. The user issues a command, allowing the software or program to learn. Through a complete voice operation, the software or program remembers the operation and names it. This learning can be voice-controlled or manual. The next time, the user can simply speak a command containing the name given during the learning process to allow the device to complete the final target command. The interface at the start of this operation can be the first interface recorded during AI learning or one of the interfaces recorded during AI learning. A simple interactive element with the marker 'm' can be added to the interface—for example, a button called "m settings" can be added to configure the software or program of this design.

10. The method for controlling screen display by language as described in claim 1, characterized in that: This design can also be implemented in this way: all clickable interactive elements and page elements on the screen can be scanned and marked with a simple marker 'm', so that users can easily perform voice operations on any operable point displayed on the screen; in order to prevent the reception of erroneous or irrelevant signals during voice control, the electronic device terminal can be set to confirm whether the operation meets the user's wishes at an appropriate time.

11. A method for controlling screen display using language as described in claim 1, characterized in that: This design can also be like this: the above-mentioned contents can be freely combined under appropriate conditions, and two or more contents can exist at the same time or overlap in a certain form. When two or more contents exist at the same time, some of the contents can be added, deleted or adjusted.

12. The method for controlling screen display with a voice as described in claim 1, characterized in that the method includes: This design utilizes an electronic terminal device to scan the current screen, plan the position coordinates, display a simple marker m, and then receive the user's voice commands to generate touch click events at the corresponding position coordinates, thereby enabling operation of the current screen.