Voice control method and electronic device

By performing speech recognition and coordinate mapping on audio data in one-handed mode, the problem of inaccurate click positions in one-handed mode is solved, ensuring the accuracy of voice control and user experience.

CN119229862BActive Publication Date: 2026-01-27HONOR DEVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310809838.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2026-01-27
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

In one-handed mode, when a user performs a simulated click operation by controlling the phone via voice, the click location determined by the phone may be inaccurate, resulting in an incorrect click location and failure to accurately execute the user's voice command.

Method used

In one-handed mode, the electronic device collects audio data for speech recognition, determines the simulated click location, and performs coordinate mapping to ensure the accuracy of the click location.

Benefits of technology

In one-handed mode, the user's voice commands can accurately execute simulated click operations, avoiding incorrect click positions and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119229862B_ABST
    Figure CN119229862B_ABST
Patent Text Reader

Abstract

The application provides a voice control method and an electronic device, relates to the technical field of electronics, and is used for enabling the electronic device to perform a simulated click operation at an accurate position in response to a voice instruction of a user in a single-hand mode. The method is applied to an electronic device with a voice control function and a single-hand operation function, and includes the following steps: in response to a single-hand mode triggering event, switching from a full-screen display picture to a small-screen display picture in a single-hand mode display state; after the voice control function is started, performing voice recognition on collected audio data based on the voice control function; if a result of the voice recognition indicates that an operation corresponding to the audio data is a simulated click operation, determining a simulated click position of the simulated click operation in the small-screen display picture. If the simulated click operation is an element click operation, a task triggered by an interface element corresponding to the element click operation is executed; if the simulated click operation is a sliding operation, the display picture is slid according to a simulated click start position to a simulated click end position.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of electronic technology, and in particular to a voice control method and electronic device. Background Technology

[0002] In many scenarios, users may want to operate electronic devices with one hand, such as using a mobile phone. However, some mobile phones are too large, making it impossible for users to tap all areas of the screen with one hand. In such cases, some phones offer a one-handed mode, which shrinks the display area on the screen to allow users to more easily tap any area of ​​the display with one hand.

[0003] In addition, many mobile phones allow users to control electronic devices via voice. In some scenarios, the user's voice command is to control the phone to perform a simulated click. In order for the phone to respond to the voice command and perform the simulated click, the location to be clicked is usually determined first, and then a simulated click is performed at that location.

[0004] Because the position of displayed content changes in one-handed mode, the phone may not accurately determine the simulated click location when the user performs a simulated click operation via voice control. This can lead to incorrect click placement when the phone responds to voice commands. Summary of the Invention

[0005] This application provides a voice control method and an electronic device, which, in one-handed mode, can respond to a user's voice command and perform a simulated click operation at a precise location.

[0006] To achieve the above objectives, the embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, a voice control method is provided, which applies an electronic device with voice control and one-handed operation capabilities, the method comprising:

[0008] In response to a one-handed mode trigger event, the electronic device switches from a full-screen display to a smaller one-handed mode display. After voice control is activated, audio data can be collected and processed by the electronic device through speech recognition. If the speech recognition result indicates that the operation corresponding to the audio data is a simulated click, the electronic device can determine the simulated click position on the currently displayed smaller screen. This simulated click can be either a click or a swipe. Specifically, when the simulated click is a swipe, the simulated click position can be divided into a start position and an end position. Finally, based on the simulated click position, the simulated click operation is executed.

[0009] In this scheme, if the simulated click operation is an element click, the electronic device executes the task triggered by the corresponding interface element. If the simulated click operation is a swipe, the electronic device swipes the displayed screen from the start to the end of the simulated click. This ensures that in one-handed mode, when the user uses voice control to perform simulated click operations, the electronic device responds to the voice command and executes the simulated click operation at the accurate location, triggering the corresponding response and avoiding errors in simulated clicks. Provided the electronic device accurately recognizes the user's voice commands, it can respond to the simulated click voice command and accurately execute the simulated click operation.

[0010] In one possible implementation of the first aspect, determining the simulated click position corresponding to the simulated click operation on the small screen display may specifically include: determining the first coordinates of the simulated click position in the screen coordinate system. In this implementation, before executing the simulated click operation based on the simulated click position, the method further includes: performing a first mapping process on the first coordinates to obtain a second coordinate in the screen coordinate system; wherein the first mapping process is used to map the coordinates of the point from the full-screen display to the small screen display.

[0011] Furthermore, in this embodiment, the electronic device performs a simulated click operation based on the simulated click position, which specifically includes: generating a simulated click event based on the second coordinates; sending the simulated click event to the multi-modal input frame of the electronic device; and the multi-modal input frame performing a second mapping process on the second coordinates in response to the simulated click event. The second mapping process maps the coordinates of the point from the small-screen display to the full-screen display; thus, the second mapping process and the first mapping process are opposite. Therefore, the multi-modal input frame can still obtain the first coordinates by performing the second mapping process on the second coordinates. Finally, the electronic device can perform a simulated click operation based on the position of the first coordinates in the multi-modal input frame. This ensures that the electronic device ultimately performs the simulated click operation at the position of the simulated click on the small-screen display, avoiding the problem of incorrect position execution of the simulated click operation.

[0012] In one possible implementation of the first aspect, the first mapping process applied to the first coordinates to obtain the second coordinates in the screen coordinate system may specifically include: obtaining the full-screen display size of the electronic device in a full-screen display, the small-screen display size in a small-screen display, and the first mapping relationship for transforming the full-screen display size into the small-screen display size. Then, the coordinates of the point can be mapped from the full-screen display to the small-screen display based on the first mapping relationship.

[0013] Before performing a simulated click operation in one-handed mode, the multi-mode input frame of the electronic device automatically defaults to the coordinates of the small-screen display and maps them to the coordinates of the full-screen display. Finally, it performs the simulated click operation based on these mapped coordinates. Therefore, in this embodiment, to ensure that the final position where the electronic device performs the simulated click operation is still the coordinates on the small-screen display, the first coordinates of the simulated click position on the small-screen display are pre-mapped to obtain a second coordinate. Then, the multi-mode input frame of the electronic device obtains the second coordinate, uses it as the coordinate on the small-screen display, performs a second mapping on it, and performs the simulated click operation on the coordinate obtained from the second mapping. Since the second coordinate is obtained from the first coordinate through the first mapping, and the second mapping is the reverse of the first mapping, the coordinate obtained after the multi-mode input frame of the electronic device performs the second mapping is still the first coordinate, which is the position of the simulated click position on the small-screen display. This ensures that the electronic device can perform the simulated click operation at the accurate position.

[0014] In one possible implementation of the first aspect, obtaining the full-screen display size of the electronic device in a full-screen display mode, the small-screen display size in a small-screen display mode, and the first mapping relationship for transforming the full-screen display size into the small-screen display size may specifically include: obtaining a first width and a first height of the full-screen display size, and a second width and a second height of the small-screen display size; determining a height reduction ratio for transforming the full-screen display size into the small-screen display size based on the first height and the second height; determining a width reduction ratio for transforming the full-screen display size into the small-screen display size based on the first width and the second width; and finally, determining the first mapping relationship based on the height reduction ratio and the width reduction ratio.

[0015] In one possible implementation of the first aspect, the above-mentioned second mapping processing of the second coordinates to obtain the first coordinates may specifically include: obtaining a second mapping relationship for the electronic device to change from a small screen display size to a full-screen display size. Based on the second mapping relationship, the second coordinates are processed again to obtain the first coordinates.

[0016] In one possible implementation of the first aspect, when the simulated click operation is an element click operation, determining the simulated click position corresponding to the simulated click operation on the small screen display can specifically include: analyzing the preset element corresponding to the element click operation. Then, finding the position corresponding to the preset element on the small screen display as the simulated click position. In this solution, after determining the element click operation of the preset element corresponding to the simulated click operation, finding the position of the preset element on the small screen display as the simulated click position ensures that the simulated click position is within the specified location on the small screen display.

[0017] In one possible implementation of the first aspect, when the simulated click operation is an element click operation, determining the simulated click position on the small screen display can specifically include: acquiring all elements contained in the small screen display and the position information corresponding to each element; analyzing the preset element corresponding to the element click operation and finding the position corresponding to the preset element from the position information. In this scheme, all elements in the small screen display and their corresponding positions can be acquired first. After determining that the operation corresponding to the audio data is an element click operation, the position of the preset element is then found from the position information of all elements to ensure that the simulated click position is the position on the small screen display.

[0018] In one possible implementation of the first aspect, the method further includes: displaying the simulated click trajectory at the location corresponding to the simulated click operation while performing the simulated click operation. This allows the user to more intuitively see the location where the simulated click operation was performed.

[0019] In one possible implementation of the first aspect, the aspect ratio of the full-screen display is the same as that of the small-screen display.

[0020] In a second aspect, an electronic device is provided, comprising: a processor, a memory, a display screen, and a recording device; the memory, the display screen, and the recording device are respectively coupled to the processor. The display screen is used to display the interface of the electronic device, the recording device is used to acquire audio data, and the memory is used to store computer execution instructions. When the electronic device is running, the processor executes the computer execution instructions stored in the memory to cause the electronic device to perform the voice control method as described in any of the first aspects above.

[0021] Thirdly, a computer-readable storage medium is provided that stores instructions which, when executed on a computer, enable the computer to perform any of the voice control methods described in the first aspect above.

[0022] Fourthly, a computer program product containing instructions is provided, which, when run on an electronic device, enables the electronic device to execute any of the voice control methods described in the first aspect above.

[0023] Fifthly, an apparatus (e.g., a system-on-a-chip) is provided, comprising a processor for supporting an electronic device in performing the functions described in the first aspect above. In one possible design, the apparatus further comprises a memory for storing program instructions and data necessary for the electronic device. When the apparatus is a system-on-a-chip, it may be composed of chips or may include chips and other discrete devices.

[0024] The technical effects of any of the design methods in aspects two through five can be found in the technical effects of different design methods in aspect one, and will not be repeated here. Attached Figure Description

[0025] Figure 1 A schematic diagram illustrating a scenario where a user uses voice to control an electronic device, as provided in an embodiment of this application;

[0026] Figure 2A A schematic diagram of a mobile phone settings interface provided in an embodiment of this application;

[0027] Figure 2B A schematic diagram of a voice-controlled mobile phone interface for performing operations, provided as an embodiment of this application;

[0028] Figure 2C A schematic diagram of a voice-controlled mobile phone interface for performing operations, provided as an embodiment of this application;

[0029] Figure 3AA schematic diagram of a mobile phone settings interface provided in an embodiment of this application;

[0030] Figure 3B This application provides a schematic diagram illustrating a process for triggering a one-handed mode in an embodiment of the present application.

[0031] Figure 3C A schematic diagram of an interface for voice control of a mobile phone in one-handed mode, provided as an embodiment of this application;

[0032] Figure 3D A schematic diagram of an interface for voice control of a mobile phone in one-handed mode, provided as an embodiment of this application;

[0033] Figure 4A A schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application;

[0034] Figure 4B A software structure block diagram of an electronic device provided in an embodiment of this application;

[0035] Figure 5A A flowchart illustrating a voice control method provided in an embodiment of this application;

[0036] Figure 5B A schematic diagram of a mobile phone interface in one-handed mode provided in an embodiment of this application;

[0037] Figure 5C A schematic diagram of a coordinate mapping provided for an embodiment of this application;

[0038] Figure 6 A schematic diagram of a coordinate mapping provided for an embodiment of this application;

[0039] Figure 7 A flowchart illustrating a voice control method provided in an embodiment of this application;

[0040] Figure 8 This is a structural block diagram of a chip system provided in an embodiment of this application. Detailed Implementation

[0041] With the development of technology, it is now possible for users to control electronic devices via voice.

[0042] For example, in scenarios like cooking and watching videos in the kitchen, or eating and watching TV, a user's hands may be occupied. In these situations, the user can typically control the electronic device by speaking voice commands. The electronic device uses a microphone to capture the user's voice, analyzes and recognizes it, and executes the corresponding voice commands, enabling the user to control the electronic device via voice. Figure 1As shown, user 10 can control mobile phone 1 to perform corresponding operations by speaking voice commands when within a certain range of mobile phone 1.

[0043] Generally, to conserve power and prevent accidental triggering of electronic devices, voice control functions need to be enabled before use. For example, a user can wake up an electronic device by inputting a preset word (called a wake-up word) through voice. Once awakened, the electronic device can execute the corresponding voice command, thus activating the voice control function. Alternatively, users can enable or disable the voice control function by toggling a preset switch on or off through the electronic device's user interface.

[0044] Voice control functions may have different names in different electronic devices, such as "voice control," "intelligent voice," "voice assistant," "see and speak," "voice command," "command at will," and "intelligent AI." The specific implementation of voice control functions with different names may also differ.

[0045] The following is a brief explanation of information that is visible and can be said:

[0046] It can be seen that this is achieved locally by electronic devices and does not require a network connection.

[0047] The "Visible and Speakable" feature is controlled by a preset switch. Users can turn on the preset switch to enable "Visible and Speakable" or turn it off to disable it.

[0048] Taking a mobile phone as an example, for instance, such as Figure 2A As shown, a user can open the settings function of mobile phone 1; for example, the user clicks the application icon of the "Settings" app on the desktop. In response to the user's click on the "Settings" app icon, mobile phone 1 displays the "Settings" interface 101. The "Settings" interface 101 includes a "Smart Voice" option 102, which is used to set up the smart voice function. For example, in response to the user's click on the "Smart Voice" option 102, mobile phone 1 displays the "Smart Voice" interface 103, which includes a "See and Speak" option 104. The user can click on the "See and Speak" option 104 to set options related to the "See and Speak" function. For example, refer to... Figure 2AIn response to a user's click on the "Speak When You See It" option 104, mobile phone 1 displays the "Speak When You See It" interface 105. Optionally, the "Speak When You See It" interface 105 includes a prompt message 106 to instruct the user on how to use the "Speak When You See It" function. The "Speak When You See It" interface 105 also includes a "Speak When You See It" switch 107. The user can click the "Speak When You See It" switch 107 to turn the "Speak When You See It" switch on or off. In one example, in response to receiving a user's click on the "Speak When You See It" switch 107, mobile phone 1's "Speak When You See It" switch turns on, enabling the "Speak When You See It" function. Optionally, the "Speak When You See It" interface 105 displays a prompt message 108 to instruct the user that the "Speak When You See It" function has been successfully enabled.

[0049] In one implementation, after the "See and Say" function is enabled, mobile phone 1 displays a first notification icon, which indicates that the "See and Say" function is enabled. For example, such as... Figure 2A As shown, after the "See and Say" switch 107 is turned on, the status bar of the mobile phone 1 displays a prompt icon 10a, indicating that the "See and Say" function has been enabled.

[0050] In one scenario, the preset switch for "See and Say" on the electronic device is not turned on, and the device's microphone is not activated. When the preset switch for "See and Say" is turned on, the electronic device activates the microphone and starts the corresponding recording channel in the system and driver. In this way, the voice (audio stream) captured by the microphone can be sent to the "See and Say" application for processing through the corresponding recording channel, enabling the user to control the electronic device via voice.

[0051] In another scenario, the preset switch for "Visible & Talkable" on the electronic device is not turned on. The device's microphone operates in a power-saving mode (e.g., searching for signals at low power) to pick up ambient sound. When the preset switch for "Visible & Talkable" is turned on, the corresponding recording channel for "Visible & Talkable" is activated in the device's system and drivers. In this way, the voice (audio stream) captured by the microphone can be sent to the "Visible & Talkable" application for processing through the corresponding recording channel, enabling the user to control the electronic device via voice.

[0052] Once "Visible and Talkable" is enabled, the corresponding recording channel is activated in the electronic device's system and drivers. The electronic device enters a "long-continuous recording" state, continuously capturing ambient sound. Users can issue commands to the electronic device at any time via voice, without needing to input a wake-up word.

[0053] In one implementation, after any function on the electronic device activates its voice recording function (turns on the microphone and establishes a recording channel), the electronic device will send a prompt message to the user, indicating that the device is in voice recording mode. This helps prevent the leakage of user privacy. For example, after the "see and speak" function is activated, the electronic device enters continuous voice recording mode, and a second prompt icon is displayed on the device's screen. This second prompt icon indicates that the recording channel is open, indicating to the user that the microphone is recording voice. For example, such as... Figure 2A As shown, the status bar of mobile phone 1 displays a prompt icon 10b, indicating that the recording channel is on, which is used to prompt the user that the microphone is capturing voice.

[0054] Once the "See It, Speak It" function is enabled, the electronic device establishes a recording channel and continuously captures the user's voice through the microphone. The user can input voice into the electronic device at any time. The electronic device parses and recognizes the user's voice input and executes the commands corresponding to the user's voice.

[0055] In some embodiments, the voice commands supported by the "See and Say" function can be divided into system commands and application commands. System commands can be used on any interface, while application commands are supported on the application interface displayed after some applications are launched. For example, system commands may include: return to the desktop, back, open an application, swipe up, swipe down, swipe left, swipe right, adjust brightness, and adjust volume, etc. Application commands may include: play, pause, full screen, like, favorite, next, previous, fast forward, and rewind, etc. It should be noted that the above commands are only some examples of the commands supported by the "See and Say" voice control function. That is, when the "See and Say" voice control function is activated, the user can control the phone 1 to perform the corresponding operation by speaking the corresponding voice command.

[0056] For controlling phone 1, different users may use different voice commands for the same purpose. Therefore, each of the above voice commands can correspond to multiple different expressions. For example, for the system voice command to return to the home screen, the user can say any of the following: "home screen," "return to home screen," "return to home screen," or "return to home screen" to control phone 1 to return to the home screen. Correspondingly, when phone 1 detects the above voice expression, it can recognize it as a voice command to return the current page to the home screen.

[0057] In one example, such as Figure 2B As shown, mobile phone 1 displays the main screen interface 201. The user inputs the voice command "Open Calendar" into mobile phone 1. In response to receiving the user's voice command "Open Calendar", mobile phone 1 launches the calendar application and displays the calendar application's user interface 202.

[0058] In another example, such as Figure 2C As shown, mobile phone 1 displays the interface 203 of a short video application. The user inputs the voice command "pause" into mobile phone 1. In response to receiving the user's voice command "pause," mobile phone 1 executes the corresponding command and displays interface 204. A "play" button 205 is displayed on interface 204 to indicate to the user that the short video has been paused. The user can click the "play" button 205 to continue playing the short video. Alternatively, the user can input the voice command "play" into mobile phone 1 to continue playing the short video.

[0059] Furthermore, on the aforementioned display interface 203, users can also use the visible-to-speak voice control function to control actions such as liking and saving the currently playing short video.

[0060] In some embodiments, the specific implementation process of the mobile phone 1 responding to voice commands to perform corresponding operations can be referred to the description in related technologies, and will not be repeated in the embodiments of this application. For example, after determining that the audio data includes voice commands supported by the visible and speakable voice control function, the specific operation corresponding to the voice command can be executed through the multi-mode input framework of the mobile phone 1.

[0061] In some embodiments, as can be seen, after the voice control function is activated, mobile phone 1 will obtain the text data displayed on the current page in real time, analyze and obtain the words contained therein, and the positions of each word and its related content (such as the application icon of the corresponding application) on the current page. For example, as shown below... Figure 2B The main screen interface 201 shown allows mobile phone 1 to obtain and analyze the text data displayed on the current page to identify words such as "clock," "calendar," "gallery," "memo," "file management," "email," and "music." It should be noted that the specific implementation process of mobile phone 1 obtaining and analyzing the text data displayed on the current page to identify the included words can be found in related technical descriptions and will not be elaborated upon in this embodiment. The location of the word "calendar" on the current page can be the location of the word "calendar" itself; or, the location of the word "calendar" on the current page can also be the location of display content related to "calendar." Figure 2B In the main screen interface 201 shown, the display content related to "Calendar" can be the location of the application icon for the "Calendar" application; such as... Figure 2B The area circled by the dashed line in the main screen interface 201 shown.

[0062] It should be noted that the above-described voice control function is merely an example of one type of voice control function. In other embodiments, the voice control function may also be a voice control function. For example, the voice control function may also be: a voice assistant function that is activated by a preset wake-up word, a voice command function that is activated by a preset scenario, and so on.

[0063] Additionally, in many scenarios, users may want to use their phones with one hand. However, some phones are too large, making it impossible for users to tap all areas of the screen with one hand. In such cases, some phones offer one-handed operation features, which shrink the displayed area on the screen to allow users to more easily tap any area of ​​the screen with one hand. It should be noted that the aforementioned one-handed operation feature can correspond to an entire application or a function within an application.

[0064] Please refer to Figure 3A The settings interface also includes accessibility settings option 300. In response to the user's activation of accessibility settings option 300, access can be made... Figure 3A The settings interface 301 for one-handed operation is shown. The settings interface 301 allows setting a switch 302 to enable one-handed operation. This switch 302 is used to turn on or off whether users are allowed to trigger one-handed mode via gestures. Specifically, after switch 302 is turned on, the user can trigger the one-handed mode of phone 1 and enter a small-screen display state by performing a preset gesture on the screen. The preset gesture can be set according to actual conditions and is not limited in this embodiment. It should be noted that the above-mentioned one-handed operation function can correspond to an application or a function within an application.

[0065] For example, taking switch 302 in the open state as an example, please refer to... Figure 3B The system displays the main screen interface 303 of mobile phone 1. When the user performs a preset gesture on the main screen interface 303 of mobile phone 1, it will trigger the one-handed mode of mobile phone 1. That is, mobile phone 1 will bring up a smaller screen, such as... Figure 3B The scaled-down home screen interface 304 is shown. Furthermore, in Figure 3B In area 305 outside the shrunken main screen interface 304, the phone displays the message: "One-handed mode, tap outside the small screen to exit." The user can exit one-handed mode by tapping outside this area of ​​the shrunken main screen interface 304, restoring the display to full-screen size.

[0066] Understandably, the scaled-down home screen interface 304 is closer to the right side of the screen, making it more suitable for right-handed users. In other embodiments, after the user performs a preset gesture on the home screen interface 303, the phone 1 can also enter... Figure 3B The scaled-down home screen interface 306 shown is more suitable for users operating the phone with their left hand. It should be noted that in other embodiments, after the phone 1 is triggered to enter one-handed mode, the size and location of the small screen display area can be adjusted according to the user's wishes to display it at a different size or in a different location.

[0067] Furthermore, in some embodiments, the process of switching from the main screen interface 303 to the scaled-down main screen interface 304, or from the main screen interface 303 to the scaled-down main screen interface 306, where the display size changes from full-screen to small-screen, is scaled down proportionally. That is, the aspect ratio of the display in the small-screen display state is the same as that in the full-screen display state.

[0068] It should be noted that the scaled-down home screen interface 304 and scaled-down home screen interface 306 described above are two display examples of the one-handed mode display state. In other embodiments, in the one-handed mode display state, the mobile phone 1 can also display other screens, such as a half-screen display near the bottom of the screen.

[0069] Furthermore, the following explanation uses the example of a user speaking "Open [Calendar]" on the minimized home screen 304. Phone 1 analyzes this voice message and determines that it corresponds to the voice command "Open [Application]" supported by the visible-to-speak voice control function, and that it contains the word "[Calendar]". Simultaneously, by acquiring the words contained in the minimized home screen 304, the phone can confirm that the minimized home screen 304 contains the word "[Calendar]". At this point, in response to the voice command "Open [Calendar]", Phone 1 performs a simulated click operation on the word "[Calendar]" and its corresponding control at the location on the current smaller screen display.

[0070] Specifically, in some embodiments, in response to the voice command "Open [Calendar]", mobile phone 1 injects a simulated click event into the system (such as a multi-modal input frame). This simulated click event carries the position D1 of [Calendar] in the current small-screen display, such as... Figure 3C The area circled in solid circles is 307.

[0071] However, in one-handed mode, for position D1, the system will automatically convert it to position D2 where the [Calendar] would appear in full-screen display. For example, D2 could be as follows: Figure 2B The dotted circle area shown in the main screen interface 201 (simulating the click location), or D2 can also be in... Figure 3CThis location corresponds to the dotted circle area 308. Then, the system performs a simulated click operation on this location D2. For example... Figure 3C As shown, the area 308 circled by the dotted line does not correspond to the location of the calendar application in one-handed mode. Therefore, if phone 1 performs a simulated click operation in area 308 circled by the dotted line, it will not be able to accurately open the calendar application.

[0072] Furthermore, in this example, the area 308 circled by the dotted line is outside the scaled-down main screen interface 304. Therefore, in response to the voice command, when phone 1 performs a simulated click operation at the corresponding position in the area 308 circled by the dotted line, it may cause phone 1 to exit one-handed mode and enter a state similar to... Figure 3C The home screen interface shown is 309. In other words, for the user, after saying "Open [Calendar]" with voice command, what they see is phone 1 exiting one-handed mode, not phone 1 opening the calendar application.

[0073] Alternatively, in other embodiments, if the area 308 circled by the dashed line is located elsewhere on the scaled-down main screen interface 304, then after the simulated click operation is performed in response to the voice command, the display interface of phone 1 may not change at all, or the click may hit the icon of another application, causing phone 1 to open another application. For the user, after saying "Open [Calendar]", no change is observed in the display interface of phone 1. Alternatively, after the user says "Open [Calendar]", they see that phone 1 has opened another application.

[0074] In other embodiments, taking the user's spoken voice command "swipe down" in the context of the scaled-down home screen interface 304 as an example, mobile phone 1 can determine from the voice content that it corresponds to the "swipe down" voice command supported by the visible-and-speakable voice control function. This voice command corresponds to a simulated click operation. During the simulated click operation in response to this voice command, mobile phone 1 needs to determine the simulated click start position and the simulated click end position. Then, based on the simulated click start position and the simulated click end position, a simulated swipe operation is performed. It can be understood that the swipe operation can be understood as a series of click operations. Specifically, mobile phone 1 can first perform a simulated press operation at the simulated click start position, then perform a simulated swipe operation from the simulated click start position to the simulated click end position, and finally perform a simulated release operation at the simulated click end position to complete the simulated swipe operation. It should be noted that the specific implementation process of mobile phone 1 determining the simulated click start position and the simulated click end position can be referred to in the description of related technologies, and will not be repeated in this embodiment.

[0075] As explained above, in the small-screen display state of one-handed mode, the system of phone 1 automatically converts the position of the simulated click operation to its position within the full-screen display size before performing the simulated click operation. Similarly, during the simulated swipe operation performed by phone 1, the system also automatically converts the positions D3 (the starting position of the simulated click) and D4 (the ending position of the simulated click) determined by phone 1 to their positions within the full-screen display size of phone 1. For example... Figure 3D As shown, after receiving the voice command "swipe down," phone 1 can determine the simulated tap start position D3 and simulated tap end position D4. Before executing the simulated swipe operation, the system will convert these positions D3 and D4 to... Figure 3D Positions D5 and D6 are shown. According to... Figure 3D As shown, D5 is located outside the shrunk main screen area. Therefore, when phone 1 performs a simulated swipe operation, if it performs a simulated press operation at the simulated click starting position, the system might mistakenly identify it as a click operation outside the shrunk screen, causing phone 1 to exit one-handed mode. For example... Figure 3D As shown, after a user says the voice command "swipe down", the phone 1 may change from a minimized home screen display 310 to a full-screen home screen display 311. However, the effect of this response operation on the phone 1 does not correspond to the user's voice command "swipe down".

[0076] The above scenarios all indicate that in one-handed mode, when a user performs a simulated click operation by controlling phone 1 via voice, the click location may be incorrect. When the simulated click location is incorrect, the operation performed by the phone may not match the user's intention, causing inconvenience. It should be noted that the voice commands "open [calendar]" and "swipe down" in the above embodiments are only two examples of voice commands corresponding to simulated click operations. In other embodiments, voice commands corresponding to simulated click operations may also include other voice commands, such as swiping up, swiping left, and swiping right. The operation to be performed corresponding to the voice command "open [calendar]" can be categorized as an element click operation, while the operations to be performed corresponding to the voice commands "swipe up," "swipe down," "swipe left," and "swipe right" can be categorized as swipe operations.

[0077] Based on this, this application proposes a voice control method applied to scenarios where users use voice to control electronic devices to perform operations in one-handed mode. For example, the electronic device can be the aforementioned mobile phone 1. In the technical solution provided in this application embodiment, when mobile phone 1 detects that the operation corresponding to the audio data is a simulated click operation, if it is determined that mobile phone 1 is currently in a small-screen display in one-handed mode, the simulated click position in the small-screen display is obtained. Then, mobile phone 1 executes the simulated click operation according to the simulated click position. Specifically, when the simulated click operation is an element click operation (such as "open [Calendar]" as mentioned above), the task triggered by the interface element corresponding to the element click operation is executed (such as opening the [Calendar] application and displaying the calendar interface). When the simulated click operation is a swipe operation (such as "swipe down" as mentioned above), the start position and end position of the simulated click are obtained, and then the display screen is swiped from the start position to the end position. This ensures that in scenarios where users use voice to control electronic devices to perform simulated click operations in one-handed mode, mobile phone 1 responds to the voice command for the simulated click and executes the simulated click operation at the accurate position, triggering the corresponding response operation and avoiding the problem of simulated click errors. Provided that the user's voice commands are accurately recognized by mobile phone 1, mobile phone 1 can respond to the simulated click voice command and accurately execute the simulated click operation.

[0078] The voice control method provided in this application embodiment can be applied to an electronic device 100 including a microphone. The aforementioned electronic device 100 may include mobile phones, tablets, laptops, personal computers (PCs), ultra-mobile personal computers (UMPCs), handheld computers, netbooks, smart home devices (e.g., smart TVs, smart screens, large screens, smart speakers, smart air conditioners, etc.), personal digital assistants (PDAs), wearable devices (e.g., smartwatches, smart bracelets, etc.), in-vehicle devices, virtual reality devices, etc., and this application embodiment does not impose any limitations on these.

[0079] like Figure 4AThe diagram shows a schematic of an electronic device 100 according to an embodiment of this application. For example, the electronic device 100 can be the aforementioned mobile phone 1. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, antenna 1, antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a sensor module 180, buttons 190, a motor 191, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a touch sensor 180B, etc.

[0080] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0081] Processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. For example, processor 110 is used to execute the voice control method in the embodiments of this application.

[0082] The controller can be the nerve center and command center of the electronic device 100. The controller can generate operation control signals according to the instruction opcode and timing signals to complete the control of fetching and executing instructions.

[0083] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0084] USB interface 130 is an interface that conforms to the USB standard specification, specifically it can be a Mini USB interface, Micro USB interface, USB Type C interface, etc. USB interface 130 can be used to connect a charger to charge electronic device 100, and it can also be used for data transfer between electronic device 100 and peripheral devices.

[0085] The external storage interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0086] Internal memory 121 can be used to store executable program code, which includes instructions. Processor 110 executes various functional applications and data processing of electronic device 100 by running the instructions stored in internal memory 121. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function (such as sound playback, image playback, etc.).

[0087] In addition, the internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0088] The charging management module 140 is used to receive charging input from the charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130.

[0089] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc.

[0090] In some other embodiments, the power management module 141 may also be located within the processor 110. In other embodiments, the power management module 141 and the charging management module 140 may also be located in the same device.

[0091] The wireless communication function of electronic device 100 can be realized through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0092] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with tuning switches.

[0093] The mobile communication module 150 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1.

[0094] The wireless communication module 160 can provide solutions for wireless communication applications on the electronic device 100, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0095] In some embodiments, antenna 1 of electronic device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, so that electronic device 100 can communicate with networks and other devices through wireless communication technology.

[0096] Electronic device 100 can implement audio functions through audio module 170 and application processor, such as music playback and recording.

[0097] The audio module 170 is used to convert digital audio signals into analog audio signals for output, and also to convert analog audio inputs into digital audio signals. The audio module 170 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 170 may be located in the processor 110, or some functional modules of the audio module 170 may be located in the processor 110. The audio module 170 may include a speaker 170A and a microphone 170B.

[0098] The loudspeaker 170A, also known as a "speaker", is used to convert audio electrical signals into sound signals. Electronic device 100 can listen to music or conduct video conferences through the loudspeaker 170A.

[0099] Microphone 170B, also known as a "microphone," is used to convert sound signals into electrical signals. During video calls, video conferences, or when using a voice assistant, the user can speak by bringing their mouth close to microphone 170B, inputting sound signals into microphone 170B. Electronic device 100 may have at least one microphone 170B. In some embodiments, electronic device 100 may have two microphones 170B, which, in addition to acquiring sound signals, can also perform noise reduction. In other embodiments, electronic device 100 may also have three, four, or more microphones 170B, enabling sound signal acquisition, noise reduction, sound source identification, and directional recording functions, etc.

[0100] Pressure sensor 180A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 180A can be disposed on display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When a force is applied to pressure sensor 180A, the capacitance between the electrodes changes. Electronic device 100 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 194, electronic device 100 detects the touch operation intensity based on pressure sensor 180A. Electronic device 100 can also calculate the touch position based on the detection signal from pressure sensor 180A.

[0101] Touch sensor 180B, also known as a "touch panel," can be located on display screen 194. The touch sensor 180B and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180B is used to detect touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180B may also be located on the surface of electronic device 100, in a different position than display screen 194.

[0102] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Electronic device 100 can receive button input and generate key signal inputs related to user settings and function control of electronic device 100.

[0103] Motor 191 can generate vibration alerts. Motor 191 can be used for incoming call vibration alerts or for touch vibration feedback.

[0104] Electronic device 100 implements display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0105] The display screen 194 is used to display images, videos, etc. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0106] Camera 193 is used to capture still images or videos. In some embodiments, electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0107] The SIM card interface 195 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 195 to achieve contact and separation with the electronic device 100. The electronic device 100 can support one or N SIM card interfaces, where N is a positive integer greater than 1.

[0108] The voice control methods described in the following embodiments can all be implemented in the electronic device 100 with the above-described hardware structure.

[0109] The software system of electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 100.

[0110] Figure 4B This is a software structure block diagram of the electronic device 100 according to an embodiment of this application.

[0111] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: the application layer, the application framework layer, the Android runtime (ART) and native C / C++ libraries, the Hardware Abstraction Layer (HAL), and the kernel layer.

[0112] The application layer can include a series of application packages.

[0113] like Figure 4B As shown, the application package may include applications such as camera, gallery, calendar, map, music, navigation, SMS, call, video, voice control module, and one-handed mode control module.

[0114] The voice control module responds to user-spoked voice commands and executes corresponding operations. The one-handed mode control module responds to a one-handed mode trigger event, enters the one-handed mode small screen display state, and displays the small screen screen.

[0115] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0116] like Figure 4B As shown, the application framework layer may include a window manager, content provider, view system, resource manager, notification manager, activity manager, input manager, and multimodal input framework, etc.

[0117] The window manager provides Window Manager Service (WMS), which can be used for window management, window animation management, surface management, and as a relay station for the input system.

[0118] Content providers store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, phone calls made and received, browsing history and bookmarks, phone books, etc.

[0119] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0120] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0121] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0122] The Activity Manager Service (AMS) can be used to start, switch, and schedule system components (such as activities, services, content providers, and broadcast receivers), as well as manage and schedule application processes.

[0123] The Input Manager Service (IMS) provides input management services, which can be used to manage system inputs such as touchscreen input, keypad input, and sensor input. IMS retrieves events from input device nodes and, through interaction with the WMS (Windows Management System), distributes these events to appropriate windows.

[0124] The multimodal input framework encapsulates various interfaces for other subsystems and applications to call. In this embodiment, after determining the operation to be performed by the voice command based on the voice control function, the functional module corresponding to the voice control function can send the command to the multimodal input framework. Subsequently, the multimodal input framework responds to the command and can execute the corresponding operation, thereby achieving the effect of voice-controlled electronic device to perform operations.

[0125] The Android runtime consists of the core libraries and the Android runtime itself. The Android runtime is responsible for converting source code into machine code. The Android runtime primarily employs ahead-of-time (AOT) compilation and just-in-time (JIT) compilation techniques.

[0126] The core library primarily provides basic Java class library functionalities, such as libraries for fundamental data structures, mathematics, I / O, tools, databases, and networking. It also provides APIs for users to develop Android applications.

[0127] Native C / C++ libraries can include multiple functional modules. Examples include: surface manager, media framework, libc, OpenGL ES, SQLite, Webkit, etc.

[0128] The Surface Manager manages the display subsystem and provides fusion of 2D and 3D layers for multiple applications. The Media Framework supports playback and recording of various common audio and video formats, as well as still image files. The Media Library supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG. OpenGL ES provides drawing and manipulation of 2D and 3D graphics in applications. SQLite provides a lightweight relational database for applications on the electronic device.

[0129] The Hardware Abstraction Layer (HAL) runs in user space, encapsulates kernel-level drivers, and provides calling interfaces to the upper layers.

[0130] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0131] The following is a brief introduction to the technical terms that may be involved in the embodiments of this application.

[0132] Automatic Speech Recognition (ASR) aims to convert the lexical content of human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.

[0133] Natural Language Understanding (NLU) is a general term for all methods, models, or tasks that support machines in understanding text content.

[0134] Dialogue Management (DM) is used to control the process of human-computer dialogue and determine the current response to the user based on dialogue history information.

[0135] The voice control method provided in this application is mainly applied to electronic devices. Specifically, it can be applied to scenarios where users control electronic devices via voice in a small-screen display state under one-handed mode. For simulated click voice commands spoken by the user in the above scenario, simulated clicks can be executed at accurate locations.

[0136] Taking a mobile phone as an example, the voice control method provided in this application embodiment may specifically include: after entering a small screen display and activating the voice control function, the mobile phone collects audio data. The mobile phone performs voice recognition on the audio data. The mobile phone determines whether the operation corresponding to the audio data is a simulated click operation. If so, the mobile phone obtains the simulated click position corresponding to the voice command of the simulated click on the small screen display. Then, the mobile phone executes the simulated click operation according to this simulated click position. Specifically, if the simulated click operation is an element click operation (such as "open [Calendar]" mentioned above), the task triggered by the interface element corresponding to the element click operation is executed (such as opening the [Calendar] application and displaying the calendar interface). If the simulated click operation is a swipe operation (such as "swipe down" mentioned above), the simulated click start position and simulated click end position are obtained, and then the display screen is swiped from the simulated click start position to the simulated click end position.

[0137] The voice control method provided in this application embodiment ensures that, in one-handed mode, when a user uses voice control to perform a simulated click operation on an electronic device, the phone can execute the simulated click operation at the accurate location, avoiding errors in the simulated click. Provided the phone accurately recognizes the user's voice command, it can accurately execute the simulated click operation based on the user's voice command. This avoids the problem of inconsistency between the click operation performed by the phone and the intent of the user's spoken voice command.

[0138] Furthermore, for the simulated click position in the small-screen display obtained from the audio data, the phone first maps it to the corresponding second coordinate in the full-screen display. Specifically, this mapping process involves mapping the point from the full-screen display size to the small-screen display size. Then, a simulated click event is generated based on the second coordinate and sent to the phone's multimodal input frame. This causes the multimodal input frame to respond to the simulated click event by performing a second mapping operation, mapping the second coordinate to the full-screen display size.

[0139] Since the first coordinate in the small-screen display has already undergone the first mapping process to obtain the second coordinate, when the multi-modal input frame performs a simulated click operation, it will automatically default to the second coordinate being a point in the small-screen display. Subsequently, it will also automatically map this second coordinate to the full-screen display; that is, the multi-modal input frame will perform a second mapping process on the second coordinate. Because the first mapping process maps from the full-screen display size to the small-screen display size, and the second mapping process maps from the small-screen display size to the full-screen display size, the second mapping process based on the multi-modal input frame yields the first coordinate. Finally, performing a simulated click operation on the first coordinate based on the multi-modal input frame ensures that the simulated click operation still corresponds to the coordinates of the simulated click position in the audio data within the small-screen display. In other words, it ensures that the phone can accurately perform the corresponding simulated click operation at the precise location according to the audio data.

[0140] The voice control method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0141] Please refer to Figure 5A This is a flowchart illustrating a voice control method provided in some embodiments of this application. In this embodiment, the method includes steps S500-S513:

[0142] S500. Phone starts up.

[0143] S501. When the one-handed operation function is enabled, in response to the first event, the small screen display corresponding to the one-handed mode is displayed.

[0144] The first event can be a one-handed mode trigger event. For example, a user can open a... Figure 3A After the switch 302 is activated, a preset gesture is executed on the phone to trigger the phone to enter one-handed mode. In this embodiment, the aforementioned first event may specifically be the detection of a preset gesture. For example, the preset gesture may specifically be a horizontal swipe and pause along the bottom of the screen, a swipe down from the bottom of the screen, a swipe into the screen from the bottom left or right corner, etc.

[0145] Understandably, in some other embodiments, if the user has not enabled the one-handed operation function in advance, the phone will not trigger the one-handed mode after detecting the preset gesture.

[0146] As explained above, entering one-handed mode reduces the size of the display on the phone screen. Therefore, upon detecting the first event, the phone responds by changing the display size from full-screen to a smaller screen size, such as... Figure 3B As shown, in response to the first event, the phone switches from display interface 303 to display interface 304, or switches to display interface 306.

[0147] S502. In response to the second event, activate the voice control function.

[0148] In some embodiments, the second event described above may specifically correspond to the user manually activating the voice control function. For example, the user can trigger a voice control function by... Figure 2A The switch 107 in interface 105 shown is used to activate the voice control function. In other embodiments, the aforementioned second event may also correspond to the user activating the intelligent human-computer interaction function. In this embodiment, after activating the intelligent human-computer interaction function, the phone will automatically activate the first voice control function to better provide the user with an interaction method with the phone. In other embodiments, the aforementioned second event may also be other events.

[0149] S503. Collect audio data based on voice control function.

[0150] As described above, in some embodiments, after the voice control function is activated, the mobile phone is in a state of continuously collecting audio data. In this embodiment, after the mobile phone activates the voice control function in S502, the mobile phone can collect audio data based on the voice control function. In some embodiments, S503 may specifically include: after activating the voice control function, the mobile phone calls the recording device through the voice control module corresponding to the voice control function to start collecting audio data.

[0151] Furthermore, in some embodiments, after the voice control module invokes the recording device to begin collecting audio data, if audio data is collected, the corresponding speech recognition engine of the voice control module needs to perform speech recognition to determine whether it contains voice commands supported by the voice control function. Therefore, in some embodiments, after the voice control module invokes the recording device, the above method further includes: activating the speech recognition engine corresponding to the voice control module. It should be noted that the specific implementation process of activating the speech recognition engine can be referred to the description in related technologies, and will not be repeated in the embodiments of this application.

[0152] In other embodiments, the voice control function may also be a command-based voice control function. After being activated, the command-based voice control function only triggers audio data acquisition and voice recognition when a preset scenario is triggered. In this embodiment, S503 may specifically include: after detecting that the mobile phone has triggered a preset scenario, acquiring audio data based on the voice control function.

[0153] Alternatively, in other embodiments, the voice control function can also be a voice assistant or other voice control function. After the mobile phone activates the voice assistant function, specifically, upon detecting a preset wake-up word, the voice assistant function can be activated to begin collecting audio data and performing speech recognition. In this embodiment, the above-mentioned S503 may specifically include: after detecting the preset wake-up word, collecting audio data based on the voice control function.

[0154] It should be noted that the above-described method of triggering the voice control function to start collecting audio data and performing voice recognition is only an example. In other embodiments, the voice control function can also trigger the mobile phone to start collecting audio data based on the voice control function in other ways.

[0155] S504. Perform speech recognition on the audio data to determine whether the audio data contains preset voice commands.

[0156] As explained above, the voice control function can support multiple voice commands. The preset voice command can be any of the voice commands supported by the voice control function. In some embodiments, the preset voice command can be a system command, such as: return to desktop, back, open application, swipe up, swipe down, swipe left, swipe right, adjust brightness or volume, etc. The preset voice command can also be an application command, such as: play, pause, full screen, like, favorite, next, previous, fast forward or rewind, etc.

[0157] Alternatively, in other embodiments, the preset voice commands may include different voice expressions corresponding to different commands. For example, different voice expressions for different commands: desktop, return to desktop, back, previous page, open [application name], swipe up, swipe down, swipe down, etc.

[0158] If the judgment result of S504 is negative, meaning the phone did not recognize any voice commands supported by the voice control function in the audio data, it indicates that the user may not have intended to control the phone via voice. The phone can then continue collecting audio data and performing voice recognition to determine if the user has uttered a voice command supported by the voice control function. In other words, if the judgment result of S504 is negative, the phone can return to executing S503 to continue collecting audio data.

[0159] Alternatively, if S504's judgment result is negative, it's possible that the phone misrecognizes the audio data and fails to accurately identify the user's voice command. In this case, the user needs to repeat the accurate voice command. In some embodiments, if S504's judgment result is negative, the phone can display a prompt message on the interface, prompting the user to try repeating the voice command.

[0160] If the S504 judgment result is yes, it means that the user wishes to control the phone to perform the corresponding operation via voice. Therefore, if the S504 judgment result is yes, the phone can perform the corresponding operation in response to the preset voice command.

[0161] S505. Determine whether the operation corresponding to the preset voice command is a simulated click operation.

[0162] As described above, in some embodiments, the simulated click operation may specifically include element click operations and swipe operations. Combined with... Figure 2A The voice control function shown supports the following voice commands: voice commands for element click operations can include application commands to open applications, and voice commands for swipe operations can include system commands such as swipe up, swipe down, swipe left, and swipe right.

[0163] In some embodiments, for the aforementioned application commands such as play, pause, full screen, like, favorite, previous, next, fast forward, and rewind, voice control can be performed by the mobile phone through simulated click operations. In other embodiments, when the mobile phone executes the aforementioned application commands via voice control, it can also be achieved through interaction between the mobile phone and the server of the corresponding application.

[0164] Therefore, in some embodiments, the preset voice commands corresponding to simulated click operations may include any of the following: system commands such as open application, swipe up, swipe down, swipe left, and swipe right; and application commands such as play, pause, full screen, like, favorite, previous, next, fast forward, and rewind. Further, the simulated click operations can be divided into two categories: element click operations and swipe operations. Element click operations specifically include: open application, play, pause, full screen, like, favorite, previous, next, fast forward, and rewind. Swipe operations include swipe up, swipe down, swipe left, and swipe right.

[0165] If the S505's judgment result is yes, it means that the user wants to control the phone to perform a simulated click operation through the voice control function.

[0166] S506. Obtain the simulated click position in the small screen display.

[0167] In some embodiments, S506 may specifically include: obtaining the first coordinates of the simulated click position in the small screen display in the screen coordinate system. Taking the simulated click operation as an element click operation of "open [Calendar]" as an example, its corresponding simulated click position in the small screen display could be the position of the [Calendar] application icon in the small screen display. This simulated click position has a unique coordinate in the screen coordinate system, namely the aforementioned first coordinate. Taking the simulated click operation as a swipe operation as an example, its corresponding simulated click position in the small screen display may specifically include the simulated click start position and simulated click end position in the small screen display. This simulated click start position and simulated click end position also each correspond to a unique coordinate in the screen coordinate system.

[0168] In some embodiments, when the simulated click operation is an element click operation, the mobile phone can first analyze all elements in the currently displayed screen and their corresponding position information. If it is determined that the operation corresponding to the audio data is an element click operation, the preset element corresponding to that element click operation is analyzed. Then, the position of the preset element is found from the position information of all elements, which is the simulated click position on the small screen display screen corresponding to that element click operation. In this way, it can be ensured that the simulated click position is the position on the small screen display screen.

[0169] As explained above, the mobile phone can pre-obtain the text data displayed on the current page and analyze the text data to obtain all the words contained on the current page, as well as the positions of each word and its corresponding icon on the current page. For example, taking S504 and S505 as determining that the preset voice command is "Open [Calendar]", and the words on the current page include the word "[Calendar]", the mobile phone can obtain the position of the word "[Calendar]" on the current page. It can be understood that the simulated click position corresponding to the voice command "Open [Calendar]" is the position of the word "[Calendar]" on the current page. Taking the position of the word "[Calendar]" on the current page including the position of the application icon of the [Calendar] application as an example, in this embodiment, in S506 above, the simulated click position obtained by the mobile phone can specifically be the position of the application icon of the [Calendar] application on the current page. For example, in the small-screen display state of the mobile phone in one-handed mode, the position of the word "[Calendar]" is... Figure 3C The area shown is 307.

[0170] In other embodiments, when the simulated click operation is an element click operation, the mobile phone can first determine the preset element corresponding to the element click operation. Then, it can find the location of the preset element in the currently displayed page, which is used as the simulated click position on the small screen display. This ensures that the simulated click position is within the small screen display.

[0171] As can be seen from the above embodiments, the preset voice command corresponding to the simulated click operation, besides "open [application]", can also be a swipe command, such as swipe up, swipe down, swipe left, and swipe right. In this embodiment, the simulated click position can include two positions: the simulated click start position and the simulated click end position. Taking S504 and S505 determining that the preset voice command is "swipe down" as an example, in S506, the simulated click position corresponding to the simulated click operation obtained by the mobile phone can specifically include, for example, Figure 3D D5 and D6 are shown.

[0172] In other embodiments, if the audio data contains other instructions, S506 can also obtain the simulated click position corresponding to the simulated click operation in other ways.

[0173] S507. Obtain the full-screen display size in full-screen mode, the small-screen display size in one-handed mode, and the first mapping relationship between the full-screen display size and the small-screen display size.

[0174] In some embodiments, the mobile phone can obtain the full-screen display size when the phone is in full-screen mode. After the phone enters the small-screen display mode in one-handed mode, the small-screen display size can be obtained. Then, based on the full-screen display size and the small-screen display size, a first mapping relationship between the phone's transformation from the full-screen display size to the small-screen display size can be determined.

[0175] The full-screen display size may specifically include a first width and a second height, while the small-screen display size may include a second width and a second height. Further, based on the aforementioned first width and second height, the width reduction ratio when transforming the first width to the second width, and the height reduction ratio when transforming the first height to the second height, can be determined. Then, combining the coordinates 1 in the full-screen display and 2 in the small-screen display of the same interface element, along with the width reduction ratio and height reduction ratio, a first mapping relationship from the full-screen display size to the small-screen display size can be determined through coordinate conversion. In some embodiments, the first mapping relationship from the full-screen display size to the small-screen display size can also be represented as a mapping relationship from coordinate 1 to coordinate 2.

[0176] As explained above, in different embodiments, the small screen display size of the phone in one-handed mode can be adjusted according to the user's needs. In other words, the small screen display size of the phone may differ in different embodiments.

[0177] S508. Based on the first mapping relationship, perform the first mapping process on the first coordinates corresponding to the simulated click position to obtain the second coordinates.

[0178] The first mapping process is used to map the coordinates of a point from the full-screen display to the smaller screen display. For example, the first mapping process is used to... Figure 3C The region 308 shown is mapped to region 307. Specifically, each coordinate included in region 308 can be mapped to each coordinate in region 307, or the coordinates of the center position of region 308 can be mapped to the coordinates of the center position of region 307. Alternatively, the first mapping process is used to... Figure 3D Position D5 is mapped to position D3, and position D6 is mapped to D4.

[0179] As explained above, the first coordinate is the coordinate of the simulated click position in the small-screen display, and it is referenced to the screen coordinate system. Therefore, the second coordinate obtained by performing the first mapping process on the first coordinate is still a coordinate referenced to the screen coordinate system. However, this second coordinate has no practical meaning in either the full-screen or small-screen display. Taking the preset voice command "Open [Application]" as an example, the first coordinate of the simulated click operation in the small-screen display is the corresponding position of the [Calendar] application in the small-screen display, which is... Figure 5B The location corresponding to area 521 in the scaled-down main screen interface 520 shown.

[0180] Furthermore, as explained above, in one-handed mode, before the phone performs a simulated click operation, the system (such as a multi-modal input framework) automatically assigns the simulated click location (e.g., the location of the simulated click carried by the injected click event) to the input. Figure 3C The area shown is 307, and Figure 3D The positions of D3 and D4 shown are automatically calculated in the full-screen display (e.g., ...). Figure 3C The area shown is 308, and Figure 3D (As shown in D5 and D6). Then, perform simulated clicks according to the positions in the full-screen display. Combined with Figure 3C and Figure 3D As shown, in one-handed mode, simulating a click at the location shown in the full-screen display is incorrect. The correct location for the simulated click should be the location shown in the smaller screen, which is the first coordinate.

[0181] Therefore, in order to ensure that the position when the phone performs a simulated click is the same as the position on the small screen display, the simulated click position needs to undergo a first mapping process, that is, converting it from the first coordinate to a virtual second coordinate. In other words, it involves... Figure 5B The region 521 shown is subjected to the first mapping process (such as...). Figure 5B The L1 mapping process shown yields a virtual region 522. (Refer to...) Figure 5B It is known that the second coordinate obtained by performing the first mapping process on the first coordinate does not correspond to the position of the [Calendar] application in the full-screen or small-screen display. Therefore, when the system performs a simulated click operation, the first coordinate can be obtained by automatically mapping the second coordinate (from the small-screen display to the full-screen display). Finally, the phone can perform a simulated click operation at the location of the first coordinate.

[0182] S509. Generate simulated click events based on the second coordinate.

[0183] S510. Sends a simulated click event to the phone's multi-mode input frame.

[0184] As explained above, in one-handed mode, when the phone's system performs a simulated click, it uses the coordinates (second coordinates) obtained from the simulated click event as the coordinates in the small-screen display, and automatically performs a second mapping process to transfer these coordinates to the full-screen display. Finally, it performs the simulated click operation at the position corresponding to the coordinates in the full-screen display obtained after this second mapping process.

[0185] S511. Based on the multi-modal input framework, perform a second mapping process on the second coordinate to obtain the first coordinate.

[0186] The second mapping process maps the coordinates of a point from the small-screen display to the full-screen display. It can be understood that the second mapping process and the first mapping process are opposites. For example, taking the mapping process from point A to point C as the first mapping process, the second mapping process is the mapping process from point C to point A. Therefore, when the phone's system performs the second mapping process on the second coordinates and converts them to coordinates in the full-screen display, it obtains the aforementioned first coordinates. Please refer to... Figure 5B As shown, S511 above performs a second mapping process on the region 522 corresponding to the second coordinate (such as...). Figure 5B The L2 mapping process shown can be used to obtain the region 521 corresponding to the first coordinate.

[0187] Taking the example that the phone enters right-handed one-handed mode when in one-handed mode, please refer to... Figure 5CConstruct the phone screen coordinate system XOY and the window coordinate system X`O`Y`. Let K be the scaling factor on the X-axis between the small screen display size and the full screen display size. x The scaling factor on the Y-axis is denoted as K. y It should be noted that the screen coordinate system is the coordinate system corresponding to the full-screen display of the phone in full-screen mode, while the window coordinate system is the coordinate system corresponding to the small-screen display of the phone in small-screen mode.

[0188] C(X0,Y0) represents the simulated click position obtained by the mobile phone, which is its location within the small screen display size. The coordinates of point C are still referenced to the origin O of the screen coordinate system. For example, point C corresponds to the first coordinate mentioned above, such as... Figure 3C The area shown is 307. In some embodiments, the coordinates of point C are absolute coordinates relative to the screen coordinate system, that is, the values ​​of X0 and Y0 are described with the origin O of the screen coordinate system as the reference point.

[0189] A(X1, Y1) represents the position of the simulated click position C after mapping to the full-screen display size. For example, point C corresponds to... Figure 3C The area shown is 308.

[0190] As shown in the diagram, the intersection of the first auxiliary line passing through point A and perpendicular to the X-axis, and the second auxiliary line passing through point C and perpendicular to the Y-axis, is B(X1,Y0). C`(X`,Y`) is the position of the lower right corner of the screen. B`(X1,Y`) is the projection point of point B onto DC`.

[0191] As explained above, the aspect ratio of a full-screen phone display is the same as that of a small-screen display. In other words, the ratio of width reduction of the small-screen display to width reduction of the full-screen display is the same as the ratio of height reduction. This is based on the principle of similarity.

[0192]

[0193]

[0194] Right now,

[0195]

[0196]

[0197] As explained above, in the small-screen display mode of one-handed operation, the phone system automatically converts the simulated click position from its position in the small-screen display mode to its position in the full-screen display mode. That is, if the simulated click position is C, the phone system will automatically convert point C to point A and perform the simulated click operation at point A. Understandably, point A is an incorrect simulated click position.

[0198] Understandably, the phone's system automatically converts the simulated click location from point C in the small-screen display state to point A in the full-screen display state. This conversion process can be represented by the following two formulas. The actual click location A(X1,Y1) changes from C(X0,Y0):

[0199]

[0200]

[0201] To ensure that the simulated click is executed accurately at point C in one-handed mode, the Z-point can be calculated first. This allows the phone's system to automatically map the input coordinates from Z to C during the simulated click, ultimately executing the click at point C. The coordinates of the Z-point are denoted as (x, y).

[0202] As explained above, formula (1-5) represents the change in the X-axis coordinate during the automatic conversion process when the mobile phone system performs a simulated click operation, and formula (1-6) represents the change in the Y-axis coordinate during the automatic conversion process. Therefore, to enable the mobile phone system to convert point Z(x, y) to point C(X0, Y0) during the automatic conversion process, we can use X0 = x and X1 = X0 in formula (1-5) to deduce the value of x.

[0203] x = K x (X0-X')+X' (1-7)

[0204] Similarly, in formula (1-6), let Y0 = y and Y1 = Y0, and we can calculate the value of y:

[0205] y = K y (Y0-Y')+Y' (1-8)

[0206] Understandably, formulas (1-7) and (1-8) correspond to the first mapping process described above. That is, if the phone generates a simulated click event with Z(x,y), the system can obtain the coordinates of the simulated click location on the small screen display, i.e., point C(X0,Y0), after performing an automatic conversion step. The values ​​of x and y can be calculated from C(X0,Y0) using formulas (1-7) and (1-8). For example, Z(x,y) represents the second coordinate mentioned above.

[0207] It should be noted that the above Figure 5B and Figure 5C The illustrated coordinate mapping process is only a scaled-down example of a right-handed one-handed mode. In other embodiments, the phone displays other one-handed mode displays (such as left-handed one-handed mode, e.g., ...). Figure 3B When the main screen interface shown is 306 (reduced to a smaller size), or when the screen is half-screen (e.g., half-screen display), the conversion process between the first mapping process and the second mapping process can be other formulas.

[0208] S512. Based on the multi-modal input framework, perform a simulated click operation at the position corresponding to the first coordinate.

[0209] Taking the simulated click operation as an example of an element click operation, the above S512 can specifically include: executing the task triggered by the interface element corresponding to the element click operation based on the multi-modal input framework. For example, for the voice command "Open [Calendar]", the phone will control the simulated click to be performed on the [Calendar] application icon, thereby opening the [Calendar] application.

[0210] Taking the simulated click operation as a sliding operation as an example, the above S512 may specifically include, based on a multi-modal input framework, sliding the displayed screen from the start position of the simulated click to the end position of the simulated click.

[0211] It should be noted that the specific implementation process of the mobile phone performing a simulated click operation at a certain position based on the multi-mode input framework can be referred to the description in the relevant technology, and will not be repeated in the embodiments of this application.

[0212] In some embodiments, while performing a simulated click operation in response to a voice command, the mobile phone can also indicate the location of the simulated click operation by displaying the simulated click trajectory on the screen. This allows users to easily see the location of the simulated click, making the voice-controlled simulated click process more intuitive. Please refer to... Figure 6In small-screen display mode, the phone displays a scaled-down home screen interface 620. On this scaled-down interface 620, the user speaks the voice command "Open [Calendar]". In response to this voice command, the phone displays a simulated click trajectory 621 at the location corresponding to the word "[Calendar]". Then, at the location corresponding to this simulated click trajectory 621, a simulated click operation is performed to open the Calendar application, entering... Figure 6 The scaled-down calendar application interface 622 is shown. Similarly, when a swipe operation is executed in response to a user's voice command, the phone can also display the simulated click trajectory sequentially along the simulated swipe operation path, or display the simulated click trajectory at the start and end points of the simulated swipe operation respectively.

[0213] Furthermore, it should be noted that in some embodiments, if the judgment result of S505 is negative, it indicates that the user wishes to control the operation performed by the phone via voice control, which is not a simulated click operation. In this case, in response to this voice command, the phone can directly execute the corresponding operation. For example, after determining that the audio data includes the preset voice command "return to desktop," in response to this voice command, the phone can directly control the phone's display interface to return to the desktop, such as... Figure 5A S513 is shown.

[0214] S513. In response to preset voice commands, the mobile phone performs the corresponding operation.

[0215] In the technical solution provided in this application embodiment, in one-handed mode, if the user's spoken voice command is not a simulated click command, the phone will perform the corresponding operation in the same way as if it were in full-screen mode and receiving a voice control command. In other words, the phone responds to the voice command and executes the corresponding operation normally, accurately realizing the user's intention to control the phone.

[0216] In the technical solution provided in this application embodiment, when the simulated click operation is an element click operation (such as "open [Calendar]" as mentioned above), the task triggered by the interface element corresponding to the element click operation is executed (such as opening the [Calendar] application and displaying the calendar interface). When the simulated click operation is a swipe operation (such as "swipe down" as mentioned above), the start position and end position of the simulated click are obtained, and then the screen is swiped from the start position to the end position of the simulated click. In this way, it can be ensured that in the scenario where the user uses voice control to perform simulated click operations in one-handed mode, the mobile phone responds to the voice command of the simulated click and executes the simulated click operation at the accurate position, triggering the corresponding response operation and avoiding the problem of simulated click errors. Provided that the mobile phone accurately recognizes the user's voice command, the mobile phone can respond to the voice command of the simulated click and accurately execute the simulated click operation.

[0217] As described in the above embodiments, this application mainly addresses scenarios where a user performs a simulated click operation on a mobile phone via voice control in one-handed mode. Specifically, after determining that the user is in the aforementioned scenario, the determined simulated click position is mapped before executing the simulated click voice command. To ensure that the mobile phone can accurately recognize the one-handed mode scenario before executing the simulated click operation, the mobile phone can determine whether it is currently in a small-screen display state in one-handed mode after activating the voice control function and receiving audio data. If so, after performing voice recognition on the audio data and determining that it contains a simulated click voice command supported by the voice control function, the simulated click position corresponding to the simulated click voice command is mapped before executing the simulated click operation.

[0218] In other embodiments, to ensure that the phone can accurately identify the one-handed mode scenario before performing a simulated click operation, the phone can also perform voice recognition on the audio data. If the audio data contains a voice command for a simulated click supported by the voice control function, the phone can then determine whether it is currently in a small-screen display state in one-handed mode. If so, the simulated click location is mapped first, and then the simulated click operation is performed.

[0219] For example, in some embodiments, before S507, the method may further include determining whether the phone is in a small-screen display state in one-handed mode. In this embodiment, if the phone is currently in a small-screen display state, S507-S512 can be executed.

[0220] In the technical solution provided in this application embodiment, after the mobile phone receives audio data based on the voice control function, before executing the simulated click operation in response to the voice command for simulated click contained in the audio data, it first determines whether the mobile phone is currently in a small-screen display state in one-handed mode. If it is determined that the mobile phone is currently in a small-screen display state in one-handed mode, then for the simulated click operation to be executed, coordinate mapping needs to be performed first, and then the simulated click operation is executed according to the coordinates obtained after mapping. In this way, it can be ensured that when the mobile phone is in a small-screen display state in one-handed mode, it can execute the simulated click operation in response to the voice command, and thus accurately execute the corresponding simulated click operation according to the user's voice command to control the mobile phone to perform the simulated click.

[0221] In both of the above scenarios, the determination of whether the phone is in one-handed mode small-screen display is only made after the audio data is received. That is, the phone needs to determine whether the audio data is in one-handed mode small-screen display every time it collects audio data based on the voice control function. Therefore, to reduce the number of times the voice control function needs to determine whether it is in one-handed mode small-screen display, in some embodiments, the phone can also send a first notification message to the voice control module corresponding to the voice control function after detecting that it has entered one-handed mode small-screen display. This first notification message is used to inform the voice control module that it has entered one-handed mode small-screen display. In this way, when the voice control function subsequently receives and recognizes a voice command containing a simulated click, it can first map the simulated click location before executing the simulated click operation.

[0222] It should be noted that the voice control function can be turned on before the phone enters the small screen display state of one-handed mode. In this way, when the phone detects that it has entered the small screen display state of one-handed mode, it will immediately send the first notification message to the voice control module.

[0223] In other embodiments, the phone can also activate the voice control function after entering the small-screen display state of one-handed mode. In this embodiment, when entering the small-screen display state of one-handed mode, the voice control function is not enabled, so the first notification message cannot be sent to the voice control module immediately. At this time, the activation of the voice control function can be monitored. Subsequently, after detecting that the phone has activated the voice control function, the first notification message is sent to the voice control module.

[0224] In the technical solution provided in this application embodiment, the voice control module corresponding to the voice control function does not need to determine whether the phone has entered the one-handed mode small screen display state every time it receives audio data. This reduces the number of times the voice control module needs to determine the one-handed mode, thus reducing power consumption.

[0225] Furthermore, in the embodiment where the phone sends a first notification message to the voice control module after entering the small-screen display state of one-handed mode, the phone should send a second notification message to the voice control module after detecting that it has exited the small-screen display state of one-handed mode. This second notification message informs the voice control module that the phone has exited the small-screen display state of one-handed mode. It is understood that after receiving the second notification message, the voice control module will no longer map the received audio data containing simulated click commands, but can directly execute the simulated click operation according to the obtained simulated click location.

[0226] In the technical solution provided in this application embodiment, after detecting that the phone has exited the small-screen display state of one-handed mode, if a preset voice command containing a simulated click operation is received based on the voice control function, it is no longer necessary to perform the first mapping processing on the obtained simulated click position. This avoids the phone being able to execute a simulated click operation at the accurate location even after entering full-screen display mode and receiving a voice command from the user requiring a simulated click operation.

[0227] Please refer to Figure 7 This diagram illustrates the interaction between internal modules of a mobile phone during the execution of the voice control method provided in this embodiment. In this embodiment, the mobile phone includes: a voice control module, a content sensor, a view fetcher, a speech recognition engine (specifically including ASR / DM / NLU), an action execution module, a system settings module, a window manager, a one-handed mode control module, and an event manager. In some embodiments, the event manager is a multimodal input framework. In this embodiment, the voice control method specifically includes:

[0228] After the voice control function is activated, the voice control module calls the framework layer's interface across processes, causing the content perceiver to invoke the top-level activity result callback function. This allows the view crawler to obtain the text data on the current page, analyze the text data, and obtain the words contained on the current page. The view crawler then returns these words to the voice control module, returning the words contained on the current page.

[0229] The voice control module sends words to the speech recognition engine. Furthermore, after acquiring the audio data, it sends the audio data back to the speech recognition engine.

[0230] The speech recognition engine determines whether the audio data contains a preset voice command, and whether the operation corresponding to the preset voice command simulates a click operation. Furthermore, if the speech recognition engine determines that the audio data contains a preset voice command that simulates a click operation, it notifies the command execution module to execute the simulated click operation.

[0231] The instruction execution module registers a one-handed mode listener with the system settings module. This listener monitors whether the device is currently in one-handed mode. If so, the one-handed mode control module obtains the phone's full-screen display size from the window management module and the scaling ratio for one-handed mode from the system settings. This scaling ratio represents the scaling of the small-screen display size relative to the full-screen display size in the current small-screen display state. Based on the first coordinate of the simulated click position corresponding to the preset voice command and the scaling ratio, the one-handed mode control module calculates the second coordinate of the target position (e.g., ...). Figure 5C (The coordinates of point Z shown).

[0232] Next, a simulated click event is generated based on the second coordinate. Then, the instruction execution module injects the simulated click event into the event manager. Upon receiving the simulated click event, the event manager parses it to obtain the second coordinate carried within the event. In response to this simulated click event, the second coordinate is mapped from the small-screen display size to the full-screen display size. That is, the second coordinate is mapped to obtain the first coordinate of the simulated click location on the small-screen display size. Finally, the phone can perform a simulated click operation at the location of this first coordinate.

[0233] In the technical solution provided in this application embodiment, because the simulated click position has been mapped to the small screen display in one-handed mode before the simulated click operation is executed in response to the simulated click voice command, the position of the simulated click operation can be accurately ensured. This avoids the problem of inconsistency between the click operation performed by the phone and the user's intention in speaking the voice command.

[0234] Other embodiments of this application provide an electronic device (such as a mobile phone). The electronic device may include a memory, a display screen, a recording device, and one or more processors. The memory, display screen, and recording device are each coupled to a processor. The display screen is used to display the interface of the electronic device, and the recording device is used to acquire audio data. The memory is also used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device can perform various functions or steps performed by the mobile phone in the above method embodiments. The structure of the electronic device can be referred to... Figure 4A The structure of the electronic device 100 shown.

[0235] This application also provides a chip system, such as... Figure 8As shown, the chip system 80 includes at least one processor 801 and at least one interface circuit 802. The processor 801 and the interface circuit 802 are interconnected via lines. For example, the interface circuit 802 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 802 can be used to send signals to other devices (e.g., the processor 801). Exemplarily, the interface circuit 802 can read instructions stored in the memory and send those instructions to the processor 801. When the instructions are executed by the processor 801, the electronic device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete components, which are not specifically limited in this application embodiment.

[0236] This application also provides a computer-readable storage medium including computer instructions that, when executed on the aforementioned electronic device (such as a mobile phone), cause the electronic device to perform various functions or steps performed by the mobile phone in the above method embodiments.

[0237] This application also provides a computer program product that, when run on a computer, causes the computer to perform the various functions or steps performed by the mobile phone in the above method embodiments. The computer can be an electronic device, such as a mobile phone.

[0238] Through the above description of the embodiments, those skilled in the art can clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0239] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0240] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0241] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0242] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0243] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A voice control method, characterized in that, The method is applied to electronic devices with voice control and one-handed operation capabilities; the method includes: In response to a one-handed mode trigger event, the screen switches from full-screen display to a smaller screen display in one-handed mode. After the voice control function is activated, the audio data collected by the electronic device is subjected to voice recognition based on the voice control function to obtain the voice recognition result; If the speech recognition result indicates that the operation corresponding to the audio data is a simulated click operation, determine the simulated click position corresponding to the simulated click operation in the small screen display; the simulated click operation includes an element click operation or a swipe operation; when the simulated click operation is the swipe operation, the simulated click position includes a simulated click start position and a simulated click end position; Execute a simulated click operation based on the simulated click location; If the simulated click operation is an element click operation, the electronic device executes the task triggered by the interface element corresponding to the element click operation; if the simulated click operation is a swipe operation, the electronic device swipes the display screen from the simulated click start position to the simulated click end position.

2. The method according to claim 1, characterized in that, Determining the simulated click position corresponding to the simulated click operation on the small screen display includes: Determine the first coordinate of the simulated click position in the screen coordinate system within the small screen display; Before performing the simulated click operation based on the simulated click location, the method further includes: The first coordinates are subjected to a first mapping process to obtain the second coordinates in the screen coordinate system; the first mapping process is used to map the coordinates of the point from the full-screen display to the small-screen display. The step of performing a simulated click operation based on the simulated click location includes: A simulated click event is generated based on the second coordinate; Send the simulated click event to the multi-mode input frame of the electronic device; In response to the simulated click event, the multimodal input framework performs a second mapping process on the second coordinates to obtain the first coordinates; the second mapping process is used to map the coordinates of the point from the small screen display to the full screen display. Based on the multi-modal input framework, a simulated click operation is performed on the location of the first coordinate.

3. The method according to claim 2, characterized in that, The step of performing a first mapping process on the first coordinates to obtain the second coordinates in the screen coordinate system includes: The full-screen display size of the electronic device in the full-screen display screen, the small-screen display size in the small-screen display screen, and a first mapping relationship between the full-screen display size and the small-screen display size are obtained. Based on the first mapping relationship, the first coordinates are subjected to a first mapping process to obtain the second coordinates.

4. The method according to claim 3, characterized in that, The step of obtaining the full-screen display size of the electronic device in the full-screen display mode, the small-screen display size in the small-screen display mode, and the first mapping relationship between the full-screen display size and the small-screen display size includes: Obtain the first width and first height of the full-screen display size, and the second width and second height of the small-screen display size; Based on the first height and the second height, determine the height reduction ratio for transforming the full-screen display size into the small-screen display size; Based on the first width and the second width, determine the width reduction ratio by which the full-screen display size is transformed into the small-screen display size; The first mapping relationship is determined based on the height reduction ratio and the width reduction ratio.

5. The method according to any one of claims 2-4, characterized in that, The step of performing a second mapping process on the second coordinates to obtain the first coordinates includes: Obtain a second mapping relationship between the electronic device and the full-screen display size under the small-screen display screen; Based on the second mapping relationship, the second coordinates are subjected to a second mapping process to obtain the first coordinates.

6. The method according to any one of claims 2-4, characterized in that, The simulated click operation is the element click operation; determining the simulated click position corresponding to the simulated click operation in the small screen display includes: Analyze the preset element corresponding to the element click operation; locate the position corresponding to the preset element in the small screen display.

7. The method according to any one of claims 2-4, characterized in that, The simulated click operation is the element click operation; determining the simulated click position corresponding to the simulated click operation in the small screen display includes: Obtain all elements contained in the small screen display, as well as the position information corresponding to each element; Analyze the preset elements corresponding to the element click operation; The location corresponding to the preset element is found based on the location information.

8. The method according to any one of claims 1-4, characterized in that, The method further includes: Simultaneously with the execution of the simulated click operation, the simulated click trajectory is displayed at the location corresponding to the simulated click operation.

9. The method according to any one of claims 1-4, characterized in that, The aspect ratio of the full-screen display is the same as that of the small-screen display.

10. An electronic device, characterized in that, The electronic device includes: a processor, a memory, a display screen, and a recording device; the memory, the display screen, and the recording device are respectively coupled to the processor. The display screen is used to display the interface of the electronic device, and the recording device is used to collect audio data; the memory stores computer program code, which includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the method as described in any one of claims 1-9.

11. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Method and device for operating mobile terminal

    CN105955602A

  • Voice-based interactive method and apparatus, electronic device and operation system

    CN108279839A