Control method of speaking window, electronic equipment and vehicle
By obtaining the target touch point coordinates of the user's voice command in the vehicle and determining the target control based on the command position rules, the problem of the "see-and-say" function failing due to incomplete hot word registration or control not being integrated into the toolkit is solved, and more accurate control operation and function usage are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GREAT WALL MOTOR CO LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-19
AI Technical Summary
In existing technologies, the lack of access to the "See and Say" software development kit or incomplete keyword registration leads to the inability of controls to implement the "See and Say" function.
By obtaining the coordinates of the target touch point corresponding to the user's voice command, the target control is determined based on the command position rules, and the control operation is executed directly, avoiding the need to search for and confirm hot words.
It improves the accuracy of target control identification, ensures the effective use of the "what you see is what you can say" function, and avoids function failure caused by incomplete hot word registration or control not being integrated into the software development kit.
Smart Images

Figure CN122064261A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of vehicle voice interaction technology, and more particularly to a control method, electronic device and vehicle with a talkable window. Background Technology
[0002] "Speakable windows" refer to the range of functions or interaction boundaries that users can directly control via voice commands. They are typically deeply integrated with the in-vehicle system's interface, hardware functions, or services. Their core logic is: "All functions visible on the screen or supported by the vehicle can be triggered via voice commands," thus achieving a seamless interactive experience of "what you see is what you can say." However, issues frequently arise where controls fail to implement the "what you see is what you can say" function due to the lack of integration with the "what you see is what you can say" software development kit or incomplete keyword registration. Summary of the Invention
[0003] In view of this, the purpose of this application is to propose a control method, electronic device and vehicle with a talkable window, so as to solve the problem that existing controls cannot realize the talkable function because they are not connected to the talkable software development kit or the registered keywords are incomplete.
[0004] To achieve the above objectives, the first aspect of this application provides a method for controlling a view, comprising: Obtain a first user interface of a viewable window, the first user interface including multiple controls; In response to receiving a user voice command, the coordinates of the target touch point in the first user interface are determined based on the command location rules. From the plurality of controls, determine the target control corresponding to the coordinates of the target touch point, and execute the control operation corresponding to the target control.
[0005] Optionally, the method further includes: Obtain the initial voice command corresponding to each of the aforementioned controls; In response to the initial voice command meeting the preset command validity conditions, the initial touch point coordinates on the first user interface are determined for the initial touch point corresponding to the initial voice command. Based on each initial voice command and the corresponding initial touch point coordinates, command position rules are generated.
[0006] Optionally, the method further includes: Execute the control operation corresponding to the control corresponding to the initial voice command to switch the first user interface to the second user interface; In response to the fact that the second user interface is different from the first user interface, it is determined that the initial voice command meets the preset command validity conditions.
[0007] Optionally, determining the initial touch point coordinates on the first user interface corresponding to the initial voice command includes: Obtain the coordinate information of the click operation corresponding to the execution of the initial voice command on the first user interface, and determine the coordinate information as the initial touch point coordinates; Alternatively, the initial position information of the control corresponding to the initial voice command on the first user interface can be obtained, and the initial position information can be determined as the initial touch point coordinates.
[0008] Optionally, determining the target touch point coordinates in the first user interface based on the command location rules includes: Determine the user command keywords of the user's voice command, and determine multiple key derivative words based on the user command keywords; Based on multiple key derived words, a target voice instruction is determined from all the initial voice instructions; Based on the command location rules, the initial touch point coordinates corresponding to the target voice command are determined, and the initial touch point coordinates are determined as the target touch point coordinates corresponding to the user's voice command.
[0009] Optionally, determining a target voice instruction from all the initial voice instructions based on multiple key derived words includes: Determine the initial instruction keywords for each of the initial voice commands; In response to the fact that the initial instruction keyword among the plurality of key derived words includes an initial voice instruction, the initial voice instruction corresponding to the initial instruction keyword is determined to be the target voice instruction; Alternatively, in response to an initial instruction keyword that does not include any of the initial voice instructions among the plurality of key derivative words, a target voice instruction is determined from all the initial voice instructions based on the initial instruction keyword, the plurality of key derivative words, and / or the user instruction keyword.
[0010] Optionally, determining a target voice command from all the initial voice commands based on the initial command keywords, the plurality of key derived words, and / or the user command keywords includes: Determine the first relevance between each initial instruction keyword and the user instruction keyword, and determine the initial voice instruction corresponding to the initial instruction keyword with the highest first relevance as the target voice instruction; Alternatively, a second relevance degree between each initial instruction keyword and each key derivative word is determined. In response to a second relevance degree between an initial instruction keyword and each key derivative word being greater than or equal to a preset relevance degree, the initial voice instruction corresponding to the initial instruction keyword is determined as the target voice instruction.
[0011] Optionally, determining the target control corresponding to the coordinates of the target touch point from the plurality of controls and performing the control operation corresponding to the target control includes: Obtain the initial position area of each control on the first user interface; In response to the initial position region including the target touch point coordinates, the control corresponding to the initial position region is determined as the target control corresponding to the target touch point coordinates; A simulated touch operation is performed in the initial position area of the target control to execute the control operation corresponding to the target control.
[0012] Based on the same inventive concept, a second aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the method as described in any of the first aspects above.
[0013] Based on the same inventive concept, a third aspect of this application provides a vehicle that includes the electronic equipment described in the second aspect above.
[0014] As can be seen from the above, the control method, electronic device, and vehicle for the speakable window provided in this application, after receiving a user's voice command, first determines the target touch point coordinates in the first user interface corresponding to the user's voice command based on the command position rules. Then, based on the target touch point coordinates, it determines the target control corresponding to the target touch point coordinates from the plurality of controls. This target control is the target control determined by the user's voice command. Thus, in the process of determining the target control, it is only necessary to determine the target touch point coordinates in the first user interface corresponding to the user's voice command. This allows the determination of the position coordinates of the target touch point that needs to be touched or clicked on the first user interface when executing the user's voice command. Then, based on the position coordinates, the target control corresponding to the position coordinates can be determined. This ensures that the determined target control is the control that needs to be touched or clicked on the first user interface when executing the user's voice command. In this way, there is no need to search and confirm the hot words of the target control, nor is it necessary to determine the target control based on registered hot words. Instead, the target control is determined directly based on the position coordinates of the target touch point. Furthermore, since the positions of each control on the first user interface do not change, the method of determining the target control based on the target touch point coordinates improves the accuracy of the determined target control. It also avoids situations where the determined target control is inaccurate or cannot be found due to incomplete hot word registration or the control not being connected to the Visible and Speak software development kit, thus ensuring the effective use of the Visible and Speak function. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a flowchart illustrating the control method of a viewport according to an embodiment of this application; Figure 2 This is an exemplary schematic diagram of the instruction location rules in an embodiment of this application; Figure 3 This is a schematic diagram of the control device for the audible window according to an embodiment of this application; Figure 4 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0018] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0019] With the rapid development of in-vehicle systems and the increasing diversity of vehicle-to-everything (V2X) interactions, users have higher and higher demands for convenient, simple, and efficient operation, requiring more efficient, simpler, and more convenient operation. Voice interaction has become an indispensable function in current in-vehicle systems, and among them, the voice-activated interface is arguably the most important function.
[0020] In vehicle voice interaction systems, a "speakable window" (or "operable voice window") refers to the range of functions or interaction boundaries that users can directly control through voice commands. It is usually deeply integrated with the UI interface, hardware functions, or services of the in-vehicle system. Its core logic is: "All screen-visible selections or vehicle-supported functions can be triggered by voice commands," thereby achieving a seamless interactive experience of "what you see is what you can say."
[0021] In layman's terms, the "what you see is what you can say" feature allows users to click on the interface and perform corresponding actions by speaking. "What you see is what you can say" means that anything visible on the interface can be executed by speaking.
[0022] However, for the implementation of the "what is visible is what is spoken" function, two conditions must be met: Condition 1: The visible window includes multiple controls, which are elements visible on the screen, such as menus, buttons, and icons displayed on the central control screen, instrument panel, or head-up display (e.g., "navigation," "air conditioning," "seat heating"). Hotwords corresponding to all these menus, buttons, and icons need to be registered in advance. When the user issues a voice command, the vehicle's system recognizes the trigger word corresponding to the voice command, iterates through all registered hotwords, and executes the function corresponding to the hotword that matches the trigger word. This achieves the "visible and speakable" function. In other words, for the "visible and speakable" function, only by pre-registering the hotwords corresponding to the controls can the controls be accurately executed based on voice commands.
[0023] Condition 2: Each control on the viewport must be pre-connected to the Visible and Speakable software development kit; otherwise, the Visible and Speakable function of that control cannot be implemented.
[0024] However, for condition one, to achieve a more efficient and accurate "see-and-talk" functionality for controls, multiple potential hotwords need to be registered for the same semantic function of each control (e.g., the "exit" function on the interface requires registering hotwords such as "exit," "back," "close," and "back"). This makes the hotword registration process very cumbersome and cannot guarantee that the registered hotwords are comprehensive enough. If certain similar hotwords are not registered, the control will be unable to execute based on voice commands, thus losing the "see-and-talk" functionality. For example, the "exit" function on the interface has registered hotwords such as "exit," "back," "close," and "back." However, if the user's voice command is "go back one step," then because the hotword "go back" is not registered, the "exit" function on the interface cannot be executed based on the voice command "go back one step," causing the "see-and-talk" functionality to fail.
[0025] For condition two, if certain controls are not pre-integrated into the Visible-Speak software development kit, the Visible-Speak function of that control cannot be implemented.
[0026] Therefore, during the use of the "See and Say" function, the control often fails to implement the "See and Say" function because it is not connected to the "See and Say" software development kit or the registered keywords are incomplete.
[0027] Based on this, see Figure 1 This application provides a control method for a viewable window, executed by a vehicle controller, the method specifically including the following steps: Step S100: Obtain the first user interface of the viewable window, wherein the first user interface includes multiple controls; Step S200: In response to receiving a user voice command, determine the target touch point coordinates in the first user interface based on the command position rules; Step S300: Determine the target control corresponding to the coordinates of the target touch point from the plurality of controls, and execute the control operation corresponding to the target control.
[0028] Specifically, a "voiceable window" refers to the range of functions or interaction boundaries that a user can directly control through voice commands. It is usually deeply integrated with the interface, hardware functions, or services of the vehicle system. In this application, the voiceable window can be the vehicle's central control interface, head-up display interface, or other in-vehicle screen interfaces.
[0029] The first user interface of a talkable window refers to the current user interface of the talkable window, which may include the user-visible parts of the current user interface. The user-visible parts may include images, text, menus, options, icons, buttons, etc. displayed on the interface. In this embodiment, images, text, menus, options, icons, buttons, etc. on the user interface are collectively referred to as "controls". Therefore, the first user interface includes multiple controls.
[0030] When a user voice command is received, it means that the user wants to execute the function of a corresponding control by sending the user voice command. In other words, the user "speaks" the user voice command to make the vehicle controller execute the control operation corresponding to the relevant control, realizing the "what you see is what you can say" function.
[0031] Therefore, it is necessary to determine the target touch point coordinates in the first user interface based on the instruction location rules to identify the target touch point corresponding to the user's voice instruction.
[0032] The command position rule is a pre-set or pre-determined correspondence rule between voice commands and corresponding touch point coordinates, used to characterize the one-to-one correspondence rule between voice commands and touch point coordinates.
[0033] The touch point coordinates refer to the position coordinates of the touch point when a click or touch operation is required on the first user interface, corresponding to the execution of the user's voice command. These position coordinates refer to the position coordinates of the touch point on the first user interface.
[0034] The command location rule can be implemented in various ways. For example, the command location rule can be a mapping relationship between a voice command and the corresponding touch point coordinates. This mapping relationship can be any regression mapping method. Alternatively, the instruction position rule can be a data model trained based on the correspondence between voice instructions and the corresponding touch point coordinates. The model type is not limited. For example, it can be a deep learning model of the converter self-attention mechanism or a neural network model. Alternatively, the instruction location rules can be a relationship diagram, relationship table, or database generated based on the correspondence between voice instructions and the corresponding touch point coordinates.
[0035] Therefore, based on the instruction location rules, the target touch point coordinates on the first user interface corresponding to the user's voice instruction are determined. These target touch point coordinates are the position coordinates of the target touch point that needs to perform a touch operation on the first user interface corresponding to the user's voice instruction.
[0036] Then, a target control corresponding to the target touch point coordinates is determined from the plurality of controls. The target control determined at this time is the control corresponding to the target touch point coordinates and the user's voice command, that is, the control that the user wants to execute when he / she speaks the user's voice command. Therefore, the control operation corresponding to the target control is executed.
[0037] In this application, upon receiving a user's voice command, the coordinates of the target touch point in the first user interface corresponding to the user's voice command are first determined based on the command location rules. Then, based on the target touch point coordinates, the target control corresponding to the target touch point coordinates is determined from the plurality of controls. This target control is the target control determined by the user's voice command. Thus, in the process of determining the target control, it is only necessary to determine the coordinates of the target touch point in the first user interface corresponding to the user's voice command. This allows the determination of the position coordinates of the target touch point that needs to be touched or clicked on the first user interface when executing the user's voice command. Then, based on the position coordinates, the target control corresponding to the position coordinates can be determined. This ensures that the determined target control is the control that needs to be touched or clicked on the first user interface when executing the user's voice command. In this way, there is no need to search and confirm the hot words of the target control, nor is it necessary to determine the target control based on registered hot words. Instead, the target control is determined directly based on the position coordinates of the target touch point.
[0038] Furthermore, since the positions of each control on the first user interface do not change, the method of determining the target control based on the target touch point coordinates improves the accuracy of the determined target control. It also avoids situations where the determined target control is inaccurate or cannot be found due to incomplete hot word registration or the control not being connected to the Visible and Speak software development kit, thus ensuring the effective use of the Visible and Speak function.
[0039] In some embodiments, the method further includes: Obtain the initial voice command corresponding to each of the aforementioned controls; In response to the initial voice command meeting the preset command validity conditions, the initial touch point coordinates on the first user interface are determined for the initial touch point corresponding to the initial voice command. Based on each initial voice command and the corresponding initial touch point coordinates, command position rules are generated.
[0040] Specifically, the initial voice command corresponding to each of the controls is obtained. The initial voice command is also a voice command issued by the user or relevant personnel, which is used to determine the command position rules.
[0041] For each control, the initial voice command corresponding to each control is obtained, and then it is determined whether the initial voice command meets the preset command validity conditions. The preset command validity conditions are preset conditions for determining the validity of the initial voice command, that is, the conditions under which the vehicle controller can perform a valid operation on the first user interface based on the initial voice command.
[0042] When the initial voice command is determined to meet the preset command validity conditions, it is considered a valid command. The initial touch point coordinates on the first user interface are then determined. These initial touch point coordinates are the position coordinates of the initial touch point corresponding to the initial voice command, which is required to perform a touch operation on the first user interface.
[0043] Then, based on the correspondence between each initial voice command and the corresponding initial touch point coordinates, a command position rule is generated. This command position rule includes the initial voice command corresponding to each control and the initial touch point coordinates corresponding to each initial voice command. Thus, there is a one-to-one correspondence between the initial voice command and the initial touch point coordinates. Based on the command position rule, the initial touch point coordinates uniquely corresponding to a given initial voice command can be determined. These initial touch point coordinates represent the position coordinates of the initial touch point that needs to be touched or clicked on the first user interface when executing the initial voice command. Once these initial touch point coordinates are determined, the target control can be accurately located, and the control operation corresponding to the initial voice command can be executed accurately. For example, the command position rule can be as follows: Figure 2 The list shown.
[0044] In this application, the initial voice command and the initial touch point coordinates in the command position rule are in one-to-one correspondence. In this way, the initial touch point coordinates that uniquely correspond to any initial voice command can be determined. Then, the target control can be accurately determined based on the initial touch point coordinates, and the control operation corresponding to the initial voice command can be accurately executed, thereby improving the accuracy of the determined target control.
[0045] In some embodiments, the method further includes: Execute the control operation corresponding to the control corresponding to the initial voice command to switch the first user interface to the second user interface; In response to the fact that the second user interface is different from the first user interface, it is determined that the initial voice command meets the preset command validity conditions.
[0046] Specifically, while obtaining the initial voice command corresponding to each control, the control operation corresponding to the control corresponding to the initial voice command is executed, such as clicking or touching the control, to execute the control operation corresponding to the control, so that the first user interface is switched to the second user interface.
[0047] For example, for the "Back" control, while the initial voice command "Back" is input, the user touches or clicks the "Back" control to perform the operation corresponding to the "Back" control, thereby switching the first user interface to the second user interface.
[0048] Then compare whether the first user interface and the second user interface are the same. Specifically, you can compare screenshots of the first user interface and the second user interface to see if their content is the same.
[0049] Alternatively, compare all the controls contained in the first user interface and the second user interface to see if they are all the same. If they are not all the same, it means that the first user interface and the second user interface are different.
[0050] When it is determined that the second user interface is different from the first user interface, it indicates that the execution of the control operation corresponding to the control corresponding to the initial voice command has indeed changed the interface information of the window. Therefore, the initial voice command is determined to be a valid voice command and meets the preset command validity conditions.
[0051] In this application, by judging whether the initial voice command meets the preset command validity conditions, the validity of the initial voice command can be ensured, the validity of the command position rules determined based on the initial voice command can be improved, and the accuracy of the "see it, say it" function can be improved.
[0052] In some embodiments, determining the initial touch point coordinates on the first user interface corresponding to the initial touch point of the initial voice command includes: Obtain the coordinate information of the click operation corresponding to the execution of the initial voice command on the first user interface, and determine the coordinate information as the initial touch point coordinates; Alternatively, the initial position information of the control corresponding to the initial voice command on the first user interface can be obtained, and the initial position information can be determined as the initial touch point coordinates.
[0053] Specifically, there are two ways to determine the initial touch point coordinates on the first user interface corresponding to the initial voice command: Method 1: Obtain the coordinate information of the click operation corresponding to the initial voice command on the first user interface, and determine this coordinate information as the initial touch point coordinates. That is, Method 1 determines the initial touch point coordinates based on the coordinate information of the actual click operation.
[0054] It is worth noting that in this application, click operation and touch operation have the same meaning. Therefore, click operation can also be called touch operation, and touch operation can also be called click operation.
[0055] Method 2: Obtain the initial position information of the control corresponding to the initial voice command on the first user interface, and determine the initial position information as the initial touch point coordinates. That is, Method 2 determines the initial touch point coordinates based on the initial position information of the control corresponding to the initial voice command.
[0056] In this application, the initial touch point coordinates can be determined based on the coordinate information of the actual click operation, or based on the initial position information of the control corresponding to the initial voice command. This allows for the determination of the initial touch point coordinates in different ways, making the determination of the initial touch point coordinates more convenient and accurate.
[0057] In some embodiments, determining the target touch point coordinates in the first user interface based on the instruction location rule includes: Determine the user command keywords of the user's voice command, and determine multiple key derivative words based on the user command keywords; Based on multiple key derived words, a target voice instruction is determined from all the initial voice instructions; Based on the command location rules, the initial touch point coordinates corresponding to the target voice command are determined, and the initial touch point coordinates are determined as the target touch point coordinates corresponding to the user's voice command.
[0058] Specifically, user command keywords for the user's voice commands are determined, and multiple key derived words are determined based on the user command keywords. The key derived words are words that are similar to, related to, or highly relevant to the user command keywords.
[0059] For example, if the user's voice command is "return to the previous page", then the user command keyword determined based on the user's voice command is "return". Multiple key derivative words determined based on the user command keyword "return" can be "return", "exit", "back", "return", "previous", "next", etc.
[0060] The specific process for determining multiple key derivative words based on the user command keyword is not limited. For example, an existing pre-trained semantic model can be used to determine the key derivative words. For instance, inputting the user command keyword "return" into an existing pre-trained semantic model can output multiple key derivative words. Alternatively, multiple key derivative words related to the user command keyword can be determined based on a database storing synonyms or near-synonyms.
[0061] In actual use, there are many minor differences in wording between the user's spoken expression and the pre-input initial voice command. Therefore, this application determines multiple key derivative words based on the user command keywords, so that the combination of the determined key derivative words can better match the user's actual intention. This can solve the problem of minor differences between the user's spoken expression and the pre-input initial voice command, and ensure that the "see it, say it" function will not fail due to these minor differences.
[0062] Then, based on multiple key derived words, a target voice command is determined from all the initial voice commands. The determined target voice command is the initial voice command that best matches the user's voice command.
[0063] Finally, based on the command position rules, the initial touch point coordinates corresponding to the target voice command are determined, and the initial touch point coordinates are determined as the target touch point coordinates corresponding to the user voice command. In this way, the target touch point coordinates are the position coordinates of the click or touch operation that needs to be performed on the first user interface when executing the user voice command. The user voice command can be executed accurately through the target touch point coordinates to ensure the effective execution of the "see and speak" function.
[0064] In this application, the target touch point coordinates can be determined solely based on the user command keywords and command location rules of the user's voice command. These target touch point coordinates are the coordinates of the position on the first user interface where a click or touch operation needs to be performed when executing the user's voice command. Determining the target touch point coordinates does not require confirming or searching for keywords for each control on the first user interface, and is unrelated to whether the keywords for each control are registered. This ensures that the determination process of the target touch point coordinates is independent of whether keywords are registered, preventing incomplete keyword registration from affecting the determination of the target touch point coordinates. This improves the accuracy of the determined target touch point coordinates, thereby improving the accuracy of the target controls subsequently determined based on the target touch point coordinates. It ensures that incomplete keyword registration or the lack of integration with the "See and Say" software development kit will not lead to inaccurate target controls or the inability to find target controls, thus ensuring the effective use of the "See and Say" function.
[0065] In some embodiments, determining a target voice instruction from all the initial voice instructions based on a plurality of key derived words includes: Determine the initial instruction keywords for each of the initial voice commands; In response to the fact that the initial instruction keyword among the plurality of key derived words includes an initial voice instruction, the initial voice instruction corresponding to the initial instruction keyword is determined to be the target voice instruction; Alternatively, in response to an initial instruction keyword that does not include any of the initial voice instructions among the plurality of key derivative words, a target voice instruction is determined from all the initial voice instructions based on the initial instruction keyword, the plurality of key derivative words, and / or the user instruction keyword.
[0066] Specifically, in the process of determining a target voice instruction from all the initial voice instructions, the initial instruction keyword for each initial voice instruction is first determined. For example, if the initial voice instruction is "go back to the previous page", then the initial instruction keyword determined based on the initial voice instruction is "go back".
[0067] When the plurality of key derivative words include an initial instruction keyword of the initial voice instruction, the initial voice instruction corresponding to the initial instruction keyword is determined to be the target voice instruction.
[0068] For example, the user instruction keyword is "return", and the multiple key derivative words determined based on the user instruction keyword "return" are "return", "exit", "return", "return", "back", and "previous".
[0069] If the initial voice command 1 is "Go back to the previous page", then the initial command keyword determined based on the initial voice command 1 is "go back"; if the initial voice command 2 is "Enable private mode", then the initial command keyword determined based on the initial voice command 1 is "private mode"; if the initial voice command 3 is "Enter privacy policy", then the initial command keyword determined based on the initial voice command 3 is "privacy policy".
[0070] Since the initial instruction keyword "return" is included in the initial voice instruction 1 among multiple key derivative words, the initial voice instruction corresponding to the initial instruction keyword is determined to be the target voice instruction, that is, the determined target voice instruction is the initial voice instruction that best matches the user's voice instruction.
[0071] Alternatively, if the plurality of key derivative words do not include any of the initial instruction keywords of the initial voice instruction, it is necessary to further determine a target voice instruction from all the initial voice instructions based on the initial instruction keyword, the plurality of key derivative words and / or the user instruction keyword.
[0072] For example, the user instruction keyword is "return", and the multiple key derivative words determined based on the user instruction keyword "return" are "return", "exit", "return", "return", "back", and "previous".
[0073] If the initial voice command 1 is "Go to the previous page", then the initial command keyword determined based on the initial voice command 1 is "previous"; if the initial voice command 2 is "Enable private mode", then the initial command keyword determined based on the initial voice command 1 is "private mode"; if the initial voice command 3 is "Enter privacy policy", then the initial command keyword determined based on the initial voice command 3 is "privacy policy".
[0074] Since the initial instruction keyword of any of the initial voice commands is not included in the plurality of key derivative words, it is necessary to further determine a target voice command from all the initial voice commands based on the initial instruction keyword, the plurality of key derivative words and / or the user command keyword, so as to improve the accuracy of the determined target voice command and make the determined target voice command more consistent with the user's voice command and more consistent with the user's control intention.
[0075] In this application, a target voice command is determined from all the initial voice commands by comparing the initial command keyword of each initial voice command with multiple key derived words. In this way, only the initial command keyword of each initial voice command needs to be determined, without the need to register hot words for each control, and without caring about whether the hot word registration is comprehensive. This reduces the steps of registering hot words in the early stage, simplifies the steps of determining the target voice command, and improves the convenience of operation.
[0076] In some embodiments, determining a target voice instruction from all the initial voice instructions based on the initial instruction keywords, the plurality of key derived words, and / or the user instruction keywords includes: Determine the first relevance between each initial instruction keyword and the user instruction keyword, and determine the initial voice instruction corresponding to the initial instruction keyword with the highest first relevance as the target voice instruction; Alternatively, a second relevance degree between each initial instruction keyword and each key derivative word is determined. In response to a second relevance degree between an initial instruction keyword and each key derivative word being greater than or equal to a preset relevance degree, the initial voice instruction corresponding to the initial instruction keyword is determined as the target voice instruction.
[0077] Specifically, based on the initial instruction keywords, the multiple key derived words, and / or the user instruction keywords, there are two ways to determine a target voice instruction from all the initial voice instructions: The first method is to directly determine the first relevance between each initial instruction keyword and the user instruction keyword, and determine the initial voice instruction corresponding to the initial instruction keyword with the highest first relevance as the target voice instruction. In this way, the final determined target voice instruction is the initial voice instruction with the highest relevance to the user voice instruction, that is, the initial voice instruction that best matches the user voice instruction.
[0078] For example, the user instruction keyword is "return".
[0079] If the initial voice command 1 is "Go to the previous page", then the initial command keyword determined based on the initial voice command 1 is "previous"; if the initial voice command 2 is "Enable private mode", then the initial command keyword determined based on the initial voice command 1 is "private mode"; if the initial voice command 3 is "Enter privacy policy", then the initial command keyword determined based on the initial voice command 3 is "privacy policy".
[0080] The initial command keyword "previous" for initial voice command 1 has a first-order relevance of 80% to the user command keyword "return"; the initial command keyword "private mode" for initial voice command 2 has a first-order relevance of 0% to the user command keyword "return"; and the initial command keyword "privacy policy" for initial voice command 3 has a first-order relevance of 0% to the user command keyword "return".
[0081] Therefore, the initial voice command 1 with the highest relevance is determined as the target voice command.
[0082] The second method involves determining the second relevance between each initial instruction keyword and each key derivative word. In response to the fact that the second relevance between an initial instruction keyword and each key derivative word is greater than or equal to a preset relevance, the initial voice instruction corresponding to the initial instruction keyword is determined as the target voice instruction.
[0083] The preset relevance is a pre-set relevance threshold. When the second relevance is greater than or equal to the preset relevance, it indicates that the initial instruction keyword and the key derived word are highly related, and they can be considered synonyms. When the second relevance is less than the preset relevance, it indicates that the initial instruction keyword and the key derived word are not highly related, and they cannot be considered synonyms. For example, the preset relevance can be 50% or 60%.
[0084] When the second relevance of the initial instruction keyword and each of the key derivative words is greater than or equal to the preset relevance, it indicates that the initial instruction keyword and each key derivative word have a large relevance. This means that the initial voice instruction corresponding to the initial instruction keyword and the user voice instruction corresponding to the key derivative word have a large relevance. Therefore, the initial voice instruction corresponding to the initial instruction keyword is determined as the target voice instruction.
[0085] The specific process for determining the first and second relevance is not limited. For example, an existing trained neural network model can be used to determine the first and second relevance, or a pre-defined relevance rule can be used to determine the first and second relevance.
[0086] In this application, the target voice command can be determined based on the first relevance between each initial command keyword and the user command keyword, or based on the second relevance between each initial command keyword and each key derivative word, thus improving the flexibility and accuracy of determining the target voice command.
[0087] In some embodiments, determining the target control corresponding to the target touch point coordinates from the plurality of controls and performing the control operation corresponding to the target control includes: Obtain the initial position area of each control on the first user interface; In response to the initial position region including the target touch point coordinates, the control corresponding to the initial position region is determined as the target control corresponding to the target touch point coordinates; A simulated touch operation is performed in the initial position area of the target control to execute the control operation corresponding to the target control.
[0088] Specifically, the initial position area of each control on the first user interface is obtained. When the initial position area includes the target touch point coordinates, it means that the initial position area includes the position coordinates that need to be clicked on the first user interface when executing the user's voice command. Then, the control corresponding to the initial position area is determined as the target control corresponding to the target touch point coordinates. In this way, the position area of the determined target control includes the position coordinates that need to be clicked on the first user interface when executing the user's voice command. That is to say, the position coordinates that need to be clicked on the first user interface when executing the user's voice command are included in the position area of the target control.
[0089] At this time, a simulated touch operation is performed in the initial position area of the target control to execute the control operation corresponding to the target control. Thus, by performing a simulated touch operation in the initial position area of the target control, the position coordinates that need to be clicked on the first user interface when executing the user's voice command are also simulated touch operations, thereby enabling the user's voice command to be effectively executed.
[0090] In this way, simulated touch operations can be performed based on the corresponding interfaces that have been pre-set in the system to execute the control operations corresponding to the target control.
[0091] In this application, based on the initial position area of each control on the first user interface and the coordinates of the target touch point, the initial position area containing the coordinates of the target touch point can be accurately determined, thereby accurately determining the target control. Then, a simulated touch operation is performed in the initial position area of the target control to execute the control operation corresponding to the target control. In this way, the function required by the user's voice command is effectively realized based on the execution of the target control, ensuring the realization of the "seeing is speaking" function.
[0092] This application eliminates the need to integrate the voice-enabled "visible and speakable" toolkit with every control, and also eliminates the need for application developers to register keywords for each user interface, saving significant communication and labor costs. Furthermore, determining multiple key derivative words based on the user command keywords allows for better extraction of user intent commands from everyday expressions, ensuring that differences in spoken expression do not affect functional implementation. Additionally, by inputting the initial voice command while clicking the control, the initial touch point coordinates are associated with the initial voice command to generate command position rules, ensuring the accuracy of subsequently determined target touch point coordinates, and consequently, the accuracy of the determined target controls. Moreover, it can simulate user clicks based on preset interfaces at coordinate points specified on the first user interface to complete the response and achieve the "visible and speakable" function.
[0093] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.
[0094] It should be noted that some embodiments of this application have been described above. In some cases, the actions or steps described in the above embodiments can be performed in a different order than that shown in the above embodiments and the desired result can still be achieved. In addition, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0095] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides a control device for a viewable window.
[0096] refer to Figure 3 The control device for the viewable window includes: The acquisition module 100 is configured to acquire a first user interface of a viewable window, the first user interface including multiple controls; The determination module 200 is configured to, in response to receiving a user voice command, determine the target touch point coordinates in the first user interface based on the command position rules for the target touch point corresponding to the user voice command; The execution module 300 is configured to determine the target control corresponding to the coordinates of the target touch point from the plurality of controls, and to perform the control operation corresponding to the target control.
[0097] In some embodiments, the determining module 200 is further configured to: Obtain the initial voice command corresponding to each of the aforementioned controls; In response to the initial voice command meeting the preset command validity conditions, the initial touch point coordinates on the first user interface are determined for the initial touch point corresponding to the initial voice command. Based on each initial voice command and the corresponding initial touch point coordinates, command position rules are generated.
[0098] In some embodiments, the determining module 200 is further configured to: Execute the control operation corresponding to the control corresponding to the initial voice command to switch the first user interface to the second user interface; In response to the fact that the second user interface is different from the first user interface, it is determined that the initial voice command meets the preset command validity conditions.
[0099] In some embodiments, the determining module 200 is further configured to: Obtain the coordinate information of the click operation corresponding to the execution of the initial voice command on the first user interface, and determine the coordinate information as the initial touch point coordinates; Alternatively, the initial position information of the control corresponding to the initial voice command on the first user interface can be obtained, and the initial position information can be determined as the initial touch point coordinates.
[0100] In some embodiments, the determining module 200 is further configured to: Determine the user command keywords of the user's voice command, and determine multiple key derivative words based on the user command keywords; Based on multiple key derived words, a target voice instruction is determined from all the initial voice instructions; Based on the command location rules, the initial touch point coordinates corresponding to the target voice command are determined, and the initial touch point coordinates are determined as the target touch point coordinates corresponding to the user's voice command.
[0101] In some embodiments, the determining module 200 is further configured to: Determine the initial instruction keywords for each of the initial voice commands; In response to the fact that the initial instruction keyword among the plurality of key derived words includes an initial voice instruction, the initial voice instruction corresponding to the initial instruction keyword is determined to be the target voice instruction; Alternatively, in response to an initial instruction keyword that does not include any of the initial voice instructions among the plurality of key derivative words, a target voice instruction is determined from all the initial voice instructions based on the initial instruction keyword, the plurality of key derivative words, and / or the user instruction keyword.
[0102] In some embodiments, the determining module 200 is further configured to: Determine the first relevance between each initial instruction keyword and the user instruction keyword, and determine the initial voice instruction corresponding to the initial instruction keyword with the highest first relevance as the target voice instruction; Alternatively, a second relevance degree between each initial instruction keyword and each key derivative word is determined. In response to a second relevance degree between an initial instruction keyword and each key derivative word being greater than or equal to a preset relevance degree, the initial voice instruction corresponding to the initial instruction keyword is determined as the target voice instruction.
[0103] In some embodiments, the execution module 300 is further configured to: Obtain the initial position area of each control on the first user interface; In response to the initial position region including the target touch point coordinates, the control corresponding to the initial position region is determined as the target control corresponding to the target touch point coordinates; A simulated touch operation is performed in the initial position area of the target control to execute the control operation corresponding to the target control.
[0104] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0105] The apparatus of the above embodiments is used to implement the control method of the corresponding view window in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0106] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the control method of the configurable window described in any of the above embodiments.
[0107] Figure 4 This embodiment illustrates a more specific hardware structure of an electronic device. The device may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0108] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0109] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0110] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0111] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0112] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0113] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0114] The electronic devices described above are used to implement the control methods of the corresponding windows in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0115] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium that stores computer instructions for causing the computer to execute the control method of the readable window as described in any of the above embodiments.
[0116] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0117] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the control method of the window as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0118] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a computer program product, including computer program instructions, which, when run on a computer, cause the computer to execute the control method of the configurable window as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0119] Based on the same inventive concept, and corresponding to the methods of any of the above embodiments, this application also provides a vehicle, which includes the control device, electronic device, computer-readable storage medium, or computer program product with a readable window as described in any of the above embodiments. The vehicle has the technical effects corresponding to any of the above embodiments, which will not be elaborated further here.
[0120] It is understood that before using the technical solutions of the various embodiments in this application, users will be informed of the type, scope of use, and usage scenarios of the personal information involved in an appropriate manner, and user authorization will be obtained.
[0121] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose, based on the prompt message, whether to provide personal information to the software or hardware such as electronic devices, applications, servers, or storage media performing the operations described in this application.
[0122] As an optional but not limited implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0123] It is understood that the above notification and user authorization process is merely illustrative and does not limit the implementation of this application. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this application.
[0124] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application is limited to these examples; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in detail for the sake of brevity.
[0125] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0126] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0127] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of this application. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. A method for controlling a viewport, characterized in that, include: Obtain a first user interface of a viewable window, the first user interface including multiple controls; In response to receiving a user voice command, the coordinates of the target touch point in the first user interface are determined based on the command location rules. From the plurality of controls, determine the target control corresponding to the coordinates of the target touch point, and execute the control operation corresponding to the target control.
2. The method according to claim 1, characterized in that, The method further includes: Obtain the initial voice command corresponding to each of the aforementioned controls; In response to the initial voice command meeting the preset command validity conditions, the initial touch point coordinates on the first user interface are determined for the initial touch point corresponding to the initial voice command. Based on each initial voice command and the corresponding initial touch point coordinates, command position rules are generated.
3. The method according to claim 2, characterized in that, The method further includes: Execute the control operation corresponding to the control corresponding to the initial voice command to switch the first user interface to the second user interface; In response to the fact that the second user interface is different from the first user interface, it is determined that the initial voice command meets the preset command validity conditions.
4. The method according to claim 2, characterized in that, Determining the initial touch point coordinates on the first user interface corresponding to the initial voice command includes: Obtain the coordinate information of the click operation corresponding to the execution of the initial voice command on the first user interface, and determine the coordinate information as the initial touch point coordinates; Alternatively, the initial position information of the control corresponding to the initial voice command on the first user interface can be obtained, and the initial position information can be determined as the initial touch point coordinates.
5. The method according to claim 2, characterized in that, The step of determining the target touch point coordinates in the first user interface corresponding to the user's voice command based on command location rules includes: Determine the user command keywords of the user's voice command, and determine multiple key derivative words based on the user command keywords; Based on multiple key derived words, a target voice instruction is determined from all the initial voice instructions; Based on the command location rules, the initial touch point coordinates corresponding to the target voice command are determined, and the initial touch point coordinates are determined as the target touch point coordinates corresponding to the user's voice command.
6. The method according to claim 5, characterized in that, The process of determining a target voice instruction from all the initial voice instructions based on multiple key derived words includes: Determine the initial instruction keywords for each of the initial voice commands; In response to the fact that the initial instruction keyword among the plurality of key derived words includes an initial voice instruction, the initial voice instruction corresponding to the initial instruction keyword is determined to be the target voice instruction; Alternatively, in response to an initial instruction keyword that does not include any of the initial voice instructions among the plurality of key derivative words, a target voice instruction is determined from all the initial voice instructions based on the initial instruction keyword, the plurality of key derivative words, and / or the user instruction keyword.
7. The method according to claim 6, characterized in that, The step of determining a target voice command from all the initial voice commands based on the initial command keywords, the multiple key derived words, and / or the user command keywords includes: Determine the first relevance between each initial instruction keyword and the user instruction keyword, and determine the initial voice instruction corresponding to the initial instruction keyword with the highest first relevance as the target voice instruction; Alternatively, a second relevance degree between each initial instruction keyword and each key derivative word is determined. In response to a second relevance degree between an initial instruction keyword and each key derivative word being greater than or equal to a preset relevance degree, the initial voice instruction corresponding to the initial instruction keyword is determined as the target voice instruction.
8. The method according to claim 1, characterized in that, The step of determining the target control corresponding to the coordinates of the target touch point from the plurality of controls and executing the control operation corresponding to the target control includes: Obtain the initial position area of each control on the first user interface; In response to the initial position region including the target touch point coordinates, the control corresponding to the initial position region is determined as the target control corresponding to the target touch point coordinates; A simulated touch operation is performed in the initial position area of the target control to execute the control operation corresponding to the target control.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 8.
10. A vehicle, characterized in that, Includes the electronic device as described in claim 9.