Voice control method, electronic device, and storage medium
Patent Information
- Application Number
- CN202211448516.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-18
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2042-11-18
AI Technical Summary
[0003]然而,目前的语音识别方案中的语音词通常是静态的,导致电子设备的语音识别能力有限
[0008]根据本申请的一些实施方式,电子设备在监听到视图变化事件后,获取变化后的视图中显示的视图容器所绑定的模型层数据,并从中获取语音词,并将获取的语音词注册至用于语音服务的语音词集合,使得语音服务能够在视图内容变化后,自动触发语音词的注册。此外,在接收到基于语音词结合识别的语音指令后,电子设备可基于模型层数据响应语音指令,实现语音控制。
Smart Images

Figure CN117542357B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this application relate to the field of electronic technology, and more specifically, to a voice control method, an electronic device, and a storage medium. Background Technology
[0002] With the development of technology, voice interaction control has become a standard feature of electronic devices. Electronic devices obtain voice commands by recognizing user voice data and execute voice commands to respond to user needs.
[0003] However, current speech recognition solutions typically use static speech words, which limits the speech recognition capabilities of electronic devices. Summary of the Invention
[0004] The embodiments of this application provide a voice control method, electronic device, and storage medium that can at least partially solve the above-mentioned or other problems existing in the prior art.
[0005] Implementations of this application provide a voice control method, comprising: in response to listening to a view change event, acquiring first model layer data bound to a view container displayed after the view change; acquiring and registering a first voice word from the first model layer data; receiving a voice command recognized based on a set of voice words, and determining second model layer data corresponding to the voice command, the set of voice words including the first voice word; and responding to the voice command according to the second model layer data.
[0006] Embodiments of this application also provide an electronic device, including: at least one processor and a memory, the memory being communicatively connected to the at least one processor and storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the voice control method mentioned in the above embodiments.
[0007] The embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the voice control method mentioned in the above embodiments.
[0008] According to some embodiments of this application, after detecting a view change event, the electronic device obtains the model layer data bound to the view container displayed in the changed view, extracts speech words from it, and registers the extracted speech words to a speech word set used for voice services. This allows the voice service to automatically trigger the registration of speech words after the view content changes. Furthermore, upon receiving a voice command based on speech word combination recognition, the electronic device can respond to the voice command based on the model layer data, thereby achieving voice control. Attached Figure Description
[0009] Other features, objects, and advantages of this application will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. Wherein:
[0010] Figure 1 This is a schematic flowchart of a voice control method according to some embodiments of this application;
[0011] Figure 2 This is a schematic diagram of the interface view of a music application according to some embodiments of this application;
[0012] Figure 3 This is a schematic diagram of the interface view of a network radio according to some embodiments of this application;
[0013] Figure 4 These are schematic diagrams of user interface views according to some embodiments of this application; and
[0014] Figure 5 This is a schematic diagram of the structure of an electronic device according to some embodiments of this application. Detailed Implementation
[0015] To better understand this application, various aspects of this application will be described in more detail with reference to the accompanying drawings. It should be understood that these detailed descriptions are merely illustrative of exemplary embodiments of this application and are not intended to limit the scope of this application in any way. Throughout the specification, the same reference numerals refer to the same elements. The expression "and / or" includes any and all combinations of one or more of the associated listed items.
[0016] It should also be understood that expressions such as "comprising," "including," "having," "containing," and / or "comprising" are open-ended rather than closed-ended expressions in this specification, indicating the presence of the stated features, elements, and / or components, but not excluding the presence of one or more other features, elements, components, and / or combinations thereof. Furthermore, when expressions such as "at least one of..." appear after a list of listed features, they modify the entire list of features, not just individual elements in the list. Additionally, when describing embodiments of this application, the word "may" is used to mean "one or more embodiments of this application." And the term "exemplary" is intended to refer to examples or illustrations.
[0017] Unless otherwise specified, all terms used herein (including engineering and technical terms) shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. It should also be understood that, unless expressly stated herein, terms defined in common dictionaries shall be interpreted as having the meaning consistent with their meaning in the context of the relevant art, and not as having an idealized or overly formalized meaning.
[0018] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. Furthermore, unless explicitly limited or contradicted by the context, the specific steps included in the methods described in this application are not limited to the order in which they are described, but can be performed in any order or in parallel. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0019] In some implementations, various view controls (such as ListView and RecyclerView controls in Android) provide users with different views. Taking RecyclerView as an example, it binds model layer data to a view container (ViewHolder) and displays different view components (Views) within the view container. The Views provided by RecyclerView can change at any time, but current speech recognition solutions cannot easily register the speech words within the RecyclerView control, resulting in limited speech recognition capabilities and hindering user voice control.
[0020] Figure 1 This is a schematic flowchart of a voice control method according to some embodiments of this application. For example... Figure 1 As shown, the voice control method 100 includes the following steps:
[0021] S11, in response to a view change event, retrieves the first model layer data bound to the view container that is displayed after the view change.
[0022] S12, Obtain and register the first speech word from the first model layer data.
[0023] S13, receive a voice command based on a set of speech words, and determine the second model layer data corresponding to the voice command. The set of speech words includes the first speech words.
[0024] S14, respond to voice commands based on the data from the second model layer.
[0025] According to some embodiments of this application, after detecting a view change event, the electronic device obtains the model layer data bound to the view container displayed in the changed view, extracts speech words from it, and registers the extracted speech words to the speech word set used for voice services. This allows the voice service to automatically trigger the registration of speech words after the view content changes, without requiring additional operations from the developer. Furthermore, upon receiving a voice command based on speech word combination recognition, the electronic device can respond to the voice command based on the model layer data, thus achieving voice control.
[0026] To facilitate understanding, the following will be combined with... Figures 2 to 4The voice control method 100 provided in the embodiments of this application will be described by way of example.
[0027] In some embodiments of this application, the electronic device may include a view control and a voice service control. The voice service control is used to recognize the user's voice data based on a registered set of voice words to determine the voice command triggered by the user. The view control is used to display the model layer data it is bound to, such as images and text within the model layer data.
[0028] In some embodiments of this application, the electronic device may initialize the view control and load the associated model layer data upon initial view loading. The electronic device may register speech words in a speech word set based on the model layer data bound to the currently visible view container of the view control, in order to provide the speech service control with speech words for recognizing speech data.
[0029] It should be understood that, without departing from the teachings of this application, the initialization process of view controls can refer to some related technologies, which will not be elaborated here.
[0030] In some embodiments of this application, after displaying a view to a user based on model layer data, the electronic device can listen for view change events and update the voice words upon detecting a view change event. For example, the view change events listened to by the electronic device may include one of the following: a view refresh event, a view jump event, and a view scroll event. A view refresh event may refer to a change in the model layer data bound to the view container within the view, requiring the view to be refreshed to present the updated model layer data. A view jump event may refer to jumping from one interface view to another, for example, from... Figure 2 The example music application's interface view 200 jumps to... Figure 3 The example is a view 300 of the internet radio interface. View scrolling events can include horizontal scrolling events and vertical scrolling events.
[0031] In some embodiments of this application, the electronic device can obtain the model layer data bound to the view container that is displayed after the view changes through the Application Program Interface (API), that is, the data bound to each view container that is currently interacting with the user (hereinafter referred to as the first model layer data).
[0032] For example, the view control is a RecyclerView control. Electronic devices can obtain the currently interacting ViewHolder through the RecyclerView control's API. Since the ViewHolder can be seen as a combined carrier of model layer data and View, electronic devices can obtain the necessary voice words and the Views they can interact with through the ViewHolder.
[0033] In some embodiments of this application, the electronic device may obtain and register speech words (hereinafter referred to as first speech words) from the first model layer data. For example, the electronic device may determine whether a first data object that allows user operation exists in the first model layer data, and if it is determined that it exists, obtain and register the first speech word from the first data object.
[0034] For example, the view control is a RecyclerView control, and user actions can include click actions or voice actions. The electronic device can determine whether each View in the ViewHolder supports click actions or voice actions. If it is determined that it supports them, the first voice word can be obtained from the model layer data corresponding to that View.
[0035] Optionally, the process by which the electronic device acquires and registers the first speech word from the first data object may include: acquiring a third speech word from the first data object; determining whether the third speech word already exists in the speech word set; if it exists, not registering the third speech word again; if it does not exist, identifying the third speech word as the first speech word and registering the first speech word. In other words, the electronic device compares the newly acquired speech word with the already registered speech words in the speech word set to determine and register the newly added speech word.
[0036] Optionally, after acquiring the first spoken word, the electronic device can record the correspondence between the storage location of the first spoken word and the storage location of the first model layer data. For example, the view control is a RecyclerView control. The RecyclerView control creates a ViewHolder through an Adapter and binds the data at the target storage location to the ViewHolder's View. Therefore, the electronic device can obtain the storage location of the first model layer data that each ViewHolder needs to bind, and it can also obtain the first model layer data bound to the ViewHolder. After extracting the first spoken word from the first model layer data, the electronic device can store the first spoken word and the corresponding storage location of the first model layer data in the database.
[0037] In some embodiments of this application, after acquiring the user's voice data, the voice service control of the electronic device can recognize the voice data based on a set of voice words, trigger a voice command, and transmit the triggered voice command to a view control. Upon receiving the voice command, the view control searches for the model layer data corresponding to the voice command (hereinafter referred to as the second model layer data) and responds to the voice command according to the second model layer data. The second model layer data mentioned herein can refer to data bound within a view container in the view's model layer data.
[0038] In some embodiments of this application, the view control finds the second model layer data corresponding to the voice command in ways including but not limited to method one, method two, and method three.
[0039] Method 1: The view control searches for the view container corresponding to the second voice word in the voice command within the currently displayed view container. If it is found to exist, the voice command can be distributed to that view container. The view container can then determine its bound third model layer data as the second model layer data and respond to the voice command based on the second model layer data.
[0040] Optionally, if the electronic device determines that there is no view container corresponding to the voice control command, it may prompt the user that the operation has failed. It should be understood that, without departing from the teachings of this application, the electronic device may prompt the user that the operation has failed through voice or pop-up windows, and this application does not limit the prompting method of the electronic device.
[0041] Method 2: The view control can determine the storage location of the second model layer data based on the second voice word in the voice command and the correspondence between the voice word and the storage location of the model layer data, and retrieve the second model layer data based on the determined storage location. For example, as described above, the view control can record the correspondence between the retrieved voice word and the storage location of the model layer data during the process of retrieving the voice word, so that the view control can find the second model layer data corresponding to the second voice word in the voice command based on the correspondence.
[0042] Method 3: The view control searches for the view container corresponding to the second voice command in the currently displayed view container. If it is found to exist, the voice command can be distributed to that view container, which can identify the third model layer data it is bound to as the second model layer data and respond to the voice command based on the second model layer data. If it is found not to exist, the view control can determine the storage location of the second model layer data based on the second voice word in the voice command and the correspondence between the voice word and the storage location of the model layer data, and retrieve the second model layer data based on the determined storage location.
[0043] In methods two and three above, since the storage location of the model layer data corresponding to the registered voice words is recorded, even if the view container corresponding to the voice words contained in the voice command is slid out of the view, making it impossible to determine the model layer data corresponding to the voice command based on the view container, the electronic device can still find the model layer data corresponding to the voice words contained in the voice command based on the record and respond to the voice command accordingly, thereby reducing the situation where the voice command cannot be responded to due to the view container being slid out of the view.
[0044] Optionally, if the electronic device determines that none of the currently displayed view containers are bound to the second model layer data, it updates the view based on the second model layer data. In other words, if the second model layer data corresponding to a user-triggered voice command is not displayed in the current view, the electronic device can update the view based on the second model layer data, for example, by jumping to a view that displays the second model layer data. In the above scheme, the electronic device can jump to the corresponding view based on a user-triggered voice command, making it easier for the user to understand the model layer data related to the voice command and improving the user experience.
[0045] For example, in some scenarios, the electronic device is a vehicle head unit, and the view control is a RecyclerView control. The view of the vehicle head unit may include at least a music application interface view 200 and an internet radio interface view 300. The music application interface view may be as follows: Figure 2 As shown, the interface view 300 of the internet radio station can be viewed as follows: Figure 3As shown. The interface view 200 of the music application may include a first view container 210, a second view container 220, and a third view container 230. The first view container 210 may include a first view component 211 bound to data related to the settings interface, a second view component 212 bound to data related to the search box interface, and a third view component 213 bound to data related to the personal center interface. The second view container 220 may include a fourth view component 221 bound to data related to the music application's daily recommendations interface, a fifth view component 222 bound to data related to the music application's leaderboard interface, and a sixth view component 223 bound to data related to the music application's favorites interface. The third view container 230 may include a seventh view component 231 bound to a text box displaying "Recommended Playlists." The fourth view container 240 may include an eighth view component 241 bound to data related to playlist 1, a ninth view component 242 bound to data related to playlist 2, and a tenth view component 243 bound to data related to playlist 3. The view interface 300 of the internet radio station may include at least the first view container 210, the fourth view container 310, and the fifth view container 320 mentioned above. The fourth view container 310 may include an eleventh view component 311, which is associated with the daily recommendation interface of the internet radio station, and a twelfth view component 312, which is associated with the favorites interface of the internet radio station. The fifth view container 320 may include a thirteenth view component 321, which is associated with the first type of radio station list interface, and a fourteenth view component 322, which is associated with the second type of radio station list interface. Among them, some content in the first view component 211, the second view component 212, the third view component 213, the fourth view component 221, the fifth view component 222, the sixth view component 222, the eighth view component 241, the ninth view component 242, the tenth view component 243, the eleventh view component 311, the twelfth view component 312, the thirteenth view component 321, and the fourteenth view component 322 supports voice control, such as "Daily Recommendation", "Ranking List", "Playlist 1", "First Type of Radio Station", etc.
[0046] In some scenarios, the user first opens the interface view 200 of a music application. After the electronic device displays the interface view 200, it can obtain voice words based on the model layer data bound to the first to tenth view components. For example, the obtained voice words may include "Daily Recommendation," "Rankings," "Playlist 1," etc., and may also include the song titles, artist names, etc., of songs within the Daily Recommendation and Ranking interfaces. The storage location corresponding to the above voice words is recorded. For example, when the user speaks a sentence containing "Daily Recommendation," the voice service control of the electronic device can trigger a voice command instructing the opening or playing of the daily recommendations of the music application and send it to the view control. Based on the voice command, the view control recognizes that it corresponds to the second view container 220 and distributes the voice command to the second view container 220. The second view container 220 determines that it needs to respond to the voice command by simulating a click and can obtain the view component corresponding to the voice command, namely the fourth view component 221 bound to the relevant data of the daily recommendations interface of the music application. The second view container 220 simulates clicking the fourth view component 221 to display or play the relevant songs of the daily recommendations of the music application. As can be seen from the above, after using the RecyclerView control and the voice control method 100 mentioned in the embodiments of this application, the electronic device can obtain and register voice words based on the model layer data bound to each view container (ViewHolder), and when the user controls the electronic device through voice commands, it can more accurately control the view component (View) associated with it in the view container (ViewHolder).
[0047] In some scenarios, if a user clicks on an internet radio station to switch to the internet radio station's interface view, after the electronic device displays the internet radio station's interface view 300, it can obtain voice words based on the model layer data bound to the eleventh to fourteenth view components. For example, the obtained voice words may include "daily recommendation," "first type of radio station," etc., and may also include the radio station names and host names in the interface such as daily recommendations and first type of radio stations in the internet radio station, and record the storage location corresponding to the above voice words.
[0048] In some scenarios, when an electronic device displays the interface view 300 of an internet radio station, if the user triggers a voice command related to a voice word obtained from the interface view 200 of a music application, such as a voice command triggered based on the voice word "playlist 1", since the storage location corresponding to each voice word is recorded when the interface view 200 of the music application is opened, the electronic device can lock the model layer data corresponding to the voice command, respond to the voice command based on the model layer data, and display the interface view corresponding to the model layer data.
[0049] In some embodiments of this application, the process by which an electronic device responds to a voice command based on the second model layer data may include: obtaining a second data object corresponding to the voice command from the second model layer data, and executing the voice command based on the second data object.
[0050] In some embodiments of this application, the same voice word may correspond to multiple third data objects. Based on this, in the embodiments of this application, the processing method of the electronic device in triggering voice commands may include, but is not limited to, method 1, method 2 and method 3.
[0051] Method 1: The electronic device can identify the third data object associated with the voice command in the second model layer data, and determine if the number of third data objects is greater than 1. If the number of third data objects is determined to be 1, the third data object can be identified as the second data object. If the number of third data objects is determined to be greater than 1, the priority of each third data object is obtained, and the third data object with the highest priority is identified as the second data object. For example, each data object is assigned a priority, which can be set by the developer, and the priorities of data objects with the same voice word are different. After identifying the third data object associated with the voice command, the electronic device can obtain the priority of the third data object and determine the second data object based on the priority of each third data object. In the above scheme, the electronic device can identify a second data object to respond to the voice command even when there are multiple third data objects, reducing the problem of voice response errors caused by multiple data objects having the same voice word.
[0052] Method 2: The electronic device can determine the third data object associated with the voice command in the second model layer data, and determine whether the number of third data objects is greater than 1. If the number of third data objects is determined to be equal to 1, the third data object can be determined as the second data object; if the number of third data objects is determined to be greater than 1, the first or last third data object can be determined as the second data object. For example, if multiple third data objects come from the same second model layer data, and the data objects in the second model layer data are arranged in a specified order, the first or last third data object can refer to the third data object that is first or last in the second model layer data. As another example, if multiple third data objects come from different second model layer data, and the multiple second model layer data in the view are arranged in a specified order, the first or last third data object can refer to the third data object in the second model layer data that is first or last in the second model layer data. Furthermore, the first or last third data object can also refer to the third data object in the second model layer data corresponding to the earliest or latest record in the record correspondence relationship; this application does not impose any restrictions on this.
[0053] Method 3: The electronic device can identify the third data object associated with the voice command obtained from the second model layer data as the second data object. If the number of second data objects is determined to be equal to 1, the electronic device can respond to the voice command based on the second data object. If the number of second data objects is determined to be greater than 1, the electronic device can send the voice command and prompt information to each second data object respectively. The prompt information is used to inform the second data object that other data objects besides itself have received the voice command, so that the second data object can determine whether to execute the voice command based on its own priority and the priorities of other data objects. For example, if the electronic device determines that the data objects associated with the voice command include data object A and data object B, it can send the voice command and first prompt information to data object A, and send the voice command and second prompt information to data object B. The first prompt information includes the identification information or priority of data object B, and the second prompt information includes the identification information or priority of data object A. After receiving the voice command and the first prompt information, data object A determines the priority of data object B based on the first prompt information and checks whether its own priority is higher than that of data object B. If it is, it responds to the voice command; if it is not, it does not respond to the voice command. Similarly, after receiving the voice command and the second prompt, data object B determines the priority of data object A based on the second prompt and checks whether its own priority is higher than that of data object A. If it is, it responds to the voice command; if it is not, it does not respond to the voice command.
[0054] Optionally, if the second data object associated with the voice command is not found in the second model layer data, i.e. the operation of obtaining the second data object fails, the electronic device can respond to the voice command through the root data object of the second model layer data.
[0055] After providing an exemplary description of determining the second data object, the process of responding to voice commands based on the second data object will now be described in an exemplary manner.
[0056] In some embodiments of this application, the differences between voice commands and data objects lead to diverse ways in which electronic devices respond to voice commands. For ease of understanding, the following description primarily uses the process by which an electronic device determines whether to simulate clicking a view component bound to a second data object based on the second data object as an example.
[0057] In some embodiments of this application, the electronic device can determine whether it is necessary to respond to the voice command by simulating clicking on the view component corresponding to the second data object based on the voice command and the model layer data corresponding to the voice command. If it is determined to be yes, then the view component is clicked.
[0058] Optionally, the electronic device can call the ViewHolder callback method after the simulated click time is completed to implement the click event callback.
[0059] Alternatively, the view control could be, for example, a RecyclerView, and the view container could be, for example, a ViewHolder. During integration, developers can add functionality to the ViewHolder using an interface pattern to achieve the above solution. For example, interfaces extending the ViewHolder could include interfaces for retrieving spoken words, interfaces for determining if a click was triggered, interfaces for retrieving the View corresponding to the click, and interfaces for voice callbacks. The interface pattern does not break the ViewHolder's logical functionality and can improve stability.
[0060] In some scenarios, the electronic device is the vehicle's main unit, and the view currently displayed on the main unit is the user interface view of the vehicle's internal devices. This user interface view can be as follows: Figure 4 As shown. Figure 4As shown, the operation interface view 400 includes a seventh view container 410, an eighth view container 420, and a ninth view container 430. The seventh view container 410 is bound to the air conditioning control program and includes a fifteenth view component 411 for adjusting the air conditioning mode, a sixteenth view component 412 for adjusting the air conditioning temperature, and a seventeenth view component 413 for adjusting the air conditioning fan speed. The eighth view container 420 is bound to the window control program and includes an eighteenth view component 421 for closing the windows, a nineteenth view component 422 for opening the windows, and a twentieth view component 423 for locking the windows. The ninth view container 430 is bound to the door control program and includes a twenty-first view component 431 for opening the doors, a twenty-second view component 432 for closing the doors, and a twenty-third view component 433 for locking the doors. All of the above functions support voice operation. The electronic device, based on the model layer data bound to the seventh view container 410, eighth view container 420, and ninth view container 430, obtains voice words that may include at least: "air conditioning," "temperature," "fan speed," "mode," "open window," "close window," "lock window," "unlock window," "open door," "close door," "lock door," and "unlock door." These voice words are registered to a voice service control. The voice service control can recognize voice data based on the registered voice words to convert it into voice commands. Taking the voice data received by the voice service control as the voice data corresponding to "lock the window," the voice service control can trigger a voice command indicating "lock the window" based on the registered voice words. After receiving the voice command indicating "lock the window," the view control can determine that the view component corresponding to the voice command is the twentieth view component 423. The electronic device obtains the model layer data related to the window control program corresponding to the twentieth view component 423 and retrieves data objects related to the window lock status information from it. If the electronic device determines that the data within the data object indicates the window is unlocked, it can respond to the voice command by simulating a click on the twentieth view component 423 to control the vehicle to close the window. If it determines that the data within the data object indicates the window is locked, it will not simulate a click on the twentieth view component 423, and may instead selectively remind the user of "the window is locked" via voice or display screen. This solution reduces unnecessary resource waste caused by the electronic device simulating a click on the twentieth view component 423 again when the window is already in an open or locked state. Especially when the functions of locking and unlocking the window correspond to the same view component, it can reduce the possibility of accidental locking or unintended locking.
[0061] As explained above, electronic devices can determine their response to voice commands by acquiring the second data object corresponding to the voice command. This reduces the likelihood of the electronic device triggering invalid or erroneous commands, thereby reducing resource waste and false responses. An invalid command can be one that, when executed, does not change the current state of the electronic device.
[0062] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this patent.
[0063] Embodiments of this application also provide an electronic device, such as... Figure 5 As shown, the electronic device 500 may include: at least one processor and a memory, the memory being communicatively connected to the at least one processor and storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the aforementioned voice control method 100. For example, the electronic device 500 may be, for instance, a vehicle host computer.
[0064] One embodiment of this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the voice control method 100.
[0065] Figure 5 This is a schematic block diagram of an electronic device 500 according to some embodiments of this application. For example... Figure 5 As shown, the electronic device 500 includes a processor 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a memory 508 into a random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of the electronic device 500. The processor 501, ROM 502, and RAM 503 are interconnected via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0066] Multiple components in electronic device 500 are connected to I / O interface 505, including: input unit 506, such as buttons or touchscreens in a vehicle infotainment system; output unit 507, connected to various types of displays, speakers, etc., to output various forms of signals; memory 508, including any medium for storing computer-executable programs; and communication unit 509, such as a network interface card (NIC), modem, or wireless transceiver. Communication unit 509 allows electronic device 500 to exchange information / data with other devices via a local area network (LAN) or other wireless communication networks.
[0067] Processor 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 501 performs the various methods and processes described above, such as voice control method 100. For example, in some embodiments, voice control method 100 may be implemented as a computer software program tangibly contained in a computer-readable storage medium, such as memory 508. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by processor 501, one or more steps of voice control method 100 described above may be performed. Alternatively, in other embodiments, processor 501 may be configured to perform voice control method 100 by any other suitable means (e.g., by means of firmware).
[0068] Various aspects of this application have been described herein with reference to flowchart illustrations and / or timing diagrams of methods, apparatus (systems), and computer program products according to exemplary embodiments of this application. It should be understood that each step of the flowchart illustrations and / or timing diagrams, as well as combinations of steps in the flowchart illustrations and / or timing diagrams, can be implemented by computer-readable program instructions.
[0069] These computer-readable program instructions can be provided to a processor, general-purpose computer, special-purpose computer, or other programmable data processing unit in an electronic device to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing device, they create means for implementing the functions / steps specified in one or more steps of a flowchart and / or timing diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing device, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / steps specified in one or more steps of a flowchart and / or timing diagram.
[0070] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / steps specified in one or more steps of a flowchart and / or timing diagram.
[0071] The flowcharts and timing diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each step in a flowchart or timing diagram may represent a module, segment, or part of an instruction that contains one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the steps may occur in a different order than those indicated in the drawings. For example, two consecutive steps may actually be performed substantially in parallel, and they may sometimes be performed in reverse order, depending on the functions involved. It should also be noted that each step in a timing diagram and / or flowchart, and combinations of steps in timing diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0072] The above description is merely an illustration of the embodiments of this application and the technical principles employed. Those skilled in the art should understand that the scope of protection involved in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the technical concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A voice control method, characterized in that, include: In response to a view change event, retrieve the first model layer data bound to the view container that is displayed after the view change; Obtain and register the first speech word from the data of the first model layer; Receive a voice command based on a set of speech words, and determine the second model layer data corresponding to the voice command, wherein the set of speech words includes the first speech words; and The voice command is responded to based on the data from the second model layer.
2. The method according to claim 1, wherein, The data for the second model layer is determined to include: In response to the existence of a view container corresponding to the voice command in the currently displayed view container, the third model layer data bound to the view container corresponding to the voice command is determined as the second model layer data.
3. The method according to claim 2, wherein, The method further includes: If the view container corresponding to the voice command does not exist in the currently displayed view container, the user is prompted that the operation failed.
4. The method according to claim 1, wherein, The method further includes: Record the correspondence between the storage locations of the first speech word and the data of the first model layer; The data for determining the second model layer includes: The storage location of the second model layer data is determined based on the second voice word in the voice command and the corresponding relationship, and the second model layer data is obtained based on the determined storage location.
5. The method according to any one of claims 1 to 4, wherein, The step of obtaining and registering the first speech word from the first model layer data includes: In response to the existence of a first data object in the first model layer data that allows user operation, the first speech word is obtained from and registered.
6. The method according to claim 5, wherein, Obtaining and registering the first speech word from the first data object includes: Obtain the third speech word from the first data object; and In response to the absence of the third speech word in the speech word set, the third speech word is identified as the first speech word, and the first speech word is registered.
7. The method according to any one of claims 1 to 4, wherein, The step of responding to the voice command based on the data from the second model layer includes: Obtain the second data object corresponding to the voice command from the second model layer data; and The voice command is executed based on the second data object.
8. The method according to claim 7, wherein, The method further includes: In response to the fact that none of the currently displayed view containers are bound to the second model layer data, the view is updated based on the second model layer data.
9. The method according to claim 7, wherein, The step of obtaining the second data object corresponding to the voice command from the second model layer data includes: Determine the third data object in the second model layer data that is associated with the voice command; and In response to the fact that the number of the third data objects is greater than 1, the priority of each of the third data objects is obtained, and the third data object with the highest priority is determined as the second data object.
10. The method according to claim 7, wherein, The execution of the voice command based on the second data object includes: In response to the fact that the number of the second data objects is greater than 1, the voice command and prompt information are sent to each of the second data objects respectively; The prompt information is used to inform the second data object that other data objects besides itself have received the voice command, so that the second data object can determine whether to execute the voice command based on its own priority and the priorities of the other data objects.
11. The method according to claim 7, wherein, The method further includes: In response to the failure of the operation to obtain the second data object, the voice command is responded to through the root data object of the second model layer data.
12. The method according to claim 1, wherein, The view change events include at least one of the following: view refresh event, view jump event, and view scroll event.
13. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the voice control method as described in any one of claims 1 to 12.
14. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the voice control method as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Voice control method and system for smart television and smart television
CN107948698A
Voice control method and device, electronic equipment and readable storage medium
CN112581946A