A voice interaction method and device
By working together with the main device and the first device, the problem of wake word memory in multi-device scenarios is solved, and the accuracy of voice wake-up and user experience are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGBO FOTILE KITCHEN WARE CO LTD
- Filing Date
- 2026-01-08
- Publication Date
- 2026-06-02
Smart Images

Figure CN122135700A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart device technology, and in particular to a voice interaction method and apparatus. Background Technology
[0002] With the development of smart device technology, more and more smart devices are equipped with voice-activated functionality. Smart devices can be woken up by a detected wake word and respond to the corresponding voice command. In scenarios involving multiple smart devices, related technologies often assign different wake words to different devices to ensure accurate wake-up and thus effective response. However, this setup requires users to remember different wake words for different smart devices, which can negatively impact the user experience. Summary of the Invention
[0003] To address at least one of the aforementioned technical problems, this application provides a voice interaction method and apparatus: According to a first aspect of this application, a voice interaction method is provided, applied to a main device, the method comprising: Acquire target speech data, wherein the target speech data includes a preset wake word or the target speech data is subsequent speech data of the reference speech data including the preset wake word; The target voice data is sent to the first device so that the first device sends a target control command to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The target device is any device in a preset network device set, which includes the master device and multiple candidate slave devices. When the target device is a target slave device, the task execution status sent by the first device is received. The task execution status indicates that the target slave device is executing the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
[0004] According to a second aspect of this application, a voice interaction method is provided, applied to a first device, the method comprising: Receive target voice data sent by the master device, wherein the target voice data includes a preset wake-up word or the target voice data is subsequent voice data of the reference voice data including the preset wake-up word; Based on the target voice data, a target device is determined from the device set indicating a preset network, and a target control command is sent to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The device set includes the master device and multiple candidate slave devices. If the target device is a target slave device, the task execution status is sent to the master device. The task execution status indicates that the target slave device is executing the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
[0005] According to a third aspect of this application, a voice interaction device is provided, configured on a main device, the device comprising: A voice data acquisition module is used to acquire target voice data, wherein the target voice data includes a preset wake-up word or the target voice data is subsequent voice data of a reference voice data including the preset wake-up word; A voice data transmission module is used to send the target voice data to a first device, so that the first device sends a target control command to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The target device is any one of the devices in a preset network device set. The device set includes the master device and multiple candidate slave devices. The execution result receiving module is used to receive the task execution status sent by the first device when the target device is a target slave device. The task execution status represents the situation in which the target slave device executes the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
[0006] According to a fourth aspect of this application, a voice interaction device is provided, configured in a first device, the device comprising: A voice data receiving module is used to receive target voice data sent by a master device, wherein the target voice data includes a preset wake-up word or the target voice data is subsequent voice data of a reference voice data including the preset wake-up word; The control command sending module is used to determine the target device from the device set indicating the preset network based on the target voice data, and to send the target control command to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The device set includes the master device and multiple candidate slave devices. The execution result sending module is used to send task execution status to the master device when the target device is a target slave device. The task execution status represents the situation of the target slave device executing the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
[0007] According to a fifth aspect of this application, a main device is provided, including a voice interaction device as described in the third aspect.
[0008] According to a sixth aspect of this application, a first device is provided, including a voice interaction device as described in the fourth aspect.
[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this application.
[0010] Implementing this application will have the following beneficial effects: This application provides a voice interaction solution that can effectively adapt to scenarios involving multiple smart devices. This application sets up a master device for a pre-defined network of devices and introduces a first device to provide decision support. After the master device is voice-activated, it sends target voice data to the first device. The first device then determines the target device from the device set through intent matching and sends target control commands to the target device, enabling the target device to provide an effective response to the target voice data. The pre-setting of the master device ensures that the device that can be voice-activated is unique, which is beneficial to the accuracy of voice activation. Compared to related technologies, it can also reduce the reliance on users to remember multiple wake words to a certain extent, thereby improving the user experience. In the cooperation between the first device and the master device, the master device is mainly responsible for transmitting the target voice data, while the first device is mainly responsible for determining the target device that needs to respond to the target voice data and sending target control commands to it. This arrangement can support users in setting up the master device according to their own needs.
[0011] Other features and aspects of this application will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description
[0012] The objectives, technical solutions, and beneficial effects of the present invention described above can be clearly obtained through the following detailed description of specific embodiments that enable the implementation of the present invention, in conjunction with the accompanying drawings.
[0013] The same reference numerals and symbols in the accompanying drawings and the specification are used to represent the same or equivalent elements.
[0014] Figure 1 This is a flowchart illustrating a voice interaction method provided in this application; Figure 2 This is a schematic diagram of an application environment provided in this application; Figure 3 This is also a flowchart illustrating a voice interaction method provided in this application; Figure 4 This is a flowchart illustrating the process of setting up the main device via a second device, as provided in this application. Figure 5 This is a device block diagram of a voice interaction device provided in this application; Figure 6 This is also a device block diagram of a voice interaction device provided in this application. Detailed Implementation
[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0017] Various exemplary embodiments, features, and aspects of this application will now be described in detail with reference to the accompanying drawings. The same reference numerals in the drawings denote elements that have the same or similar functions. Although various aspects of the embodiments are shown in the drawings, they are not necessarily drawn to scale unless specifically indicated otherwise.
[0018] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0019] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0020] Furthermore, to better illustrate this application, numerous specific details are provided in the following detailed description. Those skilled in the art should understand that this application can be implemented without certain specific details. In some instances, methods, means, components, and circuits well-known to those skilled in the art have not been described in detail in order to highlight the main points of this application.
[0021] Figure 1 This diagram illustrates a flow chart of a voice interaction method according to an embodiment of this application, as shown below. Figure 1 As shown, the method includes: S101: Obtain target speech data, wherein the target speech data includes a preset wake-up word or the target speech data is subsequent speech data of the reference speech data including the preset wake-up word; In this embodiment of the application, steps S101-S103 provide a voice interaction method with the host device as the execution subject. For example... Figure 2 As shown, the master device can be any device in the preset network device set. The master device acquires target voice data. The voice data that wakes up the master device is voice data containing a preset wake-up word. This voice data containing the preset wake-up word can be used as the target voice data. Alternatively, subsequent voice data containing the preset wake-up word can be used as the target voice data. It can be understood that the master device performs wake-up word detection on the acquired voice data; if the detection result indicates that the voice data contains a preset wake-up word, then that voice data can be used as the target voice data. That is, after acquiring voice data, if the voice data contains a preset wake-up word, it is determined to be the target voice data. Simultaneously, this voice data can also be used as reference voice data, and subsequently acquired voice data will be the target voice data. In other words, the master device is woken up by the reference voice data, and the voice data acquired after the master device is woken up is the target voice data.
[0022] A wake-up word can be set for a set of devices, with each device having at least one corresponding candidate wake-up word. A preset wake-up word can be any one of the at least one candidate wake-up word. A wake-up word can also be set for a specific device within the device set. For each device in the device set, there is at least one corresponding candidate wake-up word. A candidate wake-up word set is obtained based on the at least one candidate wake-up word corresponding to each device in the device set. The target wake-up word set can be constructed by taking all or a portion of the wake-up words from the candidate wake-up word set. A preset wake-up word can be any one of the wake-up words in the target wake-up word set.
[0023] S102: Send the target voice data to the first device so that the first device sends a target control command to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The target device is any one of the devices in a preset network device set. The device set includes the master device and multiple candidate slave devices. In this embodiment, the master device sends target voice data to the first device. The master device is primarily responsible for transmitting the target voice data. The first device is primarily responsible for identifying the target device that needs to respond to the target voice data and sending target control commands to it. The first device can be an independent physical server, a server cluster or distributed system consisting of at least two physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server may include a network communication unit, a processor, and a memory. In practical applications, the cloud service can be responsible for identifying the target device that needs to respond to the target voice data and sending target control commands to it.
[0024] For a device set, it indicates a pre-defined network configuration. It should be understood that this involves connecting multiple independent devices directly or indirectly via wired or wireless communication to form an interconnected network system. Multiple independent devices constitute a device set, and the network system is the pre-defined network configuration. Communication within a device set can rely on personal area networks (PANs), local area networks (LANs), metropolitan area networks (MANs), wide area networks (WANs), etc. Application scenarios for device sets can include home scenarios, office scenarios, etc. When all devices in the device set have voice-activated functionality, the master device can be any device in the device set. When only some devices in the device set have voice-activated functionality, the master device can be any device in those groups. With a master device determined, the remaining devices in the device set besides the master device constitute multiple candidate slave devices.
[0025] For each device in the device set, it supports at least one preset control command, meaning that it will execute the task indicated by the preset control command when it receives it. It should be understood that the support for the at least one preset control command is independent for different devices in terms of the total number of commands and the command content. There can be an empty set or an intersection between the at least one preset control command supported by two devices. For example, device x supports control command 1 and control command 2, device y supports control command 3, and device z supports control command 1 and control command 4. A control command set can be constructed based on the at least one preset control command supported by each device in the device set. For each preset control command in the control command set, the corresponding control intent is determined to obtain a control intent set.
[0026] When performing intent matching, the first device determines the voice intent corresponding to the target voice data, and then identifies the control intent that matches the voice intent from the set of control intents. The preset control instruction corresponding to this matched control intent is the target control instruction. For the control intent, it aims to reflect the expectation that the task indicated by the preset control instruction will be executed. For the voice intent, it is the result of intent recognition of the target voice data, focusing on reflecting the expected behavior implied by the target voice data. Matching the control intent and the voice intent can mean that the expected task and the expected behavior match. For example, the expected task is to turn off the lights, and the expected behavior is to turn off the lights. The target control instruction can be a pre-used preset control instruction, or a new control instruction generated based on the preset control instruction and additional information carried by the target voice data. The introduction of additional information makes the new control instruction more specific and targeted than the task indicated by the preset control instruction. For example, if the additional information carried by the target voice data indicates song i, the expected behavior implied by the target voice data is to play music, based on the voice intent of the target voice data. If the control intent matching the voice intent corresponds to the preset control instruction j, the expected task to be executed is to play music, based on the control intent. Therefore, the preset control command j can be updated based on the additional information carried by the target voice data, so that the newly generated target control command indicates the task of playing song i.
[0027] After performing intent matching, the first device determines the target device from the device set. The target device is any device in the device set. The target device can be a master device or a candidate slave device. The target device can be determined based on whether the device supports the target control command. If there are differences among the preset control commands supported by each device, i.e., there is an empty set between at least one preset control command supported by any two devices in the device set, then 1) if the target control command is a used preset control command, the device supporting the target control command can be determined as the target device; 2) if the target control command is a new control command generated based on the preset control command and additional information carried by the target voice data, the device supporting the preset control command corresponding to the target control command can be determined as the target device. If there is at least one subset of devices in the device set, and there is an intersection between at least one preset control command supported by any two devices in the subset, then 1) if the target control command is a used preset control command and the target control command is an intersection command, the target device can be determined from the target device subset according to preset rules. The preset rules can indicate that the device receiving the most control commands within a preset time period is the target device, or the device whose control command was sent closest to the current time is the target device. The target device subset is any subset of devices within at least one device subset. Intersection commands are preset control commands supported by all devices in the target device subset. 2) If the target control command is a new control command generated based on preset control commands and additional information carried by the target voice data, and the preset control command corresponding to the target control command is an intersection command, the target device can be determined from the target device subset according to the preset rules.
[0028] In practical applications, considering the correspondence between devices of the same type and specific control commands, such as devices supporting specific control command x being of type y, the first device can store the specific control commands supported by each device in the device set according to the device type dimension. After intent matching, the corresponding device type can be determined based on the matched specific control command, and then a device matching that device type can be selected as the target device. If there is more than one device matching that device type, one can be randomly selected as the target device, or the device with the most control commands sent within a preset time period can be selected as the target device, or the device whose control command was sent closest to the current time can be selected as the target device. Of course, the first device can also store the default control commands corresponding to each device type according to the device type dimension. After intent matching, the corresponding device type can be determined based on the matched default control command, and then a device matching that device type can be selected as the target device. However, there may be cases where the device set does not contain devices matching that device type. In this case, the first device can send a notification to the master device indicating that the device does not support it. The master device can output third-type feedback information based on this notification and the third rule. For details on the third rule and the third-type feedback information, please refer to the relevant description in step S103 below, which will not be repeated here.
[0029] As a possible implementation, after sending the target voice data to the first device, the method may further include the following steps: First, if the target device is the master device, receive a target control command sent by the first device; then, execute the task indicated by the target control command. The target device can be either a candidate slave device or a master device. In this case, the master device receives the target control command sent by the first device and then executes the task indicated by the target control command. Accordingly, the master device uses the task execution result or task execution progress as the task execution status, and then outputs a first type of feedback information based on the task execution status and a first rule. For details on the first rule and the first type of feedback information, please refer to the relevant description in step S103 described later; it will not be repeated here.
[0030] As a possible implementation, after sending the target voice data to the first device, the method may further include the following steps: First, receiving matching failure information sent by the first device, the matching failure information indicating that the control intentions corresponding to the preset control commands supported by each device in the device set do not match the voice intention; then, outputting a second type of feedback information based on the matching failure information and a second rule. This takes into account the occurrence of intention matching failures, and the output of the second type of feedback information allows the user to be aware of this situation from the master device side, so that the user can adjust the input in a timely manner.
[0031] The second rule can specify the form in which the user is informed of the matching failure. Candidate forms may include at least one of the following: visual presentation or audio presentation. Visual presentation can involve displaying visual elements indicating the matching failure on the screen. For example, displaying text, images, or videos indicating the matching failure, such as displaying default text. Visual presentation can also involve using a light source to display a specific color or flash according to a specific pattern to indicate the matching failure. Audio presentation can involve using a speaker to play the audio indicating the matching failure, such as playing the default audio. The main device can output the second type of feedback information based on its supported formats. For example, if the main device lacks a screen and speaker, it can use a light source to display a specific color or flash according to a specific pattern to output the second type of feedback information. The main device can also output the second type of feedback information based on user preferences. For example, if the user prefers an audio presentation, the main device can use a speaker to play the audio indicating the matching failure to output the second type of feedback information.
[0032] S103: If the target device is a target slave device, receive the task execution status sent by the first device. The task execution status indicates that the target slave device is executing the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
[0033] In this embodiment, when the target device is a target slave device, the master device receives the task execution status sent by the first device. Since the target device is a target slave device, the task indicated by the target control command is executed by the target slave device. The target slave device can send the task execution result or task execution progress to the first device. The first device can generate task execution status based on the received task execution result or task execution progress (e.g., using the task execution result or task execution progress as task execution status), and then send it to the master device.
[0034] As a possible implementation, after receiving the task execution status sent by the first device, the method may further include the following steps: outputting a first type of feedback information based on the task execution status and a first rule.
[0035] The first rule can specify the form in which the user is informed of the task's progress. Candidate forms may include at least one of the following: visual presentation or audio presentation. Visual presentation can involve displaying visual elements indicating task progress on a screen. For example, this could be text, images, or videos indicating task progress. Visual presentation can also involve using a light source to display a specific color or flash according to a specific pattern to indicate task progress. Audio presentation can involve playing voice messages indicating task progress through a speaker. The main device can output the first type of feedback information based on its supported formats. For example, if the main device lacks a screen and speaker, it can use a light source to display a specific color or flash according to a specific pattern to output the first type of feedback information. The main device can also output the first type of feedback information based on user preferences. For example, if the user prefers an audio presentation, the main device can use a speaker to play voice messages indicating task progress to output the first type of feedback information.
[0036] On the one hand, the output of the first type of feedback information facilitates the user's understanding of the task execution status from the master device. The master device is both the device that acquires the target voice data and the device that outputs the response status to the target voice data. This allows the same device to receive user input and output feedback, guiding the user to achieve orderly voice interaction through the master device in scenarios involving multiple smart devices. On the other hand, considering the potential differences in distance between the target slave device and the master device and the user, and since the master device receiving user input is generally closer to the user, having the closer master device output the first type of feedback information makes it easier for the user to know the task execution status of the farther target slave device. Taking a home scenario as an example, if the master device is a refrigerator in the dining room and the target device is an air conditioner in the study, and the target control command indicates that the task is to activate the air conditioner in cooling mode, the air conditioner can send the task execution result indicating successful execution to the first device. The first device then sends the received task execution result as the task execution status to the master device. As can be seen from the technical solutions provided in the embodiments of this application above, the embodiments of this application provide a voice interaction solution that can effectively adapt to scenarios containing multiple smart devices. The embodiments of this application configure a master device for a pre-defined network of devices and introduce a first device to provide decision support. After the master device is awakened by voice, the master device sends target voice data to the first device. The first device then determines the target device from the device set through intent matching and sends a target control command to the target device, enabling the target device to provide an effective response to the target voice data. The pre-configuration of the master device ensures that the device that can be awakened by voice is unique, which is beneficial to the accuracy of voice wake-up. Compared with related technologies, it can also reduce the reliance on users to remember multiple wake words to a certain extent, thereby improving the user experience. In the cooperation between the first device and the master device, the master device is mainly responsible for transmitting the target voice data, while the first device is mainly responsible for determining the target device that needs to respond to the target voice data and sending the target control command to it. This arrangement can provide support for users to configure the master device according to their own needs.
[0037] Figure 3 This diagram illustrates a flowchart of a method for determining cooking parameters according to an embodiment of this application. Figure 3 As shown, the method includes: S301: Receive target voice data sent by the master device, wherein the target voice data includes a preset wake-up word or the target voice data is subsequent voice data of the reference voice data including the preset wake-up word; S302: Based on the target voice data, determine the target device from the device set indicating the preset network, and send a target control command to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The device set includes the master device and multiple candidate slave devices. S303: If the target device is a target slave device, send the task execution status to the master device. The task execution status indicates the status of the target slave device executing the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
[0038] The aforementioned steps S101-S103 provide a voice interaction method with the master device as the execution subject. Correspondingly, steps S301-S303 here provide a voice interaction method with the first device as the execution subject. After the master device is awakened by voice, the master device sends target voice data to the first device. Then, the first device determines the target device from the device set through intent matching and sends a target control command to the target device, so that the target device provides an effective response to the target voice data.
[0039] As one possible implementation, such as Figure 4 As shown, the method further includes: S401: Send a device configuration request to the second device. The device configuration request includes the device identifier and device association information of each device in the device set. The device association information includes one of the following: device type, device status, device location, and device historical usage. S402: Receive device configuration results sent by the second device, wherein the device configuration results include the device identifier of the device specified in the device set, and the device configuration results are generated by the second device based on the selection instruction provided by the received target object; S403: If the specified device is not the master device, perform a master device update on the device set.
[0040] This provides a method for setting the master device via a second device, which helps guide users to configure the master device according to their own needs. By sending device configuration requests to the second device, the device identifiers and device association information of each device in the device set are transmitted, allowing users to better understand the devices in the device set and ensuring that the set master device accurately matches the user's needs. Simultaneously, updating the master device in cases where the specified device is not the current master device ensures that the specified device, as the new master device, can be activated by voice and adapt to the transmission of target voice data.
[0041] The second device can be any device in the device set. The second device can also be a device outside the device set. The second device can be a mobile phone, computer (such as a desktop computer, tablet computer, or laptop computer), augmented reality (AR) / virtual reality (VR) device, digital assistant, smart voice interaction device (such as a smart speaker), smart wearable device, smart home appliance, in-vehicle terminal, etc.
[0042] Device type can be determined based on classification rules. Taking a home scenario as an example, the devices in the device set are often household appliances, and selectable device types could include refrigerators, steam ovens, range hoods, washing machines, air conditioners, microwave ovens, etc. Taking an office scenario as an example, the devices in the device set are often office equipment, and selectable device types could include printers, projectors, paper shredders, etc. Device status can characterize the device's data processing capabilities and can include at least one of the following: CPU information, load status, and storage performance. Device location can refer to the device's location. Taking a home scenario as an example, selectable device locations could include the kitchen, dining room, bedroom, study, etc. Device historical usage can refer to the number of times the device was used within a preset time period, which can be represented by the number of control commands sent.
[0043] The target object can refer to the target account or the user using the target account. The second device can display the device identifiers and associated information of each device in the device set through an interactive interface. Users can select a specific device by triggering controls provided in the interactive interface, thereby generating a selection command. Control triggering can be manual operation or specific voice commands. If the specified device is the current master device, there is no need to update the master device in the device set. If the specified device is not the current master device, a master device update is required in the device set to make the specified device the unique and latest master device.
[0044] The first device can send a device configuration request to the second device after learning that a new device has been added to the preset network. In practical applications, the user can send the device identifier and device association information of the new device to the first device through the second device. If a device set indicating the preset network exists, the device set can be updated by adding the new device to the preset network. If a device set indicating the preset network does not exist, the new device and multiple existing devices can be directly or indirectly connected via wired or wireless communication to form an interconnected network system, thereby realizing the creation of a new device set indicating the preset network. If the new device does not have the function of being voice-activated, the first device will not initiate the master device setup process. If the new device has the function of being voice-activated, the first device will initiate the master device setup process. The initiated master device setup process includes: 1) Determining whether a master device exists in the updated device set or the newly created device set. 2) If it exists, sending a first reminder notification to the second device whether to switch the master device; if it does not exist, sending a second reminder notification to the second device whether to set the new device as the master device. 3) If the second device responds to the first reminder notification or the second reminder notification and sets the new master device, it sends the device identifier of the new master device to the first device. 4) The first device generates voice wake-up guidance information based on the device identifier of the new master device and sends it to each device in the updated device set. The voice wake-up guidance information will be described later and will not be repeated here. The second device can be a new device, any one of the multiple original devices, or a device other than a new device and the multiple original devices.
[0045] Furthermore, the master device update for the device set may include the following steps: sending voice wake-up guidance information to each device in the device set. This voice wake-up guidance information instructs devices matching the device identifier of the specified device to enable the voice wake-up function, and devices not matching the device identifier of the specified device to disable the voice wake-up function. Guided by the voice wake-up guidance information, the specified device with the matching identifier ensures that the voice wake-up function is enabled, while non-specified devices with mismatched identifiers ensure that the voice wake-up function is disabled. This ensures that the specified device becomes the unique and up-to-date master device. This improves the accuracy of the master device update and guarantees the uniqueness of devices that can be woken up by voice.
[0046] As a possible implementation, the master device and the first device communicate based on a preset liveness detection mechanism. The method may further include the following steps: if the duration between the current time and the time of the last report received from the master device is greater than a preset duration, an associated device is determined to be the new master device, and the associated device is the recipient of the control command most recently issued by the first device. The communication based on the preset liveness detection mechanism may be heartbeat communication. If the duration between the current time and the time of the last report received from the master device is greater than the preset duration, it indicates that the communication between the first device and the master device is interrupted and the master device is offline. The content reported by the master device may be an agreed-upon heartbeat packet or target voice data. If the master device and the first device have agreed on periodic heartbeat packet reporting, then if a heartbeat packet is not received for more than a preset number of periodic intervals, it can be determined that the communication with the master device is interrupted. For example, if the master device periodically reports heartbeat packets to the first device at 30-second intervals, and the first device does not receive a heartbeat packet for more than 3 periodic intervals, it is determined that the communication with the master device is interrupted and the master device is offline. In this scenario, the master device may malfunction, such as no longer supporting voice wake-up, or even if it does, failing to effectively transmit target voice data. Simultaneously, the criteria for selecting associated devices ensure that the selected associated devices are functioning correctly. This allows the associated device to promptly replace the current master device as the sole and most up-to-date master, effectively guaranteeing the voice interaction validity in scenarios involving multiple smart devices. Of course, each candidate slave device in the device set can also communicate with the first device based on a preset liveness detection mechanism. In this case, for the selected associated device, the time between the current time and the time previously received from the associated device is less than or equal to a preset duration. Furthermore, when some devices in the device set have voice wake-up capabilities, the associated device can be any device within that set. The criteria for selecting associated devices can also indicate that the associated device is the recipient of the task execution results or progress messages received by the first device.
[0047] This application also provides a voice interaction device, such as... Figure 5 As shown, the voice interaction device 50 is configured on the main device, and the voice interaction device 50 includes: The voice data acquisition module 501 is used to acquire target voice data, wherein the target voice data includes a preset wake-up word or the target voice data is subsequent voice data of the reference voice data including the preset wake-up word; The voice data sending module 502 is used to send the target voice data to the first device, so that the first device sends a target control command to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The target device is any one of the devices in a preset network device set. The device set includes the master device and multiple candidate slave devices. The execution result receiving module 503 is used to receive the task execution status sent by the first device when the target device is a target slave device. The task execution status represents the situation in which the target slave device executes the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
[0048] In some embodiments, the device further includes a first type of feedback information output module; The first type of feedback information output module is used to output the first type of feedback information based on the task execution status and the first rule.
[0049] In some embodiments, the apparatus further includes a task execution module; The task execution module is configured to receive a target control command sent by the first device when the target device is the master device, and execute the task indicated by the target control command.
[0050] In some embodiments, the device further includes a second type of feedback information output module; The second type of feedback information output module is used to receive matching failure information sent by the first device, wherein the matching failure information indicates that the control intent corresponding to the preset control command supported by each device in the device set does not match the voice intent; and outputs second type of feedback information based on the matching failure information and the second rule.
[0051] It should be noted that the apparatus and method embodiments described in the device embodiments are based on the same inventive concept.
[0052] This application also provides a voice interaction device, such as... Figure 6 As shown, the voice interaction device 60 is configured in the first device, and the voice interaction device 60 includes: The voice data receiving module 601 is used to receive target voice data sent by the master device, wherein the target voice data includes a preset wake-up word or the target voice data is subsequent voice data of the reference voice data including the preset wake-up word; The control command sending module 602 is used to determine a target device from a set of devices indicating a preset network based on the target voice data, and to send a target control command to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The set of devices includes the master device and multiple candidate slave devices. The execution result sending module 603 is used to send task execution status to the master device when the target device is a target slave device. The task execution status indicates the status of the target slave device in executing the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
[0053] In some embodiments, the apparatus further includes: The device configuration request sending module is used to send a device configuration request to the second device. The device configuration request includes the device identifier and device association information of each device in the device set. The device association information includes one of the following: device type, device status, device location, and device historical usage. The device configuration result receiving module is used to receive the device configuration result sent by the second device. The device configuration result includes the device identifier of the specified device in the device set. The device configuration result is generated by the second device based on the selection instruction provided by the received target object. The master device update module is used to perform master device updates on the device set when the specified device is not the master device.
[0054] In some embodiments, the master device update module is further configured to send voice wake-up guidance information to each device in the device set, wherein the voice wake-up guidance information is used to instruct devices that match the device identifier of the specified device to enable the voice wake-up function and devices that do not match the device identifier of the specified device to disable the voice wake-up function.
[0055] In some embodiments, the master device and the first device communicate based on a preset liveness detection mechanism. The device is further configured to determine that the associated device is a new master device if the duration between the current time and the time when the master device was last received is greater than a preset duration. The associated device is the recipient of the control command most recently issued by the first device.
[0056] It should be noted that the apparatus and method embodiments described in the device embodiments are based on the same inventive concept.
[0057] This application also provides an electronic device, including the aforementioned voice interaction device 50 or 60. The electronic device may also include a memory and a processor, wherein the memory stores a computer program, which is loaded and executed by the processor to implement the aforementioned method.
[0058] When the electronic device includes the aforementioned voice interaction device 50, or supports the voice interaction method provided in steps S101-S103, the electronic device can be a voice interaction device applied to smart home, smart home appliance, or kitchen appliance scenarios, such as a refrigerator, steam oven, or range hood.
[0059] It should be noted that the devices and methods described in the device embodiments are based on the same inventive concept.
[0060] This application also provides a computer-readable storage medium storing a computer program, which is loaded and executed by a processor to implement the above-described method. The computer-readable storage medium may be a non-volatile computer-readable storage medium.
[0061] This application also provides a computer program product, which includes a computer program that is loaded and executed by a processor to implement the above-described method.
[0062] The embodiments described above are merely examples of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A voice interaction method, characterized in that, Applied to a master device, the method includes: Acquire target speech data, wherein the target speech data includes a preset wake word or the target speech data is subsequent speech data of the reference speech data including the preset wake word; The target voice data is sent to the first device so that the first device sends a target control command to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The target device is any device in a preset network device set, which includes the master device and multiple candidate slave devices. When the target device is a target slave device, the task execution status sent by the first device is received. The task execution status indicates that the target slave device is executing the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
2. The method according to claim 1, characterized in that, After receiving the task execution status sent by the first device, the method further includes: Based on the task execution status and the first rule, the first type of feedback information is output.
3. The method according to claim 1, characterized in that, After sending the target voice data to the first device, the method further includes: If the target device is the master device, receive the target control command sent by the first device; Execute the task indicated by the target control command.
4. The method according to claim 1, characterized in that, After sending the target voice data to the first device, the method further includes: The system receives a matching failure message sent by the first device, wherein the matching failure message indicates that the control intent corresponding to the preset control commands supported by each device in the device set does not match the voice intent. Based on the matching failure information and the second rule, a second type of feedback information is output.
5. A voice interaction method, characterized in that, Applied to a first device, the method includes: Receive target voice data sent by the master device, wherein the target voice data includes a preset wake-up word or the target voice data is subsequent voice data of the reference voice data including the preset wake-up word; Based on the target voice data, a target device is determined from the device set indicating a preset network, and a target control command is sent to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The device set includes the master device and multiple candidate slave devices. If the target device is a target slave device, the task execution status is sent to the master device. The task execution status indicates that the target slave device is executing the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
6. The method according to claim 5, characterized in that, The method further includes: Send a device configuration request to the second device. The device configuration request includes the device identifier and device association information of each device in the device set. The device association information includes one of the following: device type, device status, device location, and device historical usage. The device configuration result sent by the second device is received. The device configuration result includes the device identifier of the device specified in the device set. The device configuration result is generated by the second device based on the selection instruction provided by the received target object. If the specified device is not the master device, the device set is updated to a master device.
7. The method according to claim 6, characterized in that, The process of performing a master device update on the device set includes: Voice wake-up guidance information is sent to each device in the device set. The voice wake-up guidance information is used to instruct devices that match the device identifier of the specified device to enable the voice wake-up function, and devices that do not match the device identifier of the specified device to disable the voice wake-up function.
8. The method according to any one of claims 5-7, characterized in that, The master device and the first device communicate based on a preset liveness detection mechanism, and the method further includes: If the time interval between the current time and the time of the last time the master device reported is greater than a preset time interval, the associated device is determined to be the new master device, and the associated device is the recipient of the control command most recently issued by the first device.
9. A voice interaction device, characterized in that, Configured in the main device, the device includes: A voice data acquisition module is used to acquire target voice data, wherein the target voice data includes a preset wake-up word or the target voice data is subsequent voice data of a reference voice data including the preset wake-up word; A voice data transmission module is used to send the target voice data to a first device, so that the first device sends a target control command to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The target device is any one of the devices in a preset network device set. The device set includes the master device and multiple candidate slave devices. The execution result receiving module is used to receive the task execution status sent by the first device when the target device is a target slave device. The task execution status represents the situation in which the target slave device executes the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.
10. A voice interaction device, characterized in that, Configured in a first device, the means comprising: A voice data receiving module is used to receive target voice data sent by a master device, wherein the target voice data includes a preset wake-up word or the target voice data is subsequent voice data of a reference voice data including the preset wake-up word; The control command sending module is used to determine the target device from the device set indicating the preset network based on the target voice data, and to send the target control command to the target device. The target control command is any one of at least one preset control command supported by the target device. The control intent corresponding to the target control command matches the voice intent corresponding to the target voice data. The device set includes the master device and multiple candidate slave devices. The execution result sending module is used to send task execution status to the master device when the target device is a target slave device. The task execution status represents the situation of the target slave device executing the task indicated by the target control command. The target slave device is any one of the plurality of candidate slave devices.