Interaction method, apparatus and medium

By acquiring voiceprint feature codes and calculating response priorities on the server side, the device response is dynamically set, solving the problem of multiple devices waking up simultaneously. This is applicable to various network environments and improves interaction efficiency and user experience.

CN113870868BActive Publication Date: 2025-12-30BEIJING XIAOMI MOBILE SOFTWARE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010615933.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-06-30
Publication Date
2025-12-30
Estimated Expiration
2040-06-30

AI Technical Summary

Technical Problem

Existing technologies struggle to address the issue of multiple devices waking up simultaneously in multi-device voice interaction, especially in non-LAN environments. Furthermore, existing methods may prevent low-priority devices from interacting properly.

Method used

By acquiring the voiceprint feature code, it is sent to the server to determine the response priority. The server calculates the response index based on factors such as the voiceprint feature code, loudness, device level, and voice interaction frequency, and dynamically sets the device response priority to ensure that only one device responds.

Benefits of technology

It solves the problem of multiple devices waking up simultaneously, is suitable for both LAN and cellular environments, simplifies the interaction steps, shortens the response time, improves the user experience, and avoids situations where devices are never able to respond.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113870868B_ABST
    Figure CN113870868B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an interaction method, device and medium. The method is applied to a device and includes: obtaining a voiceprint feature code based on a monitored sound; sending the voiceprint feature code to a server; receiving a response priority related to the voiceprint feature code from the server; and determining whether to respond to the monitored sound based on the response priority. The method is not limited to the online mode of the device and is suitable for devices that access the Internet through various channels such as a local area network and a cellular network. Moreover, the method simplifies the interaction steps and shortens the response time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of smart device technology, and in particular to interaction methods, devices and media. Background Technology

[0002] Voice interaction falls under the category of human-computer interaction and represents one of the most advanced interaction methods in human-computer interaction development to date. It involves using natural human language to give commands to machines to achieve one's goals. With the development of the smart device industry and the increasing prevalence of smart home devices, voice interaction, as an important interaction method for these devices, is a crucial operating procedure for users.

[0003] When designing voice interaction, manufacturers select one or more wake-up keywords. When there are multiple devices from the same manufacturer in a household, multiple devices are often woken up and respond simultaneously when a voice keyword is used for wake-up, which greatly affects normal use.

[0004] Current solutions to the problem of multiple devices waking up simultaneously rely on service discovery based on service registration information. This allows for timely detection of new devices joining the system, followed by negotiation to determine a master-slave relationship. Decision-making is then based on this self-negotiated master-slave relationship. However, this method has two major problems:

[0005] First, the interaction between multiple devices is limited to a local area network (LAN). In many cases, multiple devices in the same location are not in a LAN but may be connected through a cellular network, so the applicability of this method is relatively narrow.

[0006] Second, this method involves a simple pre-response priority setting and determines master and slave devices, which directly leads to the inability of some smart devices to function. That is, when there are multiple smart voice interaction devices, due to the preset priorities, some low-priority devices will be unable to perform normal voice interaction.

[0007] Therefore, a method is needed that can both avoid the above problems and solve the problem of multiple devices being woken up at the same time. Summary of the Invention

[0008] To overcome the problems existing in related technologies, this disclosure provides an interaction method, apparatus and medium.

[0009] According to a first aspect of the present disclosure, an interaction method is provided, the method being applied to a device, the method comprising:

[0010] Based on the monitored sound, obtain the voiceprint feature code;

[0011] Send the voiceprint feature code to the server;

[0012] Priority of receiving responses related to the voiceprint feature code from the server;

[0013] Based on the response priority, it is determined whether to respond to the detected sound.

[0014] The priority of receiving the response related to the voiceprint feature code from the server includes:

[0015] For the detected sounds, determine the response index;

[0016] The response index is sent to the server so that the server can determine the response priority of the device;

[0017] The response priority is received from the server.

[0018] The step of determining the response index for the monitored sound includes:

[0019] Obtain index correlation information, which includes the loudness of the monitored sound, the distance weight related to the loudness, the device level, the device weight related to the device level, the device voice interaction frequency, and the frequency weight related to the device voice interaction frequency.

[0020] The response index is determined by weighted calculation based on the index correlation information.

[0021] The priority of receiving the response related to the voiceprint feature code from the server further includes:

[0022] The loudness, device level, and device voice interaction frequency of the monitored sound are sent to the server, so that the server can calculate the response index and determine the response priority of the device based on the monitored sound.

[0023] The response priority is received from the server.

[0024] According to a second aspect of the present disclosure, an interaction method is provided, the method being applied to a server, the method comprising:

[0025] Receive the voiceprint signature from the device;

[0026] Determine the response priority of the devices associated with the voiceprint feature code;

[0027] The response priority is sent to the device.

[0028] The step of determining the response priority of the device associated with the voiceprint feature code includes:

[0029] Receive response index from the device;

[0030] Among devices that send the same voiceprint signature, the device with the highest response index is identified as having a high response priority, while devices with other response indices are identified as having a low response priority.

[0031] The step of determining the response priority of the device associated with the voiceprint feature code further includes:

[0032] Receive from the device the loudness of the monitored sound, device level, and device voice interaction frequency;

[0033] Determine the response index of the device based on the detected sound;

[0034] Among devices that send the same voiceprint signature, the device with the highest response index is identified as having a high response priority, while devices with other response indices are identified as having a low response priority.

[0035] The step of calculating the device's response index for the monitored sound includes:

[0036] Obtain the distance weight related to the loudness, the device weight related to the device level, and the frequency weight related to the device's voice interaction frequency;

[0037] The response index is determined by weighting the loudness, distance weight, device level, device weight, device voice interaction frequency, and frequency weight.

[0038] According to a third aspect of the present disclosure, an interactive device is provided, the device being applied to a device, the device comprising:

[0039] The voiceprint acquisition module is configured to obtain voiceprint feature codes based on the monitored sounds.

[0040] The first sending module is configured to send the voiceprint feature code to the server;

[0041] The first receiving module is configured to receive responses related to the voiceprint feature code from the server with a priority.

[0042] The response determination module is configured to determine whether to respond to the detected sound based on the response priority.

[0043] The device further includes:

[0044] The first index determination module is configured to determine the response index for the monitored sound.

[0045] The first sending module is also configured to send the response index to the server so that the server can determine the response priority of the device.

[0046] The first index determination module is further configured as follows:

[0047] Obtain index correlation information, which includes the loudness of the monitored sound, the distance weight related to the loudness, the device level, the device weight related to the device level, the device voice interaction frequency, and the frequency weight related to the device voice interaction frequency.

[0048] The response index is determined by weighted calculation based on the index correlation information.

[0049] The first sending module is further configured to send the loudness, device level, and device voice interaction frequency of the monitored sound to the server, so that the server can calculate the response index and determine the response priority of the device based on the monitored sound.

[0050] According to a fourth aspect of the present disclosure, an interactive device is provided, the device being applied to a server, the device comprising:

[0051] The second receiving module is configured to receive the voiceprint feature code from the device.

[0052] The priority determination module is configured to determine the response priority of devices associated with the voiceprint feature code;

[0053] The second sending module is configured to send the response priority to the device.

[0054] The second receiving module is further configured to receive a response index from the device.

[0055] The priority determination module is further configured to: among devices that send the same voiceprint feature code, determine the device with the highest response index as having a high response priority, and determine the devices with other response indices as having a low response priority.

[0056] The second receiving module is further configured to receive from the device the loudness of the monitored sound, the device level, and the device voice interaction frequency;

[0057] The device also includes a second index determination module, configured to determine the device's response index for the monitored sound;

[0058] The priority determination module is further configured to: among devices that send the same voiceprint feature code, determine the device with the highest response index as having a high response priority, and determine the devices with other response indices as having a low response priority.

[0059] The second index determination module is further configured as follows:

[0060] Obtain the distance weight related to the loudness, the device weight related to the device level, and the frequency weight related to the device's voice interaction frequency;

[0061] The response index is determined by weighting the loudness, distance weight, device level, device weight, device voice interaction frequency, and frequency weight.

[0062] According to a fifth aspect of the present disclosure, an interactive device is provided, comprising:

[0063] processor;

[0064] Memory used to store processor-executable instructions;

[0065] The processor is configured to perform the following steps when executing the executable instructions:

[0066] Based on the monitored sound, obtain the voiceprint feature code;

[0067] Send the voiceprint feature code to the server;

[0068] Priority of receiving responses related to the voiceprint feature code from the server;

[0069] Based on the response priority, it is determined whether to respond to the detected sound.

[0070] According to a sixth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of a device, the device is enabled to perform an interactive method, the method comprising:

[0071] Based on the monitored sound, obtain the voiceprint feature code;

[0072] Send the voiceprint feature code to the server;

[0073] Priority of receiving responses related to the voiceprint feature code from the server;

[0074] Based on the response priority, it is determined whether to respond to the detected sound.

[0075] According to a seventh aspect of the present disclosure, an interactive device is provided, comprising:

[0076] processor;

[0077] Memory used to store processor-executable instructions;

[0078] The processor is configured to perform the following steps when executing the executable instructions:

[0079] Receive the voiceprint signature from the device;

[0080] Determine the response priority of the devices associated with the voiceprint feature code;

[0081] The response priority is sent to the device.

[0082] According to an eighth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein when instructions in the storage medium are executed by a processor of a device, the device is enabled to perform an interactive method, the method comprising:

[0083] Receive the voiceprint signature from the device;

[0084] Determine the response priority of the devices associated with the voiceprint feature code;

[0085] The response priority is sent to the device.

[0086] This disclosure proposes an interaction method in which a device obtains a voiceprint feature code based on the monitored sound and sends the voiceprint feature code to a server; after receiving the voiceprint feature code from the device, the server determines the response priority of the device associated with the voiceprint feature code and sends the response priority to the device; the device receives the response priority associated with the voiceprint feature code from the server and determines whether to respond to the monitored sound based on the response priority.

[0087] This method does not restrict the device's internet access method and is applicable to devices accessing the internet through various channels such as LAN and cellular networks. Furthermore, it simplifies the interaction process and shortens response time by identifying devices located in close proximity using voiceprint signatures. The response priority is set dynamically, preventing situations where only one device can respond while others remain unresponsive. This method effectively identifies the device the user truly wants to interact with, improving the user experience.

[0088] In addition, this method can effectively solve the problem of strangers waking up their devices in public places by identifying users through voiceprint feature codes.

[0089] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0090] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0091] Figure 1 This is a flowchart illustrating an interaction method according to an exemplary embodiment.

[0092] Figure 2 This is a flowchart illustrating an interaction method according to an exemplary embodiment.

[0093] Figure 3 This is a flowchart illustrating an interaction method according to an exemplary embodiment.

[0094] Figure 4 This is a block diagram illustrating an interactive device according to an exemplary embodiment.

[0095] Figure 5 This is a block diagram illustrating an interactive device according to an exemplary embodiment.

[0096] Figure 6 This is a block diagram illustrating an apparatus according to an exemplary embodiment.

[0097] Figure 7 This is a block diagram illustrating an apparatus according to an exemplary embodiment. Detailed Implementation

[0098] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention as detailed in the appended claims.

[0099] When multiple voice-interactive smart devices are present in the same environment, they may be simultaneously woken up upon receiving a wake-up keyword, thus affecting user experience. Current solutions for simultaneous wake-up of multiple devices are limited to devices within a local area network (LAN). However, in many cases, multiple devices in a similar location are not necessarily on a LAN; they may also be connected via cellular networks, thus limiting the applicability of this solution. Furthermore, current solutions, which require master-slave priority settings before response, may prevent some lower-priority devices from performing normal voice interactions.

[0100] This disclosure proposes an interaction method in which a device obtains a voiceprint feature code based on the detected sound and sends the voiceprint feature code to a server. After receiving the voiceprint feature code from the device, the server determines the response priority of the device associated with the voiceprint feature code and sends the response priority to the device. The device receives the response priority associated with the voiceprint feature code from the server and determines whether to respond to the detected sound based on the response priority. This method effectively solves the problem of multiple devices responding to the same wake word simultaneously and can be applied to devices accessed via cellular networks and local area networks, while avoiding situations where response is impossible due to priority settings.

[0101] The devices disclosed herein are devices that can interact with people via voice, such as mobile phones, tablets, wearable devices, etc.

[0102] Figure 1 This is a flowchart illustrating an interaction method according to an exemplary embodiment, the method being applied to a device, such as... Figure 1 As shown, the method includes the following steps:

[0103] Step 101: Obtain the voiceprint feature code based on the monitored sound;

[0104] Step 102: Send the voiceprint feature code to the server;

[0105] Step 103: Receive a response priority related to the voiceprint feature code from the server;

[0106] Step 104: Based on the response priority, determine whether to respond to the detected sound.

[0107] In step 101, when the device detects a user's voice, such as keywords, it extracts a voiceprint feature code from the detected sound. Each person's voiceprint feature code is unique. The method for extracting the voiceprint feature code from the sound can be implemented using techniques known to those skilled in the art, and will not be elaborated upon here.

[0108] In step 102, after extracting the voiceprint feature code, the device sends the voiceprint feature code to the server. The server needs to obtain the response index of devices that send the same voiceprint feature code and determine the response priority based on the response index of each device. Therefore, each device needs to send its extracted voiceprint feature code to the server. It should be noted that devices that send the same voiceprint feature code are those located in close proximity and that are listening to the same user's voice.

[0109] In step 103, after determining the response priority associated with the voiceprint feature code sent by the device, the server sends the response priority to the corresponding device. For example, if the server determines that there are three devices sending the same voiceprint feature code based on the voiceprint feature codes received simultaneously, the server will determine the response priority for each of these three devices and send the determined response priority to each device respectively.

[0110] In step 104, the device determines whether to respond to the detected sound based on the received response priority. For example, if the device receives a high response priority (represented by "1"), it determines to respond to the detected sound; if the device receives a low response priority (represented by "0"), it determines not to respond to the detected sound. To prevent multiple devices from responding simultaneously, the server sets the response priority of only one device to high and the response priority of the other devices to low.

[0111] The above method does not restrict the device's internet access method and is applicable to devices accessing the internet through various channels such as LAN and cellular networks. Furthermore, identifying devices located in close proximity using voiceprint signatures simplifies the interaction process and shortens response time. Additionally, the response priority is set dynamically, preventing situations where only one device can respond while some devices never respond.

[0112] In an optional implementation, the priority of receiving the response related to the voiceprint feature code from the server includes:

[0113] For the detected sounds, determine the response index;

[0114] The response index is sent to the server so that the server can determine the response priority of the device;

[0115] The response priority is received from the server.

[0116] As described above, the server determines the response priority of each device based on its response index. The response index indicates the likelihood that a device will respond to the detected sound. A higher response index indicates a greater likelihood of response; a lower response index indicates a lower likelihood of response. The response index can be determined by the device itself or by the server. In this implementation, the device determines its own response index and sends the determined response index to the server.

[0117] The following details how the device determines its response index.

[0118] In an optional implementation, determining the response index for the monitored sound includes:

[0119] Obtain index correlation information, which includes the loudness of the monitored sound, the distance weight related to the loudness, the device level, the device weight related to the device level, the device voice interaction frequency, and the frequency weight related to the device voice interaction frequency.

[0120] The response index is determined by weighted calculation based on the index correlation information.

[0121] In this implementation, a weighted calculation is performed based on the index correlation information, and the result is the response index. The correlation information for each index is explained below:

[0122] Loudness of the sound: Determined by the loudness of the sound being monitored, with a maximum of 10 and a minimum of 1. The distance between the device and the speaker can also be determined from the loudness of the sound, and this distance value is used when calculating the response index.

[0123] Distance weight: typically set to 50%.

[0124] Device priority: Set based on device type. For example, mobile phones have a priority of 9, TVs have a priority of 8, and speakers have a priority of 7.

[0125] Device weight: typically set to 25%.

[0126] Device voice interaction frequency: This is obtained by accumulating the frequency of user interactions with the device. The higher the interaction frequency, the greater the probability of further interaction. Interaction frequency can be determined based on the number of interactions per week, or per month, etc.

[0127] Frequency weighting; typically set to 25%.

[0128] Device level, distance weight, device weight, and frequency weight are all preset and can be set according to specific application scenarios. The specific values ​​given above are just examples.

[0129] In an optional implementation, the step of receiving the response priority related to the voiceprint feature code from the server further includes:

[0130] The loudness, device level, and device voice interaction frequency of the monitored sound are sent to the server, so that the server can calculate the response index and determine the response priority of the device based on the monitored sound.

[0131] The response priority is received from the server.

[0132] In this implementation, the server determines the device's response index. As mentioned above, determining the response index requires obtaining index-related information. Therefore, the device needs to send the loudness of the detected sound, the device level, and the device's voice interaction frequency to the server. The server then performs a weighted calculation by combining its own stored (or accessed via the device or other means) distance weight, device weight, and frequency weight. The specific method for this weighted calculation is as described above and will not be repeated here.

[0133] Figure 2 This is an example of an interaction method, applied to a server, such as... Figure 2 As shown, the method includes:

[0134] Step 201: Receive the voiceprint feature code from the device;

[0135] Step 202: Determine the response priority of the device associated with the voiceprint feature code;

[0136] Step 203: Send the response priority to the device.

[0137] As described above, the server receives voiceprint signatures extracted from the devices based on the monitored sound. The server receives these voiceprint signatures because it needs to identify devices located nearby that are monitoring the same user's voice. Then, it determines a response priority for devices with the same voiceprint signature and sends this priority to the appropriate device.

[0138] This method does not restrict the device's internet access method and is applicable to devices accessing the internet through various channels such as LAN and cellular networks. Furthermore, it simplifies the interaction process and shortens response time by identifying devices located in close proximity using voiceprint signatures. Additionally, the response priority is set dynamically, preventing situations where only one device can respond while some devices never respond.

[0139] In an optional implementation, determining the response priority of the device associated with the voiceprint feature code includes:

[0140] Receive response index from the device;

[0141] Among devices that send the same voiceprint signature, the device with the highest response index is identified as having a high response priority, while devices with other response indices are identified as having a low response priority.

[0142] The server determines the response priority of each device based on its response index. The response index can be determined by the device itself or by the server. In this implementation, the device determines its response index, and the server receives the response index from the device.

[0143] When determining response priorities, the device with the highest response index is assigned a high priority, while devices with other response indices are assigned a low priority. Here, only one device is assigned a high priority so that only one device responds to the monitored sound, thus avoiding simultaneous responses from multiple devices. A high response priority can be marked as "1", and a low response priority as "0", or other marking methods can be used.

[0144] In an optional implementation, determining the response priority of the device associated with the voiceprint feature code further includes:

[0145] Receive from the device the loudness of the monitored sound, device level, and device voice interaction frequency;

[0146] Determine the response index of the device based on the detected sound;

[0147] Among devices that send the same voiceprint signature, the device with the highest response index is identified as having a high response priority, while devices with other response indices are identified as having a low response priority.

[0148] In this implementation, the server determines the device's response index. The loudness of the monitored sound, the device class, and the frequency of device voice interactions required to calculate the response index are obtained from the device. See the description of the implementation below for details.

[0149] In an optional implementation, calculating the device's response index for the monitored sound includes:

[0150] Obtain the distance weight related to the loudness, the device weight related to the device level, and the frequency weight related to the device's voice interaction frequency;

[0151] The response index is determined by weighting the loudness, distance weight, device level, device weight, device voice interaction frequency, and frequency weight.

[0152] The distance weights, device weights, and frequency weights required to calculate the response index can be stored internally or retrieved from the device. The specific method for weighted calculation is as described above and will not be repeated here.

[0153] The following describes a specific embodiment based on this disclosure, using a specific application scenario as an example. In this embodiment, four devices are placed in the same room A: a television, a speaker, mobile phone 1, and mobile phone 2.

[0154] Step 301: The user shouts out the keyword to wake up the devices in room A, and the four devices listen to the sound.

[0155] Step 302: The four devices extract voiceprint feature codes based on the monitored sounds.

[0156] Step 303: The four devices respectively send the extracted voiceprint feature code, sound loudness, device level, and device voice interaction frequency to the server.

[0157] Step 304: The server selects devices with the same voiceprint feature code from the voiceprint feature codes received simultaneously, namely TV, speaker, mobile phone 1, and mobile phone 2.

[0158] Step 305: The server obtains a distance weight of 50%, a device weight of 25%, and a frequency weight of 25%.

[0159] Step 306: The server calculates the response index of each device based on the above information, as shown in Table 1.

[0160] Step 307: The server sets the response priority of the device with the highest response index (phone 1) to 1, and sets the response priority of other devices to 0.

[0161] Step 308: The server sends its response priority to each device.

[0162] In step 309, mobile phone 1 responds, while other devices do not.

[0163] Table 1

[0164] Sound loudness Equipment level Voice interaction frequency Response Index Response Priority television 6 8 4 6 0 Speakers 6 7 4 5.75 0 Mobile Phone 1 8 9 7 8 1 Mobile Phone 2 8 9 1 6.5 0

[0165] It should be noted that after this process is completed, each device can clear its response priority; or it can leave the current response priority unchanged and let the new response priority received next time overwrite the current response priority.

[0166] This disclosure also provides an interactive device, which is applied to a device, such as... Figure 4 As shown, the device includes:

[0167] The voiceprint acquisition module 401 is configured to acquire voiceprint feature codes based on the monitored sound.

[0168] The first sending module 402 is configured to send the voiceprint feature code to the server;

[0169] The first receiving module 403 is configured to receive responses related to the voiceprint feature code from the server with a priority.

[0170] The response determination module 404 is configured to determine whether to respond to the detected sound based on the response priority.

[0171] In an optional embodiment, the apparatus further includes:

[0172] The first index determination module is configured to determine the response index for the monitored sound.

[0173] The first sending module 402 is also configured to send the response index to the server so that the server can determine the response priority of the device.

[0174] In an optional implementation, the first index determination module is further configured as follows:

[0175] Obtain index correlation information, which includes the loudness of the monitored sound, the distance weight related to the loudness, the device level, the device weight related to the device level, the device voice interaction frequency, and the frequency weight related to the device voice interaction frequency.

[0176] The response index is determined by weighted calculation based on the index correlation information.

[0177] In an optional embodiment, the first sending module 402 is further configured to send the loudness, device level, and device voice interaction frequency of the monitored sound to the server, so that the server can calculate the response index for the monitored sound and determine the response priority of the device.

[0178] This disclosure also provides an interactive device, which is applied to a server, such as... Figure 5 As shown, the device includes:

[0179] The second receiving module 501 is configured to receive the voiceprint feature code from the device.

[0180] Priority determination module 502 is configured to determine the response priority of devices associated with the voiceprint feature code;

[0181] The second sending module 503 is configured to send the response priority to the device.

[0182] In an optional embodiment, the second receiving module 501 is further configured to receive a response index from the device;

[0183] The priority determination module 502 is further configured to: among devices that send the same voiceprint feature code, determine that the device with the highest response index has a high response priority, and determine that the devices with other response indices have a low response priority.

[0184] In an optional embodiment, the second receiving module 501 is further configured to receive from the device the loudness of the monitored sound, the device level, and the device voice interaction frequency;

[0185] The device also includes a second index determination module, configured to determine the device's response index for the monitored sound;

[0186] The priority determination module 502 is further configured to: among devices that send the same voiceprint feature code, determine that the device with the highest response index has a high response priority, and determine that the devices with other response indices have a low response priority.

[0187] In an optional implementation, the second index determination module is further configured as follows:

[0188] Obtain the distance weight related to the loudness, the device weight related to the device level, and the frequency weight related to the device's voice interaction frequency;

[0189] The response index is determined by weighting the loudness, distance weight, device level, device weight, device voice interaction frequency, and frequency weight.

[0190] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0191] The method implemented by the aforementioned device does not restrict the device's internet access method and is applicable to devices accessing the internet through various channels such as local area networks and cellular networks. Furthermore, identifying devices located in close proximity using voiceprint signatures simplifies the interaction process and shortens response time. The response priority is set dynamically, preventing situations where only one device can respond while others remain unresponsive. This method can determine the device the user truly wants to interact with, improving the user experience. Additionally, identifying the user through voiceprint signatures effectively solves the problem of strangers waking up their devices in public places.

[0192] Figure 6 This is a block diagram illustrating an interactive device 600 according to an exemplary embodiment.

[0193] Reference Figure 6 The device 600 may include one or more of the following components: a processing component 602, a memory 604, a power component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.

[0194] Processing component 602 typically controls the overall operation of device 600, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 602 may include one or more modules to facilitate interaction between processing component 602 and other components. For example, processing component 602 may include a multimedia module to facilitate interaction between multimedia component 608 and processing component 602.

[0195] Memory 604 is configured to store various types of data to support the operation of device 600. Examples of this data include instructions for any application or method operating on device 600, contact data, phonebook data, messages, pictures, videos, etc. Memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0196] The power supply component 606 provides power to the various components of the device 600. The power supply component 606 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device 600.

[0197] Multimedia component 608 includes a screen that provides an output interface between the device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 608 includes a front-facing camera and / or a rear-facing camera. When the device 600 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0198] Audio component 610 is configured to output and / or input audio signals. For example, audio component 610 includes a microphone (MIC) configured to receive external audio signals when device 600 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 604 or transmitted via communication component 616. In some embodiments, audio component 610 also includes a speaker for outputting audio signals.

[0199] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0200] Sensor assembly 614 includes one or more sensors for providing status assessments of various aspects of device 600. For example, sensor assembly 614 may detect the on / off state of device 600, the relative positioning of components such as the display and keypad of device 600, changes in the position of device 600 or a component of device 600, the presence or absence of user contact with device 600, the orientation or acceleration / deceleration of device 600, and temperature changes of device 600. Sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 614 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 614 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0201] Communication component 616 is configured to facilitate wired or wireless communication between device 600 and other devices. Device 600 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 616 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 616 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0202] In an exemplary embodiment, the apparatus 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0203] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, which can be executed by a processor 620 of the device 600 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0204] A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform an interaction method, the method comprising: acquiring a voiceprint feature code based on a detected sound; sending the voiceprint feature code to a server; receiving a response priority associated with the voiceprint feature code from the server; and determining whether to respond to the detected sound based on the response priority.

[0205] Figure 6 This is a block diagram illustrating an interactive device 600 according to an exemplary embodiment. For example, device 600 may be provided as a server. (Refer to...) Figure 6 The device 600 includes a processing component 622, which further includes one or more processors, and memory resources represented by memory 632 for storing instructions executable by the processing component 622, such as application programs. The application programs stored in memory 632 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 622 is configured to execute instructions to perform the methods described above: receiving a voiceprint signature from a device; determining a response priority of the device associated with the voiceprint signature; and sending the response priority to the device.

[0206] Device 700 may also include a power supply component 726 configured to perform power management of device 700, a wired or wireless network interface 750 configured to connect device 700 to a network, and an input / output (I / O) interface 758. Device 700 may operate on an operating system stored in memory 732, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or similar.

[0207] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the invention are indicated by the following claims.

[0208] It should be understood that the present invention is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. An interaction method, the method being applied to a device, characterized in that, The method comprises: Based on the monitored sound, the voiceprint feature code is obtained; The voiceprint feature code is sent to the server; The response priority related to the voiceprint feature code is received from the server; the response priority is determined by the server according to the voiceprint feature codes received from multiple devices; wherein the response priority of the device with the highest response index is high, and the response priority of the device with other response indexes is low; Based on the response priority, it is determined whether to respond to the monitored sound; The response priority related to the voiceprint feature code is received from the server, further comprising: The loudness of the monitored sound, the device level, and the device voice interaction frequency are sent to the server for weighted calculation of the response index and determination of the response priority of the device by the server in combination with the distance weight, device weight, and frequency weight stored in the server for the monitored sound.

2. The method of claim 1, wherein, The response priority related to the voiceprint feature code is received from the server, comprising: The response index is determined for the monitored sound; The response index is sent to the server to determine the response priority of the device; The response priority is received from the server.

3. The method of claim 2, wherein, The response index is determined for the monitored sound, comprising: Obtain index correlation information, the index correlation information includes the loudness of the monitored sound, the distance weight related to the loudness, the device level, the device weight related to the device level, the device voice interaction frequency, and the frequency weight related to the device voice interaction frequency; Based on the index correlation information, the response index is determined by weighted calculation.

4. The method of claim 1, wherein, The response priority related to the voiceprint feature code is received from the server, further comprising: The response priority is received from the server.

5. An interaction method, the method being applied to a server, characterized in that, The method comprises: Receive voiceprint feature codes from multiple devices; Determine the response priority of the device related to the voiceprint feature code; wherein the response priority of the device with the highest response index is high, and the response priority of the device with other response indexes is low; The response priority is sent to the device; Wherein, the determination of the response priority of the device related to the voiceprint feature code further comprises: The loudness of the monitored sound, the device level, and the device voice interaction frequency are received from the device; The response index of the device is determined for the monitored sound; The response index of the device is calculated for the monitored sound, comprising: Obtain the distance weight related to the loudness, the device weight related to the device level, and the frequency weight related to the device voice interaction frequency, wherein the distance weight, the device weight, and the frequency weight are stored in the server; Based on the loudness, the distance weight, the device level, the device weight, the device voice interaction frequency, and the frequency weight, the response index is determined by weighted calculation.

6. The method of claim 5, wherein, The response priority of the device related to the voiceprint feature code is determined, comprising: The response index is received from the device; In the devices sending the same voiceprint feature code, the response priority of the device with the highest response index is determined to be high, and the response priority of the device with other response indexes is determined to be low.

7. The method of claim 5, wherein, The method further includes: In the devices sending the same voiceprint feature code, the response priority of the device with the highest response index is determined to be high, and the response priority of the device with other response indexes is determined to be low.

8. An interactive device, said device being applied to an apparatus, characterized in that The device includes: The voiceprint acquisition module is configured to acquire a voiceprint feature code based on the monitored sound; The first sending module is configured to send the voiceprint feature code to a server; The first receiving module is configured to receive a response priority related to the voiceprint feature code from the server; the response priority is determined by the server according to the voiceprint feature codes received from multiple devices; wherein the response priority of the device with the highest response index is high, and the response priority of the device with other response indexes is low; The response determination module is configured to determine whether to respond to the monitored sound based on the response priority; The first sending module is further configured to send the loudness of the monitored sound, the device level, and the device voice interaction frequency to the server, so that the server performs weighted calculation on the response index and determines the response priority of the device in combination with the distance weight, the device weight, and the frequency weight stored in the server for the monitored sound.

9. The apparatus of claim 8, wherein, The device further includes: The first index determination module is configured to determine the response index for the monitored sound; The first sending module is further configured to send the response index to the server to determine the response priority of the device.

10. The apparatus of claim 9, wherein, The first index determination module is further configured to: Obtain index association information, the index association information including the loudness of the monitored sound, the distance weight related to the loudness, the device level, the device weight related to the device level, the device voice interaction frequency, and the frequency weight related to the device voice interaction frequency; Determine the response index by weighted calculation based on the index association information.

11. An interactive device, the device applied to a server, characterized in that, The device includes: The second receiving module is configured to receive voiceprint feature codes from multiple devices; The priority determination module is configured to determine the response priority of the device related to the voiceprint feature code; wherein the response priority of the device with the highest response index is high, and the response priority of the device with other response indexes is low; The second sending module is configured to send the response priority to the device; The second receiving module is further configured to receive the loudness of the monitored sound, the device level, and the device voice interaction frequency from the device; The device further includes a second index determination module configured to determine the response index of the device for the monitored sound; The second index determination module is further configured to: obtaining a distance weight related to the loudness, a device weight related to the device level, and a frequency weight related to the device voice interaction frequency, wherein the distance weight, the device weight, and the frequency weight are stored in the server; performing a weighted calculation based on the loudness, the distance weight, the device level, the device weight, the device voice interaction frequency, and the frequency weight to determine the response index.

12. The apparatus of claim 11, wherein, The second receiving module is further configured to receive the response index from the device. The priority determination module is further configured to determine that the response priority of the device with the highest response index is high and the response priority of the device with other response indexes is low among the devices sending the same voiceprint feature code.

13. The apparatus of claim 11, wherein, The priority determination module is further configured to determine that the response priority of the device with the highest response index is high and the response priority of the device with other response indexes is low among the devices sending the same voiceprint feature code.

14. An interactive device, characterized by comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the following steps when running the executable instructions: obtaining a voiceprint feature code based on the listened sound; sending the voiceprint feature code to a server; receiving a response priority related to the voiceprint feature code from the server; the response priority is determined by the server according to the voiceprint feature codes received from multiple devices; wherein the response priority of the device with the highest response index is high and the response priority of the device with other response indexes is low; determining whether to respond to the listened sound based on the response priority; The receiving of the response priority related to the voiceprint feature code from the server further comprises: sending the loudness, device level, and device voice interaction frequency of the listened sound to the server for the server to perform a weighted calculation of a response index and determine the response priority of the device in combination with the distance weight, device weight, and frequency weight stored in the server for the listened sound.

15. A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of an apparatus, enabling the apparatus to perform an interaction method, the method comprising: obtaining a voiceprint feature code based on the listened sound; sending the voiceprint feature code to a server; receiving a response priority related to the voiceprint feature code from the server; The response priority is determined by the server according to the voiceprint feature codes received from multiple devices; wherein the response priority of the device with the highest response index is high and the response priority of the device with other response indexes is low; determining whether to respond to the listened sound based on the response priority; The receiving of the response priority related to the voiceprint feature code from the server further comprises: The loudness of the monitored sound, the device level, the device voice interaction frequency are sent to a server to be combined with distance weight, device weight, frequency weight stored in the server by the server to calculate a response index and determine the response priority of the device.

16. An interactive device, characterized by Comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the following steps when running the executable instructions: receiving voiceprint feature codes from a plurality of devices; determining the response priority of the device related to the voiceprint feature code; wherein the response priority of the device with the highest response index is high, and the response priority of the device with other response indexes is low; sending the response priority to the device; wherein the determination of the response priority of the device related to the voiceprint feature code further comprises: receiving the loudness of the monitored sound, the device level, the device voice interaction frequency from the device; determining the response index of the device for the monitored sound; the calculation of the response index of the device for the monitored sound comprises: obtaining distance weight related to the loudness, device weight related to the device level, and frequency weight related to the device voice interaction frequency, wherein the distance weight, the device weight, and the frequency weight are stored in a server; based on the loudness, the distance weight, the device level, the device weight, the device voice interaction frequency, and the frequency weight, performing weighted calculation to determine the response index.

17. A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of an apparatus, enabling the apparatus to perform an interaction method, the method comprising: receiving voiceprint feature codes from a plurality of devices; determining the response priority of the device related to the voiceprint feature code; wherein the response priority of the device with the highest response index is high, and the response priority of the device with other response indexes is low; sending the response priority to the device; wherein the determination of the response priority of the device related to the voiceprint feature code further comprises: receiving the loudness of the monitored sound, the device level, the device voice interaction frequency from the device; determining the response index of the device for the monitored sound; the calculation of the response index of the device for the monitored sound comprises: obtaining distance weight related to the loudness, device weight related to the device level, and frequency weight related to the device voice interaction frequency, wherein the distance weight, the device weight, and the frequency weight are stored in a server; based on the loudness, the distance weight, the device level, the device weight, the device voice interaction frequency, and the frequency weight, performing weighted calculation to determine the response index.

Citation Information

Patent Citations

  • Interaction method, robot, server and storage medium

    CN109841207A

  • Voice control method and device, storage medium and control system

    CN109920416A