Method, device and voice device for distributed voice equipment participating in elections
By self-evaluating candidate audio parameters by distributed voice devices and sending election requests when conditions are met, the problem of large amount of computational volume of voice decision-making devices is solved, and response efficiency and user experience are improved.
Patent Information
- Application Number
- CN202210607131.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2042-05-31
AI Technical Summary
In a distributed home voice control system, voice decision-making equipment needs to process audio parameters sent by multiple voice devices, resulting in large amounts of calculations and low efficiency, affecting response time and user experience.
The distributed voice device self-evaluates whether the candidate audio parameters meet the preset election participation conditions, and only sends candidate audio parameters to the voice decision device when the conditions are met, so as to reduce unnecessary election requests and reduce the calculation pressure of the voice decision device.
Improve the computing efficiency of voice command response, shorten the response time, and enhance the user experience.
Smart Images

Figure CN115019795B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of distributed intelligent voice devices, and in particular to a method, device and voice device for distributed voice devices to participate in elections. Background Art
[0002] With the development of smart technology and the Internet of Things (IoT), distributed home voice control systems have become an integral part of people's lives. These systems include multiple voice devices, each capable of recognizing user voice commands. When a user issues a voice command, multiple devices in the system may recognize it simultaneously. The decision about which device will respond is determined through an election process.
[0003] In the prior art, after each voice device receives a voice command issued by the user, it determines the audio parameters such as the sound intensity and angle of the audio corresponding to the voice command, and sends these audio parameters to the voice decision device. The voice decision device determines which voice device responds to the voice command based on the audio parameters sent by each voice device.
[0004] Since for each voice command, the voice decision device will receive audio parameters sent by each voice device that recognizes the voice command, the voice decision device has a large amount of calculation and low calculation efficiency, resulting in a long response time to the voice command, affecting the user's interactive experience. Summary of the Invention
[0005] The present invention provides a method, apparatus, and voice device for distributed voice devices to participate in elections. The technical solution of the present invention is as follows:
[0006] In a first aspect, a method for distributed voice devices to participate in an election is provided, comprising:
[0007] Get the voice command audio recorded when the user issues the current voice command;
[0008] Obtaining candidate audio parameters of the voice command audio;
[0009] Determining whether the candidate audio parameters meet the preset election participation conditions;
[0010] If the candidate audio parameters meet the preset election participation conditions, the candidate audio parameters will be sent to the voice decision device in the same local area network, and the voice decision device will determine the target voice device that responds to the current voice command based on the candidate audio parameters sent by each voice device.
[0011] Optionally, the candidate audio parameters include at least candidate sound intensity and candidate sound angle;
[0012] The determining whether the candidate audio parameters meet the preset election participation conditions includes:
[0013] According to the preset sound intensity threshold and sound angle threshold range respectively, it is judged whether the candidate sound intensity and the candidate sound angle meet the preset election participation conditions.
[0014] Optionally, before determining whether the candidate sound intensity and the candidate sound angle meet the preset election participation conditions based on the preset sound intensity threshold and sound angle threshold range respectively, the method further includes:
[0015] Obtain a historical voice interaction audio set, where the historical voice interaction audio set includes audio recorded each time the user issues a voice command during the historical voice interaction process;
[0016] Filtering a target audio set from the historical voice interaction audio set, the target audio set including the audio recorded each time the target audio device is selected as the target audio device by the voice decision device during the historical voice interaction process;
[0017] Acquire the sound intensity and sound angle of each audio in the target audio set to form a sound intensity set and a sound angle set;
[0018] A sound intensity threshold and a sound angle threshold range are determined according to the sound intensity set and the sound angle set, respectively.
[0019] Optionally, after determining the sound intensity threshold and the sound angle threshold range according to the sound intensity set and the sound angle set respectively, the method further includes:
[0020] The sound intensity threshold and the sound angle threshold range are updated regularly.
[0021] Optionally, judging whether the candidate sound intensity and the candidate sound angle meet preset election participation conditions based on preset sound intensity thresholds and sound angle threshold ranges respectively includes:
[0022] obtaining a current election participation principle, wherein the election participation principle includes a sound angle priority and a sound intensity priority;
[0023] If the current election participation principle is sound angle priority, determining whether the candidate sound angle is within the sound angle threshold range;
[0024] When the candidate sound angle is within the sound angle threshold range, determining whether the candidate sound intensity exceeds the sound intensity threshold;
[0025] If the candidate voice intensity does not exceed the voice intensity threshold, determining whether the candidate voice intensity exceeds the election participation voice intensity floating threshold;
[0026] If the candidate sound intensity exceeds the election participation sound intensity floating threshold, it is determined that the candidate sound intensity and the candidate sound angle meet the preset election participation conditions.
[0027] Optionally, judging whether the candidate sound intensity and the candidate sound angle meet preset election participation conditions based on preset sound intensity thresholds and sound angle threshold ranges respectively includes:
[0028] obtaining a current election participation principle, wherein the election participation principle includes a sound angle priority and a sound intensity priority;
[0029] If the current election participation principle is voice intensity priority, determining whether the candidate voice intensity exceeds the voice intensity threshold;
[0030] When the candidate sound intensity exceeds the sound intensity threshold, determining whether the candidate sound angle is within the sound angle threshold range;
[0031] If the candidate sound angle is not within the sound angle threshold range, determining whether the candidate sound angle is within the election participation sound angle floating range;
[0032] If the candidate sound angle is within the election participation sound angle floating range, it is determined that the candidate sound intensity and candidate sound angle meet the preset election participation conditions.
[0033] Optionally, determining a sound intensity threshold and a sound angle threshold range according to the sound intensity set and the sound angle set, respectively, includes:
[0034] filtering a source lower limit sound intensity from the sound intensity set, calculating a target lower limit sound intensity for participating in the election based on the source lower limit sound intensity and a preset sound intensity floating range, and using the target lower limit sound intensity as a sound intensity threshold;
[0035] The edge sound angles are filtered from the sound angle set, and the target sound angle range for participating in the election is calculated according to the edge sound angles and the preset sound angle floating range, and the target sound angle range is used as the sound angle threshold range.
[0036] In a second aspect, an apparatus for distributed voice devices to participate in an election is provided, comprising:
[0037] A first acquisition unit is configured to acquire the voice command audio recorded when the user issues the current voice command;
[0038] a second acquiring unit, configured to acquire candidate audio parameters of the voice command audio;
[0039] a judging unit configured to judge whether the candidate audio parameters meet its own preset election participation conditions;
[0040] The sending unit is configured to send the candidate audio parameters to the voice decision device in the same local area network if the candidate audio parameters meet its own preset election participation conditions, and the voice decision device determines the target voice device that responds to the current voice command based on the candidate audio parameters sent by each voice device.
[0041] Optionally, the candidate audio parameters include at least candidate sound intensity and candidate sound angle;
[0042] The judgment unit is configured to judge whether the candidate sound intensity and the candidate sound angle meet its own preset election participation conditions according to the preset sound intensity threshold and sound angle threshold range respectively.
[0043] Optionally, the apparatus for the distributed voice device to participate in the election further includes:
[0044] a third acquiring unit configured to acquire a historical voice interaction audio set, wherein the historical voice interaction audio set includes audio recorded each time the user issues a voice command during the historical voice interaction process;
[0045] a screening unit configured to screen a target audio set from the historical voice interaction audio set, wherein the target audio set includes audio recorded each time the target audio device is selected as a target voice device by the voice decision device during the historical voice interaction process;
[0046] a fourth acquiring unit, configured to acquire the sound intensity and sound angle of each audio in the target audio set to form a sound intensity set and a sound angle set;
[0047] The determining unit is configured to determine a sound intensity threshold and a sound angle threshold range according to the sound intensity set and the sound angle set, respectively.
[0048] Optionally, the apparatus for the distributed voice device to participate in the election further includes:
[0049] An updating unit is configured to periodically update the sound intensity threshold and the sound angle threshold range.
[0050] Optionally, the judging unit includes:
[0051] a first acquisition module configured to acquire current election participation principles, wherein the election participation principles include sound angle priority and sound intensity priority;
[0052] a first judgment module configured to, if the current election participation principle is sound angle priority, determine whether the candidate sound angle is within the sound angle threshold range;
[0053] a second judgment module configured to judge whether the candidate sound intensity exceeds the sound intensity threshold when the candidate sound angle is within the sound angle threshold range;
[0054] a third judgment module configured to judge whether the candidate voice intensity exceeds the election participation voice intensity floating threshold if the candidate voice intensity does not exceed the voice intensity threshold;
[0055] The first determination module is configured to determine whether the candidate sound intensity and the candidate sound angle meet its own preset election participation conditions if the candidate sound intensity exceeds the election participation sound intensity floating threshold.
[0056] Optionally, the judging unit includes:
[0057] a second acquisition module configured to acquire current election participation principles, wherein the election participation principles include sound angle priority and sound intensity priority;
[0058] a fourth judgment module configured to judge whether the candidate voice intensity exceeds the voice intensity threshold if the current election participation principle is voice intensity priority;
[0059] a fifth judgment module, configured to, when the candidate sound intensity exceeds the sound intensity threshold, judge whether the candidate sound angle is within the sound angle threshold range;
[0060] a sixth judgment module configured to judge whether the candidate sound angle is within the election participation sound angle floating range if the candidate sound angle is not within the sound angle threshold range;
[0061] The second determination module is configured to determine whether the candidate sound intensity and candidate sound angle meet its own preset election participation conditions if the candidate sound angle is within the election participation sound angle floating range.
[0062] Optionally, the determining unit includes:
[0063] a first screening and calculation module configured to screen a source lower limit sound intensity from the sound intensity set, calculate a target lower limit sound intensity for participating in the election based on the source lower limit sound intensity and a preset sound intensity floating range, and use the target lower limit sound intensity as a sound intensity threshold;
[0064] The second screening and calculation module is configured to screen edge sound angles from the sound angle set, calculate the target sound angle range participating in the election based on the edge sound angle and the preset sound angle floating range, and use the target sound angle range as the sound angle threshold range.
[0065] In a third aspect, a voice device is provided, comprising: at least one memory and at least one processor;
[0066] The at least one memory is configured to store a machine-readable program;
[0067] The at least one processor is configured to call the machine-readable program to execute the method described in the first aspect.
[0068] In a fourth aspect, a computer-readable medium is provided, on which computer instructions are stored. When the computer instructions are executed by a processor, the processor executes the method described in the first aspect.
[0069] According to the method and device provided in the embodiments of the specification, after obtaining the candidate audio parameters, it is first determined whether the candidate audio parameters meet the preset election participation conditions, and when it is determined that the candidate audio parameters meet the preset election participation conditions, the candidate audio parameters are sent to the voice decision device to initiate an election request, so that when many voice devices in the same local area network recognize the voice command, only the voice devices that meet the preset election participation conditions will participate in the election, and the voice devices that do not meet the preset election participation conditions will not initiate an election request, thereby reducing the computing pressure of the voice decision server, thereby improving computing efficiency, shortening the response time to voice commands, and increasing user stickiness. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0071] Figure 1 This is a schematic diagram of an implementation environment of a method for distributed voice devices to participate in elections provided by an embodiment of the present invention.
[0072] Figure 2 This is a flow chart of a method for distributed voice devices to participate in elections provided by an embodiment of the present invention.
[0073] Figure 3 This is a block diagram of an apparatus for distributed voice devices to participate in elections provided by one embodiment of the present invention.
[0074] Figure 4 It is a structural diagram of a voice device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0075] like Figure 1 As shown, it is a schematic diagram of the implementation environment of the method for distributed voice devices to participate in elections provided by an embodiment of the present invention, and the implementation environment is a distributed home voice control system. The distributed home voice control system includes multiple voice devices with voice recognition functions distributed in different locations in the room, and these voice devices are located in the same local area network. These voice devices can be some large home appliances such as smart refrigerators 101, smart range hoods 102, smart TVs 103, smart air conditioner cabinets 104, smart wall-mounted air conditioners 105, smart washing machines 106, desktop computers 107, etc., and can also be small home appliances such as smart tea bar machines, smart speakers, smart sockets, etc. ( Figure 1 (Not shown) Among these voice devices is a voice decision-making device. When a user issues a voice command, each voice device that recognizes the voice command records the voice command audio and performs a self-evaluation to determine whether it meets its own preset election participation conditions and participates in the election of the voice decision-making device. The specific method for distributed voice devices to participate in the election is detailed below:
[0076] Combine Figure 1 In the implementation environment shown, the embodiment of the present invention provides a method for distributed voice devices to participate in elections. The embodiment of the present invention takes a voice device in a distributed home voice control system that recognizes a voice command issued by a user to execute the method provided by the embodiment of the present invention as an example to explain the method provided by the embodiment of the present invention in detail. Figure 2 As shown, the method may include the following steps:
[0077] Step 201: Obtain the voice command audio recorded when the user issues the current voice command.
[0078] Among them, the current voice command is the command for the user to control the voice device to perform a certain action, for example, the current voice command is "play music", "tell time", "weather forecast" and so on.
[0079] When the user issues a voice command, if the voice device recognizes the voice command (i.e., hears the voice command), it will record the voice command audio through its own microphone or other recording accessories. Based on this, when obtaining the voice command audio, it will obtain it from its own memory.
[0080] Step 203: Acquire candidate audio parameters of the voice command audio.
[0081] The candidate audio parameters are used to represent the characteristics of the voice command audio. The candidate audio parameters can be one or more. In order to comprehensively identify the characteristics of the voice command audio, the candidate audio parameters usually include multiple parameters.
[0082] After recording the voice command audio, candidate audio parameters are obtained by analyzing and calculating the voice command audio. The specific analysis and calculation method is related to the content of the candidate audio parameters.
[0083] Specifically, since sound intensity and sound angle can usually reflect the location where the user issues the voice command, the candidate audio parameters in the embodiment of the present invention include at least candidate sound intensity and candidate sound angle. The candidate sound intensity can reflect the distance between the user and the voice device when the current voice command is issued. The candidate sound angle can reflect the angle between the user and the voice device when the current voice command is issued, for example, facing the voice device, facing away from the voice device, facing the side of the voice device, etc. Furthermore, the candidate audio parameters can also include high-frequency reverberation and low-frequency reverberation, etc., to reflect the extent to which the voice command audio is affected by noise.
[0084] When the candidate audio parameters include candidate sound intensity and candidate sound angle, the voice device analyzes the waveform (sine or cosine wave) of the voice command audio and determines the candidate sound intensity based on the amplitude of the waveform. The candidate sound angle is calculated by analyzing the time difference between different microphones when collecting the voice command audio.
[0085] Step 205: Determine whether the candidate audio parameters meet the preset election participation conditions.
[0086] Among them, the preset election participation conditions can be determined based on factors such as the response rules and response audio parameters to voice commands during the historical voice interaction process of the voice device.
[0087] Step 207: If the candidate audio parameters meet the preset election participation conditions, the candidate audio parameters are sent to the voice decision device in the same local area network. The voice decision device determines the target voice device that responds to the current voice command based on the candidate audio parameters sent by each voice device.
[0088] Among them, the voice decision-making device can be a voice device in the local area network that has certain computing capabilities and is in a constant standby state, for example, a smart TV, a smart refrigerator, etc. can serve as the voice decision-making device.
[0089] Specifically, when the voice decision device determines the target voice device based on the candidate audio parameters sent by each voice device, it can score the candidate audio parameters sent by each voice device and determine the target voice device based on the score of each voice device.
[0090] Furthermore, when multiple candidate audio parameters are included, the voice decision device pre-assigns different weights to the different candidate audio parameters. For any given voice device's candidate audio parameters, the voice decision device first calculates a score for each candidate audio parameter, then multiplies each candidate audio parameter's score by its corresponding weight, and then adds the scores together to obtain the voice device's score. For example, if the audio parameters include sound intensity and sound angle, with weights of 0.6 and 0.4, respectively, and a voice device's candidate sound intensity and candidate sound angle scores are 68 and 72, respectively, then the voice decision device's score for that voice device is 0.6*68+0.4*72=69.6.
[0091] The method provided by the embodiment of the present invention, after obtaining the candidate audio parameters, first determines whether the candidate audio parameters meet the preset election participation conditions, and when it is determined that the candidate audio parameters meet the preset election participation conditions, sends the candidate audio parameters to the voice decision-making device to initiate an election request, so that when many voice devices in the same local area network all recognize the voice command, only the voice devices that meet the preset election participation conditions will participate in the election, and the voice devices that do not meet the preset election participation conditions will not initiate the election request, thereby reducing the computing pressure of the voice decision-making device, and then improving computing efficiency, shortening the response time to voice commands, and increasing user stickiness.
[0092] Described below Figure 2 The specific implementation method of each step shown.
[0093] Specifically, when the candidate audio parameters include at least candidate sound intensity and candidate sound angle, step 205 determines whether the candidate audio parameters meet the preset voting conditions by determining whether the candidate sound intensity and candidate sound angle meet the preset voting conditions. In one embodiment, a sound intensity threshold and a sound angle threshold range can be preset. In this case, whether the candidate sound intensity and candidate sound angle meet the preset voting conditions can be determined based on the sound intensity threshold and the sound angle threshold range, respectively.
[0094] Before determining whether the candidate sound intensity and candidate sound angle meet the preset election participation conditions based on the sound intensity threshold and sound angle threshold range, it is necessary to first determine the sound intensity threshold and sound angle threshold range. In the embodiment of the present invention, determining the sound intensity threshold and sound angle threshold range includes, but is not limited to, the following steps:
[0095] Step A: Obtain a historical voice interaction audio set, which includes audio recorded each time the user issues a voice command during the historical voice interaction process.
[0096] During historical voice interaction, every time a user issues a voice command and the voice device recognizes it, audio is recorded. The audio recorded each time a user issues a voice command during historical voice interaction is combined to form a historical voice interaction audio set. The audio in the historical voice interaction audio set can be recorded from the time the voice device is enabled to the time the sound intensity threshold and sound angle threshold range are determined, or it can be recorded within a selected time period. The goal is to ensure that the historical voice interaction audio set contains a sufficient number of samples to ensure that the sound intensity threshold and sound angle threshold range are determined accurately.
[0097] Step B: Filter a target audio set from the historical voice interaction audio set. The target audio set includes the audio recorded each time the target audio device is selected as the target voice device by the voice decision device during the historical voice interaction process.
[0098] During the historical voice interaction process, although the voice device recognized the voice command, it may not be selected as the target voice device by the voice decision device due to poor recorded audio quality. This part of the audio does not have a reference value when determining the sound intensity threshold and the sound angle threshold range. However, during the historical voice interaction process, the audio recorded when the voice device was selected as the target voice device by the voice decision device is the audio selected by the voice decision device as suitable for responding to the user's voice command. Its audio parameters can meet the requirements of responding to the user's voice command. This part of the audio has a reference value when determining the sound intensity threshold and the sound angle threshold range. Therefore, the embodiment of the present invention selects the target audio set from the historical voice interaction audio set to form sample data for determining the sound intensity threshold and the sound angle threshold range.
[0099] Step C: Acquire the sound intensity and sound angle of each audio in the target audio set to form a sound intensity set and a sound angle set.
[0100] Each audio in the target audio set corresponds to at least two audio parameters, namely, sound intensity and sound angle. The sound intensities corresponding to each audio in the target audio set are combined to obtain a sound intensity set. The sound angles corresponding to each audio in the target audio set are combined to obtain a sound angle set.
[0101] Step D: determining a sound intensity threshold and a sound angle threshold range according to the sound intensity set and the sound angle set respectively.
[0102] There are many ways to determine the sound intensity threshold and the sound angle threshold range based on the sound intensity set and the sound angle set, respectively. For example, the average sound intensity value of each element in the sound intensity set can be calculated and used as the sound intensity threshold; the upper and lower sound angle limits of each element in the sound angle set can be queried and the angle between the lower and upper sound angle limits can be determined as the sound angle threshold range. However, in order to make the determined sound intensity threshold and sound angle threshold range more accurate, the embodiment of the present invention can implement step D through the following preferred embodiment.
[0103] Preferably, step D of the embodiment of the present invention, when determining the sound intensity threshold and the sound angle threshold range according to the sound intensity set and the sound angle set respectively, includes the following steps:
[0104] Step D1: Filter the source lower limit sound intensity from the sound intensity set, calculate the target lower limit sound intensity for participating in the election based on the source lower limit sound intensity and the preset sound intensity floating range, and use the target lower limit sound intensity as the sound intensity threshold.
[0105] Among them, the source lower limit sound intensity is the sound intensity with the weakest sound intensity in the sound intensity set, which represents the lowest sound intensity of the audio recorded when the voice device is used as the target voice device during the historical voice interaction process. Audio below the source lower limit sound intensity is not selected as the target voice device during the historical interaction process. Therefore, the source lower limit sound intensity can also be directly used as the sound intensity threshold. However, in order to ensure that the sound intensity of the recorded audio will not be affected due to some accidental or special reasons during the recording of the audio, thereby causing the voice device to be unable to participate in the election, the embodiment of the present invention also sets a preset sound intensity floating range, so that the voice command audio that is lower than the source lower limit sound intensity but within the preset sound intensity floating range can also participate in the election. The preset sound intensity floating range can be set as needed, for example, set to 10% of the source lower limit sound intensity, or 5 decibels, 10 decibels, etc.
[0106] By setting a preset sound intensity floating range, the target lower limit sound intensity can fluctuate relative to the source lower limit sound intensity to a certain extent, and the limit of the sound intensity threshold is appropriately expanded. This can ensure that in subsequent elections, some uncertain factors can be avoided, which may cause the candidate's voice intensity to not exceed the sound intensity threshold and thus be unable to participate in the election.
[0107] Step D2: Filter edge sound angles from the sound angle set, calculate the target sound angle range for participating in the election based on the edge sound angle and the preset sound angle floating range, and use the target sound angle range as the sound angle threshold range.
[0108] Among them, the edge sound angle is the maximum and minimum sound angles in the sound angle set, representing the angle range of the audio recorded by the voice device as the target voice device during the historical voice interaction process. Audio below the minimum angle and above the maximum angle is not selected as the target voice device. Therefore, the range of the minimum angle and above the maximum angle can also be directly used as the sound angle threshold range. However, in order to ensure that the sound angle of the recorded audio may be affected by some accidental or special reasons during the recording of audio, thereby causing the voice device to be unable to participate in the election, the embodiment of the present invention also sets a preset sound angle floating range, so that voice command audio that is not within the minimum angle and above the maximum angle range but within the preset sound angle floating range can also participate in the election. The preset sound intensity floating range refers to the floating range of the minimum angle and the maximum angle. The two can use the same preset sound intensity floating range, and the preset sound intensity floating range can be set as needed, for example, set to 5°, 10°, etc., or set to 10% of the minimum angle and 5% of the maximum angle.
[0109] By setting a preset sound angle floating range, the determined target sound angle range can fluctuate relative to the edge sound angle, and the limit of the sound angle threshold range is appropriately expanded. This can ensure that in subsequent elections, some uncertain factors (such as calculation errors or microphone failures) can be avoided, causing the candidate's voice angle to be not within the sound angle threshold range and unable to participate in the election.
[0110] Furthermore, when determining whether the candidate voice intensity and candidate voice angle meet the preset election participation conditions based on the preset voice intensity threshold and voice angle threshold range respectively, there are three methods, including but not limited to the following:
[0111] The first method includes the following steps:
[0112] Step 2051: Determine whether the candidate sound intensity exceeds the sound intensity threshold.
[0113] Specifically, the judgment is made by comparing the candidate sound intensity with a sound intensity threshold.
[0114] Step 2053: Determine whether the candidate sound angle is within the sound angle threshold range.
[0115] Step 2055: If the candidate sound intensity exceeds the sound intensity threshold and the candidate sound angle is within the sound angle threshold range, it is determined that the candidate sound intensity and the candidate sound angle meet the preset election participation conditions.
[0116] In this method, any voice command audio whose candidate sound intensity does not exceed the sound intensity threshold, or whose candidate sound angle does not fall within the sound angle threshold, is excluded from the election. This stricter approach can filter out more voice devices from participating in the election, minimizing the number of voice devices participating.
[0117] In addition to the first method, the method provided by the embodiment of the present invention also supports the election participation principle of the user-preset voice device, and the election participation principle includes sound angle priority and sound intensity priority. Sound angle priority means that the sound angle meets the conditions, and the sound intensity is within a certain floating range, and the election can still be participated in. An example of a use scenario for sound angle priority is that the user whispers a voice command to a certain voice device. Sound intensity priority means that the sound intensity meets the conditions, and the sound angle is within a certain floating range, and the election can still be participated in. An example of a use scenario for sound intensity priority is that the user loudly sends a voice command to a certain voice device, but the angle between the location and the voice device is relatively large. On this basis, when judging whether the candidate sound intensity and candidate sound angle meet the preset election participation conditions, there are the following two methods:
[0118] The second method includes the following steps:
[0119] Step 205.1: Obtain the current election participation principle, and if the current election participation principle is sound angle priority, determine whether the candidate sound angle is within the sound angle threshold range.
[0120] Among them, the current election participation principle can be obtained from its own configuration information.
[0121] Step 205.3: When the candidate sound angle is within the sound angle threshold range, determine whether the candidate sound intensity exceeds the sound intensity threshold.
[0122] Specifically, when the sound angle is prioritized, the candidate sound angle is within the sound angle threshold range, and the candidate sound intensity is within the floating range. Therefore, when the candidate sound angle is within the sound angle threshold range, it is necessary to determine the specific situation of the candidate sound intensity, that is, whether the candidate sound intensity exceeds the sound intensity threshold. If the candidate sound intensity exceeds the sound intensity threshold, the first method described above is used.
[0123] Step 205.5: If the candidate's voice intensity does not exceed the voice intensity threshold, determine whether the candidate's voice intensity exceeds the election participation voice intensity floating threshold.
[0124] The voting participation voice intensity floating threshold refers to the sound intensity level obtained by floating the voice intensity threshold downward by a certain value. The downward floating amount of the voice intensity threshold can be set as needed, for example, to 10% of the voice intensity threshold, or to 5 decibels, 10 decibels, etc. The voting participation voice intensity floating threshold is the difference between the voice intensity threshold and the downward floating amount.
[0125] Step 205.7: If the candidate sound intensity exceeds the election participation sound intensity floating threshold, determine whether the candidate sound intensity and candidate sound angle meet the preset election participation conditions.
[0126] This method sets a floating threshold for the sound intensity of election participation, and determines whether the candidate's voice intensity meets the election participation conditions based on the floating threshold. It appropriately expands the sound intensity range for self-evaluation, and can avoid accidents when recording voice command audio that cause the candidate's voice intensity to fail to meet the sound intensity threshold. This method is more reasonable to determine whether the voice device meets its own preset election participation conditions.
[0127] The third method includes the following steps:
[0128] Step 205 - 1 : Obtain the current election participation principle, and if the current election participation principle is voice intensity priority, determine whether the candidate voice intensity exceeds a voice intensity threshold.
[0129] Step 205-3: When the candidate sound intensity exceeds the sound intensity threshold, determine whether the candidate sound angle is within the sound angle threshold range.
[0130] Specifically, when sound intensity is prioritized, it is sufficient for the candidate sound intensity to exceed the sound intensity threshold and the candidate sound angle to be within the floating range. Therefore, when the candidate sound intensity exceeds the sound angle threshold, it is necessary to determine the specific situation of the candidate sound angle intensity, that is, whether the candidate sound angle is within the sound angle threshold range. If the candidate sound angle is within the sound angle threshold range, the first method described above is used.
[0131] Step 205-5: If the candidate sound angle is not within the sound angle threshold range, determine whether the candidate sound angle is within the election participation sound angle floating range.
[0132] The "voice angle fluctuation range" for election participation refers to the sound angle range obtained by floating the minimum angle of the sound angle threshold range downward by a certain value and the maximum angle by a certain angle upward by a certain angle. The upward or downward fluctuation angle can be set as needed, for example, to 5°, 10°, etc., or to 10% of the minimum angle and 5% of the maximum angle.
[0133] Step 205-7: If the candidate sound angle is within the election participation sound angle floating range, determine whether the candidate sound intensity and candidate sound angle meet the preset election participation conditions.
[0134] This method sets a floating range of election participation sound angles and determines whether the candidate voice angles meet the election participation conditions based on the floating range of election participation sound angles. It appropriately expands the sound angle range for self-evaluation and can calculate the situation where the candidate voice angles do not meet the sound angle threshold range due to accidental factors such as angle errors. It is more reasonable to determine whether the voice device meets its own preset election participation conditions in this way.
[0135] Furthermore, since the number of elements in the historical voice interaction audio set will increase with the increase in the number of voice interactions, the number of elements in the target audio set will increase with the increase in the number of voice interactions, and new voice equipment may be added to the local area network, in order to make the determined sound intensity threshold and sound angle threshold range applicable to the current voice interaction needs, the sound intensity threshold and sound angle threshold range can also be updated regularly. Therefore, after determining the sound intensity threshold and sound angle threshold range, the method provided by the embodiment of the present invention also includes regularly updating the sound intensity threshold and sound angle threshold range. The specific time of the regular period can be set as needed, for example, set to update once a week, or once a month, etc. The specific way to update the sound intensity threshold and sound angle threshold range includes but is not limited to the method of adopting the above steps A to D.
[0136] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0137] The embodiment of the present invention provides a device for distributed voice equipment to participate in elections. Figure 3 A schematic block diagram of an apparatus for distributed voice equipment to participate in elections according to an embodiment is shown. The apparatus for distributed voice equipment to participate in elections has a voice recognition function. Figure 3 As shown, the distributed voice device participating in the election includes:
[0138] The first acquisition unit 301 is configured to acquire the voice command audio recorded when the user issues the current voice command;
[0139] The second acquiring unit 303 is configured to acquire candidate audio parameters of the voice command audio;
[0140] The judging unit 305 is configured to judge whether the candidate audio parameters meet the preset election participation conditions;
[0141] The sending unit 307 is configured to send the candidate audio parameters to the voice decision device in the same local area network if the candidate audio parameters meet its own preset election participation conditions. The voice decision device determines the target voice device that responds to the current voice command based on the candidate audio parameters sent by each voice device.
[0142] Optionally, the candidate audio parameters include at least candidate sound intensity and candidate sound angle;
[0143] The judgment unit 305 is configured to judge whether the candidate sound intensity and the candidate sound angle meet its own preset election participation conditions according to the preset sound intensity threshold and sound angle threshold range respectively.
[0144] Optionally, the apparatus for the distributed voice device to participate in the election further includes:
[0145] A third acquiring unit is configured to acquire a historical voice interaction audio set, where the historical voice interaction audio set includes audio recorded each time the user issues a voice command during the historical voice interaction process;
[0146] a screening unit configured to screen a target audio set from a historical voice interaction audio set, where the target audio set includes audio recorded each time the target audio device is selected as a target voice device by the voice decision device during the historical voice interaction process;
[0147] A fourth acquisition unit is configured to acquire the sound intensity and sound angle of each audio in the target audio set to form a sound intensity set and a sound angle set;
[0148] The determining unit is configured to determine a sound intensity threshold and a sound angle threshold range according to the sound intensity set and the sound angle set, respectively.
[0149] Optionally, the apparatus for the distributed voice device to participate in the election further includes:
[0150] An updating unit is configured to periodically update the sound intensity threshold and the sound angle threshold range.
[0151] Optionally, the judging unit 305 includes:
[0152] a first acquisition module configured to acquire current election participation principles, the election participation principles including sound angle priority and sound intensity priority;
[0153] A first judgment module is configured to judge whether the candidate voice angle is within a voice angle threshold range if the current election participation principle is voice angle priority;
[0154] a second judgment module configured to judge whether the candidate sound intensity exceeds a sound intensity threshold when the candidate sound angle is within the sound angle threshold range;
[0155] a third judgment module configured to judge whether the candidate's voice intensity exceeds the election participation voice intensity floating threshold if the candidate's voice intensity does not exceed the voice intensity threshold;
[0156] The first determination module is configured to determine whether the candidate sound intensity and the candidate sound angle meet its own preset election participation conditions if the candidate sound intensity exceeds the election participation sound intensity floating threshold.
[0157] Optionally, the judging unit 305 includes:
[0158] a second acquisition module configured to acquire current election participation principles, the election participation principles including sound angle priority and sound intensity priority;
[0159] a fourth judgment module configured to judge whether the candidate's voice intensity exceeds a voice intensity threshold if the current election participation principle is voice intensity priority;
[0160] a fifth judgment module configured to, when the candidate sound intensity exceeds the sound intensity threshold, determine whether the candidate sound angle is within a sound angle threshold range;
[0161] a sixth judgment module configured to judge whether the candidate voice angle is within the election participation voice angle floating range if the candidate voice angle is not within the voice angle threshold range;
[0162] The second determination module is configured to determine whether the candidate sound intensity and the candidate sound angle meet its own preset election participation conditions if the candidate sound angle is within the election participation sound angle floating range.
[0163] Optionally, the determining unit includes:
[0164] A first screening and calculation module is configured to screen a source lower limit sound intensity from the sound intensity set, calculate a target lower limit sound intensity for participating in the election based on the source lower limit sound intensity and a preset sound intensity floating range, and use the target lower limit sound intensity as the sound intensity threshold;
[0165] The second screening and calculation module is configured to screen edge sound angles from the sound angle set, calculate the target sound angle range for participating in the election based on the edge sound angle and the preset sound angle floating range, and use the target sound angle range as the sound angle threshold range.
[0166] The apparatus for distributed voice devices to participate in elections provided by an embodiment of the present invention, after obtaining candidate audio parameters, first determines whether the candidate audio parameters meet its own preset election participation conditions, and when it is determined that the candidate audio parameters meet its own preset election participation conditions, then sends the candidate audio parameters to the voice decision device to initiate an election request, so that when many voice devices in the same local area network all recognize the voice command, only the voice devices that meet their own preset election participation conditions will participate in the election, and the voice devices that do not meet their own preset election participation conditions will not initiate an election request, thereby reducing the computing pressure of the voice decision server, and further improving computing efficiency, shortening the response time to voice commands, and increasing user stickiness.
[0167] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the apparatus for distributed voice device participation in elections. In other embodiments of the present invention, the apparatus for distributed voice device participation in elections may include more or fewer components than illustrated, or may combine or separate certain components, or employ different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of both.
[0168] The information interaction, execution process, etc. between the units in the above-mentioned device are based on the same concept as the embodiment of the method of the present invention. For specific contents, please refer to the description in the embodiment of the method of the present invention and will not be repeated here.
[0169] like Figure 4 As shown, an embodiment of the present invention further provides a voice device, comprising: at least one memory and at least one processor;
[0170] The at least one memory is configured to store a machine-readable program;
[0171] The at least one processor is configured to call the machine-readable program to execute the method for distributed voice device participation in an election in any embodiment of the present invention.
[0172] An embodiment of the present invention further provides a computer-readable medium having computer instructions stored thereon. When executed by a processor, the computer instructions cause the processor to execute the method for distributed voice device participation in an election according to any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided. The storage medium stores software program code that implements the functions of any of the above-described embodiments, and the computer (or CPU or MPU) of the system or device can read and execute the program code stored in the storage medium.
[0173] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute part of the present invention.
[0174] Examples of storage media for providing program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.
[0175] In addition, it should be clear that the functions of any of the above embodiments can be achieved not only by executing the program code read by the computer, but also by enabling the operating system operating on the computer to complete part or all of the actual operations based on the instructions of the program code.
[0176] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer, and then based on the instructions of the program code, the CPU installed on the expansion board or expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above embodiments.
[0177] It should be noted that not all steps and modules in the above processes and system structure diagrams are required, and certain steps or modules can be omitted according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or may be implemented by certain components in multiple independent devices.
[0178] In the above embodiments, the hardware unit can be realized by mechanical means or electrical means. For example, a hardware unit can include permanent dedicated circuits or logic (such as special processors, FPGA or ASIC) to complete the corresponding operations. The hardware unit can also include programmable logic or circuits (such as general-purpose processors or other programmable processors), which can be temporarily set up by software to complete the corresponding operations. Concrete implementation (mechanical means or dedicated permanent circuits or temporarily set circuits) can be determined based on the consideration on cost and time.
[0179] The present invention has been shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the scope of protection of the present invention.
Claims
1. A method for distributed voice devices to participate in elections, characterized in that: include: The voice device obtains the voice command audio recorded when the user issues the current voice command; The voice device obtains candidate audio parameters of the voice command audio; The voice device determines whether the candidate audio parameters meet its own preset election participation conditions and thus determines whether it is necessary to initiate an election request to the voice decision device; If the candidate audio parameters meet the preset election participation conditions, the voice device initiates an election request to the voice decision device, thereby sending the candidate audio parameters to the voice decision device in the same local area network; if the candidate audio parameters do not meet the preset election participation conditions, the voice device does not initiate an election request to the voice decision device; The voice decision device determines a target voice device that responds to the current voice command based on candidate audio parameters sent by each voice device; in, The candidate audio parameters include candidate sound intensity and candidate sound angle; The determining whether the candidate audio parameters meet the preset election participation conditions includes: Determine whether the candidate sound intensity and the candidate sound angle meet the preset election participation conditions based on the preset sound intensity threshold and sound angle threshold range respectively; Before determining whether the candidate sound intensity and the candidate sound angle meet the preset election participation conditions according to the preset sound intensity threshold and sound angle threshold range respectively, the method further includes: Obtain a historical voice interaction audio set, where the historical voice interaction audio set includes audio recorded each time the user issues a voice command during the historical voice interaction process; Filtering a target audio set from the historical voice interaction audio set, the target audio set including the audio recorded each time the target audio device is selected as the target audio device by the voice decision device during the historical voice interaction process; Acquire the sound intensity and sound angle of each audio in the target audio set to form a sound intensity set and a sound angle set; Determining the sound intensity threshold and the sound angle threshold range according to the sound intensity set and the sound angle set, respectively; The determining of the sound intensity threshold and the sound angle threshold range according to the sound intensity set and the sound angle set, respectively, includes: filtering a source lower limit sound intensity from the sound intensity set, calculating a target lower limit sound intensity for participating in the election based on the source lower limit sound intensity and a preset sound intensity floating range, and using the target lower limit sound intensity as a sound intensity threshold; screening edge sound angles from the sound angle set, calculating a target sound angle range for participating in the election based on the edge sound angles and a preset sound angle floating range, and using the target sound angle range as a sound angle threshold range; in, The source lower limit sound intensity is the sound intensity with the weakest sound intensity in the sound intensity set, representing the lowest sound intensity of the audio recorded when the voice device was used as the target voice device during the historical voice interaction process. Audio with a sound intensity lower than the source lower limit was not selected as the target voice device during the historical interaction process. The edge sound angle is the maximum and minimum sound angles in the sound angle set, representing the angle range of the audio recorded when the voice device is used as the target voice device during the historical voice interaction process. Audio below the minimum angle and above the maximum angle is not selected as the target voice device.
2. The method according to claim 1, characterized in that After determining the sound intensity threshold and the sound angle threshold range according to the sound intensity set and the sound angle set respectively, the method further includes: The sound intensity threshold and the sound angle threshold range are updated regularly.
3. The method according to claim 1, characterized in that The determining whether the candidate sound intensity and the candidate sound angle meet the preset election participation conditions according to the preset sound intensity threshold and sound angle threshold range respectively includes: obtaining a current election participation principle, wherein the election participation principle includes a sound angle priority and a sound intensity priority; If the current election participation principle is sound angle priority, determining whether the candidate sound angle is within the sound angle threshold range; When the candidate sound angle is within the sound angle threshold range, determining whether the candidate sound intensity exceeds the sound intensity threshold; If the candidate voice intensity does not exceed the voice intensity threshold, determining whether the candidate voice intensity exceeds the election participation voice intensity floating threshold; If the candidate sound intensity exceeds the election participation sound intensity floating threshold, it is determined that the candidate sound intensity and the candidate sound angle meet the preset election participation conditions.
4. The method according to claim 1, wherein The determining whether the candidate sound intensity and the candidate sound angle meet the preset election participation conditions according to the preset sound intensity threshold and sound angle threshold range respectively includes: obtaining a current election participation principle, wherein the election participation principle includes a sound angle priority and a sound intensity priority; If the current election participation principle is voice intensity priority, determining whether the candidate voice intensity exceeds the voice intensity threshold; When the candidate sound intensity exceeds the sound intensity threshold, determining whether the candidate sound angle is within the sound angle threshold range; If the candidate sound angle is not within the sound angle threshold range, determining whether the candidate sound angle is within the election participation sound angle floating range; If the candidate sound angle is within the election participation sound angle floating range, it is determined that the candidate sound intensity and candidate sound angle meet the preset election participation conditions.
5. A device for distributed voice equipment to participate in elections, characterized in that: include: A first acquisition unit is configured to acquire the voice command audio recorded when the user issues the current voice command; a second acquiring unit, configured to acquire candidate audio parameters of the voice command audio; a judgment unit configured to judge whether the candidate audio parameters meet its own preset election participation conditions and thus judge whether it is necessary to initiate an election request to the voice decision device; a sending unit configured to, if the candidate audio parameters satisfy its own preset election participation conditions, initiate an election request to the voice decision device, thereby sending the candidate audio parameters to the voice decision device in the same local area network; if the candidate audio parameters do not satisfy its own preset election participation conditions, the voice device does not initiate an election request to the voice decision device; The voice decision device determines a target voice device that responds to the current voice command based on candidate audio parameters sent by each voice device; The candidate audio parameters include candidate sound intensity and candidate sound angle; The judgment unit is configured to: judge whether the candidate sound intensity and candidate sound angle meet its own preset election participation conditions based on the preset sound intensity threshold and sound angle threshold range respectively; The apparatus for the distributed voice device to participate in the election further includes: A third acquiring unit is configured to acquire a historical voice interaction audio set, where the historical voice interaction audio set includes audio recorded each time the user issues a voice command during the historical voice interaction process; a screening unit configured to screen a target audio set from a historical voice interaction audio set, where the target audio set includes audio recorded each time the target audio device is selected as a target voice device by the voice decision device during the historical voice interaction process; A fourth acquisition unit is configured to acquire the sound intensity and sound angle of each audio in the target audio set to form a sound intensity set and a sound angle set; a determining unit configured to determine the sound intensity threshold and the sound angle threshold range according to a sound intensity set and a sound angle set, respectively; The determining of the sound intensity threshold and the sound angle threshold range according to the sound intensity set and the sound angle set, respectively, includes: filtering a source lower limit sound intensity from the sound intensity set, calculating a target lower limit sound intensity for participating in the election based on the source lower limit sound intensity and a preset sound intensity floating range, and using the target lower limit sound intensity as a sound intensity threshold; screening edge sound angles from the sound angle set, calculating a target sound angle range for participating in the election based on the edge sound angles and a preset sound angle floating range, and using the target sound angle range as a sound angle threshold range; in, The source lower limit sound intensity is the sound intensity with the weakest sound intensity in the sound intensity set, representing the lowest sound intensity of the audio recorded when the voice device was used as the target voice device during the historical voice interaction process. Audio with a sound intensity lower than the source lower limit was not selected as the target voice device during the historical interaction process. The edge sound angle is the maximum and minimum sound angles in the sound angle set, representing the angle range of the audio recorded when the voice device is used as the target voice device during the historical voice interaction process. Audio below the minimum angle and above the maximum angle is not selected as the target voice device.
6. A voice device, characterized in that include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 4.
7. A computer-readable medium, characterized in that The computer-readable medium stores computer instructions, which, when executed by a processor, enable the processor to perform the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Information processing method and electronic device
CN107195305A
Voice awakening method and electronic equipment
CN111369988A