Voice activation threshold adjustment method and electronic equipment

By obtaining historical data of voice wake-up events, using machine learning models or time distribution functions to predict probability, dynamically adjusting the voice activation threshold, solving the problem of insufficient sensitivity or frequent false triggering of voice devices, and improving the user experience.

CN120279904APending Publication Date: 2025-07-08HANGZHOU EZVIZ SOFTWARE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510333245.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, improper setting of voice activation threshold results in insufficient sensitivity of the device or frequent false triggering, affecting the user experience.

Method used

By obtaining historical data of speech wake-up events, using machine learning models or time distribution functions to predict probability, dynamically adjust the voice activation threshold to adapt to user usage habits, improve sensitivity and reduce the probability of false wake-up.

Benefits of technology

It realizes dynamic adjustment of voice activation threshold according to user usage habits, improves the sensitivity of the device during the high-frequency usage period and reduces the probability of false wake-up in the low-frequency usage period, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279904A_ABST
    Figure CN120279904A_ABST
Patent Text Reader

Abstract

The invention discloses a method for adjusting a voice activation threshold value, and the method comprises the steps: obtaining the historical data of voice wake-up events of an electronic device with a voice interaction function, the historical data of the voice wake-up events comprises the historical time information of each voice wake-up event, and the historical time information of each voice wake-up event is obtained based on the historical data of the voice wake-up events; and determining a prediction probability of the voice wake-up event in a set prediction time period, and adjusting a voice activation threshold value in the prediction time period by using the prediction probability, so that the voice activation threshold value is reduced along with the increase of the prediction probability. According to the invention, the adjustment of the voice activation threshold of any time granularity is realized, the accuracy of the voice activation threshold is improved, and the user experience of the electronic equipment with the voice interaction function is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of voice interaction, and in particular, to a method for adjusting a voice activation threshold. Background Art

[0002] Voice wake-up is one of the core modules of voice interaction. For a device with voice interaction function, when the user does not perform voice interaction, it needs to be in a sleep state. When the user wants to perform voice interaction, the device is woken up by a specific wake-up word.

[0003] Generally, the voice needs to be greater than the set voice activation threshold to be recognized by the device. The voice activation threshold is used to represent the minimum audio energy level required to trigger the voice recognition or voice control function, which directly affects the wake-up sensitivity. If the voice activation threshold is too high, the user needs to increase the volume or repeat several times to trigger the device response, increasing the user's operation difficulty and fatigue. If the voice activation threshold is too low, the device may be too sensitive to minor sounds in the environment, resulting in false triggers or frequent interruptions of the user's normal communication. Summary of the Invention

[0004] The present invention provides a method for adjusting a voice activation threshold to achieve dynamic adjustment of the voice activation threshold.

[0005] In a first aspect of the present invention, a method for adjusting a voice activation threshold is provided, the method comprising:

[0006] Obtaining historical data of voice wake-up events of an electronic device with voice interaction function, wherein the historical data of voice wake-up events includes: historical time information of each voice wake-up event,

[0007] Based on the historical data of voice wake-up events, determining the prediction probability of voice wake-up events within a set prediction time period,

[0008] Using the prediction probability to adjust the voice activation threshold within the prediction time period, so that the voice activation threshold decreases as the prediction probability increases.

[0009] As a possible implementation manner, the determining the prediction probability of voice wake-up events within a set prediction time period based on the historical data of voice wake-up events includes:

[0010] Based on the historical data of voice wake-up events, obtaining the time distribution function of voice wake-up events,

[0011] Determining the prediction probability of voice wake-up events within a set prediction time period according to the time distribution function of voice wake-up events;

[0012] Or,

[0013] Input the historical data of voice wake-up events into the trained machine learning model to obtain the predicted probability output by the machine learning model, which represents the probability of the predicted time of the voice wake-up event.

[0014] Among them,

[0015] The machine learning model is trained with the historical data of voice wake-up events as sample data, and the model parameters are adjusted according to the loss function between the predicted sample probability and the expected sample probability during the training process.

[0016] As a possible implementation manner, obtaining the time distribution function of the voice wake-up event based on the historical data of the voice wake-up event includes:

[0017] Perform statistical analysis on the historical data of voice wake-up events according to the set statistical time periods to obtain the voice wake-up event distribution functions of the statistical time periods.

[0018] Or,

[0019] Perform clustering analysis on the historical data of voice wake-up events according to the historical time information to obtain a clustering result, where the clustering result represents the cluster center to which each historical data of the voice wake-up event belongs, and the cluster center represents the historical time information.

[0020] Based on the clustering result, determine the voice wake-up event distribution of each cluster center to obtain the time distribution function of the voice wake-up of each cluster center.

[0021] As a possible implementation manner, the voice wake-up distribution function is fitted according to the normal distribution function, where the mean of the normal distribution function represents the concentrated time of the voice wake-up event, and the standard deviation of the normal distribution function represents the degree of deviation of the time of the voice wake-up event from the concentrated time.

[0022] As a possible implementation manner, determining the predicted probability of voice wake-up within the prediction time period according to the time distribution function of voice wake-up includes:

[0023] Derive based on the time distribution function of voice wake-up to obtain the probability distribution density function.

[0024] Calculate the integral of the probability distribution density within the prediction time period to obtain the predicted probability within the prediction time period.

[0025] As a possible implementation manner, using the predicted probability to adjust the voice activation threshold within the prediction time period includes:

[0026] Convert the predicted probability within the prediction time period into a voice activation adjustment amount.

[0027] Based on the voice activation system threshold, reduce the voice activation adjustment amount to obtain the adjusted voice activation threshold.

[0028] As a possible implementation, the converting the prediction probability within the prediction time period into the voice activation adjustment amount includes:

[0029] Calculate the product of the prediction probability and the set decrease coefficient to obtain the voice activation adjustment amount.

[0030] Wherein, the decrease coefficient is used to characterize the dimension of the prediction probability converted into the voice activation system threshold, and the value is a real number greater than 0.

[0031] As a possible implementation, the method further includes:

[0032] An electronic device with voice interaction function collects voice wake-up events in real time and records the time information of the collected voice wake-up events.

[0033] Store the voice wake-up events and their time information as voice wake-up event historical data locally in the electronic device with voice interaction function, and / or upload the voice wake-up events and their time information, as well as the identification information of the electronic device, to the server side accessed by the electronic device for storage.

[0034] Wherein, the time information is the time stamp information of the moment when the voice wake-up event occurs.

[0035] As a possible implementation, the obtaining the voice wake-up event historical data of the electronic device with voice interaction function includes:

[0036] Obtain the voice wake-up event historical data from the local of the electronic device with voice interaction function, and / or obtain the voice wake-up event historical data of the target electronic device from the server side according to the identification information of the target electronic device.

[0037] As a possible implementation, the prediction time period is a time period composed of any moment as the minimum unit.

[0038] The voice wake-up event historical data is processed in the following manner:

[0039] Extract the voice wake-up events of each statistical time period from the stored voice wake-up event historical data according to each statistical time period as random variables.

[0040] Wherein, the statistical time period is a time period with hours as the minimum unit.

[0041] In a second aspect of the present invention, an electronic device is provided, which includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to implement the method for adjusting the voice activation threshold as described in any one of claims 1 to 9.

[0042] The method for adjusting the voice activation threshold provided in the embodiments of the present application utilizes historical data of voice wake-up events in which an electronic device with voice interaction function is woken up by voice, obtains the prediction probability of voice wake-up events within a prediction time period, and uses the prediction probability to adjust the voice activation threshold within the prediction time period, such that the voice activation threshold decreases as the prediction probability increases. In this way, the electronic device with voice interaction function is equivalent to knowing that it is very likely to be woken up within the prediction time period and waits for voice wake-up instructions with high sensitivity, improving the sensitivity during historical high-frequency usage time periods and reducing the probability of false wake-up during historical low-frequency usage time periods. The present application realizes the adjustment of the voice activation threshold at any time granularity, improves the accuracy of the voice activation threshold, and improves the user experience of the electronic device with voice interaction function. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic flowchart of the method for adjusting the voice activation threshold in the embodiments of the present application.

[0044] Figure 2 It is a schematic flowchart of the method for adjusting the voice activation threshold in an electronic device with voice interaction function in Embodiment 1 of the present application.

[0045] Figure 3 It is the distribution of voice wake-up events with each hour of each day as the statistical time period.

[0046] Figure 4 It is a schematic diagram of the clustering result of the historical data of each day.

[0047] Figure 5 It is a schematic flowchart of the method for adjusting the voice activation threshold of the target electronic device on the server side accessing the electronic device with voice interaction function in Embodiment 2 of the present application.

[0048] Figure 6 It is a schematic diagram of an electronic device in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] In order to make the objectives, technical means, and advantages of the present application clearer and more understandable, the following further describes the present application in detail with reference to the accompanying drawings.

[0050] The applicant has found that in the related technologies for adjusting the voice activation threshold, through the historical voice wake-up threshold of the user for dynamic adjustment, setting a personalized threshold, which is equivalent to becoming a personalized voice activation threshold; or, determining the difficulty of user activation through automatic speech recognition (ASR) to adjust the user's personalized voice activation threshold. In practical applications, the device is often fixed at different time points during the day. For example, User 1 will ask the device about the weather in the morning, and User 2 will use the device to control the lights at home to turn on or off at night.

[0051] In view of this, the embodiments of the present application utilize the historical data of voice wake-up events when the device is woken up by voice, obtain the prediction probability of the voice wake-up event within the prediction time period, and use the prediction probability to adjust the voice activation threshold within the prediction time period, so that the voice activation threshold decreases as the prediction probability increases. In this way, the device knows that it is very likely to be woken up within the prediction time period and waits for the voice wake-up instruction with high sensitivity.

[0052] See Figure 1 shown in Figure 1 is a schematic flowchart of a method for adjusting the voice activation threshold according to an embodiment of the present application. The method includes:

[0053] Step 101, obtain the historical data of voice wake-up events of an electronic device with voice interaction function, where the historical data of voice wake-up events includes: the historical time information of each voice wake-up event.

[0054] As an example, the historical data of voice wake-up events can be formed in the following way:

[0055] An electronic device with voice interaction function collects voice wake-up events in real time and records the time information of the collected voice wake-up events. The time information is the timestamp information of the moment when the voice wake-up event occurs, and this timestamp information can be the system time of the electronic device.

[0056] Store the voice wake-up event and its time information as the historical data of voice wake-up events locally in the electronic device with voice interaction function, and / or upload the voice wake-up event and its time information, as well as the identification information of the electronic device, to the server side accessed by the electronic device for storage, so as to store the historical data of voice wake-up events of each electronic device on the server side.

[0057] To improve the quality of historical data of voice wake-up events, as an implementation, user identification is performed on the voice, and different confidence levels are marked for voice wake-up events from different users, so as to clean the historical data of voice wake-up events according to the confidence level. For example, voice wake-up events from children are marked with a low confidence level, and voice wake-up events from the elderly are marked with a high confidence level; as another implementation, the confidence level is marked for voice wake-up events according to the effectiveness of the interaction response after voice wake-up to exclude invalid voice wake-up events. For example, when the interaction response is successful after voice wake-up, the voice wake-up event is marked with a high confidence level, otherwise, the voice wake-up event is marked with a low confidence level.

[0058] When the historical data of voice wake-up events is stored locally in the electronic device, the electronic device or the server obtains the historical data of voice wake-up events from the local of the electronic device.

[0059] When the historical data of voice wake-up events is stored on the server side, the electronic device obtains its historical data of voice wake-up events from the server side according to its identification information, or the server side obtains the historical data of voice wake-up events of the target electronic device according to the identification information of the specified target electronic device.

[0060] Step 102, based on the historical data of voice wake-up events, determine the prediction probability of voice wake-up events within a set prediction time period.

[0061] As an example, based on the historical data of voice wake-up events, obtain the time distribution function of voice wake-up events, and according to the time distribution function of voice wake-up events, determine the prediction probability of voice wake-up events within a set prediction time period.

[0062] Among them, the time distribution function of voice wake-up events can be obtained in the following way:

[0063] According to each set statistical time period, perform statistical analysis on the historical data of voice wake-up events to obtain the voice wake-up event distribution function of each statistical time period, or perform clustering analysis on the historical data of voice wake-up events according to historical time information to obtain a clustering result, where the clustering result characterizes the clustering center to which each historical data of voice wake-up events belongs, and the clustering center characterizes the historical time information. Based on the clustering result, determine the voice wake-up event distribution of each clustering center to obtain the time distribution function of voice wake-up of each clustering center.

[0064] For example, according to each statistical time period, extract the voice wake-up events of each statistical time period from the stored historical data of voice wake-up events as random variables, where the statistical time period is a time period with the hour as the smallest unit.

[0065] The time distribution function can be a normal distribution function. The mean of the normal distribution function represents the concentrated time of the voice wake-up event, and the standard deviation of the normal distribution function represents the degree of deviation of the historical data of the voice wake-up event from the concentrated time. The time distribution function can also be a distribution related to the normal distribution, including but not limited to chi-square distribution, t-distribution, F-distribution, Rayleigh distribution, Cauchy distribution, and lognormal distribution, etc. The time distribution function can also be determined by methods such as maximum likelihood estimation and data fitting.

[0066] The prediction probability can be determined in the following way:

[0067] Derive based on the time distribution function of voice wake-up to obtain the probability distribution density function.

[0068] Calculate the integral of the probability distribution density within the prediction time period to obtain the prediction probability within the prediction time period.

[0069] Among them, the prediction time period is a time period composed of any moment as the minimum unit.

[0070] As another example, input the historical data of voice wake-up events into the trained machine learning model to obtain the prediction probability output by the machine learning model. This prediction probability represents the probability of the predicted time of the voice wake-up event.

[0071] Among them,

[0072] The machine learning model is trained with the historical data of voice wake-up events as sample data, and the model parameters are adjusted according to the loss function between the predicted sample probability and the expected sample probability during the training process.

[0073] Step 103, use the prediction probability to adjust the voice activation threshold within the prediction time period, so that the voice activation threshold decreases as the prediction probability increases.

[0074] As an example, convert the prediction probability within the prediction time period into a voice activation adjustment amount, and on the basis of the voice activation system threshold, reduce the voice activation adjustment amount to obtain the adjusted voice activation threshold.

[0075] Among them, the voice activation adjustment amount is determined in the following way:

[0076] Calculate the product of the prediction probability and the set decrease coefficient to obtain the voice activation adjustment amount.

[0077] Among them, the decrease coefficient is used to characterize the dimension of the prediction probability converted into the voice activation system threshold, and the value is a real number greater than 0.

[0078] The method for adjusting the voice activation threshold in the embodiments of the present application realizes reducing the voice activation threshold according to the predicted probability in the predicted time period. In the predicted time period with high historical frequency of user use, the device greatly reduces the voice activation threshold to improve sensitivity. In the predicted time period with low frequency of use, the device slightly reduces the voice activation threshold to reduce the probability of false wake-up.

[0079] For the convenience of understanding the embodiments of the present application, the following takes the side of an electronic device with a voice interaction function or the server side as an example for illustration.

[0080] Embodiment 1

[0081] See Figure 2 as shown in Figure 2 FIG. 13 is a schematic flowchart of a method for adjusting the voice activation threshold in an electronic device with a voice interaction function according to Embodiment 1 of the present application. The method includes: on the side of an electronic device with a voice interaction function,

[0082] Step 201, in response to a voice wake-up input, generate a voice wake-up event, and record the voice wake-up event and its time information locally in the electronic device to form historical data of the voice wake-up event.

[0083] See Table 1 shown below. Table 1 is an example of historical data of voice wake-up events. The historical data of voice wake-up events includes each voice wake-up event and its corresponding time information.

[0084] Table 1

[0085]

[0086]

[0087] Preferably, perform user identification on the voice wake-up input, mark the confidence level of the voice wake-up event according to the user identification result, and / or mark the confidence level of the voice wake-up event according to the situation of the interaction response after voice wake-up.

[0088] See Table 2 shown below. Table 2 is another example of historical data of voice wake-up events. The historical data of voice wake-up events includes each voice wake-up event and its corresponding time information, as well as the confidence level.

[0089] Table 2

[0090] Voice wake-up event Time information of voice wake-up event Confidence level Voice wake-up event 1 Timestamp 1 Confidence level 1 Voice wake-up event 2 Timestamp 2 Confidence level 2 … … …

[0091] Step 202, the electronic device obtains the historical data of the voice wake-up event from local,

[0092] As an example, obtain the historical data of the voice wake-up event according to the specified time information. For example, the historical data within n days before the current time.

[0093] As another example, historical data of voice wake-up events that meet a desired confidence level are filtered out to improve data quality.

[0094] Step 203: Based on the acquired historical data, a time distribution function of the voice wake-up event is acquired, and the time distribution function is derived to obtain a probability density function.

[0095] To facilitate statistical analysis of historical data, the time information of the historical data of the voice wake-up event is extracted according to the set statistical time period to obtain the voice wake-up events of each statistical time period. For example, the statistical time period uses hours as the smallest unit, and the timestamp information of the voice wake-up event is h hours, x minutes, and s seconds, then the statistical time extracted is h hours, x minutes, and s seconds.

[0096] As an example, statistical analysis is performed on historical data of each statistical time period to obtain a time distribution function of a voice wake-up event, where the time distribution function represents the probability of occurrence of a voice wake-up event as a random variable at each time point.

[0097] Given the general rules of voice interaction devices, we interact with the devices at a fixed time every day. According to the central limit theorem, we can assume that the distribution of voice wake-up events will gradually show a normal distribution trend, denoted by X~N(μ,σ 2 ), where X represents a voice wake-up event, which is a random variable, μ is the mean, which determines the center position of the normal distribution and represents the concentrated time period of daily use, and σ is the standard deviation, which can indicate the degree of deviation of the time of a random voice wake-up event from the center position.

[0098] For example, when you wake up every day, you ask your smart device about the weather. Assuming that the user wakes up at 8 o'clock every day, the mean μ will most likely be close to 8 o'clock. σ is the standard deviation, which can indicate the degree of deviation of the user's wake-up time from 8 o'clock.

[0099] By taking the derivative of the normal distribution function, we can get the probability density function, which is expressed mathematically as follows:

[0100]

[0101] Where f(x) is the probability density function.

[0102] If users have multiple habitual usage time points every day, a normal distribution of multiple statistical time periods will be formed. Figure 3 As shown, Figure 3 The figure shows the distribution of voice wake-up events with each hour of the day as the statistical time period, from 0 o'clock to 24 o'clock. Two different normal distribution functions are shown in the figure.

[0103] As another example, historical data for each statistical time period is clustered to obtain a clustering result, which characterizes the clustering center of each historical data. The clustering center corresponds to time information. Refer to Figure 4 as shown in Figure 4 FIG. Figure 4 is a schematic diagram of the clustering result of historical data for each day. Different clustering centers correspond to different time information.

[0104] For each clustering center, obtain the normal distribution function of the clustering center, and take the derivative of the normal distribution function to obtain the probability density function of the clustering center.

[0105] Step 204: Determine the predicted probability of the voice wake-up event within the predicted time period according to the probability density function, and determine the adjustment degree of the voice activation threshold according to the predicted probability to obtain the adjusted voice activation threshold.

[0106] In view of obtaining the probability density function in step 203, substitute the set predicted time period into the probability density function for integration, and the predicted probability value of the voice wake-up event within the predicted time period can be calculated as the probability of the device being used. Since the predicted probability is obtained by integrating the probability density function, in this way, the predicted time period can be a time period composed of any moment, which is beneficial to improving the adjustment granularity of the voice activation threshold and also beneficial to improving the adjustment accuracy of the voice activation threshold.

[0107] For example, if the user is used to getting up at 8 o'clock, the central value of the normal distribution calculated in the previous step is probably at 8 o'clock. Divide the predicted time period discretely by minutes, and calculate the predicted probability from 8 o'clock sharp to 8:01. It can be calculated using the following formula:

[0108]

[0109] where f(x) is the probability density distribution function, a and b are the predicted time points where the predicted time period is located, for example, 8 o'clock and 8:01. P is the predicted probability within the predicted time period. According to the normal distribution function, the predicted probability of the predicted time period near the mean value is the largest.

[0110] After obtaining the predicted probability of the predicted time period, for the voice activation threshold within the predicted time period, it is adjusted to:

[0111] th' = th - alpha × P (3)

[0112] In Formula 3, th' is the adjusted voice activation threshold, th is the voice activation system threshold configured by the system, and alpha is the decay coefficient, which is used to represent the dimension of converting the prediction probability into the voice activation threshold. The value of alpha is a real number greater than 0, which can be a fixed value or a non-fixed value. This embodiment does not limit this. P is the prediction probability within the prediction time period, and alpha×P is the adjustment amount.

[0113] The adjustment of the voice activation threshold can make the voice activation threshold decrease correspondingly as the prediction probability increases. The higher the prediction probability, the greater the degree of decrease. For example, if the voice activation threshold is 0.5, alpha is 10, and the probability from 8:00 to 8:01 is 0.02, then the adjusted voice activation threshold is 0.3.

[0114] The characteristic of the normal distribution is that the probability is higher at the time points closer to the mean. Then, the voice activation threshold will show that the closer to the mean position, the lower the voice activation threshold, that is, it is easier to be awakened. It is equivalent to that the electronic device knows that it is very likely to be awakened during this central position time period, and then it will concentrate on waiting for the wake-up instruction.

[0115] In this embodiment, the adjustment of the voice activation threshold can be realized on the side of the electronic device with voice interaction function, which is beneficial to improving the independent electronic device with voice interaction function, thereby improving the sensitivity of the electronic device with voice interaction function and reducing the false wake-up probability.

[0116] Embodiment 2

[0117] See Figure 5 as shown in Figure 5 This is a schematic flowchart of a method for adjusting the voice activation threshold of a target electronic device on the server side connected to an electronic device with voice interaction function in the second embodiment of the present application. The method includes: on the server side,

[0118] Step 501, receiving the historical data of the voice wake-up event uploaded by the electronic device, and storing it according to the identification information of the electronic device from which the historical data of the voice wake-up event comes.

[0119] See Table 3 shown below. Table 3 is an example of the historical data of the voice wake-up event stored on the server side. The historical data of the voice wake-up event includes: each voice wake-up event, its corresponding time information, and the identification information of the corresponding electronic device.

[0120] Table 3

[0121]

[0122]

[0123] To reduce the amount of data uploaded, improve the quality of historical data, and reduce the occupation of server storage resources, the electronic device can screen voice wake-up events according to the confidence level to screen out the historical data of voice wake-up events that meet the expected confidence level for uploading.

[0124] Step 502: According to the identification information of the target electronic device for which the voice activation threshold is to be adjusted, obtain the historical data of the voice wake-up events of the target electronic device from the stored historical data of voice wake-up events.

[0125] As an example, in response to a request from the target electronic device for voice activation threshold adjustment, the request carries the identification information of the target electronic device, and the historical data of voice wake-up events is obtained according to the specified time information.

[0126] Step 503: Based on the obtained historical data, determine the predicted probability value of the voice wake-up event.

[0127] As an example, the obtained historical data is input into the trained machine learning model, and the predicted probability value is obtained from the output of the machine learning model. This predicted probability value represents the probability within the prediction time. For example, the historical data input into the machine learning model is several data at different hours, minutes, and seconds of each day, and the machine learning model outputs the probability of each prediction time of each day, where each prediction time includes any time with the smallest unit among hours, minutes, and seconds.

[0128] Among them, the machine learning model can be trained with the historical data of voice wake-up events as sample data, and the model parameters are adjusted according to the loss function between the predicted sample probability and the expected sample probability during the training process.

[0129] Different electronic devices can correspond to different trained machine learning models or share the same trained machine learning model. The embodiments of the present application do not limit this.

[0130] Step 504: Send the predicted probability value to the target electronic device, so that the target electronic device determines the adjustment degree of the voice activation threshold according to the predicted probability and obtains the adjusted voice activation threshold.

[0131] In this step, the target electronic device can obtain the adjusted voice activation threshold according to Formula 3.

[0132] In this embodiment, the adjustment of the voice activation threshold is implemented through the server side, which is beneficial to the management and maintenance of the voice activation thresholds of the electronic devices accessing the server.

[0133] See Figure 6 as shown Figure 6A schematic diagram of an electronic device according to an embodiment of the present application. The electronic device may be an electronic device with voice interaction function or a server device. The electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of the method for adjusting the voice activation threshold according to the embodiment of the present application.

[0134] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory (Non-Volatile Memory, NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0135] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0136] An embodiment of the present invention also provides a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the method for adjusting the voice activation threshold according to the embodiment of the present application.

[0137] For the device / network-side device / storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, please refer to the partial description of the method embodiment.

[0138] In this article, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0139] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for adjusting a voice activation threshold, characterized in that The method includes: Obtaining historical data of voice wake-up events of an electronic device with voice interaction function, where the historical data of voice wake-up events includes historical time information of each voice wake-up event. Based on the historical data of voice wake-up events, determining the prediction probability of voice wake-up events within a set prediction time period. Using the prediction probability to adjust the voice activation threshold within the prediction time period, such that the voice activation threshold decreases as the prediction probability increases.

2. The adjustment method according to claim 1, wherein The determining the prediction probability of voice wake-up events within a set prediction time period based on the historical data of voice wake-up events includes: Based on the historical data of voice wake-up events, obtaining the time distribution function of voice wake-up events. According to the time distribution function of voice wake-up events, determining the prediction probability of voice wake-up events within a set prediction time period. Or, Inputting the historical data of voice wake-up events into a trained machine learning model to obtain the prediction probability output by the machine learning model, where the prediction probability represents the probability of the predicted time of voice wake-up events. Wherein, The machine learning model is trained with the historical data of voice wake-up events as sample data, and the model parameters are adjusted according to the loss function between the predicted sample probability and the expected sample probability during the training process.

3. The adjustment method according to claim 2, wherein The obtaining the time distribution function of voice wake-up events based on the historical data of voice wake-up events includes: According to set statistical time periods, performing statistical analysis on the historical data of voice wake-up events to obtain the voice wake-up event distribution function of each statistical time period. Or, Performing clustering analysis on the historical data of voice wake-up events according to the historical time information to obtain a clustering result, where the clustering result represents the clustering center to which each historical data of voice wake-up events belongs, and the clustering center represents the historical time information. Based on the clustering result, determining the voice wake-up event distribution of each clustering center to obtain the time distribution function of voice wake-up of each clustering center.

4. The adjustment method according to claim 3, characterized in that The voice wake-up distribution function is fitted according to a normal distribution function, where the mean of the normal distribution function represents the concentrated time of voice wake-up events, and the standard deviation of the normal distribution function represents the degree of deviation of the time of voice wake-up events relative to the concentrated time.

5. The adjustment method according to any one of claims 2 to 4, characterized in that The determining the prediction probability of voice wake-up within a prediction time period according to the time distribution function of voice wake-up includes: Taking the derivative based on the time distribution function of voice wake-up to obtain the probability distribution density function. Calculating the integral of the probability distribution density within the prediction time period to obtain the prediction probability within the prediction time period.

6. The adjustment method according to claim 1, characterized in that The using the prediction probability to adjust the voice activation threshold within the prediction time period includes: Converting the prediction probability within the prediction time period into a voice activation adjustment amount. On the basis of the voice activation system threshold, reducing the voice activation adjustment amount to obtain the adjusted voice activation threshold.

7. The adjustment method according to claim 6, wherein The converting the prediction probability within the prediction time period into a voice activation adjustment amount includes: Calculating the product of the prediction probability and a set decrease coefficient to obtain the voice activation adjustment amount. Wherein, the decrease coefficient is used to represent the dimension of converting the prediction probability into the voice activation system threshold, and the value is a real number greater than 0.

8. The adjustment method according to claim 1, wherein The method further includes: An electronic device with voice interaction function collects voice wake-up events in real time and records the time information of the collected voice wake-up events. The voice wake-up events and their time information are stored as voice wake-up event historical data locally in the electronic device with voice interaction function, and / or the voice wake-up events and their time information, as well as the identification information of the electronic device, are uploaded to the server side accessed by the electronic device for storage. Among them, the time information is the time stamp information at the moment when the voice wake-up event occurs. The obtaining of the voice wake-up event historical data of the electronic device with voice interaction function includes: Obtaining the voice wake-up event historical data from the local of the electronic device with voice interaction function, and / or obtaining the voice wake-up event historical data of the target electronic device from the server side according to the identification information of the target electronic device.

9. The adjustment method according to claim 8, characterized in that The prediction time period is a time period composed of any moment as the minimum unit. The voice wake-up event historical data is processed in the following manner: According to each statistical time period, the voice wake-up events of each statistical time period are extracted from the stored voice wake-up event historical data as random variables. Among them, the statistical time period is a time period with an hour as the minimum unit.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to implement the method for adjusting the voice activation threshold as described in any one of claims 1 to 9.