Method and apparatus for determining custom wake-up word threshold, storage medium, and electronic device
Patent Information
- Application Number
- CN202211009180.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-22
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2042-08-22
AI Technical Summary
[0005]本申请的主要目的在于提供一种确定自定义唤醒词阈值的方法以及装置、存储介质、电子装置,以解决语音唤醒时误唤醒率高的问题
[0019] The method, apparatus, storage medium, and electronic device for determining a custom wake-up word threshold in this application embodiment score the user-defined wake-up parameters based on the posterior results generated by the test set of the preset acoustic model to obtain a false wake-up index. This achieves the purpose of determining the custom wake-up word threshold based on the false wake-up index, thereby realizing the technical effect of determining the custom wake-up word threshold based on the false wake-up index and solving the technical problem of high false wake-up rate during voice wake-up.
Smart Images

Figure CN115376493B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech processing, and more specifically, to a method, apparatus, storage medium, and electronic device for determining a custom wake word threshold. Background Technology
[0002] With the development of smart electronic devices, more and more devices are beginning to support voice control. However, when developing voice wake-up or command word recognition, the wake-up words set by different manufacturers and different devices are not the same, which is not user-friendly. Therefore, it is necessary to support users to customize wake-up words.
[0003] In related technologies, settings are often based on empirical values, making it difficult to precisely control wake-up / false wake-up performance, resulting in difficulty waking up devices or excessively high false wake-ups. For example, common custom wake words work well, but uncommon wake words may lead to high false wake-ups or low wake-up rates.
[0004] There is currently no effective solution to the problem of high false wake-up rate during voice wake-up in related technologies. Summary of the Invention
[0005] The main objective of this application is to provide a method, apparatus, storage medium, and electronic device for determining a custom wake-up word threshold, in order to solve the problem of high false wake-up rate during voice wake-up.
[0006] To achieve the above objectives, according to one aspect of this application, a method for determining a custom wake word threshold is provided for use on a user terminal.
[0007] The method for determining a custom wake word threshold according to this application includes: scoring user-defined wake-up parameters based on the posterior results generated by a test set of a preset acoustic model to obtain a false wake-up index, wherein the preset acoustic model is used to store the forward calculation in the posterior; and determining the custom wake word threshold based on the false wake-up index.
[0008] Furthermore, the method also includes: filtering the posterior results generated by the test set of the preset acoustic model, using the sparsity of non-space frames in the posterior output of the CTC algorithm, training the preset acoustic model with the CTC loss function, and filtering frames whose posterior probability of space exceeds a preset threshold by setting a space threshold, wherein the space refers to the empty node of the CTC.
[0009] Furthermore, the preset acoustic model is trained using the CTC loss function, and the posterior results generated by the test set are compressed after filtering frames whose posterior probability of the space exceeds the preset threshold by setting a space threshold.
[0010] Furthermore, the preset acoustic model is trained using the CTC loss function to store the peak posterior time series, and frames with posterior probabilities of spaces exceeding the preset space probability are filtered out by setting a space threshold.
[0011] Further, determining the custom wake word threshold based on the false wake-up index includes: the false wake-up index includes at least one of the following: a custom wake-up word, a false wake-up count; and calculating the false wake-up count of a new custom wake-up word in the test set using the stored posterior results based on the false wake-up count in the false wake-up index.
[0012] Furthermore, determining the custom wake word threshold based on the false wake-up index further includes: filtering the custom wake word threshold based on the number of false wake-ups in the false wake-up index; and notifying the user of a message indicating that the wake word threshold has been successfully selected.
[0013] To achieve the above objectives, according to one aspect of this application, a method for determining a custom wake word threshold is provided for use on a server.
[0014] The method for determining a custom wake word threshold according to this application includes: receiving audio data of a preset test set; scoring user-defined wake-up parameters based on the posterior results generated by the test set of a preset acoustic model to obtain a false wake-up index, wherein the preset acoustic model is used to store the forward calculation in the posterior; determining a custom wake word threshold based on the false wake-up index and sending it to the user terminal.
[0015] To achieve the above objectives, according to another aspect of this application, an apparatus for determining a custom wake word threshold is provided.
[0016] The apparatus for determining a custom wake-up word threshold according to this application includes: a preprocessing module, used to score user-defined wake-up parameters based on posterior results generated by a test set of a preset acoustic model to obtain a false wake-up index, wherein the preset acoustic model is used to store the forward calculation in the posterior; a calculation module, used to score user-defined wake-up parameters based on posterior results generated by the test set to obtain a false wake-up index; and a threshold determination module, used to determine a custom wake-up word threshold based on the false wake-up index.
[0017] To achieve the above objectives, according to another aspect of this application, a storage medium is also provided, wherein the storage medium stores a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when running.
[0018] To achieve the above objectives, according to another aspect of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0019] The method, apparatus, storage medium, and electronic device for determining a custom wake-up word threshold in this application embodiment score the user-defined wake-up parameters based on the posterior results generated by the test set of the preset acoustic model to obtain a false wake-up index. This achieves the purpose of determining the custom wake-up word threshold based on the false wake-up index, thereby realizing the technical effect of determining the custom wake-up word threshold based on the false wake-up index and solving the technical problem of high false wake-up rate during voice wake-up. Attached Figure Description
[0020] The accompanying drawings, which form part of this application, are used to provide a further understanding of the application and to make other features, objects, and advantages of the application more apparent. The illustrative embodiments and descriptions of this application are used to explain the application and do not constitute an undue limitation of the application. In the drawings:
[0021] Figure 1 This is a schematic diagram of the hardware structure of a method for determining a custom wake word threshold according to an embodiment of this application;
[0022] Figure 2 This is a flowchart illustrating a method for determining a custom wake word threshold according to an embodiment of this application;
[0023] Figure 3 This is a schematic diagram of the device structure for determining a custom wake word threshold according to an embodiment of this application;
[0024] Figure 4 This is a flowchart illustrating a method for determining a custom wake word threshold according to an embodiment of this application;
[0025] Figure 5 This is a schematic diagram of the user threshold determination process in the method for determining a custom wake word threshold according to an embodiment of this application;
[0026] Figure 6 This is a schematic diagram of the initialization process in the method for determining a custom wake word threshold according to an embodiment of this application. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] In this application, the terms "upper," "lower," "left," "right," "front," "rear," "top," "bottom," "inner," "outer," "middle," "vertical," "horizontal," "lateral," and "longitudinal" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are primarily for the purpose of better describing this application and its embodiments, and are not intended to limit the indicated device, element, or component to having a specific orientation, or to be constructed and operated in a specific orientation.
[0030] Furthermore, in addition to indicating location or positional relationship, some of the aforementioned terms may also have other meanings. For example, the term "above" may also be used in some cases to indicate a certain dependency or connection relationship. Those skilled in the art can understand the specific meaning of these terms in this application based on the specific circumstances.
[0031] Furthermore, the terms "installation," "setup," "equipped with," "connection," "linking," and "socketing" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral structure; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium, or an internal connection between two devices, components, or parts. Those skilled in the art can understand the specific meaning of these terms in this application based on the specific circumstances.
[0032] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0033] like Figure 1 As shown, it mainly includes voice wake-up 100, a preset acoustic model 200, a test set 300, and a false wake-up index 400. Voice wake-up 100 is input into the preset acoustic model 200 to determine the custom wake-up word threshold. The wake-up threshold is determined based on the false wake-up index 400 in the test set 300. From a user experience perspective, the threshold automatically determined by the test set 300 is actually calculated for the preset acoustic model. This avoids problems such as excessively high false wake-ups or difficulty in waking up, thus preventing a poor user experience.
[0034] like Figure 2 As shown, the method includes the following steps S201 to S203:
[0035] Step S201: The user-defined wake-up parameters are scored based on the posterior results generated by the test set of the preset acoustic model to obtain the false wake-up index, wherein the preset acoustic model is used to store the forward calculation in the posterior.
[0036] Step S202: Determine the custom wake-up word threshold based on the false wake-up index.
[0037] As can be seen from the above description, this application achieves the following technical effects:
[0038] By pre-acquiring audio data and inputting it into a preset acoustic model, the preset acoustic model is used to store the posterior result of the forward computation. The posterior result generated by the test set of the preset acoustic model is stored in time series. This achieves the goal of scoring user-defined wake-up parameters based on the posterior result generated by the test set, obtaining a false wake-up index. This realizes the technical effect of determining the threshold of a custom wake-up word based on the false wake-up index, thereby solving the technical problem of high false wake-up rate during voice wake-up.
[0039] The preset acoustic model mentioned in step S201 above is a neural network model based on CTC.
[0040] As an optional implementation, the preset acoustic model is used to store the forward computation as a posterior. This is a preprocessing step of storing the forward computation as a posterior.
[0041] As an optional implementation, the posterior results generated by the test set of the preset acoustic model are stored in a time series.
[0042] In the above steps, a 16-hour chat scenario false wake-up test set can be used. The posterior generated by the test set is stored in time series by the forward calculation of the acoustic model. For the customized wake-up words, a scoring algorithm is used to score the stored posterior. Based on the false wake-up index, the threshold is automatically determined.
[0043] Furthermore, the user-defined wake-up parameters are scored based on the posterior results generated from the test set to obtain a false wake-up index. In other words, for customized wake-up words, a scoring algorithm is used to score them on the stored posterior, and a threshold is automatically determined based on the false wake-up index.
[0044] As an optional implementation, user-defined wake-up parameters include the number of wake-ups.
[0045] As an optional implementation, the user-defined wake-up parameters also include the number of false wake-ups.
[0046] In step S202 above, a custom wake-up word threshold is determined based on the false wake-up index.
[0047] As an optional implementation, the custom wake-word threshold is determined based on the actual output of a preset acoustic model. That is, it is variable and corresponds to the user-defined false wake-up index.
[0048] As an optional implementation, the custom wake-word threshold does not lead to excessively high false wake-ups or difficulty in waking up. In other words, it can match most scenarios.
[0049] As a preferred embodiment, the method further includes: filtering the posterior results generated by the test set of the preset acoustic model, utilizing the sparsity of non-space frames in the posterior output of the CTC algorithm, training the preset acoustic model using the CTC loss function, and filtering frames whose posterior probability of space exceeds a preset threshold by setting a space threshold, wherein the space refers to the empty node of the CTC.
[0050] In practical implementation, directly using the above method typically generates 600 minutes of posterior results on a one-hour test set, while commonly used test sets are only a few dozen hours long, rendering this method unusable. To optimize this, the sparsity of non-Blank space frames in the CTC posterior output is utilized. The model is trained using the CTC loss function, and a Blank space threshold is set to filter out frames with a Blank space posterior probability exceeding the threshold. Preferably, for example, a Blank threshold of 0.9 can filter 98% of the posterior results, compressing the storage results to just over ten megabytes. This makes the method practical.
[0051] Preferably, the preset acoustic model is trained using the CTC loss function, and the posterior results generated by the test set are compressed after filtering frames whose posterior probability of the space exceeds the preset threshold by setting a space threshold.
[0052] Preferably, the preset acoustic model is trained using the CTC loss function to store the peak posterior time series, and frames with a space posterior probability exceeding a preset space probability are filtered out by setting a space threshold.
[0053] Specifically, the posterior time series of the spikes is stored, and frames with high Blank probability are filtered out. Since the number of non-Blank spikes is small, the storage space and computing resources required for the generated posterior file are sufficient to meet the requirements of various devices.
[0054] In a preferred embodiment, determining the custom wake word threshold based on the false wake-up index includes: the false wake-up index includes at least one of the following: a custom wake-up word, a false wake-up count; and calculating the new custom wake-up word false wake-up count using stored posterior results in the test set based on the false wake-up count in the false wake-up index.
[0055] To obtain better results and allow users to select the number of false wake-ups themselves, the number of false wake-ups for the new custom wake-up word is calculated using stored posterior results in the test set, based on the number of false wake-ups in the false wake-up metric.
[0056] In a preferred embodiment, determining the custom wake-up word threshold based on the false wake-up index further includes: filtering the custom wake-up word threshold based on the number of false wake-ups in the false wake-up index; and notifying the user of a message indicating that the custom wake-up word threshold has been successfully selected.
[0057] Based on the number of false wake-ups, the custom wake-up word threshold is filtered, and a message is sent to the user indicating that the custom wake-up word threshold has been successfully selected. The user can then directly use the corresponding custom wake-up word based on the result.
[0058] like Figure 4 As shown, in another embodiment of this application, a method for determining a custom wake word threshold is also provided for use on a server side, the method comprising:
[0059] Step S401: Receive audio data from a preset test set;
[0060] Step S402: The user-defined wake-up parameters are scored based on the posterior results generated by the test set of the preset acoustic model to obtain the false wake-up index, wherein the preset acoustic model is used to store the forward calculation in the posterior.
[0061] Step S403: Determine the custom wake-up word threshold based on the false wake-up index and send it to the user terminal.
[0062] In step S401 above, the server receives the acquired audio data from the test set.
[0063] In step S402 above, the audio data is input into a preset acoustic model. The preset acoustic model is a neural network model based on CTC.
[0064] As an optional implementation, the preset acoustic model is used to store the forward computation as a posterior. This is a preprocessing step of storing the forward computation as a posterior.
[0065] As an optional implementation, the posterior results generated by the test set of the preset acoustic model are stored in a time series.
[0066] The server scores the user-defined wake-up parameters based on the posterior results generated from the test set, thus obtaining a false wake-up index. In other words, for customized wake-up words, a scoring algorithm is used to score them on the stored posterior, and a threshold is automatically determined based on the false wake-up index.
[0067] In the above steps, a 16-hour chat scenario false wake-up test set can be used. The posterior of the acoustic model forward calculation is stored in time sequence. For customized wake-up words, a scoring algorithm is used to score the stored posterior. Based on the false wake-up index, the threshold is automatically determined.
[0068] As an optional implementation, user-defined wake-up parameters include the number of wake-ups.
[0069] As an optional implementation, the user-defined wake-up parameters also include the number of false wake-ups.
[0070] In step S403 above, the server determines the custom wake-up word threshold based on the false wake-up index.
[0071] As an optional implementation, the custom wake-word threshold is determined based on the actual output of a preset acoustic model. That is, it is variable and corresponds to the user-defined false wake-up index.
[0072] As an optional implementation, the custom wake-word threshold does not lead to excessively high false wake-ups or difficulty in waking up. In other words, it can match most scenarios.
[0073] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0074] According to embodiments of this application, an apparatus for determining a custom wake-up word threshold for implementing the above method is also provided, such as... Figure 3 As shown, the device includes:
[0075] Preprocessing module 301 is used to score user-defined wake-up parameters based on the posterior results generated by the test set of the preset acoustic model to obtain false wake-up index, wherein the preset acoustic model is used to store the forward calculation in the posterior.
[0076] The calculation module 302 is used to score the user-defined wake-up parameters based on the posterior results generated by the test set to obtain a false wake-up index;
[0077] The threshold determination module 303 is used to determine the custom wake-up word threshold based on the false wake-up index.
[0078] The preset acoustic model in the preprocessing module 301 of this application embodiment is a neural network model based on CTC.
[0079] As an optional implementation, the preset acoustic model is used to store the forward computation as a posterior. This is a preprocessing step of storing the forward computation as a posterior.
[0080] As an optional implementation, the posterior results generated by the test set of the preset acoustic model are stored in a time series.
[0081] In the above steps, a 16-hour chat scenario false wake-up test set can be used. The posterior generated by the test set is stored in time series by the forward calculation of the acoustic model. For the customized wake-up words, a scoring algorithm is used to score the stored posterior. Based on the false wake-up index, the threshold is automatically determined.
[0082] Furthermore, the user-defined wake-up parameters are scored based on the posterior results generated from the test set to obtain a false wake-up index. In other words, for customized wake-up words, a scoring algorithm is used to score them on the stored posterior, and a threshold is automatically determined based on the false wake-up index.
[0083] As an optional implementation, user-defined wake-up parameters include the number of wake-ups.
[0084] As an optional implementation, the user-defined wake-up parameters also include the number of false wake-ups.
[0085] In the threshold determination module 303 of this application embodiment, a custom wake-up word threshold is determined based on the false wake-up index.
[0086] As an optional implementation, the custom wake-word threshold is determined based on the actual output of a preset acoustic model. That is, it is variable and corresponds to the user-defined false wake-up index.
[0087] As an optional implementation, the custom wake-word threshold does not lead to excessively high false wake-ups or difficulty in waking up. In other words, it can match most scenarios.
[0088] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device, or fabricating them separately as individual integrated circuit modules, or fabricating multiple modules or steps as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0089] To better understand the above method for determining the custom wake word threshold, the following explanation of the technical solution is provided in conjunction with preferred embodiments, but it is not intended to limit the technical solution of the embodiments of the present invention.
[0090] The method for determining the custom wake-up word threshold in this embodiment automatically determines the threshold through a test set, which is actually calculated for a preset acoustic model. This avoids problems such as excessively high false wake-ups or difficulty in waking up, thus preventing a poor user experience. Furthermore, it allows users to select the number of false wake-ups themselves.
[0091] like Figure 5 The diagram shown is a flowchart illustrating a method for determining a custom wake word threshold according to an embodiment of this application. The specific implementation process includes the following steps:
[0092] Step S501, Begin.
[0093] Step S502: The user defines the wake word and the number of false wake-ups.
[0094] Optionally, the default is 1 time / 12 hours.
[0095] Step S503: Calculate the number of false wake-ups using the stored posterior data.
[0096] The user-defined wake-up parameters are scored based on the posterior results generated from the test set to obtain a false wake-up index. In other words, for customized wake-up words, a scoring algorithm is used to score them on the stored posterior, and a threshold is automatically determined based on the false wake-up index.
[0097] Step S504: Filter the threshold based on the number of false wake-ups.
[0098] Based on the false wake-up index, determine the custom wake-up word threshold.
[0099] As an optional implementation, the custom wake-word threshold is determined based on the actual output of a preset acoustic model. That is, it is variable and corresponds to the user-defined false wake-up index.
[0100] As an optional implementation, the custom wake-word threshold does not lead to excessively high false wake-ups or difficulty in waking up. In other words, it can match most scenarios.
[0101] Step S505: Select the threshold and notify the user that the setting was successful.
[0102] Based on the number of false wake-ups, the custom wake-up word threshold is filtered, and a message is sent to the user indicating that the custom wake-up word threshold has been successfully selected. The user can then directly use the corresponding custom wake-up word based on the result.
[0103] Using a test set of chat scenarios with known durations for false wake-up, the posterior calculated by the preset acoustic model is stored in time sequence. For customized wake-up words, a scoring algorithm is used to score the stored posterior. Based on the false wake-up index, the threshold is automatically determined.
[0104] To optimize the dataset, the sparsity of non-Blank frames in the CTC posterior output is utilized. The model is trained using the CTC loss function, and a Blank threshold is set to filter out frames with Blank posterior probabilities exceeding the threshold.
[0105] It is understandable that CTC stands for Connection Timing Classification, and Blank represents an empty node in CTC.
[0106] In addition, such as Figure 6 The initialization process shown also includes:
[0107] Step S601, Begin.
[0108] Step S602, input audio.
[0109] Step S603, forward calculation.
[0110] Step S604, store the post-verification.
[0111] Audio data is acquired in advance and then input into a preset acoustic model. The preset acoustic model is a neural network model based on CTC.
[0112] As an optional implementation, the preset acoustic model is used to store the forward computation as a posterior. This is a preprocessing step of storing the forward computation as a posterior.
[0113] As an optional implementation, the posterior results generated by the test set of the preset acoustic model are stored in a time series.
[0114] In the above steps, a 16-hour chat scenario false wake-up test set can be used. The posterior of the acoustic model forward calculation is stored in time sequence. For customized wake-up words, a scoring algorithm is used to score the stored posterior. Based on the false wake-up index, the threshold is automatically determined.
[0115] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for determining a custom wake word threshold, characterized in that, For use on the user end, the method includes: The user-defined wake-up parameters are scored based on the posterior results generated by the test set of the preset acoustic model to obtain the false wake-up index, wherein the preset acoustic model is used to store the forward calculation in the posterior. The specific steps for scoring user-defined wake-up parameters based on the posterior results generated from the test set of the preset acoustic model to obtain the false wake-up index include: S1 pre-acquires speech data, trains the preset acoustic model using the CTC loss function to perform forward computation, generates posterior results, and stores the posterior results in a time series. S2 utilizes the sparsity of non-space frames in the CTC algorithm output, sets a space threshold, filters out frames with a space posterior probability exceeding the space threshold, and generates a compressed posterior result. S3 scores the user-defined wake-up parameters based on the compressed posterior result to obtain the false wake-up index; Based on the false wake-up index, determine the custom wake-up word threshold.
2. The method according to claim 1, characterized in that, The step of determining the custom wake-up word threshold based on the false wake-up index includes: The false wake-up indicators include at least one of the following: custom wake word, number of false wake-ups; Based on the number of false wake-ups in the false wake-up index, the number of false wake-ups for the new custom wake-up word is calculated using the stored posterior results in the test set.
3. The method according to claim 2, characterized in that, The step of determining the custom wake-up word threshold based on the false wake-up index further includes: Based on the number of false wake-ups in the false wake-up metric, the custom wake-up word threshold is selected; The user will be notified of a message indicating that the custom wake word threshold has been successfully selected.
4. A method for determining a custom wake word threshold, characterized in that, For use on the server side, the method includes: Receive audio data from a preset test set; The user-defined wake-up parameters are scored based on the posterior results generated by the test set of the preset acoustic model to obtain the false wake-up index, wherein the preset acoustic model is used to store the forward calculation in the posterior. The specific steps for scoring user-defined wake-up parameters based on the posterior results generated from the test set of the preset acoustic model to obtain the false wake-up index include: S1 pre-acquires speech data, trains the preset acoustic model using the CTC loss function to perform forward computation, generates posterior results, and stores the posterior results in a time series. S2 utilizes the sparsity of non-space frames in the CTC algorithm output, sets a space threshold, filters out frames with a space posterior probability exceeding the space threshold, and generates a compressed posterior result. S3 scores the user-defined wake-up parameters based on the compressed posterior result to obtain the false wake-up index; Based on the false wake-up index, a custom wake-up word threshold is determined and sent to the user terminal as described in claim 1.
5. An apparatus for determining a custom wake word threshold, characterized in that, include: The preprocessing module is used to score user-defined wake-up parameters based on the posterior results generated by the test set of the preset acoustic model to obtain a false wake-up index, wherein the preset acoustic model is used to store the forward calculation in the posterior. The specific steps for scoring user-defined wake-up parameters based on the posterior results generated from the test set of the preset acoustic model to obtain the false wake-up index include: Speech data is acquired in advance, and the preset acoustic model is trained using the CTC loss function to perform forward computation, generate posterior results, and store the posterior results in a time series. By utilizing the sparsity of non-space frames in the CTC algorithm output, a space threshold is set, and frames with a space posterior probability exceeding the space threshold are filtered out to generate a compressed posterior result. The calculation module is used to score the user-defined wake-up parameters based on the posterior results generated by the test set to obtain the false wake-up index; The threshold determination module is used to determine the custom wake-up word threshold based on the false wake-up index.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program is configured to perform the method as described in any one of claims 1 to 4 when executed.
7. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Customizable voice wake-up method and system
CN106098059A
Awakening method and device of intelligent equipment, electronic equipment and medium
CN111554288A
Voice interaction method, electronic equipment and readable storage medium
CN113555016A