Voice wake-up method, electronic equipment and computer storage medium

By integrating Pinyin and energy confidence into voice wake-up technology to calculate wake-up confidence and using multiple threshold groups to determine the wake-up status, the problem of false wake-up and wake-up failure caused by the single wake-up threshold in the existing technology is solved, achieving a higher wake-up success rate and a lower false wake-up rate.

CN121506104APending Publication Date: 2026-02-10MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511896669.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-15
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In existing voice wake-up technologies, the single wake-up threshold leads to a high false wake-up rate or a high wake-up failure rate, making it difficult to balance the wake-up success rate and the false wake-up rate.

Method used

By obtaining the wake-up confidence of the wake-up flag frame in the speech frame sequence, and combining it with the pinyin confidence and energy confidence for fusion calculation, multiple threshold groups are set to determine whether the wake-up is successful or enters the pre-wake state, thereby reducing the false wake-up rate and wake-up failure rate.

Benefits of technology

It effectively reduces the false wake-up rate and wake-up failure rate, improves user experience, and balances the accuracy and sensitivity of wake-up.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506104A_ABST
    Figure CN121506104A_ABST
Patent Text Reader

Abstract

The invention provides a voice wake-up method, electronic equipment and a computer storage medium. The voice wake-up method comprises the following steps: acquiring a wake-up confidence coefficient of a wake-up flag frame in a voice frame sequence; determining whether wake-up succeeds based on the wake-up confidence; wherein the wake-up confidence of the voice frame of the voice frame sequence is obtained by calculating the pinyin confidence and the energy confidence of the voice frame. The wake-up failure rate can be effectively reduced, the false wake-up rate is reduced, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the voice wake-up technical field, in particular to a voice wake-up method, an electronic device and a computer storage medium. BACKGROUND

[0002] In the prior art, the voice wake-up technology is the core entrance for the intelligent device to realize natural interaction, and the response sensitivity and the recognition accuracy directly determine the pros and cons of the user experience. In the prior art, since the wake-up threshold is usually a single threshold, if the wake-up threshold is reduced in order to improve the wake-up success rate, the false wake-up situation will occur frequently; if the wake-up threshold is increased in order to reduce the false wake-up rate, the wake-up difficulty will be increased, and the wake-up failure rate is high. How to improve the voice wake-up method to reduce the false wake-up rate and the wake-up failure rate is one of the important problems to be solved by the technical personnel in the field. SUMMARY

[0003] The present application provides a voice wake-up method, an electronic device and a computer storage medium, which can reduce the false wake-up rate and the wake-up failure rate.

[0004] To solve the above technical problems, the present application provides a voice wake-up method, which comprises: obtaining a wake-up confidence of a wake-up flag frame in a voice frame sequence; determining whether the wake-up is successful based on the wake-up confidence; wherein the wake-up confidence of a voice frame of the voice frame sequence is calculated from the pinyin confidence and the energy confidence of the voice frame.

[0005] To solve the above technical problems, the present application provides an electronic device, which comprises a memory and a processor, the memory is used to store program data, the program data can be executed by the processor to realize the voice wake-up method described above.

[0006] To solve the above technical problems, the present application further provides a computer storage medium. The computer storage medium stores program instructions, and the program instructions are executed by the processor to realize the voice wake-up method described above.

[0007] To solve the above technical problems, the present application further provides a computer program product. The computer program product comprises computer program instructions, and the computer program instructions enable the computer to realize the voice wake-up method described above.

[0008] The beneficial effects of this application are as follows: The voice wake-up method of this application includes: obtaining the wake-up confidence of the wake-up flag frame in the voice frame sequence; determining whether wake-up is successful based on the wake-up confidence; wherein, the wake-up confidence of the voice frame in the voice frame sequence is calculated from the pinyin confidence and energy confidence of the voice frame. By using the above method to perform fusion calculation using the pinyin confidence and energy confidence of the voice frame to obtain the wake-up confidence, it can take into account both the high wake-up accuracy of pinyin confidence and the high wake-up success rate of energy confidence. Therefore, the above method can effectively reduce the false wake-up rate, reduce the wake-up failure rate, and improve the user experience. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a flowchart illustrating an embodiment of the voice wake-up method of this application; Figure 2 yes Figure 1 The steps in this embodiment are a detailed flowchart of S14; Figure 3 yes Figure 2 A detailed flowchart of step S23 in the embodiment; Figure 4 yes Figure 3 A detailed flowchart of step S34 in the embodiment is shown below; Figure 5 This is a flowchart illustrating another embodiment of the voice wake-up method of this application; Figure 6 yes Figure 5 A schematic diagram of the specific process of step S52 in the embodiment; Figure 7 This is a flowchart illustrating yet another embodiment of the voice wake-up method of this application; Figure 8 This is a flowchart illustrating another embodiment of the voice wake-up method of this application; Figure 9 This is a flowchart illustrating another embodiment of the voice wake-up method of this application; Figure 10 This is a flowchart illustrating another embodiment of the voice wake-up method of this application; Figure 11 yes Figure 10 A schematic diagram of the specific process of step S102 in the embodiment; Figure 12 This is a flowchart illustrating another embodiment of the voice wake-up method of this application; Figure 13 This is a flowchart illustrating another embodiment of the voice wake-up method of this application; Figure 14 This is a flowchart illustrating another embodiment of the voice wake-up method of this application; Figure 15 This is a flowchart illustrating another embodiment of the voice wake-up method of this application; Figure 16 This is a schematic diagram of the structure of an embodiment of the computer storage medium of this application. Detailed Implementation

[0010] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0011] The terms "first," "second," etc., used in this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. It should be understood that, when used in this specification, the term "comprising" indicates the presence of the described feature, integral, step, operation, element, and / or component, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof. It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this specification, unless the context clearly indicates otherwise, the singular forms "a," "an," and "the" are intended to include the plural forms. It should also be further understood that the term "and / or," as used in this specification, refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0012] As used in this specification, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determination" or "if [the described condition or event] is detected" may be interpreted, depending on the context, as "once determination," "in response to determination," "once [the described condition or event] is detected," or "in response to detection of [the described condition or event]."

[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0014] In existing technologies, voice wake-up technology serves as the core entry point for natural interaction in smart devices, and its response sensitivity and recognition accuracy directly determine the quality of the user experience. However, existing technologies typically use a single wake-up threshold. Lowering the threshold to increase the success rate leads to frequent false wake-ups; conversely, raising the threshold to reduce false wake-ups increases the difficulty of wake-up and results in a high failure rate. Therefore, improving voice wake-up methods to reduce false wake-ups and failure rates is a crucial problem that urgently needs to be solved by those skilled in the art.

[0015] This application first proposes a voice wake-up method, such as Figure 1 As shown, the voice wake-up method includes steps S11 to S14.

[0016] Step S11: Obtain the wake-up confidence of the wake-up flag frame in the speech frame sequence.

[0017] The speech frame sequence consists of multiple speech frames arranged in time sequence.

[0018] The wake-up flag frame is the position in the speech frame sequence where a candidate wake-up word appears. Typically, the last speech frame (i.e., the tail frame) of a candidate wake-up word is the wake-up flag frame.

[0019] In some application scenarios, wake-up models such as wake-up algorithm libraries can be used to analyze speech frame sequences and output wake-up confidence frame sequences corresponding to the speech frame sequences.

[0020] The higher the wake-up confidence of a speech frame, the higher the probability of a wake-up word appearing. Wake-up confidence can be used to determine whether wake-up was successful. Typically, among multiple speech frames corresponding to a candidate wake-up word, the last frame of the candidate wake-up word has the highest wake-up confidence; therefore, the wake-up flag frame is usually the speech frame corresponding to the last frame of the candidate wake-up word.

[0021] In some embodiments, prior to step S11, the voice wake-up method further includes: determining the corresponding voice frame as a wake-up flag frame in response to a wake-up confidence level of a voice frame in the voice frame sequence being greater than or equal to a candidate threshold. This method can determine the location of the wake-up standard frame.

[0022] Step S12: Obtain the maximum threshold in the preset threshold group.

[0023] The preset threshold group contains multiple thresholds. Under different conditions, different thresholds within the preset threshold group can be selected as wake-up thresholds and compared with wake-up confidence to determine whether wake-up is successful. The preset threshold group can be a set of thresholds expressed by a threshold formula, or a set of thresholds stored in a pre-stored database; there is no limitation.

[0024] Step S13: Wake-up is successful in response to a wake-up confidence level greater than or equal to the maximum threshold.

[0025] When a user's initial wake-up confidence is high, the wake-up can be successful directly, improving response efficiency. Setting the maximum threshold to the default threshold for successful wake-up can effectively limit the probability of false wake-ups and reduce the false wake-up rate.

[0026] Step S14: In response to the wake-up confidence being less than the maximum threshold, determine whether to enter the first pre-wake-up state based on the wake-up confidence and the preset threshold group.

[0027] Other thresholds in the preset threshold group can be used to determine whether to enter the first pre-wake state.

[0028] Setting a first pre-wake state prevents a direct transition to a wake-up failure state when the initial wake-up attempt fails, providing a margin of error for low-energy wake-up scenarios and reducing the difficulty of wake-up for users. Typically, if the device doesn't respond during the initial wake-up attempt, the user will attempt a second wake-up within a short period. Specifically, if the wake-up confidence of the user's initial wake-up attempt doesn't reach the maximum threshold, it's considered a valid wake-up attempt, thus determining whether the conditions for entering the first pre-wake state are met. If the conditions for entering the first pre-wake state are met, other judgment logic can be used within this state to reduce the difficulty of a second wake-up attempt within a certain timeframe. For example, using other thresholds within the wake-up threshold group to determine whether the user's subsequent second wake-up attempt was successful can reduce the difficulty of the second wake-up attempt; or, using other judgment logic or other voice features to determine whether the user's subsequent second wake-up attempt was successful can effectively reduce the wake-up difficulty.

[0029] The settings in steps S11 to S14 can limit the probability of successful false wake-up under false wake-up conditions using a maximum threshold, thereby reducing the false wake-up rate. Furthermore, when the wake-up confidence does not meet the maximum threshold requirement, it does not directly determine wake-up failure, but instead determines whether to enter the first pre-wake state based on a preset threshold group and wake-up confidence, effectively reducing the wake-up failure rate under low-energy wake-up conditions. Therefore, this embodiment can effectively reduce the wake-up failure rate and the false wake-up rate, improving the user experience.

[0030] In some embodiments, the wake-up confidence of a speech frame can be obtained based on the pinyin confidence and energy confidence of the speech frame. For example, when a pinyin confidence frame sequence and an energy confidence frame sequence corresponding to the speech frame sequence are available, the pinyin confidence and energy confidence can be fused to calculate the wake-up confidence. Specific implementation methods can be found in the embodiments related to expressions 1-1 or 1-2 below, and will not be repeated here. For another example, when only a pinyin confidence frame sequence corresponding to the speech frame sequence is available, the pinyin confidence of the speech frame can be used as the wake-up confidence of the speech frame; similarly, when only an energy confidence frame sequence corresponding to the speech frame sequence is available, the energy confidence of the speech frame can be used as the wake-up confidence of the speech frame.

[0031] In some embodiments, the speech frame sequence and the corresponding wake-up confidence frame sequence are from a third-party wake-up algorithm library; or the speech frame sequence and the corresponding pinyin confidence frame sequence and energy confidence frame sequence are from a third-party wake-up algorithm library, and the corresponding wake-up confidence frame sequence is calculated through the embodiments related to Expression 1-1 or Expression 1-2 below; or the speech frame sequence and the corresponding pinyin score frame sequence and energy confidence frame sequence are from a third-party wake-up algorithm library, and the wake-up confidence frame sequence is calculated through... Figure 10 , Figure 11 The illustrated embodiment obtains a pinyin confidence frame sequence, and then obtains the corresponding wake-up confidence frame sequence. In some embodiments, the speech frame sequence comes from a third-party wake-up algorithm library, and the pinyin confidence frame sequence or energy confidence frame sequence comes from a third-party wake-up algorithm library, thus determining the pinyin confidence frame sequence or energy confidence frame sequence as the wake-up confidence frame sequence.

[0032] In one application scenario, a computer storage medium or electronic product storing the voice wake-up method described in any embodiment of this application can interface with any third-party wake-up algorithm library. The voice data output by the third-party wake-up algorithm library includes at least a voice frame sequence and a confidence frame sequence. This application can optimize the voice data output by the third-party wake-up algorithm library to improve its wake-up accuracy and wake-up success rate.

[0033] In some embodiments, step S14 can be performed as follows: Figure 2 The method shown is implemented in the following way, specifically including steps S21 to S22.

[0034] Step S21: In response to the wake-up confidence being less than the maximum threshold, determine the wake-up segment where the candidate wake-up word corresponding to the wake-up flag frame is located.

[0035] Further analysis of the corresponding features of the wake-up segment can be performed to determine whether there is a possibility of successful wake-up. For example, if so, step S23 can be executed to compare the wake-up confidence with the energy threshold less than the maximum threshold in the preset threshold group to determine whether a successful wake-up is possible. This avoids directly entering the first pre-wake state, provides a margin of error for the candidate wake-up word, and reduces the difficulty of wake-up for the user. If not, step S22 can be executed.

[0036] Step S22: In response to the failure of the wake-up segment determination, determine whether to enter the first pre-wake-up state based on the wake-up confidence and the preset threshold group.

[0037] The decision to enter the first pre-wake state is only made when the wake-up segment cannot be determined. If the wake-up segment can be determined, other judgment logic can be executed in subsequent steps to utilize the relevant characteristics of the wake-up segment to determine whether a direct wake-up can be successful. That is, the settings of steps S21 to S22 can prevent the "determine whether to enter the first pre-wake state" step from being executed directly when the wake-up confidence is less than the maximum threshold, providing fault tolerance for the candidate wake-up word and reducing the difficulty of the user's first wake-up.

[0038] In some embodiments, Figure 2 The method shown may further include step S23.

[0039] Step S23: In response to the successful determination of the wake-up segment, determine whether the wake-up was successful based on the wake-up confidence, the wake-up segment, and the preset threshold group.

[0040] In some embodiments, step S21 can be implemented by steps B1 to B2.

[0041] Step B1: In response to the wake-up confidence being less than the maximum threshold, obtain the pinyin confidence of the wake-up flag frame.

[0042] The pinyin confidence score of a speech frame is calculated using the pinyin score of the speech frame. The pinyin score represents the pinyin matching probability of the speech frame. In some embodiments, the method for obtaining the pinyin confidence score from the pinyin score can be referred to... Figure 11 The illustrated embodiment.

[0043] Step B2: In response to the successful acquisition of Pinyin confidence, determine the wake-up segment where the candidate wake-up word is located based on the Pinyin confidence.

[0044] Wake words are typically composed of specific pinyin combinations. A high pinyin confidence level indicates a high degree of matching between the candidate wake word and the original wake word's pinyin. Based on pinyin confidence level, the wake segment containing the candidate wake word can be accurately located. In one application scenario, please refer to the following text. Figure 11In the illustrated embodiment, the position of the candidate wake-up word is obtained by inversely deducing the method of calculating the pinyin confidence frame sequence from the pinyin score frame sequence.

[0045] For example, each speech frame has a corresponding pinyin confidence and pinyin score; the wake-up word is "Xiaomei Xiaomei", which includes a total of four wake-up characters, and the last frame of the character "mei" is the wake-up flag frame; each wake-up character corresponds to a group of candidate character frame sequences. For example, each group of candidate character frame sequences may include 10 speech frames; the speech frame with the highest pinyin score in each group of candidate character frame sequences is usually near the last frame of the group of candidate character frame sequences. Therefore, referring to Figure 11 the calculation method of the illustrated embodiment, based on the pinyin confidence and pinyin score of the last frame of the character "mei", the position of the last frame of each wake-up character can be estimated in turn; solve the time difference δt between the position of the last frame of the first character "xiao" and the wake-up flag frame, and 4δt / 3 is the duration of the wake-up segment; based on the duration of the wake-up segment and the position of the wake-up flag frame, the position of the wake-up segment can be determined. For example, if the total duration of the speech frame sequence is known to be 2000 ms, and if it has been determined that 1000 - 1900 ms is the wake-up segment, then it is determined that 0 - 1000 ms is the leading segment and 1900 - 2000 ms is the subsequent segment.

[0046] In some embodiments, when the acquisition of the pinyin confidence fails, it is also possible to determine whether to enter the first pre-wake-up state based on the wake-up confidence and a preset threshold group. The specific implementation manner of determining whether to enter the first pre-wake-up state based on the wake-up confidence and the preset threshold group can refer to step S22.

[0047] In some embodiments, step S23 can be implemented in the manner as Figure 3 illustrated, and specifically includes steps S31 to S34.

[0048] Step S31: Based on the wake-up segment, determine the leading segment before the wake-up segment and the subsequent segment after the wake-up segment in the speech frame sequence.

[0049] For example, if the total duration of the speech frame sequence is known to be 2000 ms, and if it has been determined that 1000 - 1900 ms is the wake-up segment, then it is determined that 0 - 1000 ms is the leading segment and 1900 - 2000 ms is the subsequent segment.

[0050] Step S32: Calculate the first average decibel value of the leading segment, the second average decibel value of the wake-up segment, and the third average decibel value of the subsequent segment.

[0051] The average decibel value can represent the speech energy of the corresponding speech segment to a certain extent.

[0052] Step S33: Calculate the first decibel difference between the second average decibel value and the first average decibel value, and the second decibel difference between the second average decibel value and the third average decibel value.

[0053] For example, if the second average decibel value is 50 decibels, the first average decibel value is 45 decibels, and the third average decibel value is 40 decibels, then the first decibel difference is 5 decibels and the second decibel difference is 10 decibels.

[0054] Step S34: In response to the fact that both the first decibel difference and the second decibel difference are greater than or equal to the decibel difference threshold, determine whether the wake-up is successful based on the wake-up confidence and the preset threshold group.

[0055] In some applications, the decibel difference threshold is 3 dB. It is also possible to set the decibel difference to 2 dB, 4 dB, 5 dB, 6 dB, etc.

[0056] By comparing the average decibel values, the difference in speech energy between the wake-up segment and the preceding and following segments can be obtained. Typically, when a user issues a wake-up command, the speech energy of the wake-up word is significantly higher than the ambient noise. For example, if both the first and second decibel differences are greater than or equal to a decibel difference threshold, it can be determined that the conditions for entering the energy judgment are met. Therefore, the wake-up confidence can be compared with the energy threshold less than the maximum threshold in a preset threshold group to determine whether wake-up is successful.

[0057] Steps S31 to S34 determine whether the speech energy of the wake-up segment is significantly higher than that of the preceding and subsequent segments by comparing the average decibel values. If it is determined to be higher, it indicates that the probability of the candidate wake-up word being the wake-up word is relatively high. Therefore, subsequent steps can be executed to re-determine whether the wake-up was successful by using wake-up confidence and a preset threshold group. This setting can effectively improve the accuracy of successful wake-up.

[0058] In some embodiments, Figure 3 The method shown may further include step S35.

[0059] Step S35: In response to the first decibel difference or the second decibel difference being less than the decibel difference threshold, determine whether to enter the first pre-wake state based on the wake-up confidence and the preset threshold group.

[0060] If the voice energy of the wake-up segment is not significantly more prominent than that of the preceding and following segments, it indicates that it does not meet the condition for executing the step of "determining whether to successfully wake up based on wake-up confidence and a preset threshold group". In this case, although the candidate wake-up word does not have the possibility of directly waking up successfully, it still has a certain possibility of being a valid wake-up attempt by the user. Therefore, it can be determined whether it meets the condition for entering the first pre-wake state, and the step of "determining whether to enter the first pre-wake state based on wake-up confidence and a preset threshold group" can be executed.

[0061] In the above settings, if the decibel threshold condition is not met, the wake-up is not directly considered a failure. Instead, it is judged whether the conditions for entering the first pre-wake state are met. If the conditions for entering the first pre-wake state are met, other judgment logic can be used in the first pre-wake state to reduce the difficulty of the user's second wake-up within a certain period of time. For example, in the first pre-wake state, other thresholds within the wake-up threshold group can be used to judge whether the user's subsequent second wake-up is "successful," reducing the difficulty of the second wake-up; or in the first pre-wake state, other judgment logic or other voice features can be used to judge whether the user's subsequent second wake-up is "successful," which can effectively reduce the wake-up difficulty.

[0062] In some embodiments, step S34 can be performed as follows: Figure 4 The method shown is implemented in the following way, specifically including steps S41 to S42.

[0063] Step S41: Obtain the energy threshold of the preset threshold group, where the energy threshold is less than the maximum threshold.

[0064] Step S42: In response to the wake-up confidence being greater than or equal to the energy threshold, determine that the wake-up was successful.

[0065] Specifically, if the second average decibel value of the wake-up segment containing the candidate wake-up word is significantly greater than the first average decibel value of the preceding segment and the third average decibel value of the following segment, the probability of determining the candidate wake-up word as a wake-up word is relatively high. Therefore, the wake-up threshold is lowered to the energy threshold. The wake-up confidence is compared with the energy threshold. If the judgment criteria are met, the wake-up is determined to be successful.

[0066] The above settings can reduce the difficulty of initial wake-up for users by utilizing pinyin confidence, decibel value, and energy threshold. When the user's wake-up confidence is low and does not meet the maximum threshold condition, the probability of a candidate wake-up word being a wake-up word can be determined from the perspective of the decibel energy of the wake-up word using pinyin confidence and decibel value. After meeting the decibel energy requirement, the wake-up threshold can be reduced, and the "whether wake-up was successful" judgment can be performed again, which can effectively reduce the difficulty of initial wake-up for users.

[0067] In some embodiments, Figure 4 The method shown may further include step S43.

[0068] Step S43: In response to the wake-up confidence being less than the energy threshold, determine whether to enter the first pre-wake-up state based on the wake-up confidence and the preset threshold group.

[0069] If the wake-up confidence is less than the energy threshold, the wake-up confidence corresponding to the wake-up flag frame is determined to be low, and the conditions for executing "wake-up successful" are not met, meaning there is no possibility of direct wake-up success. In this case, although the candidate wake-up word does not have the possibility of direct wake-up success, it still has a certain possibility of being a valid wake-up attempt by the user. Therefore, it can be determined whether it meets the conditions for entering the first pre-wake state, and the step of "determining whether to enter the first pre-wake state based on wake-up confidence and preset threshold group" is executed.

[0070] In the first pre-wake state, other judgment logic can be used to reduce the difficulty of waking up the user for a second time within a certain period of time. For example, in the first pre-wake state, other thresholds within the wake-up threshold group can be used to judge whether the user's subsequent second wake-up is successful, thus reducing the difficulty of second wake-up; or in the first pre-wake state, other judgment logic or other voice features can be used to judge whether the user's subsequent second wake-up is successful, which can effectively reduce the difficulty of waking up.

[0071] In some embodiments, the step of determining whether to enter the first pre-wake state based on wake-up confidence and a preset threshold group in any of steps S14, S22, S35, and S43 can be achieved by, for example... Figure 5 The method shown is implemented in the following way, specifically including steps S51 to S52.

[0072] Step S51: Obtain the minimum threshold of the preset threshold group.

[0073] Step S52: In response to the wake-up confidence being greater than or equal to the minimum threshold, enter the first pre-wake-up state and continue for a first preset duration.

[0074] When the wake-up confidence is greater than or equal to the minimum threshold, it means that although the current wake-up confidence has not reached the higher requirements, it meets the conditions for entering the first pre-wake state. It is determined that the candidate wake-up word has a certain probability of being a valid wake-up attempt by the user, and therefore enters the first pre-wake state.

[0075] In the first pre-wake state, the device maintains a relatively sensitive state, such as lowering the wake-up threshold, to wait for the user to wake it up a second time within a certain period of time. If the user fails to wake up successfully on the first attempt due to unclear pronunciation, low volume, or other reasons resulting in low wake-up confidence, it does not mean that the user did not intend to wake up. Therefore, setting a first pre-wake state can reduce the difficulty of waking up the user.

[0076] After the first preset time period has elapsed, the device will exit the first pre-wake state and revert to the normal wake-up judgment logic. That is, it will again use the initially set maximum threshold and other conditions to determine whether wake-up was successful. This setting reduces the difficulty of waking up for users and avoids wasting resources and the risk of false wake-ups caused by the device remaining in the first pre-wake state for an extended period.

[0077] In some embodiments, Figure 5 The method shown may further include step S53.

[0078] Step S53: Wake-up fails in response to wake-up confidence being less than the minimum threshold.

[0079] A wake-up confidence score below the minimum threshold indicates an extremely low score, failing to meet the conditions for entering the first pre-wake state. Therefore, the candidate wake-up word is deemed an invalid wake-up attempt, and the wake-up attempt is directly considered a failure. This setting prevents the device from unnecessarily entering the first pre-wake state when the wake-up confidence is extremely low, reducing device resource consumption and invalid responses.

[0080] If the wake-up fails or the first pre-wake-up state is exited, step S11 can be executed again to perform the next wake-up judgment when the wake-up flag frame appears next time.

[0081] In other embodiments, other features such as speech duration, intonation, and pinyin confidence can be further combined to determine whether to enter the first pre-wake state.

[0082] In some embodiments, step S52 can be achieved through... Figure 6 The method shown is used to achieve this, specifically including steps S61 to S62.

[0083] Step S61: In response to a wake-up confidence level greater than or equal to the minimum threshold, obtain the normal threshold of the preset threshold group. The normal threshold is less than the maximum threshold and greater than the minimum threshold.

[0084] For example, in one application scenario, the maximum threshold is 0.85, the normal threshold is 0.70, and the minimum threshold is 0.65.

[0085] Step S62: In response to the wake-up confidence of any voice frame within the first preset duration after the wake-up flag frame being greater than or equal to the normal threshold, wake-up is successful.

[0086] In some embodiments, the first preset duration is 5s, 6s, 7s, 8s, 9s, 10s, 11s, 12s, 13s, or 14s, etc. In some embodiments, the first preset duration can be adjusted based on historical wake-up data. In some embodiments, the first preset duration is set to 10s to accommodate common situations.

[0087] For example, in a quiet environment, a user whispers a wake-up word. Due to volume issues, the wake-up confidence level doesn't reach the maximum threshold, but it's greater than or equal to the minimum threshold. The system then enters the first pre-wake state. During this state, the system lowers the requirements for a second wake-up, for example, by using a lower threshold within the wake-up threshold group. Assuming the maximum threshold is 0.85, the normal threshold is 0.7, the minimum threshold is 0.65, and the first preset duration is 10 seconds; the user's initial wake-up confidence level is 0.7, meeting the conditions for entering the first pre-wake state. After entering this state, the system modifies the wake-up threshold to the normal threshold of 0.7. When the user issues a wake-up command again within the next 10 seconds, it will be easier to successfully wake up.

[0088] In the first pre-wake state, a normal threshold is used as the new wake-up threshold to determine whether wake-up was successful, and the system waits for the user to attempt wake-up again within a first preset time period. When the wake-up confidence of a certain voice frame is detected to be greater than or equal to the normal threshold within this time period, it indicates that the user's second wake-up meets the conditions for successful wake-up, and the wake-up is determined to be successful. This mechanism can take into account various unfavorable situations that may occur during the user's first wake-up, and by providing the user with more wake-up opportunities through the first pre-wake state and threshold adjustment, it can effectively reduce the difficulty of repeated wake-up attempts.

[0089] In some embodiments, Figure 6 The method shown may further include step S63.

[0090] Step S63: In response to the wake-up confidence of any voice frame within a first preset time period after the wake-up flag frame being less than the normal threshold, exit the first pre-wake-up state and wake-up fails.

[0091] This setting reduces resource waste on the device. Instead of waiting indefinitely in the first pre-wake state, monitoring is only performed for a first preset duration. If a wake-up confidence level meeting a normal threshold is not detected after this time, the device will stop waiting, thus reducing resource waste caused by prolonged periods in the first pre-wake state.

[0092] In some embodiments, utilizing Figure 6 Improved method shown Figure 4 In step S43, the preset threshold group includes an energy threshold, which is greater than the normal threshold and less than the maximum threshold.

[0093] In some embodiments, the wake-up confidence value ranges from [0,1], and the maximum threshold value within the preset threshold group ranges from [0.85,0.90], for example, the maximum threshold is 0.85, 0.86, 0.87, 0.875, 0.88, 0.89, or 0.90. The energy threshold value within the preset threshold group ranges from [0.78,0.83], for example, the energy threshold is 0.78, 0.79, 0.80, 0.81, 0.82, or 0.83. The normal threshold value within the preset threshold group ranges from [0.65,0.75], for example, the normal threshold is 0.65, 0.66, 0.67, 0.68, 0.69, 0.70, 0.71, 0.72, 0.73, or 0.75. The minimum threshold value within the preset threshold group ranges from [0.60, 0.65], for example, the minimum threshold is 0.60, 0.61, 0.62, 0.63, 0.64 or 0.65, etc.

[0094] The following example illustrates the situation with a maximum threshold of 0.85, an energy threshold of 0.80, a normal threshold of 0.70, a minimum threshold of 0.65, a decibel difference threshold of 3 dB, and a first preset duration of 10 seconds.

[0095] In a wake-up attempt, the user's initial wake-up confidence score is 0.9, which is greater than the maximum threshold of 0.85, thus confirming a successful wake-up.

[0096] In a wake-up attempt, the user's initial wake-up confidence was 0.80. The first decibel difference between the wake-up segment and the preamble, and the second decibel difference between the wake-up segment and the subsequent segment, were both 5 dB. Since 0.80 is less than the maximum threshold of 0.85, and both the first and second decibel differences are greater than 3 dB, the wake-up confidence was compared to the energy threshold. Because 0.80 equals the energy threshold of 0.80, the wake-up was successful.

[0097] In a wake-up attempt, the user's initial wake-up confidence is 0.70. The first decibel difference between the wake-up segment and the preamble, and the second decibel difference between the wake-up segment and the subsequent segment, are both 5 dB. Because 0.70 is less than the maximum threshold of 0.85, and both the first and second decibel differences are greater than 3 dB, the wake-up confidence is compared to the energy threshold. Since 0.70 is less than the energy threshold of 0.80, the wake-up confidence is compared to the minimum threshold to determine whether to enter the first pre-wake state. Since 0.70 is greater than the minimum threshold of 0.65, the user enters the first pre-wake state. If the user attempts a second wake-up within 10 seconds, the wake-up confidence for the second wake-up is 0.72. Since 0.72 is greater than the normal threshold of 0.70, the wake-up is successful.

[0098] In a wake-up attempt, the user's initial wake-up confidence is 0.60. The first decibel difference between the wake-up segment and the preamble, and the second decibel difference between the wake-up segment and the subsequent segment, are both 5 dB. Since 0.60 is less than the maximum threshold of 0.85, and both the first and second decibel differences are greater than 3 dB, the wake-up confidence is compared to an energy threshold. Because 0.60 is less than the energy threshold of 0.80, the wake-up confidence is compared to a minimum threshold to determine whether to enter the first pre-wake state. Since 0.60 is less than the minimum threshold of 0.65, the wake-up attempt fails.

[0099] In a wake-up attempt, the user's initial wake-up confidence is 0.60. The first decibel difference between the wake-up segment and the preamble, and the second decibel difference between the wake-up segment and the subsequent segment, are both 2 dB. Because 0.60 is less than the maximum threshold of 0.85, and both the first and second decibel differences are less than 3 dB, the wake-up confidence is compared with the minimum threshold to determine whether to enter the first pre-wake state. Since 0.60 is less than the minimum threshold of 0.65, the wake-up attempt fails.

[0100] In a wake-up attempt, the user's initial wake-up confidence score is 0.70. Because the pinyin confidence score acquisition failed, the decibel difference between the wake-up segment and the preceding and subsequent segments cannot be determined. Since 0.70 is less than the maximum threshold of 0.85, and the pinyin confidence score acquisition failed, the wake-up confidence score is compared with the minimum threshold to determine whether to enter the first pre-wake state. Because 0.70 is greater than the minimum threshold of 0.65, the user enters the first pre-wake state. If the user attempts a second wake-up within 10 seconds, the wake-up confidence score for the second wake-up is 0.72. Since 0.72 is greater than the normal threshold of 0.70, the wake-up is successful.

[0101] To reduce the difficulty of waking up multiple times, any embodiment of this application can be improved to optimize the wake-up threshold after successful wake-up and improve user experience. For example, after step S13, step S42 or step S62 is completed, the voice wake-up method may further include step A1.

[0102] Step A1: In response to successful wake-up, enter the second pre-wake-up state and continue for the second preset duration.

[0103] In the second pre-wake state, the device maintains a relatively sensitive state, such as lowering the wake-up threshold, to wait for the user to wake it up again within a certain period of time. Once the user has successfully woken up the device, there may be further voice command requests. By entering the second pre-wake state, the device can respond to subsequent possible wake-up operations with more lenient wake-up conditions, reducing the difficulty for the user to wake it up again in a short period. For example, after the user successfully wakes up the device and performs a query, if they want to ask for other relevant information, the device in the second pre-wake state can be woken up more easily without the user needing to increase their volume or speak more clearly.

[0104] In some embodiments, a suitable second preset duration can be determined based on historical wake-up data. After the second preset duration is exceeded, the device will exit the second pre-wake state and revert to normal wake-up logic. That is, it will again use the initially set maximum threshold and other conditions to determine whether wake-up was successful. This setting satisfies users' needs for secondary wake-ups within a short period while reducing resource waste and the risk of false wake-ups caused by the device remaining in the second pre-wake state for extended periods.

[0105] In some embodiments, step A1 can be achieved through... Figure 7 The steps shown are implemented, specifically including steps S71 to S72.

[0106] Step S71: In response to a successful wake-up, obtain the normal threshold of the preset threshold group. The normal threshold is less than the maximum threshold and greater than the minimum threshold.

[0107] Step S72: In response to the wake-up confidence of any voice frame within the second preset time after the wake-up flag frame being greater than or equal to the normal threshold, wake-up is successful.

[0108] In some embodiments, the second preset duration is 15s, 20s, 25s, 30s, 35s, or 40s, etc. In some embodiments, the second preset duration can be adjusted based on historical wake-up data. In some embodiments, the second preset duration is set to 30s to accommodate common situations.

[0109] Taking a maximum threshold of 0.85, a normal threshold of 0.70, and a second preset duration of 30 seconds as an example. For instance, in one application scenario, when a user first wakes up the device, the wake-up confidence is 0.9, which is greater than the maximum threshold of 0.85, indicating a successful wake-up. Twenty seconds after the wake-up flag frame appears, the user wakes up the device again. Due to environmental noise or the user's pronunciation being somewhat random, the wake-up confidence of this wake-up flag frame is 0.72. Since this wake-up flag frame is within 30 seconds of the previous successful wake-up flag frame, the wake-up confidence of this wake-up flag frame is compared with the normal threshold. Because 0.72 is greater than 0.7, this wake-up is still successful. This setting provides users with more lenient wake-up conditions within a certain timeframe, reducing the difficulty of waking up the device again after a period of successful wake-up.

[0110] In some embodiments, Figure 7 The illustrated embodiment may further include step S73.

[0111] Step S73: In response to the wake-up confidence of any voice frame within a second preset time period after the wake-up flag frame being less than the normal threshold, exit the second pre-wake-up state.

[0112] After the second preset time period has elapsed, the device will exit the second pre-wake state and revert to the normal wake-up logic. That is, it will again use the initially set maximum threshold and other conditions to determine whether wake-up was successful. This setting satisfies users' needs for a second wake-up within a short period while reducing resource waste and the risk of false wake-ups caused by the device remaining in the second pre-wake state for extended periods.

[0113] In some embodiments, utilizing Figure 7 Improved method shown Figure 4 In the embodiment shown, the preset threshold group includes an energy threshold, which is greater than the normal threshold and less than the maximum threshold.

[0114] The following example illustrates the setting with a maximum threshold of 0.85, an energy threshold of 0.80, a normal threshold of 0.70, a minimum threshold of 0.65, a first preset duration of 10 seconds, and a second preset duration of 30 seconds.

[0115] In a wake-up attempt, the user's initial wake-up confidence score was 0.9, which is greater than the maximum threshold of 0.85, thus confirming a successful wake-up. Twenty seconds after the successful wake-up, a new wake-up flag frame was acquired, with a wake-up confidence score of 0.75. Since the new wake-up flag frame appeared after the previous wake-up flag frame and within the second preset duration of 30 seconds, its wake-up confidence score was compared to the normal threshold. Because 0.75 is greater than the normal threshold of 0.70, the wake-up was confirmed as successful.

[0116] In a wake-up attempt, the user's initial wake-up confidence score was 0.9, which is greater than the maximum threshold of 0.85, indicating a successful wake-up. Twenty seconds after the successful wake-up, a new wake-up flag frame was acquired, with a wake-up confidence score of 0.65. Since the new wake-up flag frame appeared after the previous wake-up flag frame and within the second preset duration of 30 seconds, its wake-up confidence score was compared with the normal threshold. Because 0.65 is less than the normal threshold of 0.70, this wake-up flag frame was determined to be an invalid wake-up.

[0117] In some embodiments, the second preset duration is longer than the first preset duration.

[0118] The first pre-wake state is entered when the initial wake-up attempt fails but certain conditions are met. It primarily addresses potential issues with the initial wake-up, giving the user a second chance to wake up. Its duration is relatively short, preventing the device from wasting resources by remaining in a waiting state for extended periods. The second pre-wake state is entered after a successful wake-up. At this point, the user has successfully interacted with the device and is more likely to initiate a second wake-up to issue new commands. Therefore, setting the second preset duration longer than the first better meets the user's need to use voice commands multiple times within a short period, further enhancing the user experience.

[0119] Combining the various judgment logics and speech features mentioned in previous embodiments, such as pinyin confidence and average decibel value, the voice wake-up method can form a multi-level, multi-dimensional judgment system. (See also...) Figure 15 End means completing this interaction process. After a long period of inactivity, obtain the wake-up confidence F of the current wake-up flag frame, and input... Figure 15 The judgment logic shown first compares the wake-up confidence F with the maximum threshold Th. If the wake-up confidence F ≥ the maximum threshold Th, the wake-up is successful, equivalent to step S13.

[0120] In response to the wake-up confidence F being less than the maximum threshold Th, the system uses the pinyin confidence and decibel threshold to determine whether to proceed to the energy judgment, which is equivalent to steps S31 to S33.

[0121] In response to meeting the conditions for entering the energy judgment, the wake-up confidence level F and the energy threshold Te are compared to determine whether wake-up can be successful, which is equivalent to step S34. If F ≥ energy threshold Te, wake-up is determined to be successful, which is equivalent to steps S41 to S42. If F < energy threshold Te, the wake-up confidence level F and the minimum threshold Tp are compared to determine whether the first pre-wake state can be entered, which is equivalent to step S43.

[0122] In response to the failure to meet the conditions for entering the energy judgment, the wake-up confidence F and the minimum threshold Tp are compared to determine whether the first pre-wake state can be entered, which is equivalent to step S35.

[0123] In response to a wake-up confidence level F ≥ minimum threshold Tp, the system enters the first pre-wake-up state (equivalent to steps S51 to S52). In the first pre-wake-up state, it is determined whether a new wake-up flag frame with a wake-up confidence level F ≥ normal threshold Tn exists within the next 10 seconds. If it exists, the wake-up is successful, equivalent to steps S61 to S62; if it does not exist, the current interaction ends, equivalent to step S63. When a new wake-up flag frame appears again, a new round of logical judgment will begin.

[0124] If the wake-up confidence F is less than the minimum threshold Tp, the wake-up fails, the judgment logic ends, and the current interaction ends, which is equivalent to step S53.

[0125] After a successful wake-up, the system enters the second pre-wake-up state and determines whether there is a new wake-up flag frame with a wake-up confidence F ≥ the normal threshold Tn within the next 30 seconds. If it exists, the wake-up is successful; otherwise, the interaction ends, which is equivalent to steps S71 to S73.

[0126] From comparing wake-up confidence with the initial maximum threshold, to acquiring and utilizing pinyin confidence, to comparing average decibel values ​​and setting and applying other different thresholds, the voice wake-up method can more accurately identify the user's wake-up intent in different voice environments and user wake-up situations, improve the success rate and accuracy of wake-up, and provide users with a more convenient and efficient voice wake-up experience.

[0127] In one application scenario, using the above implementation method, setting a maximum threshold can ensure that the device's wake-up threshold remains at a high level during prolonged periods of inactivity, reducing the false wake-up rate from 5-10 times / 72 hours to 0-1 times / 72 hours. In the first pre-wake state, the success rate of repeated wake-ups after an initial wake-up failure increases from 85-90% to 95%-100%. In the second pre-wake state, the success rate of secondary wake-ups increases from less than 85-90% to 95%-100%. This approach effectively reduces the false wake-up rate and improves the wake-up success rate in various scenarios, including fast-paced speech, dialectal accents, and high-noise environments.

[0128] This application further proposes a voice wake-up method, such as Figure 8 As shown, the voice wake-up method includes steps S81 to S82.

[0129] Step S81: Obtain the wake-up confidence of the wake-up flag frame in the speech frame sequence.

[0130] The speech frame sequence consists of multiple speech frames arranged in time sequence.

[0131] The wake-up flag frame is the position in the speech frame sequence where a candidate wake-up word appears. Typically, the last speech frame (i.e., the tail frame) of a candidate wake-up word is the wake-up flag frame.

[0132] In some application scenarios, wake-up models such as wake-up algorithm libraries can be used to analyze speech frame sequences and output wake-up confidence frame sequences corresponding to the speech frame sequences.

[0133] The higher the wake-up confidence of a speech frame, the higher the probability of a wake-up word appearing. Wake-up confidence can be used to determine whether wake-up was successful. Typically, among multiple speech frames corresponding to a candidate wake-up word, the last frame of the candidate wake-up word has the highest wake-up confidence; therefore, the wake-up flag frame is usually the speech frame corresponding to the last frame of the candidate wake-up word.

[0134] Step S82: Determine whether the wake-up was successful based on the wake-up confidence.

[0135] The wake-up confidence of a speech frame in a speech frame sequence is calculated from the pinyin confidence and energy confidence of the speech frame.

[0136] The pinyin confidence score of a speech frame is derived from its pinyin score. The pinyin score represents the pinyin matching probability of the corresponding speech frame, while the pinyin confidence score represents the pinyin matching probability of the corresponding word-level speech frame sequence relative to the wake word. Therefore, typically, among multiple speech frames corresponding to a candidate wake word, the last frame of the candidate wake word has the highest pinyin confidence score; the higher the score, the greater the probability that the corresponding candidate wake word is indeed the wake word. Pinyin confidence scores have high accuracy and low false wake-up rate. For example, when there are some similar-sounding but not wake words, pinyin confidence scores can accurately determine that they are not valid wake-up commands. However, pinyin confidence scores have weaker adaptability in fast-paced speech scenarios and conversational wake word scenarios.

[0137] The energy confidence score of a speech frame is a confidence score obtained based on word-level speech feature matching, resulting in a high wake-up success rate. Typically, among multiple speech frames corresponding to a candidate wake-up word, the energy confidence score is highest near the last frame of the candidate wake-up word; a higher score indicates a greater probability that the corresponding candidate wake-up word is indeed the wake-up word. High energy confidence scores result in a high wake-up success rate, allowing the device to respond promptly. However, energy confidence scores are sensitive to noise.

[0138] By fusing the confidence scores of the pinyin and energy in the speech frames, a wake-up confidence score can be obtained. This approach can balance the high wake-up accuracy of pinyin confidence scores with the high wake-up success rate of energy confidence scores. Therefore, the above method can effectively reduce the false wake-up rate and the wake-up failure rate.

[0139] In some embodiments, step S82 can be implemented with reference to steps S12 to S14, and will not be described again.

[0140] In some embodiments, the wake-up confidence of a voice frame is the weighted geometric mean of the pinyin confidence and the energy confidence of the voice frame.

[0141] In some embodiments, referring to Expression 1-1, the geometric mean of the coefficient-weighted sum of the pinyin confidence and the energy confidence is calculated to obtain the wake-up confidence. Where F is the wake-up confidence, P is the pinyin confidence, E is the energy confidence, a is the weight of the pinyin confidence, and b is the weight of the energy confidence. The values ​​of a and b can be determined based on the usage scenario. In one application scenario, the value range of a is (0,1), the value range of b is (0,1), and (a+b) is 1. For example, in a normal scenario, the value of a is 0.5, and the corresponding value of b is 0.5; as another example, in a noisy scenario, the value of a is adjusted to 0.6 according to the noise level, and the corresponding value of b is 0.4.

[0142] ...1-1 In some application scenarios, the value range of P is [0,1], the value range of E is [0,1], and the value range of F is [0,1].

[0143] In some embodiments, referring to expressions 1-2, the exponentially weighted geometric mean of the pinyin confidence score and the energy confidence score is calculated to obtain the wake-up confidence score. F is the wake-up confidence score, P is the pinyin confidence score, E is the energy confidence score, a is the weight of the pinyin confidence score, and b is the weight of the energy confidence score. The values ​​of a and b can be determined based on the usage scenario. In one application scenario, the value range of a is [0,2], the value range of b is [0,2], and (a+b) is 2. For example, in a normal scenario, the value of a is 1, and the value of b is 1; as another example, in a noisy scenario, the value of a is adjusted to 1.5 according to the noise level, and the value of b is 0.5.

[0144] ...1-2 In other embodiments, the wake-up confidence can also be obtained by calculating the exponentially weighted geometric mean of the pinyin confidence and the energy confidence, referring to expressions 1-3. F represents the wake-up confidence, P represents the pinyin confidence, E represents the energy confidence, a represents the weight of the pinyin confidence, and b represents the weight of the energy confidence. The values ​​of a and b can be determined based on the usage scenario. In one application scenario, the value range of a is [0, n], the value range of b is [0, n], and (a+b) is n. For example, in a normal scenario, the value of a is n / 2, and the value of b is correspondingly n / 2; as another example, in a noisy scenario, the value of a can be adjusted to 3n / 4, and the value of b can be correspondingly n / 4, depending on the noise level.

[0145] ...1-3 The weighted geometric mean fusion method can better combine the advantages of pinyin confidence and energy confidence, avoid the shortcomings of each, and make the obtained wake-up confidence more effective in representing the matching probability of candidate wake words relative to wake words, thereby improving wake-up accuracy and wake-up success rate.

[0146] In other implementations, the wake-up confidence can be obtained by calculating the arithmetic mean of the two.

[0147] In some embodiments, the pinyin confidence frame sequence and energy confidence frame sequence corresponding to the speech frame sequence can be obtained first. The wake-up confidence frame sequence can be calculated based on the pinyin confidence frame sequence and energy confidence frame sequence. Then, the wake-up confidence of the wake-up flag frame can be obtained based on the wake-up confidence frame sequence.

[0148] In some embodiments, after obtaining the wake-up confidence frame sequence, a wake-up flag frame can also be determined based on the wake-up confidence frame sequence. For example, see [link to documentation]. Figure 9 This application further proposes a voice wake-up method, specifically including steps S91 to S95.

[0149] Step S91: Obtain the pinyin confidence frame sequence and energy confidence frame sequence corresponding to the speech frame sequence.

[0150] Step S92: Calculate the wake-up confidence frame sequence based on the Pinyin confidence frame sequence and the energy confidence frame sequence.

[0151] The wake-up confidence score corresponding to a speech frame is calculated from the pinyin confidence score and energy confidence score corresponding to the speech frame. When calculating the wake-up confidence score, refer to the embodiments shown in Expression 1-1 or Expression 1-2; further details are omitted.

[0152] Step S93: Determine the speech frames corresponding to wake-up confidence scores greater than the candidate threshold in the wake-up confidence frame sequence as wake-up flag frames.

[0153] The method for determining the candidate threshold is not limited. For example, in one application scenario, the candidate threshold is preset within the device; in another application scenario, the candidate threshold can be obtained from the cloud or a third-party device, without limitation.

[0154] Step S94: Obtain the wake-up confidence of the wake-up flag frame in the speech frame sequence.

[0155] The specific implementation of step S94 can be referred to step S11 or step S81, and will not be repeated here.

[0156] Step S95: Determine whether the wake-up was successful based on the wake-up confidence.

[0157] The wake-up confidence score characterizes the similarity probability between the candidate wake-up word corresponding to the wake-up flag frame and the actual wake-up word. Generally, the higher the wake-up confidence score, the greater the probability that the candidate wake-up word is indeed the wake-up word. For example, in some embodiments, step S95 can be implemented with reference to step S82 or steps S12 to S14, which will not be described in detail here.

[0158] Steps S91 to S95 can obtain the wake-up confidence frame sequence through the pinyin confidence frame sequence and the energy confidence frame sequence, and then determine the wake-up flag frame, which helps to improve the wake-up accuracy.

[0159] In some embodiments, it may also be used Figure 10 Improved embodiments shown Figure 9 In the illustrated embodiment, steps S101 to S102 are set before step S91. When the pinyin confidence frame sequence cannot be obtained directly, the pinyin confidence frame sequence corresponding to the speech frame sequence can be obtained through the pinyin score frame sequence corresponding to the speech frame sequence.

[0160] Step S101: Obtain the pinyin score frame sequence corresponding to the speech frame sequence.

[0161] Each speech frame has a corresponding pinyin score, which represents the pinyin matching probability of the corresponding speech frame.

[0162] Step S102: Calculate the Pinyin confidence frame sequence based on the Pinyin score frame sequence.

[0163] Steps S101 to S102 utilize the pinyin score frame sequence to obtain the pinyin confidence frame sequence corresponding to the speech frame sequence, which can then be used to calculate the wake-up confidence. In one application scenario, the speech frame sequence can be analyzed using wake-up models such as wake-up algorithm libraries to output the pinyin score frame sequence and energy confidence sequence corresponding to the speech frame sequence.

[0164] In one application scenario, a computer storage medium or electronic product storing the voice wake-up method described in this embodiment can interface with any third-party wake-up algorithm library to optimize the wake-up of the voice data output by the third-party wake-up algorithm library, thereby improving its wake-up accuracy and success rate. When the voice data output by the third-party wake-up algorithm library does not include the pinyin confidence frame sequence, the pinyin confidence frame sequence can be obtained using the pinyin score frame sequence.

[0165] In some embodiments, step S102 can be performed as follows: Figure 11 The method shown is implemented in the following way, specifically including steps S111 to S114.

[0166] Step S111: Obtain a first pinyin score frame sequence whose timing is within a first preset period before the current frame.

[0167] A wake-up word is usually composed of a specific combination of pinyin and has a specific speech duration. The first preset period matches the regular speech duration of the wake-up word. Based on the training data of the relevant model in the wake-up algorithm library and user habits, the first preset period can be determined. In some application scenarios, the value of the first preset period can also be updated according to user usage habits. For example, the first preset period can be 1s, 2s, 2.5s, 3.5s, 4s, 4.5s, etc.

[0168] Step S112: Determine a candidate word frame sequence corresponding to each wake-up word based on the first pinyin score frame sequence.

[0169] The candidate word frame sequence is the pinyin score frame sequence corresponding to each wake-up word. Each wake-up word in the wake-up word usually has a specific speech duration. Based on the training data of the relevant model in the wake-up algorithm library and user habits, the specific position and number of frames corresponding to each wake-up word can be inferred, and thus the corresponding candidate word frame sequence can be determined.

[0170] Step S113: Calculate the geometric mean of the highest pinyin scores in all candidate word frame sequences to obtain the pinyin confidence of the current frame.

[0171] Illustrate with an example. In an application scenario, the wake-up word is "Xiaomei Xiaomei", each speech frame corresponds to 40ms, the speech frame sequence corresponding to the first pinyin score frame sequence is 4s before the current frame, the first pinyin score frame sequence includes a total of 100 frames, and the speech frame corresponding to the last frame in the first pinyin score frame sequence is the current frame. It is determined that the candidate word frame sequence corresponding to the first character "Xiao" in the candidate wake-up word is located at frames 10 to 25 in the first pinyin score frame sequence, including a total of 15 frames, referring to Table 1; the candidate word frame sequence corresponding to the second character "Mei" in the candidate wake-up word is located at frames 30 to 50 in the first pinyin score frame sequence, including a total of 20 frames, referring to Table 2; the candidate word frame sequence corresponding to the third character "Xiao" in the candidate wake-up word is located at frames 55 to 70 in the first pinyin score frame sequence, including a total of 15 frames, referring to Table 3; the candidate word frame sequence corresponding to the fourth character "Mei" in the candidate wake-up word is located at frames 80 to 100 in the first pinyin score frame sequence, including a total of 20 frames, referring to Table 4.

[0172] In Table 1, the highest pinyin score is 0.90 at frame 15.

[0173] In Table 2, the highest pinyin score is 0.85 at frame 20. Thus, the geometric mean S(12) of the first wake-up word "Xiao" and the second wake-up word "Mei" can be calculated to satisfy the expression 2-1.

[0174] ……2-1 In Table 3, the highest Pinyin score is 0.92 for the 15th frame. From this, the geometric mean of the first wake-up word "Xiao", the second wake-up word "Mei", and the third wake-up word "Xiao" can be calculated as satisfying Expression 2-2.

[0175] ……2-2 In Table 4, the highest Pinyin score is 0.88 for the 20th frame; from this, the geometric mean of the first wake-up word "Xiao", the second wake-up word "Mei", the third wake-up word "Xiao", and the fourth wake-up word "Mei" can be calculated satisfying Expression 2-3.

[0176] ……2-3 Therefore, it can be determined that the Pinyin confidence of the current frame is 0.89.

[0177] It should be noted that the number of wake-up words in the wake-up phrase is not limited in a specific usage scenario. For example, if there are n wake-up words in the wake-up phrase, the Pinyin confidence of the wake-up flag frame The calculation formula can refer to Expression 2-4.

[0178] ……2-4 where Sn is the highest Pinyin score in the candidate word frame sequence corresponding to the last wake-up word.

[0179] Table 1 Candidate word frame sequence corresponding to the first wake-up word "Xiao"

[0180] Table 2 Candidate word frame sequence corresponding to the second wake-up word "Mei"

[0181] Table 3 Candidate word frame sequence corresponding to the third wake-up word "Xiao"

[0182] Table 4 Candidate word frame sequence corresponding to the fourth wake-up word "Mei"

[0183] By successively calculating the geometric mean to obtain the geometric mean of the highest Pinyin scores in all candidate word frame sequences, and then obtaining the Pinyin confidence of the current frame, the calculation process can be optimized.

[0184] Step S114: Arrange the Pinyin confidences corresponding to all speech frames in chronological order to obtain a Pinyin confidence frame sequence.

[0185] Steps S111 to S114 can obtain the pinyin confidence score by taking the geometric mean of multiple pinyin scores, thereby obtaining the pinyin confidence score frame sequence, which can improve the reliability of the pinyin confidence score.

[0186] In some embodiments, by performing a backward calculation in accordance with the above calculation method, the position of the last frame of each wake-up word can be determined based on the pinyin confidence, pinyin score, and first pinyin score frame sequence of the wake-up flag frame, which helps to estimate the duration and corresponding position of the wake-up segment.

[0187] This application further proposes a voice wake-up method, such as Figure 12 As shown, the voice wake-up method includes steps S121 to S124.

[0188] Step S121: Obtain speech data including speech frame sequences and energy confidence frame sequences from a third-party wake-up algorithm library. The speech data also includes at least one of the pinyin confidence frame sequences and pinyin score frame sequences.

[0189] When the speech data includes a pinyin confidence frame sequence and an energy confidence frame sequence, the wake-up confidence frame sequence can be directly calculated. When the speech data includes a pinyin score frame sequence, it can be calculated through... Figure 11 The confidence frame sequence of the pinyin is calculated in the manner shown.

[0190] Step S122: Determine the wake-up flag frame based on voice data.

[0191] For example, the wake-up flag frame can be determined based on the wake-up confidence frame sequence, for example, by referring to... Figure 9 The implementation shown determines the wake-up flag frame, which will not be described in detail here.

[0192] Step S123: Obtain the wake-up confidence of the wake-up flag frame in the speech frame sequence.

[0193] The specific implementation of step S123 can be referred to step S94, and will not be repeated here.

[0194] Step S124: Determine whether the wake-up was successful based on the wake-up confidence.

[0195] The specific implementation of step S124 can be referred to step S95, and will not be repeated here.

[0196] Based on steps S121 to S124, this embodiment can interface with third-party wake-up algorithm libraries of different products, and use the voice data output by them to optimize their wake-up mechanism, thereby improving wake-up accuracy and wake-up success rate.

[0197] To improve wake-up accuracy, further improvements can be made by setting other steps after step S91 and before step S92. Figure 9The illustrated embodiment. For example, this application further proposes a voice wake-up method, such as... Figure 13 As shown, it specifically includes steps S131 to S137.

[0198] Step S131: Obtain the pinyin confidence frame sequence and energy confidence frame sequence corresponding to the speech frame sequence.

[0199] The specific implementation of step S131 can be referred to step S91, and will not be repeated here.

[0200] Step S132: Obtain all first peak frames in the Pinyin confidence frame sequence and all second peak frames in the energy confidence frame sequence.

[0201] The Pinyin confidence frame sequence contains multiple peak frames, each with a higher Pinyin confidence score than nearby speech frames. These peak frames are defined as the first peak frame. Similarly, the energy confidence frame sequence contains multiple peak frames, each with a higher energy confidence score than nearby speech frames. These peak frames are defined as the second peak frame.

[0202] Step S133: Based on the frame sequence number of all first peak frames, correct the frame sequence number of all second peak frames in sequence according to the time order.

[0203] Typically, the speech frame corresponding to the first peak frame of the Pinyin confidence frame sequence is a candidate wake-up flag frame, and the speech frame corresponding to the second peak frame of the energy confidence frame sequence is also a candidate wake-up flag frame. Both the Pinyin confidence frame sequence and the energy confidence frame sequence correspond to the same speech frame sequence. For the same speech frame sequence, the temporal position of the wake-up flag frame is consistent in both the Pinyin confidence frame sequence and the energy confidence frame sequence. However, in practical applications, the frame durations of the Pinyin confidence frame sequence and the energy confidence frame sequence may differ. Therefore, although the temporal position of the wake-up flag frame is fixed, there may still be cases where the frame number of the wake-up flag frame in the Pinyin confidence frame sequence differs from that in the energy confidence frame sequence.

[0204] Step S133 can correct the frame number corresponding to the second peak frame of the positive energy confidence frame sequence according to the frame number corresponding to the first peak frame of the Pinyin confidence frame sequence, so that the frame number of the first peak frame is the same as the frame number of the corresponding second peak frame. This makes it easier to directly determine the timing of the wake-up standard frame and calculate the wake-up confidence of the wake-up flag frame based on the frame number, thereby reducing the probability of calculation errors and improving accuracy.

[0205] Step S134: Calculate the wake-up confidence frame sequence based on the Pinyin confidence frame sequence and the energy confidence frame sequence.

[0206] The specific implementation of step S134 can be referred to step S92, and will not be repeated here.

[0207] Step S135: Determine the speech frames corresponding to wake-up confidence scores greater than the candidate threshold in the wake-up confidence frame sequence as wake-up flag frames.

[0208] The specific implementation of step S135 can be referred to step S93, and will not be repeated here.

[0209] Step S136: Obtain the wake-up confidence of the wake-up flag frame in the speech frame sequence.

[0210] The specific implementation of step S136 can be referred to step S94, and will not be repeated here.

[0211] Step S137: Determine whether wake-up was successful based on wake-up confidence.

[0212] The specific implementation of step S137 can be referred to step S95, and will not be repeated here.

[0213] In some embodiments, the step of determining whether wake-up was successful based on wake-up confidence in any of steps S82, S95, S124, and S137 can be achieved by, for example... Figure 14 The method shown is implemented in the following way, specifically including steps S141 to S143.

[0214] Step S141: Obtain the maximum threshold in the preset threshold group.

[0215] The specific implementation of step S141 can be referred to step S12, and will not be repeated here.

[0216] Step S142: Wake-up is successful in response to a wake-up confidence level greater than or equal to the maximum threshold.

[0217] The specific implementation of step S142 can be referred to step S13, and will not be repeated here.

[0218] Step S143: In response to the wake-up confidence being less than the maximum threshold, determine whether to enter the first pre-wake-up state.

[0219] The specific implementation of step S143 can be referred to step S14, and will not be repeated here.

[0220] For example, in some embodiments, reference may also be made to Figures 2 to 7 The improved steps S141 to S143 of the embodiment shown will not be described again.

[0221] For example, the step of determining whether to enter the first pre-wake state in response to a wake-up confidence level less than the maximum threshold includes: determining the wake-up segment containing the candidate wake-up word in the speech frame sequence, the leading segment before the wake-up segment, and the subsequent segment after the wake-up segment based on the pinyin confidence level; calculating the first average decibel value of the leading segment, the second average decibel value of the wake-up segment, and the third average decibel value of the subsequent segment; calculating the first decibel difference between the second average decibel value and the first average decibel value, and the second decibel difference between the second average decibel value and the third average decibel value; and determining whether wake-up is successful based on the wake-up confidence level and a preset threshold group in response to both the first and second decibel differences being greater than or equal to a decibel difference threshold. Specific implementation of this embodiment can refer to the above embodiment and will not be repeated here. As another example, the step of determining whether wake-up is successful based on the wake-up confidence level and a preset threshold group includes: obtaining the energy threshold of the preset threshold group, where the energy threshold is less than the maximum threshold; and determining that wake-up is successful in response to the wake-up confidence level being greater than or equal to the energy threshold. Specific implementation of this embodiment can refer to the above embodiment and will not be repeated here.

[0222] In some embodiments, the voice wake-up method of this application can be applied to multiple scenarios such as smart home appliances, portable devices, and in-vehicle terminals, including air conditioners, robot vacuums, and microwave ovens, to achieve cross-platform user experience enhancement.

[0223] In some embodiments, compared to devices that directly use energy confidence as wake-up confidence, using a fusion of Pinyin confidence and energy confidence to obtain wake-up confidence reduces the false wake-up rate from 20-30 times / 72h to 0-1 times / 72h. Compared to devices that directly use Pinyin confidence as wake-up confidence, using a fusion of Pinyin confidence and energy confidence to obtain wake-up confidence increases the wake-up success rate from 80%-85% to 92%-97%.

[0224] In some embodiments, the wake-up confidence is obtained by fusing the pinyin confidence and energy confidence, and see reference. Figure 15 By combining various judgment logics and speech features mentioned in the above embodiments, such as pinyin confidence, average decibel value, maximum threshold, energy threshold, normal threshold, minimum threshold, and decibel difference threshold, the voice wake-up method can form a multi-level, multi-dimensional judgment system. In one application scenario, the voice wake-up method of this embodiment can reduce the false wake-up rate to 0-1 times / 72h and increase the wake-up success rate to 92%-97%; when repeatedly wake-up, the wake-up success rate can reach over 90% in complex scenarios combining noise, accents, and fast speech.

[0225] In some embodiments, users can adjust the specific value of the common threshold themselves. For example, the adjustable range of the common threshold can be set to 0.65-0.75, 0.65-0.72, or 0.65-0.7, etc., to improve the user experience.

[0226] This application further proposes an electronic device including a memory and a processor. The memory is used to store program data, which can be executed by the processor to implement the voice wake-up method described in any of the above embodiments.

[0227] This application further proposes a computer program product, including computer program instructions that enable a computer to implement the voice wake-up method described in any of the above embodiments.

[0228] This application further proposes a computer storage medium. For example... Figure 16 As shown, Figure 16 This is a schematic diagram of a computer storage medium according to an embodiment of the present application. The computer storage medium 10 stores program instructions 11, which are executed by a processor to implement the above-described voice wake-up method.

[0229] Specifically, program instructions 11 can form a program file and be stored in the aforementioned storage medium as a software product, so that an electronic device (which may be a personal computer, server, or network device, etc.) or processor can execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, or terminal devices such as computers, servers, mobile phones, and tablets.

[0230] In this embodiment, the computer storage medium 10 can be, but is not limited to, a USB flash drive, SD card, PD optical drive, portable hard drive, large-capacity floppy drive, flash memory, multimedia memory card, server, etc.

[0231] In one embodiment, a computer program product or computer program is provided, comprising computer instructions stored in a computer storage medium. A processor of an electronic device reads the computer instructions from the computer storage medium and executes the computer instructions, causing the electronic device to perform the steps described in the above method embodiments.

[0232] Furthermore, if the aforementioned functions are implemented as software functions and sold or used as independent products, they can be stored in a mobile terminal-readable storage medium. That is, this application also provides a storage device storing program data, which can be executed to implement the methods of the above embodiments. This storage device can be, for example, a USB flash drive, an optical disc, or a server. In other words, this application can be embodied in the form of a software product, which includes several instructions to cause a smart terminal to execute all or part of the steps of the methods described in the various embodiments.

[0233] Unlike existing technologies, the voice wake-up method of this application includes: obtaining the wake-up confidence of a wake-up flag frame in a voice frame sequence; obtaining the maximum threshold in a preset threshold group; waking up successfully in response to the wake-up confidence being greater than or equal to the maximum threshold; and determining whether to enter a first pre-wake state based on the wake-up confidence and the preset threshold group in response to the wake-up confidence being less than the maximum threshold. Through this method, the probability of successful false wake-up in cases of false wake-up can be limited by the maximum threshold, thereby reducing the false wake-up rate. Furthermore, when the wake-up confidence does not meet the maximum threshold requirement, wake-up failure is not directly determined; instead, the determination of whether to enter the first pre-wake state is based on the preset threshold group and the wake-up confidence, effectively reducing the wake-up failure rate under low-energy wake-up conditions. Therefore, this embodiment can effectively reduce the wake-up failure rate and the false wake-up rate, improving the user experience.

[0234] In the several embodiments provided in this application, it should be understood that the disclosed methods and apparatus can be implemented in other ways. For example, the apparatus implementations described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0235] Any process or method description in the flowchart or otherwise herein can be understood as representing an apparatus, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order according to the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0236] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A voice wake-up method, characterized in that, The voice wake-up method includes: Obtain the wake-up confidence of the wake-up flag frame in the speech frame sequence; Whether the wake-up was successful is determined based on the wake-up confidence level. The wake-up confidence of the speech frames in the speech frame sequence is calculated from the pinyin confidence and energy confidence of the speech frames.

2. The voice wake-up method according to claim 1, characterized in that, The wake-up confidence of the speech frame is the weighted geometric mean of the pinyin confidence and the energy confidence of the speech frame.

3. The voice wake-up method according to claim 2, characterized in that, The weighted geometric mean is an exponentially weighted geometric mean. The sum of the weights of the pinyin confidence score and the energy confidence score is a preset value. Both the weights of the pinyin confidence score and the energy confidence score are non-negative.

4. The voice wake-up method according to claim 3, characterized in that, The weights of the pinyin confidence score and the energy confidence score are both half of the preset values.

5. The voice wake-up method according to claim 1, characterized in that, Before the step of obtaining the wake-up confidence of the wake-up flag frame in the speech frame sequence, the speech wake-up method further includes: Obtain the pinyin confidence frame sequence and energy confidence frame sequence corresponding to the speech frame sequence; The wake-up confidence frame sequence is calculated based on the Pinyin confidence frame sequence and the energy confidence frame sequence; The speech frames corresponding to wake-up confidence scores greater than the candidate threshold in the wake-up confidence frame sequence are identified as the wake-up flag frames.

6. The voice wake-up method according to claim 1, characterized in that, Before the step of obtaining the wake-up confidence of the wake-up flag frame in the speech frame sequence, the method further includes: Speech data including the speech frame sequence and the energy confidence frame sequence are obtained from a third-party wake-up algorithm library. The speech data also includes at least one of the pinyin confidence frame sequence and the pinyin score frame sequence. The wake-up flag frame is determined based on the voice data.

7. The voice wake-up method according to claim 5, characterized in that, Before the step of obtaining the pinyin confidence frame sequence and energy confidence frame sequence corresponding to the speech frame sequence, the voice wake-up method further includes: Obtain the pinyin score frame sequence corresponding to the speech frame sequence; The pinyin confidence frame sequence corresponding to the speech frame sequence is calculated based on the pinyin score frame sequence.

8. The voice wake-up method according to claim 7, characterized in that, The step of calculating the pinyin confidence frame sequence based on the pinyin score frame sequence includes: Obtain the first pinyin score frame sequence within the first preset time period preceding the current frame; Based on the first pinyin score frame sequence, determine the candidate character frame sequence corresponding to each wake-up character; Calculate the geometric mean of the highest pinyin scores in all candidate character frame sequences to obtain the pinyin confidence of the current frame; Arrange the pinyin confidence scores corresponding to all the speech frames in chronological order to obtain the pinyin confidence score frame sequence.

9. The voice wake-up method according to claim 5, characterized in that, Before the step of calculating the wake-up confidence frame sequence based on the pinyin confidence frame sequence and the energy confidence frame sequence, the voice wake-up method further includes: Obtain all first peak frames in the Pinyin confidence frame sequence and all second peak frames in the energy confidence frame sequence; Based on the frame sequence number of all the first peak frames, the frame sequence number of all the second peak frames is corrected sequentially according to the time sequence.

10. The voice wake-up method according to claim 1, characterized in that, The steps for determining whether wake-up was successful based on the wake-up confidence include: Get the maximum threshold in the preset threshold group; A wake-up is successful if the wake-up confidence level is greater than or equal to the maximum threshold. In response to the wake-up confidence being less than the maximum threshold, it is determined whether to enter the first pre-wake-up state.

11. The voice wake-up method according to claim 10, characterized in that, The step of determining whether to enter the first pre-wake state in response to the wake-up confidence being less than the maximum threshold includes: Based on the pinyin confidence level, the wake-up segment containing the candidate wake-up word in the speech frame sequence, the preceding segment before the wake-up segment, and the subsequent segment after the wake-up segment are determined. Calculate the first average decibel value of the preamble segment, the second average decibel value of the wake-up segment, and the third average decibel value of the subsequent segment; Calculate the first decibel difference between the second average decibel value and the first average decibel value, and the second decibel difference between the second average decibel value and the third average decibel value; In response to the first decibel difference and the second decibel difference both being greater than or equal to the decibel difference threshold, a determination is made as to whether the wake-up was successful based on the wake-up confidence and the preset threshold group.

12. The voice wake-up method according to claim 11, characterized in that, The step of determining whether wake-up was successful based on the wake-up confidence level and the preset threshold group includes: Obtain the energy threshold of the preset threshold group, wherein the energy threshold is less than the maximum threshold; In response to the wake-up confidence being greater than or equal to the energy threshold, a successful wake-up is determined.

13. An electronic device, characterized in that, It includes a memory and a processor, the memory being used to store program data, the program data being executable by the processor to implement the voice wake-up method according to any one of claims 1-12.

14. A computer storage medium, characterized in that, It stores program instructions that are executed by a processor to implement the voice wake-up method according to any one of claims 1 to 12.

15. A computer program product, characterized in that, It includes computer program instructions that cause a computer to implement the voice wake-up method according to any one of claims 1 to 12.