Device awakening method, device and equipment based on voiceprint information

By analyzing sound signals layer by layer and matching voiceprint features, the problems of high power consumption and recognition errors in waking up small, low-power devices have been solved, realizing a low-power and convenient device wake-up method.

CN120998207APending Publication Date: 2025-11-21SHANGHAI IND U TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511213795.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

The wake-up methods for small, low-power smart terminal devices suffer from high power consumption and poor ease of operation. Existing button and fingerprint wake-up technologies cannot effectively solve the high power consumption bottleneck of voice wake-up and are prone to recognition failure due to finger condition.

Method used

By analyzing the sound signal layer by layer, using short-time energy and zero-crossing rate to determine the presence of a speech signal, and combining simplified and complete voiceprint feature matching, the similarity threshold is dynamically adjusted to achieve low-power device wake-up.

Benefits of technology

It significantly reduces the standby power consumption of the device, increases the standby time, avoids identification errors and operational inconvenience, and improves the device's response efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120998207A_ABST
    Figure CN120998207A_ABST
Patent Text Reader

Abstract

The invention discloses a device awakening method and device based on voiceprint information, a device and a readable storage medium, and relates to the technical field of voiceprint recognition. Comprising the steps that sound signals are collected in real time, and whether voice signals exist in the sound signals or not is judged according to the short-time energy and the zero-crossing rate of the sound signals; if the voice signal exists in the sound signal, judging whether the voice signal is matched with a preset speaker or not; and if the voice signal is matched with the preset speaker, waking up target equipment. According to the method, the energy consumed by voiceprint recognition is reduced, and the standby time of the equipment is prolonged.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voiceprint recognition, in particular to a device wake-up method and device based on voiceprint information, an apparatus and a readable storage medium. BACKGROUND

[0002] In the digital era, small low-power intelligent terminal devices such as portable emergency callers and small smart door locks are increasingly popular, and there is an urgent need for convenient interaction wake-up methods. Voiceprint recognition, as a non-contact identity recognition method, has great potential in such device interactions and is widely used in security and home care scenarios.

[0003] However, in actual applications, there are problems with the wake-up of small low-power terminal devices. On the one hand, implementing a voice-based wake-up function requires continuous monitoring of external sounds, and the device has a small battery capacity, so continuous monitoring can cause serious power consumption problems, resulting in a significant reduction in battery life. Previous voice wake-up solutions have not been widely adopted due to high power consumption. On the other hand, existing wake-up methods have limitations and are mostly through key or fingerprint wake-up. For example, small smart door locks using fingerprint wake-up can easily fail to recognize due to finger conditions, and are not convenient to operate, which can result in a delayed response in critical scenarios.

[0004] Overall, the existing key and fingerprint wake-up technology is a compromise choice for the industry when it cannot break through the high power consumption bottleneck of voice wake-up. Although it meets the basic power consumption needs, it sacrifices the convenience of operation and the adaptability of scenarios. Therefore, there is an urgent need for a device wake-up method that can overcome the above-mentioned defects. SUMMARY

[0005] The present application provides a device wake-up method and device based on voiceprint information, which can significantly reduce standby power consumption by analyzing sound signals layer by layer without directly deep analyzing sound signals, thereby increasing standby time and avoiding identification errors and operational inconvenience when using fingerprints to wake up devices.

[0006] To achieve the above-mentioned purpose, the present application provides the following technical solutions: In a first aspect, the present application provides a device wake-up method based on voiceprint information, which comprises: real-time acquisition of sound signals and determination of whether there is a voice signal in the sound signals based on the short-time energy and zero-crossing rate of the sound signals; if there is a voice signal in the sound signals, determining whether the voice signal matches a preset speaker; if the voice signal matches the preset speaker, waking up the target device.

[0007] In some embodiments, determining whether the voice signal matches the preset speaker comprises: extracting simplified voiceprint features from the voice signal, and determining whether the voice signal preliminarily matches the preset speaker based on the simplified voiceprint features; if the voice signal preliminarily matches the preset speaker, extracting complete voiceprint features from the voice signal, and determining whether the voice signal matches the preset speaker based on the complete voiceprint features.

[0008] In some embodiments, determining whether the voice signal preliminarily matches the preset speaker based on the simplified voiceprint features comprises: calculating first mel-frequency cepstral coefficients of the simplified voiceprint features, and inputting the first mel-frequency cepstral coefficients into a pre-trained voiceprint matching model to obtain a first cosine similarity between the voice signal and the preset speaker; if the first cosine similarity is greater than a first similarity threshold, determining that the voice signal preliminarily matches the preset speaker.

[0009] In some embodiments, determining whether the voice signal matches the preset speaker based on the complete voiceprint features comprises: calculating second mel-frequency cepstral coefficients of the complete voiceprint features, and inputting the second mel-frequency cepstral coefficients into a pre-trained voiceprint matching model to obtain a second cosine similarity between the voice signal and the preset speaker; if the second cosine similarity is greater than a second similarity threshold, determining that the voice signal matches the preset speaker.

[0010] In some embodiments, the method further comprises: obtaining current environment information; the current environment information comprises a noisy degree of a current environment; based on the current environment information, dynamically adjusting the second similarity threshold.

[0011] In some embodiments, determining whether the voice signal exists in the sound signal according to the short-time energy and the zero-crossing rate of the sound signal comprises: extracting basic features in the sound signal; the basic features comprise an amplitude, a frame length and a frame shift of the sound signal; calculating the short-time energy and the zero-crossing rate of the sound signal according to the basic features; if the short-time energy and the zero-crossing rate are both in a preset range, determining that the voice signal exists in the sound signal.

[0012] In a second aspect, the present application further provides a device wake-up apparatus based on voiceprint information, the apparatus comprising: a voice determining module, configured to collect sound signals in real time, and determine whether voice signals exist in the sound signals according to short-time energy and zero-crossing rate of the sound signals; a person matching module, configured to determine whether the voice signals match a preset speaker if the voice signals exist in the sound signals; a device wake-up module, configured to wake up a target device if the voice signals match the preset speaker.

[0013] In a third aspect, the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the device wake-up method based on voiceprint information provided in the first aspect when executing the computer program.

[0014] In a fourth aspect, the present application further provides a computer readable storage medium, and the computer readable storage medium stores a computer program, and the computer program is executable on a processor to implement the device wake-up method based on voiceprint information provided in the first aspect.

[0015] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, and the computer program is executable on a processor to implement the device wake-up method based on voiceprint information provided in the first aspect.

[0016] The device wake-up method based on voiceprint information provided in the present application collects sound signals in real time, and determines whether voice signals exist in the sound signals according to short-time energy and zero-crossing rate of the sound signals, and determines whether the voice signals match a preset speaker if the voice signals exist in the sound signals, and wakes up a target device if the voice signals match the preset speaker. If the voice signals do not exist in the sound signals or do not match the preset speaker, the process is directly ended. The sound signals are analyzed layer by layer, and the sound signals are not deeply analyzed directly, so that the preliminary judgment of the sound signals can be realized by using less power consumption after receiving the sound signals, that is, when the sound signals are invalid signals in the environment, the sound signals can be recognized by using less power consumption, standby power consumption is significantly reduced, standby time is increased, and recognition errors and inconvenience of operation when using fingerprints to wake up the device are avoided.

[0017] The above description is only a summary of the technical solutions of the present application, in order to more clearly understand the technical means of the present application, and the content of the specification can be implemented. The following describes the preferred embodiments of the present application in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0018] Figure 1 A flowchart of a device wake-up method based on voiceprint information according to an embodiment of the present application is shown in FIG. 1. Figure 2 A flowchart of another device wake-up method based on voiceprint information according to an embodiment of the present application is shown in FIG. 2. Figure 3 A structural diagram of a device wake-up apparatus based on voiceprint information according to an embodiment of the present application is shown in FIG. 3. Figure 4 A structural diagram of another device wake-up apparatus based on voiceprint information according to an embodiment of the present application is shown in FIG. 4. Figure 5 An electronic device structure according to an embodiment of the present application is shown in FIG. 5. DETAILED DESCRIPTION

[0019] The technical solutions of the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application. It should be noted that the description of "one embodiment", "an embodiment", "example embodiment" and the like in the specification means that the described embodiment can include specific features, structures or characteristics, but not every embodiment must include these specific features, structures or characteristics. In addition, such expressions do not mean the same embodiment. Furthermore, when a specific feature, structure or characteristic is described in connection with an embodiment, it is indicated that such a feature, structure or characteristic is incorporated into other embodiments within the knowledge of those skilled in the art, whether or not it is explicitly described.

[0020] In addition, the technical features involved in different embodiments of the present application described below can be combined with each other as long as there is no conflict.

[0021] In some embodiments, as shown in FIG. 1, a device wake-up method based on voiceprint information is provided, and the specific method includes: Figure 1 S101, collecting sound signals in real time, and determining whether there is a voice signal in the sound signals according to the short-time energy and zero-crossing rate of the sound signals. S101, collecting sound signals in real time, and determining whether there is a voice signal in the sound signals according to the short-time energy and zero-crossing rate of the sound signals.

[0022] The sound signals are comprehensive signals of all sounds collected in the environment by using a microphone array, and the sound signals can contain environmental noise, human voice and other sounds.

[0023] Optionally, the basic features in the sound signal are extracted; the basic features include the amplitude, frame length and frame shift of the sound signal; the short-time energy and zero-crossing rate of the sound signal are calculated according to the basic features; if both the short-time energy and the zero-crossing rate are in the preset range, it is determined that the sound signal contains the speech signal.

[0024] In an example, since the process only deals with the basic features of the sound signal, it can be completed by using an MCU with ultra-low power consumption, and the waveform of the sound signal can satisfy the following formula (1): (1) wherein x(n) is the time-domain signal, w(n) is the window function, yi(n) is the value of a frame, n=1, 2, …, L, i=1, 2, …, fn, L is the frame length; inc is the frame shift length; fn is the total number of frames after framing.

[0025] For a discrete-time signal y(i), it is framed (the signal is intercepted by a sliding window with a length of N, and the window can be overlapped), and the short-time energy E(i) of the i-th frame signal yi(m) (m is the sample index within the frame) is: (2) The formula for calculating the zero-crossing rate is: (3) The preset range of the speech signal is set in advance, and the preset range includes the short-time energy preset range and the zero-crossing rate preset range. When the short-time energy of the sound signal is in the short-time energy preset range and the zero-crossing rate is in the zero-crossing rate preset range, it means that the short-time energy and the zero-crossing rate of the sound signal are in the preset range of the speech signal, and at this time it can be determined that the sound signal contains the speech signal; otherwise, it means that the sound signal does not contain the speech signal, and at this time S102 does not need to be executed. That is, when the sound signal does not contain the speech signal, no subsequent steps need to be executed, only the steps in S101 are executed. Since the processing in S101 is relatively simple and an MCU with ultra-low power consumption is used, when the sound signal does not contain the speech signal, only a small amount of energy needs to be consumed, which can eliminate the influence of environmental noise on the device wake-up, and thus the standby time of the device can be significantly improved.

[0026] S102, if the sound signal contains the speech signal, it is determined whether the speech signal matches the preset speaker.

[0027] Specifically, when it is determined in the previous step that the sound signal contains a voice signal, a further determination is made as to whether the voice signal is emitted by a preset speaker. The features of the voice signal can be further collected, and a feature comparison is made between the collected features and the features of the preset speaker. If the similarity between the two is higher than a certain value, it is determined that the voice signal matches the preset speaker; otherwise, it is determined that the voice signal does not match the preset speaker, i.e., the speaker is a stranger and cannot wake up the device, and step S103 is not continued.

[0028] Optionally, when determining whether the voice signal matches the preset speaker, the following steps can also be taken: simplified voiceprint features are extracted from the voice signal, and it is determined whether the voice signal preliminarily matches the preset speaker based on the simplified voiceprint features; if the voice signal preliminarily matches the preset speaker, complete voiceprint features are extracted from the voice signal, and it is determined whether the voice signal matches the preset speaker based on the complete voiceprint features.

[0029] The simplified voiceprint features can be the first 10-dimensional features in the voice signal, which include the most prominent features of the sound, such as overall pitch, energy distribution, and basic timbre.

[0030] Specifically, the first mel-frequency cepstral coefficient of the simplified voiceprint features is calculated, and the specific calculation formula is as follows: (4) where C(k) is the kth mel-frequency cepstral coefficient, k represents the index of the cepstral coefficient; S(m) is the output energy of the mth mel filter after logarithmic operation; M is the number of mel filters; and L is the number of final retained MFCC coefficients, which is usually less than or equal to M.

[0031] The first mel-frequency cepstral coefficient is then input into a pre-trained voiceprint matching model to obtain the first cosine similarity between the voice signal and the preset speaker. The process of calculating the first cosine similarity between the voice signal and the preset speaker by the pre-trained voiceprint matching model is as follows: (5) where v represents the voiceprint feature vector of the voice signal, v k is a pre-stored reference voiceprint feature vector.

[0032] If the first cosine similarity is greater than the first similarity threshold, it is determined that the voice signal matches the preset speaker; otherwise, it is determined that the voice signal does not match the preset speaker, and the subsequent steps do not need to be executed. Since only the simplified voiceprint features (for example, the first 10 features) are calculated at this time, the most prominent features (such as overall pitch, energy distribution, and basic timbre) are included, and the high-dimensional features do not need to be calculated. When the speaker is a stranger, the exclusion can be completed at this step, the power consumption is greatly reduced, and the standby time of the device is increased.

[0033] Then, the second mel-frequency cepstral coefficient of the complete voiceprint feature is calculated again, the calculation formula is referred to formula (4), and the second mel-frequency cepstral coefficient is input into the pre-trained voiceprint matching model to obtain the second cosine similarity between the voice signal and the preset speaker, which is referred to formula (5). If the second cosine similarity is greater than the second similarity threshold, it is determined that the voice signal matches the preset speaker. At this time, it is determined that the speaker is the preset speaker, and the subsequent device wake-up can be performed. Otherwise, it is determined that the voice signal does not match the preset speaker, and the subsequent steps do not need to be executed.

[0034] In S103, if the voice signal matches the preset speaker, the target device is woken up.

[0035] Specifically, when it is determined in S102 that the voice signal matches the preset speaker, the target device is woken up.

[0036] In the device wake-up method based on the voiceprint information in the above embodiment, the sound signal is first collected in real time, and it is determined whether the voice signal exists in the sound signal according to the short-time energy and the zero-crossing rate of the sound signal. If the voice signal exists in the sound signal, it is determined whether the voice signal matches the preset speaker. If the voice signal matches the preset speaker, the target device is woken up. It can be known that if the voice signal does not exist in the sound signal or if the voice signal does not match the preset speaker, the process is directly ended. The sound signal is analyzed layer by layer, and the sound signal is not directly deeply analyzed, so that the preliminary judgment of the sound signal can be realized by using less power consumption after the sound signal is received. That is, when the sound signal is an invalid signal in the environment, the sound signal can be identified by using less power consumption. The standby power consumption is significantly reduced, the standby time is increased, and the identification error and the inconvenience of operation when the device is woken up by using the fingerprint are avoided.

[0037] In another embodiment, the method in the above embodiment further includes: obtaining current environment information; the current environment information includes the noise level of the current environment; and the second similarity threshold is dynamically adjusted based on the current environment information.

[0038] Specifically, since the environment can affect the matching degree of the voice signal and the preset speaker, when the environment is relatively noisy, the second similarity threshold can be appropriately reduced, for example, to 80%, and when the environment is relatively quiet, the second similarity threshold can be appropriately increased, for example, to 90%, so as to flexibly cope with different environments.

[0039] In order to more fully show the present scheme, the present embodiment gives an optional way of a device wake-up method based on voiceprint information, as shown in Figure 2 S201, collecting a sound signal in real time and extracting a basic feature in the sound signal.

[0040] S202, calculating a short-time energy and a zero-crossing rate of the sound signal according to the basic feature.

[0041] S203, if the short-time energy and the zero-crossing rate are both in a preset range, determining that there is a voice signal in the sound signal.

[0042] S204, if there is a voice signal in the sound signal, extracting a simplified voiceprint feature from the voice signal.

[0043] S205, calculating a first mel-frequency cepstral coefficient of the simplified voiceprint feature, and inputting the first mel-frequency cepstral coefficient into a pre-trained voiceprint matching model to obtain a first cosine similarity of the voice signal and a preset speaker.

[0044] S206, if the first cosine similarity is greater than a first similarity threshold, determining that the voice signal and the preset speaker are preliminarily matched.

[0045] S207, if the voice signal and the preset speaker are preliminarily matched, extracting a complete voiceprint feature from the voice signal.

[0046] S208, obtaining current environment information; the current environment information includes a noisy degree of a current environment.

[0047] S209, dynamically adjusting a second similarity threshold based on the current environment information.

[0048] S210, calculating a second mel-frequency cepstral coefficient of the complete voiceprint feature, and inputting the second mel-frequency cepstral coefficient into the pre-trained voiceprint matching model to obtain a second cosine similarity of the voice signal and the preset speaker.

[0049] S211, if the second cosine similarity is greater than the second similarity threshold, determining that the voice signal and the preset speaker are matched.

[0050] S212, if the voice signal and the preset speaker are matched, waking up a target device.

[0051] ​The specific process of S201-S212 can be referred to the description of the method embodiments, and the implementation principle and technical effects are similar, which will not be described here.

[0052] Based on the same inventive concept, the embodiments of the present application also provide a voiceprint information based device wake-up apparatus for implementing the voiceprint information based device wake-up method. The implementation scheme for solving the problem provided by the apparatus is similar to the implementation scheme described in the method, and therefore the specific limitations in one or more voiceprint information based device wake-up apparatus embodiments provided below can be referred to the limitations of the voiceprint information based device wake-up method described above, which will not be described here.

[0053] In one embodiment, as shown in Figure 3 , a voiceprint information based device wake-up apparatus is provided, which comprises: a voice judgment module 30, configured to collect sound signals in real time, and judge whether there is a voice signal in the sound signals according to the short-time energy and zero-crossing rate of the sound signals; a person matching module 31, configured to judge whether the voice signal matches a preset speaker if there is a voice signal in the sound signals; a device wake-up module 32, configured to wake up a target device if the voice signal matches the preset speaker.

[0054] In another embodiment, as shown in Figure 4 , the person matching module 31 in the above Figure 3 comprises: a first matching unit 310, configured to extract simplified voiceprint features from the voice signal, and judge whether the voice signal preliminarily matches a preset speaker based on the simplified voiceprint features; a second matching unit 311, configured to extract complete voiceprint features from the voice signal if the voice signal preliminarily matches the preset speaker, and judge whether the voice signal matches the preset speaker based on the complete voiceprint features.

[0055] In another embodiment, the first matching unit 310 in the above Figure 4 is specifically configured to: calculate first mel-frequency cepstral coefficients of the simplified voiceprint features, and input the first mel-frequency cepstral coefficients into a pre-trained voiceprint matching model to obtain a first cosine similarity of the voice signal and the preset speaker; and if the first cosine similarity is greater than a first similarity threshold, it is determined that the voice signal preliminarily matches the preset speaker.

[0056] In another embodiment, the voiceprint matching model in the above Figure 4The second matching unit 311 is specifically used to: calculate the second Mel frequency cepstral coefficient of the complete voiceprint feature, and input the second Mel frequency cepstral coefficient into the pre-trained voiceprint matching model to obtain the second cosine similarity between the speech signal and the preset speaker; if the second cosine similarity is greater than the second similarity threshold, then the speech signal is determined to match the preset speaker.

[0057] In another embodiment, the above Figure 3 The device wake-up device based on voiceprint information is also specifically used for: acquiring current environmental information; the current environmental information includes the noise level of the current environment; and dynamically adjusting the second similarity threshold based on the current environmental information.

[0058] In another embodiment, the above Figure 3 The speech judgment module 30 is specifically used for: extracting basic features from the sound signal; calculating the short-time energy and zero-crossing rate of the sound signal based on the basic features; and determining that a speech signal exists in the sound signal if both the short-time energy and the zero-crossing rate are within a preset range.

[0059] This application also provides an electronic device, in some embodiments, referring to... Figure 5 As shown, the electronic device 700 includes an input unit 710, a memory 720, a processor 730, and an output unit 740. The memory 720 stores program instructions that can be executed on the processor 730. The processor 730 can execute the device wake-up method and / or technical solution based on voiceprint information in the foregoing embodiments by calling the program instructions. The electronic device 700 can be a mobile terminal device such as a mobile phone or a computer.

[0060] Furthermore, embodiments of this application also provide a computer-readable storage medium for storing a computer program that executes a device wake-up method based on voiceprint information. For example, computer program instructions, when executed by a computer, can invoke or provide the methods and / or technical solutions according to this application through the operation of the computer. The program instructions that invoke the methods of this application may be stored in a fixed or removable storage medium, and / or transmitted via data streams in broadcast or other signal carrying media, and / or stored in a storage medium that operates according to the program instructions.

[0061] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be realized by universal computing devices, which can be centralized on a single computing device or distributed on a network composed of multiple computing devices, and optionally, they can be realized by program codes executable by the computing devices, so that they can be stored in storage devices and executed by the computing devices, or they can be respectively manufactured into individual integrated circuit modules, or multiple modules or steps among them can be manufactured into a single integrated circuit module to realize. Thus, the present application is not limited to any specific combination of hardware and software.

[0062] The technical features of the above embodiments can be integrated in any manner. In order to make the description simple, all possible integrations of the technical features in the above embodiments are not described, however, as long as the integration of the technical features does not exist contradictions, it should be considered as the scope of the present disclosure.

[0063] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the protection scope of the present application. Therefore, the protection scope of the present application patent should be subject to the appended claims.

Claims

1. A voiceprint information-based device wake-up method, characterized in that, The method comprises: real-time acquisition of a sound signal, and determination of whether a voice signal exists in the sound signal according to short-time energy and a zero-crossing rate of the sound signal; if the voice signal exists in the sound signal, determination of whether the voice signal matches a preset speaker; if the voice signal matches the preset speaker, waking up a target device. 2.The voiceprint information-based device wake-up method of claim 1, wherein, The determination of whether the voice signal matches the preset speaker comprises: extracting simplified voiceprint features from the voice signal, and determining whether the voice signal preliminarily matches the preset speaker based on the simplified voiceprint features; if the voice signal preliminarily matches the preset speaker, extracting complete voiceprint features from the voice signal, and determining whether the voice signal matches the preset speaker based on the complete voiceprint features. 3.The voiceprint information-based device wake-up method of claim 2, wherein, The determination of whether the voice signal preliminarily matches the preset speaker based on the simplified voiceprint features comprises: calculating first mel-frequency cepstral coefficients of the simplified voiceprint features, and inputting the first mel-frequency cepstral coefficients into a pre-trained voiceprint matching model to obtain a first cosine similarity between the voice signal and the preset speaker; if the first cosine similarity is greater than a first similarity threshold, it is determined that the voice signal preliminarily matches the preset speaker.

4. The voiceprint information-based device wake-up method of claim 3, wherein, The determination of whether the voice signal matches the preset speaker based on the complete voiceprint features comprises: calculating second mel-frequency cepstral coefficients of the complete voiceprint features, and inputting the second mel-frequency cepstral coefficients into the pre-trained voiceprint matching model to obtain a second cosine similarity between the voice signal and the preset speaker; if the second cosine similarity is greater than a second similarity threshold, it is determined that the voice signal matches the preset speaker.

5. The voiceprint information-based device wake-up method of claim 4, wherein, The method further comprises: obtaining current environment information; the current environment information comprises a noisy degree of a current environment; based on the current environment information, dynamically adjusting the second similarity threshold. 6.The voiceprint information-based device wake-up method of claim 1, wherein, The determination of whether a voice signal exists in the sound signal according to short-time energy and a zero-crossing rate of the sound signal comprises: extracting basic features in the sound signal; the basic features comprise amplitude, frame length and frame shift of the sound signal; calculating the short-time energy and the zero-crossing rate of the sound signal according to the basic features; if the short-time energy and the zero-crossing rate are both in a preset range, it is determined that the voice signal exists in the sound signal.

7. A device wake-up apparatus based on voiceprint information, characterized in that, The device comprises: a voice determination module, configured to real-time acquisition of a sound signal, and determination of whether a voice signal exists in the sound signal according to short-time energy and a zero-crossing rate of the sound signal; a person matching module, configured to, if the voice signal exists in the sound signal, determination of whether the voice signal matches a preset speaker; a device wake-up module, configured to, if the voice signal matches the preset speaker, waking up a target device.

8. An electronic device comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, The processor executes the computer program to implement the device wake-up method based on voiceprint information in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer program is stored on the computer readable storage medium, and the computer program is executed by the processor to implement the device wake-up method based on voiceprint information in any one of claims 1 to 6.

10. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the voiceprint information-based device wake-up method of any one of claims 1 to 6.