Equipment awakening method, electronic equipment and computer readable storage medium

By acquiring audio streams and environmental information from terminal devices and using a preset voice wake-up model for preliminary recognition and secondary verification, the problem of low wake-up accuracy caused by limited computing resources of terminal devices is solved, and a higher device wake-up accuracy is achieved.

CN121963720APending Publication Date: 2026-05-01ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG DAHUA TECH CO LTD
Filing Date
2025-12-19
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Due to limited computing resources in terminal devices, existing device wake-up methods rely on speech recognition models with low accuracy, resulting in insufficient device wake-up accuracy.

Method used

By acquiring the current audio stream and environmental information, and using a preset voice wake-up model for initial recognition, the target wake-up range is determined based on the environmental information, and a second verification is performed to confirm the existence of the wake-up word, thereby improving the wake-up accuracy.

Benefits of technology

By using a secondary verification method for the wake word, the accuracy of device wake-up is improved, adapting to current environmental factors and enhancing wake-up accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963720A_ABST
    Figure CN121963720A_ABST
Patent Text Reader

Abstract

The invention discloses a device awakening method, an electronic device and a computer readable storage medium. The method comprises the steps of obtaining a current audio stream and information of a current environment where equipment is located; inputting the current audio stream into a preset voice wake-up model to obtain a model output result; determining a target wake-up range from a preset audio wake-up range according to the current environment information in response to the condition that the model output result represents that a preset wake-up word of at least one preset device exists in the current audio stream, and responding to the condition that a matching result between the current audio stream and a preset audio template sample is in the target wake-up range, if yes, determining that a target wake-up word for waking up equipment exists in the current audio stream; and performing wakeup processing on the equipment based on the target wakeup word. Therefore, the device wakeup accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Device wake-up methods, electronic devices and computer-readable storage media Technical Field

[0001] This invention relates to the field of audio processing technology, and more particularly to a device wake-up method, an electronic device, and a computer-readable storage medium. Background Technology

[0002] As smart devices become increasingly popular, their functions are also constantly expanding. Among these, voice wake-up functionality is gradually becoming one of the important features of smart devices. Smart devices, such as mobile phone voice assistants and smart speakers, can be woken up by voice and then enter working mode, without having to remain in working mode continuously.

[0003] Current device wake-up methods typically involve inputting collected audio into a pre-defined speech recognition model, and then directly waking the device based on the wake-up word output by the model. This method relies on the accuracy of the speech recognition model. However, to protect privacy and avoid sending private voice data acquired by the terminal to servers, some smart devices typically deploy the pre-defined speech recognition model on terminal products with limited computing resources. Because terminal products cannot accommodate computationally intensive models, their recognition accuracy is insufficient, resulting in low device wake-up accuracy. Summary of the Invention

[0004] The main technical problem addressed by this application is to provide a device wake-up method, an electronic device, and a computer-readable storage medium, thereby improving the accuracy of device wake-up.

[0005] To address the aforementioned technical problems, this application provides a device wake-up method, comprising: acquiring a current audio stream and current environmental information of the device; inputting the current audio stream into a preset voice wake-up model to obtain a model output result; responding to the model output result indicating that there is a preset wake-up word for at least one preset device in the current audio stream, determining a target wake-up range from a preset audio wake-up range based on the current environmental information; responding to the matching result between the current audio stream and a preset audio template sample being within the target wake-up range, determining that there is a target wake-up word for waking up the device in the current audio stream; and performing wake-up processing on the device based on the target wake-up word.

[0006] In one embodiment, the step of determining the target wake-up range from the preset audio wake-up range based on the current environment information includes: determining the range division boundary of the preset audio wake-up range based on the current environment information; and dividing the preset audio wake-up range according to the range division boundary to obtain the target wake-up range.

[0007] In one embodiment, the step of determining the range boundary of the preset audio wake-up range based on the current environment information includes: determining adjustment information based on the comparison result between the current environment information and preset standard environment information; adjusting the preset standard fitting line according to the adjustment information to obtain the target fitting line; and determining the target fitting line as the range boundary of the preset audio wake-up range.

[0008] In one embodiment, the sample audio includes multiple positive sample audios containing a target wake-up word and multiple negative sample audios not containing a target wake-up word. Before the step of adjusting the preset standard fitting line according to the adjustment information to obtain the target fitting line, the method further includes: constructing matching points corresponding to each positive sample audio with the matching score between each positive sample audio and the preset positive audio template sample as the abscissa and the matching score between each positive sample audio and the preset negative audio template sample as the ordinate; constructing matching points corresponding to each negative sample audio with the matching score between each negative sample audio and the preset positive audio template sample as the abscissa and the matching score between each negative sample audio and the preset negative audio template sample as the ordinate; and performing linear fitting processing on each matching point to obtain the preset standard fitting line.

[0009] In one embodiment, the current environment information includes the current time and current ambient noise when the current audio stream is acquired. The step of determining the range boundary of the preset audio wake-up range based on the current environment information includes: determining a wake-up period based on the time interval of the current time; determining a wake-up environment state based on the noise interval of the current noise; determining corresponding fitting line parameters from a preset fitting line parameter mapping table based on the wake-up period and the wake-up environment state, wherein the preset fitting line parameter mapping table includes the correspondence between the preset wake-up period, the preset wake-up environment state, and the preset fitting line parameters; constructing the target fitting line based on the fitting line parameters; and determining the target fitting line as the range boundary of the preset audio wake-up range.

[0010] In one embodiment, before the step of determining the target wake-up range from the preset audio wake-up range based on the current environment information, the method further includes: constructing a target coordinate system with the matching score between the obtained sample audio and the preset positive audio template sample as the horizontal axis and the matching score between the sample audio and the preset negative audio template sample as the vertical axis; and determining the first quadrant region in the target coordinate system as the preset audio wake-up range.

[0011] In one embodiment, the preset audio template sample includes a preset positive audio template sample and a preset negative audio template sample. Before the step of determining that a target wake-up word for waking up the device exists in the current audio stream in response to the matching result between the current audio stream and the preset audio template sample being within the target wake-up range, the method further includes: obtaining a first matching score between the current audio stream and the preset positive audio template sample and a second matching score between the current audio stream and the preset negative audio template sample; determining whether a target point with the first matching score as the abscissa and the second matching score as the ordinate is within the target wake-up range; if so, determining that the matching result between the current audio stream and the preset audio template sample is within the target wake-up range.

[0012] In one embodiment, the preset voice wake-up model includes an acoustic analysis module, a decoding module, and a confidence decision module. The step of inputting the current audio stream into the preset voice wake-up model to obtain the model output includes: inputting the current audio features of the current audio stream into the acoustic analysis module to obtain the probability distribution of each phoneme; inputting the probability distribution of each phoneme into the decoding module to obtain a word sequence; and inputting the word sequence into the confidence decision module to obtain the model output.

[0013] The above scheme involves inputting the current audio stream into a preset voice wake-up model to obtain the model's output. Responding to the model output indicating the presence of a preset wake-up word for at least one preset device in the current audio stream, a target wake-up range is determined from the preset audio wake-up range based on the current environmental information. If the matching result between the current audio stream and a preset audio template sample falls within the target wake-up range, then a target wake-up word for waking up a device is confirmed in the current audio stream. The device is then woken up based on this target wake-up word. Thus, when the model output indicates the presence of wake-up information for at least one preset device in the current audio stream, the target wake-up range determined based on the device's current environmental information is used to further verify the presence of a target wake-up word in the current audio stream. This secondary verification improves the accuracy of wake-up information and consequently, the accuracy of device wake-up. Furthermore, determining the target wake-up range based on the current environmental information ensures that the target wake-up range better reflects the actual wake-up scenario, further enhancing the accuracy of device wake-up. Attached Figure Description

[0014] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them: Figure 1 is a flowchart of an exemplary embodiment of the device wake-up method shown in this application; Figure 2 is a structural diagram of an exemplary embodiment of the preset voice wake-up model shown in this application; Figure 3 is a schematic diagram of the region division by the fitting line shown in this application; Figure 4 is a flowchart of another exemplary embodiment of the device wake-up method shown in this application; Figure 5 is a block diagram of a device wake-up device shown in an exemplary embodiment of this application; Figure 6 is a structural diagram of an embodiment of the electronic device provided in this application; Figure 7 is a structural diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation

[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only for explaining this application and not for limiting it. Furthermore, it should be noted that, for ease of description, only the parts related to this application are shown in the accompanying drawings, not all structures. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0016] First, it's important to note that as smart devices become increasingly popular, their functionality is also constantly expanding. Among these features, voice wake-up has gradually become a crucial function. Smart devices, such as mobile phone voice assistants and smart speakers, can be woken up by voice and then enter working mode, without needing to remain in working mode continuously. Current device wake-up methods typically involve inputting the collected audio into a preset speech recognition model, which then directly wakes the device based on the wake-up word output by the model. This method relies on the accuracy of the speech recognition model. However, to protect privacy and avoid sending private voice data acquired by the terminal to servers, some smart devices typically deploy the preset speech recognition model on terminal products with limited computing resources. Because terminal products cannot accommodate computationally intensive models, the model's recognition accuracy is not high enough, resulting in insufficient accuracy for device wake-up.

[0017] Based on this, this application provides a device wake-up method, an electronic device, and a computer-readable storage medium. For details, please refer to Figure 1, which is a schematic flowchart illustrating an exemplary embodiment of a device wake-up method according to this application.

[0018] The execution entity of a device wake-up method can be a terminal device, a server, or other processing device. The terminal device can be a computer, mobile device, terminal, computing device, vehicle-mounted device, etc. The execution entity of the device wake-up method can also be a device wake-up device. In some possible implementations, the device wake-up method can be implemented by a processor calling computer-readable instructions stored in memory. The execution entity of the device wake-up method can also be a big data cluster. A big data cluster is a computer system architecture formed by multiple computers connected through a network. The big data cluster can be deployed on a private cloud built with K8S (Kubernetes, a container orchestration engine).

[0019] Specifically, a device wake-up method in this embodiment includes the following steps: Step S110, obtaining the current audio stream and the current environment information of the device.

[0020] The current audio stream refers to the audio data acquired at the current moment. As an example, the device wake-up device uses a recording device to capture the audio at the current moment, obtaining the current audio stream. The recording device can be a microphone or a mobile phone, etc. As another example, the device wake-up device queries an audio database for the audio at the current moment, and determines the retrieved audio as the current audio stream. The audio database stores audio from multiple moments.

[0021] Current environmental information refers to environmental parameters at the current moment. This information may include at least one of the current noise level and the current time. Specifically, the device wake-up device queries the device's time to obtain the current moment; an environmental sensing device collects the noise at the current moment to obtain the current noise level. This environmental sensing device can be a sound level meter, a noise dosimeter, or an acoustic camera, etc.

[0022] Step S120: Input the current audio stream into the preset voice wake-up model to obtain the model output result.

[0023] The preset voice wake-up model is used to identify wake words in the audio stream. The model output includes the preset wake words for the preset device. The preset device can be the corresponding device that wakes up the target device, or it can be other devices besides the target device.

[0024] The device wake-up device inputs the current audio stream into a preset voice wake-up model and obtains the model output. As one example, the device wake-up device inputs the current audio stream into a preset voice wake-up model and obtains the model output. As another example, the device wake-up device performs feature extraction processing on the current audio stream to obtain current audio features; it then inputs these current audio features into the preset voice wake-up model and obtains the model output.

[0025] The device wake-up device performs feature extraction processing on the current audio stream to obtain the current audio features. Specifically, the device wake-up device can use methods such as averaging, specification, and clustering to perform feature extraction processing on the current audio stream to obtain the current audio features. Alternatively, it can use a preset audio template sample feature extraction network to perform feature extraction processing on the current audio stream to obtain the current audio features.

[0026] The device wake-up device inputs the current audio features into a preset voice wake-up model and obtains the model's output. As an example, the device wake-up device inputs the current audio features into the preset voice wake-up model and obtains a preset wake-up word output by the preset voice wake-up model. As another example, the device wake-up device inputs the current audio features into the preset voice wake-up model and obtains a word sequence output by the preset voice wake-up model; the preset wake-up word is determined based on the confidence level of the word sequence. For example, if the confidence level of the word sequence is greater than a preset confidence threshold, then the word sequence is determined as the preset wake-up word.

[0027] Step S130: In response to the model output result indicating that there is a preset wake-up word for at least one preset device in the current audio stream, determine the target wake-up range from the preset audio wake-up range based on the current environment information.

[0028] The preset wake-up word can be either text or wake-up audio. The wake-up text can be something like "flower" or "open the door," while the wake-up audio can be the spoken word "open." As one example, the device wake-up unit performs speech recognition on the current audio stream to obtain the recognized text. It then searches for the preset wake-up word within the recognized text. If the preset wake-up word exists, it's determined that the current audio stream contains the preset wake-up word; otherwise, it's determined that the current audio stream does not contain the preset wake-up word. As another example, the device wake-up unit extracts features from the current audio stream to obtain the current audio features. It then matches these current audio features with the audio features of the preset wake-up word to obtain a matching result. If the matching result indicates a successful match, it's determined that the current audio stream contains the preset wake-up word; if the matching result indicates an unsuccessful match, it's determined that the current audio stream does not contain the preset wake-up word.

[0029] The preset audio wake-up range includes all audio matching results between sample audio and preset audio template samples. Sample audio includes positive sample audio containing a wake-up word and negative sample audio not containing a wake-up word. Preset audio template samples include positive audio template samples containing a wake-up word and negative audio template samples not containing a wake-up word. The audio matching results can be audio matching scores. In other words, the preset audio wake-up range includes: audio matching scores between positive sample audio and positive audio template samples; audio matching scores between positive sample audio and negative audio template samples; audio matching scores between negative sample audio and negative audio template samples; and audio matching scores between negative sample audio and positive audio template samples.

[0030] The steps of the device wake-up device to obtain the audio matching results between sample audio and preset audio template samples include: extracting features from each sample audio to obtain the audio features of each sample audio; extracting features from the preset audio template samples to obtain the audio features of the preset audio template samples; and calculating the matching score between the audio features of each sample audio and the audio features of the preset audio template samples.

[0031] Audio features refer to numerical or symbolic representations extracted from audio signals that can represent certain key attributes of the audio. Audio features can be loudness, pitch, timbre, time-domain features, and frequency features, and can also be one or more of MFCC (Mel-Frequency Cepstral Coefficients), Ivector (speaker identification vector), and PLP (Perceptual Linear Predictive features), or combinations or splicing of multiple features.

[0032] Specifically, the device wake-up device can acquire audio features using methods such as averaging, specification, and clustering. It can also use a preset audio feature extraction network to extract features from each sample audio or preset positive and negative audio template samples to obtain audio features. Each audio feature can be a single feature or multiple features. The device wake-up device determines that the number of audio features is less than or equal to a preset number of features if the acquired remaining device resources are less than or equal to a preset number of resources; conversely, it determines that the number of audio features is greater than the preset number of features if the remaining device resources are greater than the preset number of resources. The preset number of features can be 1.

[0033] In one embodiment, the device wake-up device takes the average of multiple feature dimensions of the sample audio to obtain the audio features of the preset audio template sample; or, the device wake-up device selects the most representative feature from the multiple features of the sample audio as the audio features of the preset audio template sample; or, the device wake-up device uses the K-means clustering method to cluster the audio features of the sample audio and uses the cluster centers obtained by clustering as the audio features of the preset audio template sample.

[0034] The audio matching results of the sample audio include the positive sample matching score between the sample audio and the preset positive audio template sample, and the negative sample matching score between the sample audio and the preset negative audio template sample.

[0035] Specifically, the device wake-up device uses a preset matching model or preset matching algorithm to calculate the similarity between the audio features of each sample audio and the audio features of the preset audio template sample, thereby obtaining the audio similarity of each sample audio; the audio similarity of each sample audio is then converted into a score to obtain the audio matching result. The preset matching model can be a lightweight model that can measure similarity, and the preset matching algorithm can be a similarity algorithm or a feature matching algorithm.

[0036] In one embodiment, the device wake-up device uses a preset matching model or preset matching algorithm to calculate the similarity between the audio features of each sample audio and the audio features of a preset positive audio template sample, thereby obtaining the audio similarity between each sample audio and the preset positive audio template sample; the audio similarity between each sample audio and the preset positive audio template sample is converted into a score to obtain a positive sample matching score; the preset matching model or preset matching algorithm is used to calculate the similarity between the audio features of each sample audio and the audio features of a preset negative audio template sample, thereby obtaining the audio similarity between each sample audio and the preset positive audio template sample; the audio similarity between each sample audio and the preset positive audio template sample is converted into a score to obtain a negative sample matching score.

[0037] The preset audio wake-up range can also be represented as the first quadrant region in a coordinate system constructed with the matching score between the acquired sample audio and the preset positive audio template sample as the horizontal axis and the matching score between the sample audio and the preset negative audio template sample as the vertical axis.

[0038] The target wake-up range includes the range of audio matching scores between positive sample audio containing the wake-up word and preset audio template samples.

[0039] The device wake-up device determines the target wake-up range from a preset audio wake-up range based on the current environmental information. Specifically, the device wake-up device determines the corresponding dividing line from a preset dividing line mapping table based on the current environmental information. The preset dividing line mapping table includes the correspondence between preset environmental information and preset dividing lines. Based on the dividing line, the preset audio wake-up range is divided into two sub-regions. The sub-region with the audio matching score between the positive sample audio containing the wake-up word and the standard audio containing the wake-up word is determined as the target wake-up range.

[0040] Step S140: In response to the matching result between the current audio stream and the preset audio template sample being within the target wake-up range, it is determined that there is a target wake-up word for waking up the device in the current audio stream.

[0041] The matching result refers to the data characterizing the degree of matching between the current audio stream and the preset audio template sample. The matching result includes the target audio matching score. Specifically, the device wake-up device obtains the current audio features of the current audio stream and the audio feature template of the preset audio template sample; it uses a feature matching algorithm to calculate the matching value based on the current audio features and the audio feature template; and it determines the target audio matching score by taking the reciprocal of the matching value. The feature matching algorithm can be DTW (Dynamic Time Warping).

[0042] Specifically, the device wake-up device queries the audio matching results within the target wake-up range. If a match is found, it determines that the matching result is within the target wake-up range.

[0043] Step S150: Perform wake-up processing on the device based on the target wake-up word.

[0044] Specifically, the device wake-up device uses a target wake-up word to wake up the device.

[0045] The device wake-up device further includes: in response to the fact that the matching result between the current audio stream and the preset audio template sample is not within the target wake-up range, the step of returning to step S110 to obtain the current audio stream and the current environment information of the device.

[0046] As can be seen, by inputting the current audio stream into a preset voice wake-up model, the model output is obtained. Responding to the model output indicating the presence of a preset wake-up word for at least one preset device in the current audio stream, a target wake-up range is determined from the preset audio wake-up range based on the current environmental information. If the matching result between the current audio stream and the preset audio template sample falls within the target wake-up range, then a target wake-up word for waking up a device is determined to exist in the current audio stream. The device is then woken up based on this target wake-up word. Therefore, when the model output indicates the presence of wake-up information for at least one preset device in the current audio stream, the target wake-up range determined based on the device's current environmental information is used to further verify whether a target wake-up word exists in the current audio stream. This secondary verification of the wake-up word in the current audio stream improves the accuracy of the wake-up information, thereby increasing the accuracy of device wake-up. Furthermore, determining the target wake-up range based on the current environmental information takes into account current environmental factors, making the target wake-up range more closely aligned with the actual device wake-up scenario, which is beneficial for improving the accuracy of device wake-up.

[0047] The preset voice wake-up model includes an acoustic analysis module, a decoding module, and a confidence decision module. The steps of inputting the current audio stream into the preset voice wake-up model and obtaining the model output include: inputting the current audio features of the current audio stream into the acoustic analysis module to obtain the probability distribution of each phoneme; inputting the probability distribution of each phoneme into the decoding module to obtain the word sequence; and inputting the word sequence into the confidence decision module to obtain the model output.

[0048] Current audio features refer to the audio characteristics of the current audio stream.

[0049] In one embodiment, as shown in Figure 2, the preset voice wake-up model may further include a feature extraction module, an acoustic analysis module, a decoding module, and a confidence decision module connected in sequence. The device wake-up device inputs the current audio stream into the feature extraction module, which performs feature extraction processing on the current audio stream to obtain the current audio features; the current audio features are input into the acoustic analysis module to obtain the probability distribution of each phoneme; the probability distribution of each phoneme is input into the decoding module to obtain the word sequence; and the word sequence is input into the confidence decision module to obtain the preset wake-up word output by the preset voice wake-up model.

[0050] The acoustic analysis module is used to model the probabilistic relationship between audio features and phonemes. The acoustic analysis module includes a pre-alignment module and a deep learning network module. The device wake-up device inputs the current audio features into the pre-alignment module to obtain aligned current audio features; then, it inputs the aligned current audio features into the deep learning network module to obtain the probability distribution of each phoneme.

[0051] The pre-alignment module can be a GMM-HMM (Gaussian Mixture Model-Hidden Markov Model) network. The device wake-up device trains the GMM-HMM network using low-dimensional features to obtain the pre-alignment module. The GMM-HMM network is used to output the frame alignment result. Each integer within parentheses in the frame alignment result corresponds to a unique phoneme. For example, the frame alignment result is shown below:

[0052] In one embodiment, the deep learning network module is a deep neural network. The deep neural network can be a convolutional neural network (CNN), a recurrent neural network (RNN), a time-delay neural network (TDNN), etc.

[0053] The decoding module is used to obtain a word sequence based on the probability distribution of each phoneme. The decoding module optimizes the word order of phonemes and probabilities through the language module to obtain words, and performs an optimal text search through a decoding search algorithm to obtain the optimal text, that is, the word sequence. Specifically, the language model predicts the obtained word X and gets the most likely word after X. For example, if X is "今" (today), the language model predicts that the word after "今" is "天" (day).

[0054] The confidence judgment module is used to judge the word sequence according to the confidence of each word in the word sequence to obtain the wake-up word. The judgment method of the confidence judgment module can be to set a decoding path score threshold, a confidence threshold, a phoneme number threshold, etc. For example, the confidence judgment module determines the words with a confidence greater than or equal to the confidence threshold as the preset wake-up words.

[0055] In one embodiment, the word sequence output by the decoding module is "小花" (little flower), and the judgment method in the confidence judgment module is to compare the decoding path scores. The higher the decoding path score, the greater the reliability of the output result of the decoding module. For example, the decoding path score threshold is 2. If the confidence judgment module calculates the decoding path score as 2.2 and judges that 2.2 > 2, the result output by the preset voice wake-up model is "小花" (little flower). If the confidence judgment module calculates the decoding path score as 1.9 and judges that 1.9 < 2, the preset voice wake-up model does not output a result.

[0056] Before the step of the device wake-up device determining the target wake-up range from the preset audio wake-up range according to the current environmental information, the method further includes: constructing a target coordinate system with the matching score between the acquired sample audio and the preset positive audio template sample as the horizontal axis and the matching score between the sample audio and the preset negative audio template sample as the vertical axis; determining the first quadrant area in the target coordinate system as the preset audio wake-up range.

[0057] In one embodiment, the device wake-up device constructs a two-dimensional coordinate system, determines the horizontal axis as the matching score between the sample audio and the preset positive audio template sample, and determines the vertical axis as the matching score between the sample audio and the preset negative audio template sample to obtain the target coordinate system; and determines the first quadrant area in the target coordinate system as the preset audio wake-up range.

[0058] The step of the device wake-up device determining the target wake-up range from the preset audio wake-up range according to the current environmental information includes: determining the range division boundary of the preset audio wake-up range according to the current environmental information; performing a division process on the preset audio wake-up range according to the range division boundary to obtain the target wake-up range.

[0059] The range division boundary is the boundary used to divide the range. The range division boundary can be a target fitting line, and the target fitting line is used to divide the preset audio wake-up range into multiple sub-regions to obtain the target wake-up range.

[0060] As an example, the step of determining the range boundary of the preset audio wake-up range based on the current environmental information by the device wake-up device includes: determining adjustment information based on the comparison result between the current environmental information and the preset standard environmental information; adjusting the preset standard fitting line according to the adjustment information to obtain the target fitting line; and determining the target fitting line as the range boundary of the preset audio wake-up range.

[0061] The preset environmental standard information is the preset standard environmental information. For example, the preset environmental standard information is the time within the preset standard time interval and the noise within the preset standard noise interval.

[0062] The device wake-up device determines adjustment information based on a comparison between the current environmental information and preset standard environmental information. Specifically, if the current time is within the preset standard time interval and the current noise is within the preset standard noise interval, then the current environmental information is consistent with the preset standard environmental information, and the adjustment information is no adjustment. If the current time is not within the preset standard time interval or the current noise is greater than the maximum value of the preset standard noise interval, then the slope is increased by a first preset value as the adjustment information. If the current time is not within the preset standard time interval and the current noise is greater than the maximum value of the preset standard noise interval, then the slope is increased by a second preset value as the adjustment information. If the current time is within the preset standard time interval and the current noise is less than the minimum value of the preset standard noise interval, then the slope is decreased by a third preset value as the adjustment information.

[0063] Before the device wake-up device adjusts the preset standard fitting line according to the adjustment information to obtain the target fitting line, the method further includes: constructing matching points corresponding to each positive sample audio with the matching score between each positive sample audio and the preset positive audio template sample as the abscissa and the matching score between each positive sample audio and the preset negative audio template sample as the ordinate; constructing matching points corresponding to each negative sample audio with the matching score between each negative sample audio and the preset positive audio template sample as the abscissa and the matching score between each negative sample audio and the preset negative audio template sample as the ordinate; and performing linear fitting processing on each matching point to obtain the preset standard fitting line.

[0064] The sample audio includes multiple positive sample audios containing the target wake-up word and multiple negative sample audios not containing the target wake-up word. Specifically, the device wake-up device selects multiple audios containing the target wake-up word from a preset audio database as preset positive sample audios and selects multiple audios not containing the target wake-up word as preset negative sample audios.

[0065] The preset audio template samples include a preset positive audio template sample and a preset negative audio template sample, which serve as standard templates. Specifically, the device wake-up device selects the most representative sample audio from the preset positive sample audio and the preset negative sample audio as the preset audio template samples. It should be noted that the preset audio template samples can be pre-selected audio, or they can be positive sample audio and negative sample audio sent by the client.

[0066] The device wake-up device obtains the matching score between each sample audio and a preset positive audio template sample. Specifically, the device wake-up device obtains the audio features of each sample audio and the audio features of the preset positive audio template sample; it performs matching processing on the audio features of each sample audio and the audio features of the preset positive audio template sample to obtain a matching score. The sample audio can be either positive or negative sample audio.

[0067] The device wake-up device obtains the matching score between each sample audio and a preset negative audio template sample. Specifically, the device wake-up device obtains the audio features of each sample audio and the audio features of the preset negative audio template sample; it performs matching processing on the audio features of each sample audio and the audio features of the preset negative audio template sample to obtain a matching score.

[0068] As can be seen, by selecting preset audio template samples to construct the core features of different categories of samples, the difference between positive and negative sample audio is maximized. At the same time, by pre-setting preset audio template samples or receiving audio sent by the client, the requirements for personalized deployment are enhanced.

[0069] The device wake-up device performs linear fitting on each matching point to obtain a preset standard fitting line. Specifically, the device wake-up device plots the coordinates of the matching points corresponding to each sample audio in the target coordinate system; performs linear fitting on the coordinates of each matching point to obtain multiple fitting lines, and selects the preset standard fitting line from the multiple fitting lines.

[0070] As an example, the device wake-up device uses a preset linear fitting method to linearly fit the coordinates of each matching point, resulting in multiple fitted lines. The linear fitting method can be the least squares method, gradient descent method, etc.

[0071] As another example, the device wake-up device performs linear fitting on the coordinates of each matching point to obtain multiple fitting lines, and selects a preset standard fitting line from these multiple fitting lines. Specifically, the device wake-up device uses multiple fitting lines to segment the first quadrant of the two-dimensional coordinate system on which the coordinates of the matching points are plotted, obtaining a first sub-region where the positive sample audio is located and a second sub-region where the negative sample audio is located; it counts the number of matching points in the first sub-region to obtain the count results; and it selects a preset standard fitting line from the multiple fitting lines based on the count results corresponding to each fitting line. The matching score corresponding to the matching point in the first sub-region is the audio matching score included in the wake-up range.

[0072] Referring to Figure 3, the device wake-up device plots the matching point coordinates corresponding to each sample audio in the target coordinate system, using the matching score between each sample audio and a preset positive audio template sample as the x-axis and the matching score between each positive sample audio and a preset negative audio template sample as the y-axis. Circles represent matching points corresponding to positive sample audio, and crosses represent matching points corresponding to negative sample audio. In the target coordinate system with the matching point coordinates plotted, the device wake-up device draws multiple fitting lines. Each fitting line divides the first quadrant of the target coordinate system into two sub-regions: the first sub-region is the area containing positive sample audio enclosed by the fitting line and the x-axis, and the second sub-region is the area containing negative sample audio enclosed by the fitting line and the y-axis. The number of matching points for positive sample audio and the number of matching points for negative sample audio within the first sub-region are counted. The number of matching points for negative sample audio within the first sub-region is determined as the false alarm count, the ratio between the number of matching points for positive sample audio in the first sub-region and the total number of positive sample audio is determined as the recognition rate, and the slope and intercept of the fitting line are determined as the fitting line parameters.

[0073] In one embodiment, the device wake-up device calculates the positive sample matching score between 5 positive sample audios and a single positive sample audio feature, and the negative sample matching score between 5 positive sample audios and a single negative sample audio feature; it also calculates the positive sample matching score between 15 negative sample audios and a single positive sample audio feature, and the negative sample matching score between 15 negative sample audios and a single negative sample audio feature. Matching points for each sample audio are plotted in a two-dimensional coordinate system with the positive sample matching score as the x-axis and the negative sample matching score as the y-axis. The matching points for both positive and negative sample audios are then divided according to preset fitting lines, resulting in the slope, intercept, recognition rate, and false alarm count of each fitting line, as shown in Table 1.

[0074] As shown in Table 1, a slope of 1 and an intercept of 15 correspond to a recognition rate of 90% and 4 false alarms; a slope of 1 and an intercept of 5 correspond to a recognition rate of 80% and 2 false alarms; and a slope of 3 and an intercept of 1 correspond to a recognition rate of 65% and 0 false alarms. Different fitted lines correspond to different stages of recognition rate and false alarm count. The more stringent the slope and intercept of the fitted line, the lower the recognition rate may be. In this example, the parameters with (slope, intercept) of (1,5) are selected as the parameters of the preset standard fitted line, and the other groups of parameters are backup parameters for dynamic adjustment.

[0075] Table 1

[0076] The statistical results can include the number of false alarms and the recognition rate. The device wake-up device selects a preset standard fitting line from multiple fitting lines based on the statistical results corresponding to each fitting line. Specifically, the device wake-up device determines the fitting line with a recognition rate greater than a preset recognition rate threshold and a number of false alarms less than a preset false alarm threshold as the preset standard fitting line.

[0077] The device wake-up mechanism also includes: setting a correspondence between the number of false alarms and / or the recognition rate and environmental information; wherein, the number of false alarms, the recognition rate, and the fitting line parameters are correlated. Different recognition rates and false alarm counts are selected based on the device wake-up time period to obtain the corresponding fitting line. For example, for nighttime, the boundary parameters need to be tightened, meaning that the later the time, the fewer false alarms the fitting line should be selected to reduce false alarms and avoid affecting user experience. If the device has abundant available computing resources, the device wake-up environment status is incorporated to assist in selecting the fitting line parameters. For example, in a high-noise environment, the boundary is tightened, meaning that the higher the noise level, the fewer false alarms the fitting line should be selected to reduce potential false alarms caused by noise. It should be noted that when selecting fitting line parameters, the priority of the device wake-up time period is higher than the priority of the device wake-up environment status.

[0078] In one embodiment, when the device wakes up during the first half of the night and the device wakes up in a quiet environment, there is less external interference, so a lower recognition rate and a smaller number of false alarms can be selected, such as 70% and 0. When the device wakes up during the morning and the device wakes up in a noisy environment, there is more external interference, so the recognition rate requirement is increased, such as 75%.

[0079] The device wake-up device adjusts the preset standard fitting line according to the adjustment information to obtain the target fitting line. If the device wake-up device responds to the adjustment information as "no adjustment," it determines the preset standard fitting line as the target fitting line; if the adjustment information indicates an increase in slope by a first preset value, it increases the slope of the preset standard fitting line by the first preset value to obtain the target fitting line; if the adjustment information indicates an increase in slope by a second preset value, it increases the slope of the preset standard fitting line by the second preset value to obtain the target fitting line; if the adjustment information indicates a decrease in slope by a third preset value, it decreases the slope of the preset standard fitting line by the third preset value to obtain the target fitting line.

[0080] The parameters of the preset standard fitting line can include a slope of 1 and an intercept of 5.

[0081] As another example, the current environment information includes the current time and current ambient noise when the current audio stream is acquired. The step of determining the range boundary of the preset audio wake-up range based on the current environment information includes: determining the wake-up period based on the time interval of the current time; determining the wake-up environment state based on the noise interval of the current noise; determining the corresponding fitting line parameters from the preset fitting line parameter mapping table based on the wake-up period and the wake-up environment state, the preset fitting line parameter mapping table including the correspondence between the preset wake-up period, the preset wake-up environment state and the preset fitting line parameters; constructing a target fitting line based on the fitting line parameters; and determining the target fitting line as the range boundary of the preset audio wake-up range.

[0082] The device wake-up device determines the device wake-up period based on the time interval in which the current moment falls. Specifically, the device wake-up device divides a day into multiple time intervals, each time interval corresponding to a wake-up period; the wake-up period corresponding to the time interval in which the current moment falls is determined as the device wake-up period.

[0083] In one embodiment, the 24 hours are divided into four time intervals: the wake-up period corresponding to the time interval 0:01-6:00 is the second half of the night; the wake-up period corresponding to the time interval 6:01-12:00 is the morning; the wake-up period corresponding to the time interval 12:01-18:00 is the afternoon; and the wake-up period corresponding to the time interval 18:01-24:00 is the first half of the night. If the current time is 23:00, then the time interval is 18:01-24:00, and the corresponding wake-up period is the first half of the night; if the current time is 7:00, then the time interval is 6:01-12:00, and the corresponding wake-up period is the morning.

[0084] The device wake-up device determines the device wake-up environment state based on the noise range in which the current noise is located. For example, if the device wake-up device responds to the current noise being in the first noise range, it determines the device wake-up environment state to be noisy; if it responds to the current noise being in the second noise range, it determines the device wake-up environment state to be quiet. The first noise range is a noise range greater than or equal to a preset noise threshold, and the second noise range is a noise range less than the preset noise threshold.

[0085] For example, the device wake-up device compares the current noise with a preset noise threshold. If the current noise is greater than or equal to the preset noise threshold, the device wake-up environment is determined to be noisy. If the current noise is less than the preset noise threshold, the device wake-up environment is determined to be quiet.

[0086] Fitted line parameters are parameters used to characterize the properties of the fitted line. Fitted line parameters include the slope and intercept of the fitted line.

[0087] In one embodiment, if the device wake-up time is the first half of the night and the device wake-up environment is quiet, then the corresponding slope and intercept are 2 and 1, respectively; if the device wake-up time is the first half of the night and the device wake-up environment is noisy, then the corresponding slope and intercept are 3 and 1, respectively; if the device wake-up time is in the morning and the device wake-up environment is noisy, then the corresponding slope and intercept are 1 and 1, respectively.

[0088] A target fitted line is constructed based on the fitted line parameters. Specifically, the device wake-up device draws line segments based on the fitted line parameters to obtain the target fitted line. For example, the device wake-up device draws a straight line in a two-dimensional coordinate system according to the slope and intercept of the fitted line to obtain the target fitted line.

[0089] In one embodiment, the slope is 3 and the intercept is 1. Then, the coordinate point (0, 1) is determined in the two-dimensional coordinate system, and the straight line with a slope of 3 that passes through the coordinate point (0, 1) is determined as the target fitting line.

[0090] The preset audio template samples include preset positive audio template samples and preset negative audio template samples. Before the step of determining that a target wake-up word for waking up the device exists in the current audio stream in response to the matching result between the current audio stream and the preset audio template samples being within the target wake-up range, the method further includes: obtaining a first matching score between the current audio stream and the preset positive audio template samples and a second matching score between the current audio stream and the preset negative audio template samples; determining whether a target point with the first matching score as the horizontal axis and the second matching score as the vertical axis is within the target wake-up range; if so, determining that the matching result between the current audio stream and the preset audio template samples is within the target wake-up range.

[0091] The device wake-up device obtains a first matching score between the current audio stream and a preset positive audio template sample. Specifically, the device wake-up device obtains the current audio features of the current audio stream and the positive audio feature template of the preset positive audio template sample; performs similarity calculation on the current audio features and the positive audio feature template to obtain the positive audio similarity; and performs score conversion processing on the positive audio similarity to obtain the first matching score.

[0092] The device wake-up device obtains a second matching score between the current audio stream and a preset negative audio template sample. Specifically, the device wake-up device obtains the current audio features of the current audio stream and the negative audio feature template of the preset negative audio template sample; performs similarity calculation on the current audio features and the negative audio feature template to obtain a negative audio similarity; and performs score conversion processing on the negative audio similarity to obtain a second matching score.

[0093] The device wake-up device can use lightweight models, similarity algorithms, or feature matching algorithms to perform similarity calculations.

[0094] In one embodiment, the coordinate points corresponding to the current audio stream are plotted in the target coordinate system with the first matching score as the abscissa and the second matching score as the ordinate. The target coordinate system includes a target fitting line. If the coordinate points corresponding to the current audio stream are in the first sub-region enclosed by the target fitting line and the x-axis, the current audio stream is determined to be within the positive sample audio range, that is, the target audio matching result between the current audio stream and the preset audio template sample is within the target wake-up range, and the device is woken up. If the coordinate points corresponding to the current audio stream are in the second sub-region enclosed by the target fitting line and the y-axis, the current audio stream is determined to be within the negative sample audio range, and the device is not woken up.

[0095] As an example, the device wake-up device extracts features from the current audio stream to obtain the current audio features. It then extracts features from a preset audio template sample to obtain an audio feature template. The current audio features and the audio feature template are then dimension-unified to obtain a dimension-unified current audio feature sequence and an audio feature template sequence. The distance between the current audio feature sequence and the audio feature template sequence is calculated, and this distance value is determined as the audio similarity. The reciprocal of the audio similarity is determined as the matching score. The preset audio template sample can be a preset positive audio template sample or a preset negative audio template sample.

[0096] For example, the distance value satisfies the following formula:

[0097] In the above formula, Representing distance value, Characterizes the i-th dimension of the current audio feature sequence. The distance represents the i-th dimension of the audio feature template sequence, where n represents the total dimension of the current audio feature. A larger distance value indicates lower similarity.

[0098] The target audio matching score satisfies the following formula:

[0099] In the above formula, Characterizes the target audio matching score.

[0100] As another example, the device wake-up device extracts features from the current audio stream to obtain the current audio features of the current audio stream, extracts features from a preset audio template sample to obtain an audio feature template; unifies the dimensions of the current audio features and the audio feature template to obtain a dimension-unified current audio feature sequence and an audio feature template sequence; uses the DTW algorithm to calculate the dynamic time warp value between the current audio feature sequence and the audio feature template sequence, and determines the obtained dynamic time warp value as the audio similarity; the reciprocal of the dynamic time warp value is determined as the target audio matching score.

[0101] For example, the dynamic time warp value satisfies the following formula:

[0102]

[0103] In the above formula, Characterize the j-th dimension of the audio feature template sequence. The squared value representing the difference of the k-th feature. The dynamic time warping value is represented. A higher DTW result indicates lower similarity.

[0104] The target audio matching score satisfies the following formula:

[0105] Referring to Figure 4, the process of the device wake-up method provided in this embodiment is as follows: Step S410, obtain the current audio stream.

[0106] Step S420: Input the current audio stream into the preset voice wake-up model.

[0107] Step S430: Determine whether the preset voice wake-up model outputs a wake-up word. If yes, proceed to step S440; otherwise, return to step S410.

[0108] Step S440: Divide the preset audio wake-up range according to the obtained current environment information to obtain the target wake-up range.

[0109] Step S450: Obtain the matching result between the current audio stream and the preset audio template sample.

[0110] Step S460: Determine whether the matching result between the current audio stream and the preset audio template sample is within the target wake-up range; if yes, proceed to step S470; otherwise, return to step S410.

[0111] Step S470: Wake up the device and receive subsequent voice commands as command words for the device.

[0112] It can be seen that by obtaining device wake-up information in the current audio stream through the preset voice wake-up module, there is no need to retrain the model, which reduces the iteration time cost and system debugging difficulty, and the device resource consumption is low. At the same time, it achieves the goal of reducing false alarms of wake-up words.

[0113] Before the step of inputting the current audio features into a preset voice wake-up model and obtaining the model output, the device wake-up device further includes: constructing a training audio set, which includes training sample audio and actual wake-up words in the training sample audio; inputting the training sample audio and actual wake-up words in the training sample audio into the voice wake-up model to be trained, and obtaining the prediction result output by the voice wake-up model to be trained; calculating the loss value between the actual wake-up word and the prediction result; training the voice wake-up model to be trained with the goal of reducing the loss value, until the loss value meets the preset loss requirement; and determining the voice wake-up model to be trained that meets the requirement as the preset voice wake-up model.

[0114] Specifically, the device wake-up device collects multiple initial audios, including audios containing a wake-up word and noisy audios without a wake-up word; it sets text labels for each initial audio, which are used to represent the wake-up word; it performs data augmentation on the initial audios to obtain augmented initial audios; and it combines the initial audios and the augmented initial audios into a training audio set.

[0115] In one embodiment, if the wake word of the device is "Xiaohua", then audio containing "Xiaohua" and noise audio not containing "Xiaohua" are collected in various scenarios. The audio containing "Xiaohua" is taken as a positive sample and the text label is set as "Xiaohua". The noise audio not containing "Xiaohua" is taken as a negative sample and the text label is set as FREETEXT. FREETEXT represents any text.

[0116] Data augmentation methods can include adding reverberation, velocity perturbation, and noise. By augmenting the initial audio data, the diversity of the audio data can be increased, and the real-world device usage environment can be simulated, thereby increasing the accuracy of voice wake-up model training.

[0117] Figure 5 is a block diagram illustrating a device wake-up device according to an exemplary embodiment of this application. As shown in Figure 5, the exemplary device wake-up device 500 includes: an acquisition module 510, a model processing module 520, a target wake-up range determination module 530, a target wake-up word determination module 540, and a wake-up module 550. Specifically, the acquisition module 510 is used to acquire the current audio stream and the current environment information of the device.

[0118] The model processing module 520 is used to input the current audio stream into a preset voice wake-up model and obtain the model output result.

[0119] The target wake-up range determination module 530 is used to determine the target wake-up range from the preset audio wake-up range based on the current environment information, in response to the model output result indicating that there is a preset wake-up word of at least one preset device in the current audio stream.

[0120] The target wake-up word determination module 540 is used to determine that there is a target wake-up word for waking up the device in the current audio stream if the matching result between the current audio stream and the preset audio template sample is within the target wake-up range.

[0121] The wake-up module 550 is used to wake up the device based on the target wake-up word.

[0122] In this exemplary device wake-up device, the current audio stream is input into a preset voice wake-up model to obtain the model output. In response to the model output indicating the presence of a preset wake-up word for at least one preset device in the current audio stream, a target wake-up range is determined from a preset audio wake-up range based on the current environment information. If the matching result between the current audio stream and a preset audio template sample falls within the target wake-up range, then a target wake-up word for waking up the device is determined to exist in the current audio stream. The device is then woken up based on the target wake-up word. Thus, when the model output indicates the presence of wake-up information for at least one preset device in the current audio stream, the target wake-up range determined based on the device's current environment information is used to further verify whether a target wake-up word exists in the current audio stream. This secondary verification of the wake-up word in the current audio stream improves the accuracy of the wake-up information, thereby increasing the accuracy of device wake-up. Furthermore, determining the target wake-up range based on the current environment information takes into account environmental factors, making the target wake-up range more closely aligned with the actual device wake-up scenario, which further enhances the accuracy of device wake-up.

[0123] The functions of each module can be found in the device wake-up method implementation examples, and will not be repeated here.

[0124] To implement the device wake-up method of the above embodiments, this application proposes another electronic device. Please refer to FIG6 for details. FIG6 is a schematic diagram of the structure of an embodiment of the electronic device provided in this application.

[0125] Electronic device 600 includes memory 601 and processor 602, wherein memory 601 and processor 602 are coupled together.

[0126] The memory 601 is used to store program data, and the processor 602 is used to execute the program data to implement the device wake-up method of the above embodiment.

[0127] In this embodiment, processor 602 can also be referred to as CPU (Central Processing Unit). Processor 602 may be an integrated circuit chip with signal processing capabilities. Processor 602 can also be a general-purpose processor, digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor can be a microprocessor, or processor 602 can be any conventional processor.

[0128] This application also provides a computer-readable storage medium, as shown in FIG7. The computer-readable storage medium 700 is used to store program data 701. When the program data 701 is executed by the processor, it is used to implement the device wake-up method as described in the method embodiment of this application.

[0129] The methods involved in the device wake-up method embodiments of this application, when implemented as software functional units and sold or used as independent products, can be stored in a device, such as a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0130] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A device wake-up method, characterized in that, The method includes: acquiring the current audio stream and the current environment information of the device; inputting the current audio stream into a preset voice wake-up model to obtain the model output result; in response to the model output result indicating that there is a preset wake-up word for at least one preset device in the current audio stream, determining a target wake-up range from the preset audio wake-up range based on the current environment information; in response to the matching result between the current audio stream and a preset audio template sample being within the target wake-up range, determining that there is a target wake-up word for waking up the device in the current audio stream; and performing wake-up processing on the device based on the target wake-up word.

2. The method according to claim 1, characterized in that, The step of determining the target wake-up range from the preset audio wake-up range based on the current environment information includes: determining the range division boundary of the preset audio wake-up range based on the current environment information; and dividing the preset audio wake-up range according to the range division boundary to obtain the target wake-up range.

3. The method according to claim 2, characterized in that, The step of determining the range boundary of the preset audio wake-up range based on the current environment information includes: determining adjustment information based on the comparison result between the current environment information and the preset standard environment information; adjusting the preset standard fitting line according to the adjustment information to obtain the target fitting line; and determining the target fitting line as the range boundary of the preset audio wake-up range.

4. The method according to claim 3, characterized in that, The sample audio includes multiple positive sample audios containing the target wake-up word and multiple negative sample audios not containing the target wake-up word. Before the step of adjusting the preset standard fitting line according to the adjustment information to obtain the target fitting line, the method further includes: constructing matching points corresponding to each positive sample audio with the matching score between each positive sample audio and the preset positive audio template sample as the abscissa and the matching score between each positive sample audio and the preset negative audio template sample as the ordinate; constructing matching points corresponding to each negative sample audio with the matching score between each negative sample audio and the preset positive audio template sample as the abscissa and the matching score between each negative sample audio and the preset negative audio template sample as the ordinate; and performing linear fitting processing on each matching point to obtain the preset standard fitting line.

5. The method according to claim 2, characterized in that, The current environment information includes the current time and current ambient noise when the current audio stream is acquired. The step of determining the range boundary of the preset audio wake-up range based on the current environment information includes: determining a wake-up period based on the time interval of the current time; determining a wake-up environment state based on the noise interval of the current noise; determining corresponding fitting line parameters from a preset fitting line parameter mapping table based on the wake-up period and the wake-up environment state, wherein the preset fitting line parameter mapping table includes the correspondence between the preset wake-up period, the preset wake-up environment state, and the preset fitting line parameters; constructing the target fitting line based on the fitting line parameters; and determining the target fitting line as the range boundary of the preset audio wake-up range.

6. The method according to claim 1, characterized in that, Before the step of determining the target wake-up range from the preset audio wake-up range based on the current environment information, the method further includes: constructing a target coordinate system with the matching score between the obtained sample audio and the preset positive audio template sample as the horizontal axis and the matching score between the sample audio and the preset negative audio template sample as the vertical axis; and determining the first quadrant region in the target coordinate system as the preset audio wake-up range.

7. The method according to claim 1, characterized in that, The preset audio template sample includes a preset positive audio template sample and a preset negative audio template sample. Before the step of determining that a target wake-up word for waking up the device exists in the current audio stream in response to the matching result between the current audio stream and the preset audio template sample being within the target wake-up range, the method further includes: obtaining a first matching score between the current audio stream and the preset positive audio template sample and a second matching score between the current audio stream and the preset negative audio template sample; determining whether a target point with the first matching score as the abscissa and the second matching score as the ordinate is within the target wake-up range; if so, determining that the matching result between the current audio stream and the preset audio template sample is within the target wake-up range.

8. The method according to claim 1, characterized in that, The preset voice wake-up model includes an acoustic analysis module, a decoding module, and a confidence decision module. The step of inputting the current audio stream into the preset voice wake-up model to obtain the model output includes: inputting the current audio features of the current audio stream into the acoustic analysis module to obtain the probability distribution of each phoneme; inputting the probability distribution of each phoneme into the decoding module to obtain a word sequence; and inputting the word sequence into the confidence decision module to obtain the model output.

9. An electronic device, characterized in that, include: A memory and a processor, wherein the memory stores program instructions, and the processor retrieves the program instructions from the memory to perform the method as claimed in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, include: The system stores program data, which, when executed by a processor, is used to implement the method as described in any one of claims 1-8.