A voice wake-up method, apparatus, electronic device, and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-19
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]这种方法可以实现语音唤醒,但是可能因为环境的干扰而出现误触,或者在用户需要使用的时候却没有唤醒的问题
[0062]In this invention, after collecting target data for determining whether a user has performed a voice wake-up operation, a target wake-up word threshold is first determined based on the target data quality for the current wake-up scenario. Then, based on the target wake-up word threshold and the target wake-up word confidence level of the target data, it is determined whether voice wake-up is necessary. This method, by using a wake-up word threshold that varies with the environment, allows voice wake-up to be adjusted for different environments, thereby improving the accuracy of voice wake-up.
Smart Images

Figure CN116935848B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of voice wake-up, and specifically to a method, apparatus, electronic device, and medium for voice wake-up. Background Technology
[0002] As the most direct and convenient interaction method, voice has the advantage of being more convenient than physical touch interaction. Thanks to the continuous development of technology, voice wake-up is being applied to more and more fields and scenarios.
[0003] There are many implementation schemes for voice wake-up technology, which typically include: wake-up schemes based on recognition technology, which determine whether to wake up by judging whether the recognition result contains a wake word; specifically, the confidence level of the wake word can be calculated based on the posterior probability first, and then the confidence level of the wake word can be compared with a fixed threshold to determine whether to wake up.
[0004] This method can achieve voice wake-up, but it may cause accidental touches due to environmental interference, or it may fail to wake up when the user needs to use it. Summary of the Invention
[0005] In view of the above problems, a method, apparatus, electronic device, and medium for voice wake-up are proposed to overcome or at least partially solve the above problems, including:
[0006] A method for voice wake-up, the method comprising:
[0007] Acquire at least one type of target data collected from a user, the target data being used to determine whether the user performs a voice wake-up operation;
[0008] Determine the target data quality of the at least one target data, and determine the target wake word threshold corresponding to the target data quality;
[0009] Determine the target wake word confidence level of the at least one target data;
[0010] When the confidence level of the target wake-up word is greater than the threshold of the target wake-up word, voice wake-up is performed.
[0011] Optionally, the at least one target data includes target audio data and target image data, and determining the target data quality of the at least one target data includes:
[0012] Determine the target pronunciation quality score of the target audio data and the target image quality score of the target image data;
[0013] The target data quality is generated based on the target pronunciation quality score and the target image quality score.
[0014] Optionally, determining the target wake word threshold corresponding to the target data quality includes:
[0015] Obtain the weight information for image quality and pronunciation quality settings, as well as the preset initial wake word threshold;
[0016] The target wake word threshold is calculated based on the weight information, the target data quality, and the initial wake word threshold.
[0017] Optionally, determining the target pronunciation quality score of the target audio data and the target image quality score of the target image data includes:
[0018] The target audio data is input into the first prediction model to obtain the target pronunciation quality score;
[0019] Furthermore, the target image data is input into the second prediction model to obtain the target image quality score.
[0020] Optionally, inputting the target audio data into the first prediction model to obtain the target pronunciation quality score includes:
[0021] The target audio data is input into the first prediction model to obtain the posterior probability of each frame in the speech segment for the wake word in the target audio data.
[0022] The target pronunciation quality score is calculated based on the posterior probability of each frame.
[0023] Optionally, the method further includes the step of training a first prediction model:
[0024] Acquire first training data, which includes multiple audio data and pronunciation quality scores corresponding to each audio data;
[0025] The first training data is used to train the first model to be trained, and the first prediction model is obtained.
[0026] Optionally, the method further includes the step of training a second prediction model:
[0027] Acquire second training data, which includes multiple image data and image quality scores corresponding to each image data;
[0028] The second training data is used to train the second model to be trained, and the second prediction model is obtained.
[0029] Optionally, determining the target wake word confidence of the at least one target data includes:
[0030] The target audio data and the target image data are input into the third prediction model to obtain the confidence level of the target wake word.
[0031] Optionally, inputting the target audio data and the target image data into the third prediction model includes:
[0032] First feature data is obtained from the target audio data, and second feature data is obtained from the target image data;
[0033] Based on the first feature data and the second feature data, a correlation matrix between speech and image is generated;
[0034] Based on the correlation matrix, feature selection is performed on the second feature data to obtain the third feature data;
[0035] The first feature data and the third feature data are concatenated to form the fourth feature data, and the fourth feature data is input into the third prediction model.
[0036] Optionally, the method further includes the step of training a third prediction model:
[0037] The image teacher model is trained using image data from image data and audio / video data, and the first max pooling output of the trained image teacher model is determined.
[0038] The audio teacher model is trained using audio data and audio data from the audio and video data, and the second max pooling output result of the trained audio teacher model is determined.
[0039] The audio and video data are used to train the audio and video student model, and the third max pooling output result of the trained audio and video student model is determined.
[0040] Based on the third max pooling output, the relative entropy between the first max pooling output and the third max pooling output, and the relative entropy between the second max pooling output and the third max pooling output, determine the total loss of the trained audio-visual student model.
[0041] The parameters of the trained audio-visual student model are adjusted based on the total loss to obtain the third prediction model.
[0042] The present invention also provides a voice wake-up device, the device comprising:
[0043] The acquisition module is used to acquire at least one type of target data collected from the user, the target data being used to determine whether the user has performed a voice wake-up operation;
[0044] A threshold determination module is used to determine the target data quality of the at least one target data, and to determine the target wake word threshold corresponding to the target data quality;
[0045] Confidence determination module, used for the target wake word confidence of the at least one target data;
[0046] The wake-up module is used to perform voice wake-up when the confidence level of the target wake-up word is greater than the target wake-up word threshold.
[0047] Optionally, the at least one target data includes target audio data and target image data, and the threshold determination module is used to determine the target pronunciation quality score of the target audio data and the target image quality score of the target image data; and to generate the target data quality based on the target pronunciation quality score and the target image quality score.
[0048] Optionally, the threshold determination module is used to obtain weight information set for image quality and pronunciation quality, as well as a preset initial wake-up word threshold; and to calculate the target wake-up word threshold based on the weight information, the target data quality, and the initial wake-up word threshold.
[0049] Optionally, the threshold determination module is used to input the target audio data into a first prediction model to obtain the target pronunciation quality score; and to input the target image data into a second prediction model to obtain the target image quality score.
[0050] Optionally, the threshold determination module is used to input the target audio data into the first prediction model to obtain the posterior probability of each frame in the speech segment for the wake word in the target audio data; and to calculate the target pronunciation quality score based on the posterior probability of each frame.
[0051] Optionally, the device further includes:
[0052] The first training module is used to acquire first training data, which includes multiple audio data and pronunciation quality scores corresponding to each audio data; and to train the first model to be trained using the first training data to obtain the first prediction model.
[0053] Optionally, the device further includes:
[0054] The second training module is used to acquire second training data, which includes multiple image data and image quality scores corresponding to each image data; the second training data is used to train the second model to be trained to obtain the second prediction model.
[0055] Optionally, the confidence determination module is used to input the target audio data and the target image data into a third prediction model to obtain the confidence of the target wake word.
[0056] Optionally, the confidence determination module is configured to obtain first feature data from the target audio data and second feature data from the target image data; generate a speech-image correlation matrix based on the first feature data and the second feature data; perform feature selection on the second feature data based on the correlation matrix to obtain third feature data; concatenate the first feature data and the third feature data to form fourth feature data, and input the fourth feature data into the third prediction model.
[0057] Optionally, the device further includes:
[0058] The third training module is used to train an image teacher model using image data from the image data and audio / video data, and determine the first max-pooling output of the trained image teacher model; to train an audio teacher model using audio data from the audio / video data, and determine the second max-pooling output of the trained audio teacher model; to train an audio / video student model using the audio / video data, and determine the third max-pooling output of the trained audio / video student model; to determine the total loss of the trained audio / video student model based on the third max-pooling output, the relative entropy between the first and third max-pooling outputs, and the relative entropy between the second and third max-pooling outputs; and to adjust the parameters of the trained audio / video student model based on the total loss to obtain the third prediction model.
[0059] The present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the above-described voice wake-up method.
[0060] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described voice wake-up method.
[0061] The beneficial effects of this invention are:
[0062] In this invention, after collecting target data for determining whether a user has performed a voice wake-up operation, a target wake-up word threshold is first determined based on the target data quality for the current wake-up scenario. Then, based on the target wake-up word threshold and the target wake-up word confidence level of the target data, it is determined whether voice wake-up is necessary. This method, by using a wake-up word threshold that varies with the environment, allows voice wake-up to be adjusted for different environments, thereby improving the accuracy of voice wake-up. Attached Figure Description
[0063] Figure 1 This is a flowchart illustrating the steps of a voice wake-up method according to an embodiment of the present invention;
[0064] Figure 2 This is a flowchart of another voice wake-up method according to an embodiment of the present invention;
[0065] Figure 3 This is a flowchart of a voice wake-up process according to an embodiment of the present invention;
[0066] Figure 4 This is a flowchart of determining an image quality score according to an embodiment of the present invention;
[0067] Figure 5 This is a flowchart illustrating how to determine pronunciation quality scores according to an embodiment of the present invention;
[0068] Figure 6 This is a flowchart illustrating how to determine a target wake word threshold according to an embodiment of the present invention;
[0069] Figure 7 This is a flowchart of a training model according to an embodiment of the present invention;
[0070] Figure 8 This is a schematic diagram of an image teacher model according to an embodiment of the present invention;
[0071] Figure 9 This is a schematic diagram of an audio teacher model according to an embodiment of the present invention;
[0072] Figure 10 This is a flowchart of a feature fusion method according to an embodiment of the present invention;
[0073] Figure 11 This is a schematic diagram of the structure of a voice wake-up device according to an embodiment of the present invention. Detailed Implementation
[0074] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0075] In practical applications, it is possible to determine whether a user has performed a voice wake-up operation based on audio data, image data, etc. For example, audio data can be recognized and it can be determined whether the audio data contains a wake-up word, thereby determining whether the user has performed a voice wake-up operation; or, image data can be recognized and the shape of the user's lips can be identified from the image data to determine whether the user is saying a wake-up word, thereby determining whether the user has performed a voice wake-up operation; or, both image data and audio data can be recognized, and the recognition results can be used to determine whether the user has performed a voice wake-up operation.
[0076] In all the above methods, it is necessary to determine the relationship between the wake word confidence level corresponding to the data and a preset fixed wake word threshold. The wake word confidence level can refer to the probability that the currently collected data indicates that the user has performed a voice wake-up operation. The wake word threshold can be used to compare with the wake word confidence level to determine whether the user has performed a voice wake-up operation.
[0077] If the confidence level of the wake word is greater than the fixed wake word threshold, it indicates that the user has performed a voice wake-up operation; otherwise, it indicates that the user has not performed a voice wake-up operation.
[0078] In practical applications, the data used to determine whether a user has performed a voice wake-up operation may be affected by the environment, which may lead to misjudgments when determining whether the user has performed a voice wake-up operation. For example, when there is ambient noise in the environment, the wake word in the audio data may not be accurately identified, resulting in a low confidence level for the wake word in the audio data.
[0079] If a fixed, preset wake-up word threshold is still used to determine whether a user has performed a voice wake-up operation, misjudgments may occur. To address this issue, this invention provides a voice wake-up method. After collecting target data for determining whether a user has performed a voice wake-up operation, the method first determines a target wake-up word threshold for the current wake-up scenario based on the target data quality. Then, it determines whether voice wake-up is necessary based on the target wake-up word threshold and the target wake-up word confidence level of the target data. This method, by using a wake-up word threshold that varies with the environment, allows voice wake-up to be adjusted for different environments, thereby improving the accuracy of voice wake-up.
[0080] Reference Figure 1 The diagram illustrates a flowchart of a voice wake-up method according to an embodiment of the present invention, which may include the following steps:
[0081] Step 101: Obtain at least one type of target data collected from the user. The target data is used to determine whether the user performs a voice wake-up operation.
[0082] The target data may refer to the data collected by the user to determine whether the user has performed a voice wake-up operation, such as audio data, image data, etc. This embodiment of the invention does not limit this.
[0083] In practical applications, at least one type of target data can be collected from the user in real time; when it is necessary to determine whether the user has performed a voice wake-up operation on the device, the at least one type of target data can be obtained; wherein, the device can refer to a device with voice interaction function, such as: mobile phone, watch, computer, tablet, vehicle, etc., and the embodiments of the present invention do not limit this.
[0084] Step 102: Determine the target data quality of at least one type of target data, and determine the target wake word threshold corresponding to the target data quality.
[0085] After acquiring at least one type of target data collected from the user, the target data quality of this at least one type of target data can be determined first. The target data quality can be used to measure the quality of the target data. For example, if the target data is audio data, the target data quality can be used to measure the GOP (Goodness of Pronunciation) of the audio data. If the target data is image data, the target data quality can be used to measure the clarity, resolution, image completeness, etc. of the image data. This embodiment of the invention does not limit this.
[0086] As an example, the data quality of each target data can be determined first, and then the target data quality can be determined based on the data quality of each target data. The target data quality can be obtained by summing the multiple data qualities, by averaging, or by calculating through preset weights. This embodiment of the invention does not limit this.
[0087] After determining the target data quality, a target wake-word threshold corresponding to that target data quality can be determined. Specifically, one method for determining the target wake-word threshold is to pre-set different wake-word thresholds for different data qualities; the correspondence between data quality and wake-word thresholds can be pre-measured in a laboratory. Then, the target wake-word threshold corresponding to the target data quality can be determined based on the pre-determined correspondence.
[0088] Another method for determining the target wake word threshold is to pre-establish a conversion relationship between data quality and wake word threshold; this conversion relationship can be obtained in advance through laboratory measurements; then, the target data quality can be substituted into the conversion relationship to calculate the target wake word threshold. This embodiment of the invention does not limit this method.
[0089] In practical applications, if there is only one type of target data, the target data quality of that type of target data can be determined, and the target wake word threshold can be determined based on this target data quality. For example, the target data quality can be directly used as the target wake word threshold; or, a preset conversion logic can be performed based on the target data quality to obtain the target wake word threshold. The preset conversion logic can be set according to the actual situation (for example, the target wake word threshold can be obtained based on the product of the target data quality and the initial wake word threshold). This embodiment of the invention does not limit this.
[0090] Furthermore, if there are two or more types of target data, the data quality of each target data can be determined, and the target data quality can be determined based on the data quality of each target data. For example, the sum of the data quality of all target data can be calculated, and the average value can be used as the target data quality. Alternatively, the sum of the data quality of all target data can be used as the target data quality. The target data quality can also be calculated based on preset weights and the data quality of each target data. This embodiment of the invention does not impose any limitations on these methods.
[0091] Step 103: Determine the confidence level of the target wake word for at least one target data.
[0092] After obtaining at least one target data, the confidence level of the target wake word corresponding to the at least one target data can be determined first. The confidence level of the target wake word can be used to determine whether the at least one target data indicates that the user has performed a voice wake-up operation based on its relationship with the target wake word threshold.
[0093] After obtaining the target wake-up word confidence level and the target wake-up word threshold, it can be determined whether the device's voice interaction function needs to be activated based on the relationship between the target wake-up word confidence level and the target wake-up word threshold. Specifically, if the target wake-up word confidence level is not greater than the target wake-up word threshold, it indicates that the probability of the target data indicating that the user has performed a voice wake-up operation is low. In this case, it can be assumed that the user has not actually performed a voice wake-up operation; therefore, voice wake-up can be omitted.
[0094] Conversely, step 104 can be executed to activate the device's voice interaction function.
[0095] As an example, when there is only one type of target data, the confidence level of the target wake word for that type of target data can be directly determined; however, if there are two or more types of target data, the confidence level of the target wake word can be determined for each type of target data. The method for determining the confidence level of the target wake word for multiple types of target data can be found in the following examples, and will not be elaborated here.
[0096] Step 104: When the confidence level of the target wake word is greater than the target wake word threshold, perform voice wake-up.
[0097] Specifically, if the confidence level of the target wake-up word is greater than the target wake-up word threshold, the device's voice interaction function can be activated by voice, so that users can interact with the device by voice after the device's voice interaction function is activated.
[0098] In this embodiment of the invention, after collecting target data for determining whether a user has performed a voice wake-up operation, a target wake-up word threshold is first determined based on the target data quality of the target data for the current wake-up scenario. Then, based on the target wake-up word threshold and the target wake-up word confidence level of the target data, it is determined whether voice wake-up is required. This method, by using a wake-up word threshold that varies with the environment, allows voice wake-up to be adjusted for different environments, thereby improving the accuracy of voice wake-up.
[0099] Reference Figure 2 The diagram illustrates a flowchart of another voice wake-up method according to an embodiment of the present invention, which may include the following steps:
[0100] Step 201: Obtain at least one type of target data collected from the user, including target audio data and target image data.
[0101] Specifically, the target audio data can refer to the audio data collected by the user to determine whether the user has performed a voice wake-up operation; the target image data can refer to the image data collected by the user to determine whether the user has performed a voice wake-up operation. In particular, if the target audio data packet contains a wake-up word, then the target audio data can indicate that the user has performed a voice wake-up operation; if the target image data includes the lip movements of the user reciting the wake-up word, then the target image data can indicate that the user has performed a voice wake-up operation.
[0102] In practical applications, target audio data and target image data can be collected from the user in real time. When it is necessary to determine whether the user has performed a voice wake-up operation on the device, target audio data and target image data can be obtained, and voice wake-up can be performed using steps 202-207.
[0103] Step 202: Determine the target pronunciation quality score of the target audio data and the target image quality score of the target image data.
[0104] In practical applications, in order to determine the target wake-up word threshold for the current wake-up scenario, the current wake-up scenario can be determined first; specifically, the target pronunciation quality score of the target audio data and the target image quality score of the target image data can be determined separately.
[0105] The target pronunciation quality score can be used to represent the pronunciation quality of the target audio data. When the pronunciation quality of the target audio data is higher (e.g., the more informative the pronunciation), the target pronunciation quality score is higher. The target image quality score can be used to represent the image quality of the target image data. When the image quality of the target image data is higher (e.g., the higher the clarity), the target image quality score is higher. This embodiment of the invention does not impose any limitations on this.
[0106] In one embodiment of the present invention, the target pronunciation quality score and the target image quality score can be determined by the following sub-steps:
[0107] Sub-step 11: Input the target audio data into the first prediction model to obtain the target pronunciation quality score.
[0108] In practical applications, a first prediction model can be trained to predict the pronunciation quality of audio data; the specific training process can be found in the following instructions, which will not be elaborated here.
[0109] When it is necessary to determine the pronunciation quality of target audio data, the target audio data can be input into the first prediction model; the first prediction model can output the pronunciation quality of the target audio data based on the target audio data; and then obtain the target pronunciation quality score of the target audio data based on the pronunciation quality of the target audio data.
[0110] As an example, the target pronunciation quality score can be determined as follows:
[0111] The target audio data is input into the first prediction model to obtain the posterior probability of each frame in the speech segment corresponding to the wake word in the target audio data; the target pronunciation quality score is calculated based on the posterior probability of each frame.
[0112] In practical applications, the first prediction model can be used to determine the posterior probability of each frame of audio data in the audio data; the posterior probability can be used to represent the probability of a phoneme corresponding to that frame of audio data.
[0113] After obtaining the target audio data, the target audio data can be input into the first prediction model; the first prediction model can output the posterior probability of the audio in each frame of the target audio data based on the target audio data; then, based on the posterior probability of each frame, the pronunciation quality score of the audio in each frame can be determined; furthermore, based on the pronunciation quality score of the audio in each frame, the target pronunciation quality score of the entire target audio data can be calculated.
[0114] Sub-step 12: Input the target image data into the second prediction model to obtain the target image quality score.
[0115] In practical applications, a second prediction model can be trained first to predict the image quality of image data; the specific training process can be found in the following instructions, which will not be elaborated here.
[0116] When it is necessary to determine the image quality of target image data, the target image data can be input into the second prediction model; the second prediction model can output the image quality of the target image data based on the target image data; and then obtain the target image quality score of the target image data based on the image quality of the target image data.
[0117] Step 203: Generate target data quality based on the target pronunciation quality score and the target image quality score.
[0118] After obtaining the target pronunciation quality score and the target image quality score, the target data quality can be generated based on the target pronunciation quality score and the target image quality score; specifically, the set of the target pronunciation quality score and the target image quality score can be used as the target data quality.
[0119] Step 204: Obtain the weight information for image quality and pronunciation quality settings, as well as the preset initial wake word threshold.
[0120] In practical applications, corresponding weight information can be pre-set for image quality and pronunciation quality; this weight information can be set based on the importance of image quality and pronunciation quality in determining whether the user performs voice wake-up operation.
[0121] When it is necessary to determine the target wake word threshold for the current wake-up scenario, the weight information set for image quality and pronunciation quality, as well as a preset initial wake word threshold, can be obtained first. The preset initial wake word threshold can be a fixed value, which can be set according to the actual situation. For example, when it is necessary to trigger wake word wake-up more easily, the initial wake word threshold can be set smaller; when it is necessary to trigger wake word wake-up more difficultly, the initial wake word threshold can be set larger. This embodiment of the invention does not limit this.
[0122] Step 205: Calculate the target wake word threshold based on the weight information, target data quality, and initial wake word threshold.
[0123] After obtaining the weight information, target data quality, and initial wake-up word threshold, the target wake-up word threshold can be calculated based on these factors. Specifically, the target wake-up word threshold can be calculated using the following formula:
[0124] Target wake word threshold = (λ * target image quality score + (1-λ) * target pronunciation quality score) * initial wake word threshold;
[0125] Where λ represents the weight information.
[0126] Step 206: Input the target audio data and target image data into the third prediction model to obtain the target wake word confidence.
[0127] On the other hand, in embodiments of the present invention, the target audio data and target image data can also be input into a pre-trained third prediction model for determining the wake word confidence of audio and video data. The specific training method of the third prediction model can be referred to in the following description, and will not be repeated here.
[0128] In one embodiment of the present invention, the target audio data and target image data can be input into a third prediction model for prediction through the following sub-steps:
[0129] Sub-step 21: Obtain first feature data from target audio data and second feature data from target image data.
[0130] In practical applications, to ensure a linear relationship between the target audio data and target image data input to the third prediction model over time, first feature data can be extracted from the target audio data, and second feature data can be extracted from the target image data. The first feature data can be audio feature data extracted from the target audio data, and the second feature data can be image feature data extracted from the user's lip region within the target image data.
[0131] Sub-step 22: Generate a correlation matrix between speech and image based on the first feature data and the second feature data.
[0132] After obtaining the first feature data and the second feature data, the first feature data can be used as the key matrix and value matrix, and the second feature data can be used as the query matrix. Then, the query matrix is used to perform a click operation on the key matrix to generate the relevance matrix between speech and image.
[0133] Sub-step 23: Select features from the second feature data based on the correlation matrix to obtain the third feature data.
[0134] After obtaining the correlation matrix, the correlation matrix can be used to select features from the value matrix and generate lip features based on the first feature data, i.e., the third feature data.
[0135] Sub-step 24: Concatenate the first feature data and the third feature data to form the fourth feature data, and input the fourth feature data into the third prediction model.
[0136] After obtaining the third feature data, the first and third feature data can be concatenated to obtain the fourth feature data. At this point, the fourth feature data can be input into the third prediction model. The third prediction model can determine the target wake word confidence of the target data based on the fourth feature data.
[0137] Step 207: When the confidence level of the target wake word is greater than the threshold of the target wake word, perform voice wake-up.
[0138] After obtaining the target wake-up word confidence level and the target wake-up word threshold, it can be determined whether the device's voice interaction function needs to be activated based on the relationship between the target wake-up word confidence level and the target wake-up word threshold. Specifically, if the target wake-up word confidence level is not greater than the target wake-up word threshold, it indicates that the probability of the target data indicating that the user has performed a voice wake-up operation is low. In this case, it can be assumed that the user has not actually performed a voice wake-up operation; therefore, voice wake-up can be omitted.
[0139] Conversely, if the confidence level of the target wake-up word is greater than the target wake-up word threshold, it indicates that the target data suggests a high probability that the user has performed a voice wake-up operation. In this case, it can be assumed that the user has performed a voice wake-up operation. Therefore, the device's voice interaction function can be activated by voice so that the user can interact with the device by voice after the device's voice interaction function is activated.
[0140] In one embodiment of the present invention, the step of training a first prediction model may be further included:
[0141] Obtain first training data, which includes multiple audio data and the pronunciation quality score corresponding to each audio data; use the first training data to train the first model to be trained to obtain the first prediction model.
[0142] In practical applications, when training the first prediction model, the first training data can be obtained first. The first training data can include multiple audio data and the pronunciation quality score corresponding to each audio data. The pronunciation quality score corresponding to each audio data can be set in advance for each training audio data.
[0143] Then, a first model to be trained can be trained using multiple audio data and multiple pronunciation quality scores from the first training data, thereby obtaining a first prediction model.
[0144] In another embodiment of the present invention, the step of training a second prediction model may also be included:
[0145] Obtain second training data, which includes multiple image data and the image quality score corresponding to each image data; use the second training data to train the second model to be trained to obtain the second prediction model.
[0146] In practical applications, when training the second prediction model, second training data can be obtained first. This second training data may include multiple image data and the image quality score corresponding to each image data. The image quality score corresponding to each image data can be pre-set for each training image data.
[0147] Then, a second model to be trained can be trained using multiple image data and multiple image quality scores from the second training data, thereby obtaining a second prediction model.
[0148] In yet another embodiment of the present invention, the step of training a third prediction model may also be included:
[0149] The image teacher model is trained using image data from both image and audio / video data, and the first max-pooling output of the trained image teacher model is determined. The audio teacher model is trained using audio data from both audio and video data, and the second max-pooling output of the trained audio teacher model is determined. The audio / video student model is trained using audio / video data, and the third max-pooling output of the trained audio / video student model is determined. Based on the third max-pooling output, the relative entropy between the first and third max-pooling outputs, and the relative entropy between the second and third max-pooling outputs, the total loss of the trained audio / video student model is determined. The parameters of the trained audio / video student model are adjusted based on the total loss to obtain the third prediction model.
[0150] For the third prediction model, due to the limited amount of audio and video data, a training scheme can be proposed that jointly guides the audio and video student model using a visual teacher model and an audio teacher model. Specifically, the image teacher model can be trained using image data, and the audio teacher model can be trained using audio data. This results in an image teacher model that can predict the confidence level of the wake-up word corresponding to the image data, and an audio teacher model that can predict the confidence level of the wake-up word corresponding to the audio data.
[0151] As an example, in order to improve the accuracy of the audio-visual student model prediction, audio-visual data can be introduced when training the image teacher model and the audio teacher model. Specifically, after training the image teacher model with image data, the image data in the audio-visual data can be used to fine-tune the image teacher model to obtain the final image teacher model.
[0152] In addition, after training the audio teacher model using audio data, the audio teacher model can be fine-tuned using audio data from the audio and video data to obtain the final audio teacher model. This embodiment of the invention does not limit this.
[0153] In practical applications, audio and video data can also be used to train audio and video student models when training image teacher models and audio teacher models.
[0154] After training the image teacher model, audio teacher model, and audio-visual student model, the first max pooling output of the trained image teacher model, the second max pooling output of the trained audio teacher model, and the third max pooling output of the trained audio-visual student model can be determined respectively. The max pooling output can refer to the point with the largest value in the local receptive field.
[0155] After obtaining the first, second, and third max pooling outputs, the total loss of the trained audio-visual student model can be determined based on these three outputs. Specifically, the first, second, and third max pooling outputs can be weighted and the result can be used as the total loss of the trained audio-visual student model.
[0156] Then, the total loss can be used to adjust the parameters of the trained audio and video student model to obtain the third prediction model.
[0157] In this embodiment of the invention, at least one type of target data collected from a user is acquired. This target data is used to determine whether the user performs a voice wake-up operation. The at least one type of target data includes target audio data and target image data. A target pronunciation quality score is determined for the target audio data, and a target image quality score is determined for the target image data. Based on the target pronunciation quality score and the target image quality score, a target data quality is generated. Weight information set for image quality and pronunciation quality, as well as a preset initial wake-up word threshold, are acquired. A target wake-up word threshold is calculated based on the weight information, the target data quality, and the initial wake-up word threshold. The target audio data and target image data are input into a third prediction model to obtain the target wake-up word confidence score. When the target wake-up word confidence score is greater than the target wake-up word threshold, voice wake-up is performed. Through this embodiment of the invention, voice wake-up can be adjusted for different environments, thereby improving the accuracy of multimodal voice wake-up.
[0158] To further illustrate the above-mentioned voice wake-up method, the following specific example will be used for explanation:
[0159] like Figure 3 As shown, the acquired target image data can first undergo image quality assessment to obtain a target image quality score; and the acquired target audio data can undergo pronunciation quality assessment to obtain a target pronunciation quality score. Based on the target image quality score and the target pronunciation quality score, the target wake-up word threshold for the current wake-up scenario can be determined.
[0160] On the other hand, an audio-visual student model that jointly guides wake-word confidence prediction can be trained using an image teacher model and an audio teacher model; then, when it is necessary to determine the wake-word confidence of the target image data and the target audio data, the target image data and the target audio data can be input into the trained audio-visual student model.
[0161] The trained audio-visual student model can output the target wake-word confidence scores for the target image data and target audio data. Based on the target wake-word confidence scores and the target wake-word threshold, it can be determined whether the device's voice interaction function needs to be activated.
[0162] If the confidence level of the target wake word is greater than the target wake word threshold, the device's voice interaction function will be activated; otherwise, it will not be activated.
[0163] For image quality evaluation, such as Figure 4 As shown, a large amount of second training data can be constructed first. The image data of the second training data can include lip images of the wake word and any other content, and the image data of the second training data needs to contain image data of different qualities.
[0164] During the training phase, image features can be extracted from the image data and then globally normalized. Specifically, the mean and variance of the image features of all image data can be calculated first. Then, the global mean is subtracted from the image features of each image data and divided by the variance.
[0165] After global normalization of image features, the obtained features can be input into a convolutional neural network for feature abstraction. Then, the extracted features are input into a linear regression network; this regression network can be used to obtain NR-IQA (No Reference Image Quality Assessment), which characterizes image quality. During the testing phase, inputting image data yields its image quality score.
[0166] For audio quality evaluation, such as Figure 5 As shown, a large amount of initial training data can be constructed first. This initial training data can include wake word audio data and audio data that is not a wake word. Furthermore, the audio data in the initial training data needs to contain audio data of varying quality.
[0167] During the training phase, Fbank (Filter Bank) features can be extracted from the audio data in the first training data. Then, the Fbank features are used to train the TDNN (Time Delay Neural Network), and the cross-entropy loss function is used to optimize the TDNN network. After multiple rounds of training, the first prediction model is obtained when the loss converges to a certain value. The first prediction model can output the posterior probability of each frame of audio data in the audio data.
[0168] After obtaining the posterior probability, the pronunciation quality score of the audio data can be calculated; this pronunciation quality score can be the pronunciation quality score evaluated for the wake-up word. Specifically, the start time and end time of the wake-up word in a segment of audio data can be found based on the alignment information of the audio data first; in this audio segment containing several frames, each frame will correspond to multiple phonemes with a certain probability (phonemes can be simply understood as the pinyin sequence contained in the wake-up word, for example: the phoneme sequence of "nihao" for "hello"), and this probability is the posterior probability.
[0169] After that, the pronunciation quality of the current frame can be calculated through the GOP calculation formula. For example, the posterior probability that the current frame is the phoneme of the wake-up word at the corresponding position and the posterior probabilities that the current frame is each phoneme in the wake-up word can be determined; a maximum probability is determined from the posterior probabilities that the current frame is each phoneme in the wake-up word, and the ratio of the posterior probability that the current frame is the phoneme of the wake-up word at the corresponding position to the maximum probability is used as the pronunciation quality of the current frame. The pronunciation quality score can be the average value of the pronunciation quality scores of all phonemes within the pronunciation segment of the wake-up word.
[0170] Of course, the ratio of the posterior probability that the current frame is the phoneme of the wake-up word at the corresponding position to the sum of the posterior probabilities that the current frame is each phoneme in the wake-up word can also be used as the pronunciation quality of the current frame, and the embodiments of the present invention do not limit this.
[0171] As Figure 6 shown, after obtaining the target image quality score and the target pronunciation quality score, the target wake-up word threshold can be calculated according to the initial wake-up word threshold, the target image quality score, the target pronunciation quality score, and the weight information λ.
[0172] As Figure 7 shown, a process of training a model according to an embodiment of the present invention is shown:
[0173] For the image teacher model, second training data can be collected first. The second training data can include image data of lip movements containing the wake-up word and image data of lip movements not containing the wake-up word; the label of the image data of lip movements containing the wake-up word is labeled as 1, and the label of the image data of lip movements not containing the wake-up word is labeled as 0.
[0174] During the training process, the lip region in the image data can be located first, and the feature information of this region can be extracted. This feature information can be used to input into Image-Net for image classification processing; the specific structure of Image-Net can refer to Figure 8 .
[0175] Image-Net can be trained using Max-pooling as the loss function. Once the network training converges, a pre-trained image teacher model can be obtained. To enable the image teacher model to better provide reference information for the audio-visual network, image data from the audio-visual dataset can be used for model tuning. Specifically, lip features are extracted from the image data in the audio-visual dataset, and the pre-trained model is retrained using a smaller learning rate than before. The resulting image teacher model is then used to train the audio-visual student model.
[0176] For the inference phase of the image teacher model, an image data can be input; the image teacher model can output 0 or 1; where 0 indicates that the input image data belongs to the image data of lip movements that do not contain a wake word, and 1 indicates that the input image data belongs to the image data of lip movements that contain a wake word.
[0177] like Figure 8 As shown, the feature information of the lip region can be input into a CNN (Convolutional Neural Network) for feature abstraction. The CNN output is expanded into a one-dimensional vector through a fully connected layer, and then this one-dimensional vector is input into an LSTM (Long Short-Term Memory) network. LSTM has the ability to capture contextual features and can better represent the short-term changes of the image. Finally, the output of LSTM is fed into CTC (Connectionist Temporal Classification) for classification processing, ultimately obtaining the classification result of the lip image features.
[0178] The feature information of the lip region can be processed by Conv (Convolution) 1 for preliminary feature processing, then by Max Pooling & BN (Batch Normalization), then by Conv 2 for a second feature processing; then by BN, then by Conv 3 for a third feature processing; then by Max Pooling & BN, the output result is expanded into a one-dimensional vector by a fully connected layer.
[0179] like Figure 7 As shown, for the audio teacher model, the first training data can be collected first. The first training data can include audio data containing a wake word and audio data not containing a wake word; the audio data containing a wake word is labeled as 1, and the audio data not containing a wake word is labeled as 0.
[0180] During training, Fbank features can be extracted from the first training data, and then the Fbank features can be input into the first Kws (Keyword Spotting)-Net for classification processing. The specific structure of the first Kws-Net is as follows: Figure 9 As shown.
[0181] The first Kws-Net outputs audio data including and excluding the wake word probabilities. Max-pooling is used as the loss function for training, and the class with the highest score is used as the audio teacher model's score. After the network training converges, the obtained model is used as a pre-trained model, and then fine-tuned using audio data from the audio / video dataset.
[0182] Specifically, Fbank features can be extracted from the audio data of the audio-video dataset, and the pre-trained audio teacher model can be retrained using a smaller learning rate than the previous training to finally obtain the audio teacher model to train the audio-video student model.
[0183] like Figure 9 As shown, for the first Kws-Net, the extracted Fbank features are first processed by Conv for preliminary feature processing. Then, the feature processing result is input into three consecutive DS-Conv (Depthwise Separable Convolution) modules. The DS-Conv modules further abstract the feature processing results, and an Avg-pooling (general average filtering convolution operation) is connected after the last DS-Conv module to reduce the dimensionality of the features. Finally, a fully connected layer expands the vector after Avg-pooling into a one-dimensional vector, and the classification result is obtained through a decision function. An output of 0 indicates that the input audio data does not contain a wake word; an output of 1 indicates that the input audio data contains a wake word.
[0184] like Figure 7 As shown, for the audio-visual student model, a third training data set can be collected first. This third training data set can include audio and video data, namely, time-synchronized audio and image data containing a wake word, and time-synchronized audio and image data without a wake word. Synchronization here refers to the simultaneous recording of audio and image data, ensuring complete synchronization between pronunciation and mouth movements.
[0185] Lip features are extracted from the image data in the audio / video training dataset through lip localization, while F-bank features are extracted from the audio data. Lip features and F-bank features can then be fused using an attention mechanism.
[0186] The fused features are input into a second Kws-Net network for classification, yielding the classification results of the audio-video model. To better train the audio-video student network, both an image teacher network and an audio teacher network can be used to provide reference information. To ensure the audio-video student network learns the multimodal features after audio-video feature fusion more effectively, a teacher network architecture is employed.
[0187] The image teacher network provides image information references, and the audio teacher network provides speech information references. Specifically, the loss during the training of the audio-video student network comes from three parts: the Max-pooling output of the audio-video student model itself, the KL divergence between the Max-pooling outputs of the image teacher model and the audio-video student model, and the KL divergence between the Max-pooling outputs of the audio teacher model and the audio-video student model. The weighted sum of these three losses is used as the total loss for training the audio-video model.
[0188] like Figure 10 As shown, when performing feature fusion, the lip features in the image data can be further extracted by Conv, and the output of Conv can be used as the key matrix and value matrix respectively.
[0189] The Fbank features extracted from the audio data are used as the query matrix. Since the audio and video sampling frames are inconsistent, the strategy adopted here is to copy the image frames. Generally, an audio frame is 10ms long, while an image frame is 30ms long. If the same audio / video data is divided into frames, the number of audio frames will be three times the number of image frames. To meet the requirements of the attention mechanism, the number of frames needs to be standardized; specifically, this can be achieved by copying image frames.
[0190] Then, the query matrix can be used to perform a dot product operation on the key matrix to generate a score matrix (relevance matrix) for both speech and image. This score matrix is then used to perform feature selection on the value to generate lip feature extraction based on audio information.
[0191] Then, the original audio features (i.e., Fbank features) can be concatenated with the final lip features.
[0192] Specifically, because the original Fbank features and lip features have the same number of frames, but their dimensions (number of columns) may differ, the concatenation method is as follows: for any given frame, the original Fbank features are directly placed after the lip features. For example, if the original Fbank features have M rows and N1 columns, and the lip features have M rows and N2 columns, then the concatenated feature dimension will be M rows and (N1+N2) columns. To prevent large differences in feature value ranges, a normalization operation is required before concatenating the features. That is, the mean of the audio features is subtracted from the original Fbank features, and then divided by the variance of the audio features to obtain the normalized audio features; similarly, the mean of the image features is subtracted from the lip features, and then divided by the variance of the image features to obtain the normalized image features. Finally, the normalized audio features and image features are concatenated to obtain the audio-video fusion features.
[0193] In practical applications, the above method can be used in intelligent cockpit scenarios in vehicles. Feature fusion relying on visual and audio information can eliminate interference from speakers and human voices and noise inside and outside the vehicle, supporting both normal automotive noise and high-noise scenarios. Especially in high-noise scenarios (such as high-speed tire noise, multiple people talking inside the vehicle, and music playing), multi-mode voice wake-up can significantly suppress false wake-ups while maintaining a high wake-up rate, significantly improving the user experience. Furthermore, multi-mode wake-up is easily expandable; combined with sound source localization, it can adapt various "person-voice-scene" combinations based on voice information from different locations and different people within the camera's field of view.
[0194] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0195] Reference Figure 11 The diagram shows a structural schematic of a voice wake-up device according to an embodiment of the present invention, which may include the following modules:
[0196] The acquisition module 1101 is used to acquire at least one type of target data collected from the user, and the target data is used to determine whether the user performs a voice wake-up operation;
[0197] The threshold determination module 1102 is used to determine the target data quality of at least one type of target data and to determine the target wake word threshold corresponding to the target data quality.
[0198] Confidence determination module 1103 is used for the target wake word confidence of at least one target data;
[0199] The wake-up module 1104 is used to perform voice wake-up when the confidence level of the target wake-up word is greater than the target wake-up word threshold.
[0200] In one embodiment of the present invention, at least one target data includes target audio data and target image data. The threshold determination module 1102 is used to determine the target pronunciation quality score of the target audio data and the target image quality score of the target image data; and to generate target data quality based on the target pronunciation quality score and the target image quality score.
[0201] In one embodiment of the present invention, the threshold determination module 1102 is used to obtain weight information set for image quality and pronunciation quality, as well as a preset initial wake-up word threshold; and to calculate the target wake-up word threshold based on the weight information, target data quality, and initial wake-up word threshold.
[0202] In one embodiment of the present invention, the threshold determination module 1102 is used to input target audio data into a first prediction model to obtain a target pronunciation quality score; and to input target image data into a second prediction model to obtain a target image quality score.
[0203] In one embodiment of the present invention, the threshold determination module 1102 is used to input the target audio data into the first prediction model to obtain the posterior probability of each frame in the speech segment for the wake word in the target audio data; and to calculate the target pronunciation quality score based on the posterior probability of each frame.
[0204] In one embodiment of the present invention, the apparatus further includes:
[0205] The first training module is used to acquire first training data, which includes multiple audio data and the pronunciation quality score corresponding to each audio data; the first training data is used to train the first model to be trained to obtain the first prediction model.
[0206] In one embodiment of the present invention, the apparatus further includes:
[0207] The second training module is used to acquire the second training data, which includes multiple image data and the image quality scores corresponding to each image data. The second training data is used to train the second model to be trained to obtain the second prediction model.
[0208] In one embodiment of the present invention, the confidence determination module 1103 is used to input the target audio data and the target image data into the third prediction model to obtain the confidence of the target wake word.
[0209] In one embodiment of the present invention, the confidence determination module 1103 is used to obtain first feature data from target audio data and second feature data from target image data; generate a correlation matrix between speech and image based on the first feature data and the second feature data; perform feature selection on the second feature data based on the correlation matrix to obtain third feature data; concatenate the first feature data and the third feature data to form fourth feature data, and input the fourth feature data into a third prediction model.
[0210] In one embodiment of the present invention, the apparatus further includes:
[0211] The third training module is used to train the image teacher model using image data from the image data and audio / video data, and to determine the first max pooling output of the trained image teacher model; to train the audio teacher model using audio data from the audio / video data, and to determine the second max pooling output of the trained audio teacher model; to train the audio / video student model using audio / video data, and to determine the third max pooling output of the trained audio / video student model; to determine the total loss of the trained audio / video student model based on the third max pooling output, the relative entropy between the first and third max pooling outputs, and the relative entropy between the second and third max pooling outputs; and to adjust the parameters of the trained audio / video student model based on the total loss to obtain the third prediction model.
[0212] In this embodiment of the invention, after collecting target data for determining whether a user has performed a voice wake-up operation, a target wake-up word threshold is first determined based on the target data quality of the target data for the current wake-up scenario. Then, based on the target wake-up word threshold and the target wake-up word confidence level of the target data, it is determined whether voice wake-up is required. By using a wake-up word threshold that varies with the environment, voice wake-up can be adjusted for different environments, thereby improving the accuracy of voice wake-up.
[0213] This invention also provides an electronic device that may include a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the above-described voice wake-up method.
[0214] This invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above-described voice wake-up method.
[0215] As the apparatus embodiment is basically similar to the method embodiment, it is described in a relatively simple manner. For relevant details, please refer to the description of the method embodiment.
[0216] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0217] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0218] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0219] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0220] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0221] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0222] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0223] The above provides a detailed description of the method, apparatus, electronic device, and medium for voice wake-up. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method of voice wake-up, the method comprising: The method includes: Acquire target audio data and target image data collected from the user, the target audio data and the target image data being used to determine whether the user performs a voice wake-up operation; Determine the target data quality of the target audio data and the target image data, and determine the target wake word threshold corresponding to the target data quality; Determine the target wake word confidence level of the target audio data and the target image data; When the confidence level of the target wake-up word is greater than the threshold of the target wake-up word, voice wake-up is performed; Determining the target data quality of the target audio data and the target image data includes: Determine the target pronunciation quality score of the target audio data and the target image quality score of the target image data; The target data quality is generated based on the target pronunciation quality score and the target image quality score; Determining the target wake word threshold corresponding to the target data quality includes: Obtain the weight information for image quality and pronunciation quality settings, as well as the preset initial wake word threshold; The target wake word threshold is calculated based on the weight information, the target data quality, and the initial wake word threshold.
2. The method of claim 1, wherein, Determining the target pronunciation quality score of the target audio data and the target image quality score of the target image data includes: The target audio data is input into the first prediction model to obtain the target pronunciation quality score; Furthermore, the target image data is input into the second prediction model to obtain the target image quality score.
3. The method according to claim 2, characterized in that, The step of inputting the target audio data into the first prediction model to obtain the target pronunciation quality score includes: The target audio data is input into the first prediction model to obtain the posterior probability of each frame in the speech segment for the wake word in the target audio data. The target pronunciation quality score is calculated based on the posterior probability of each frame.
4. The method according to claim 2, characterized in that, The method also includes the step of training a first prediction model: Acquire first training data, which includes multiple audio data and pronunciation quality scores corresponding to each audio data; The first training data is used to train the first model to be trained, and the first prediction model is obtained.
5. The method according to claim 2, characterized in that, The method also includes the step of training a second prediction model: Acquire second training data, which includes multiple image data and image quality scores corresponding to each image data; The second training data is used to train the second model to be trained, and the second prediction model is obtained.
6. The method according to claim 1, characterized in that, Determining the target wake word confidence of the target audio data and the target image data includes: The target audio data and the target image data are input into the third prediction model to obtain the confidence level of the target wake word.
7. The method according to claim 6, characterized in that, The step of inputting the target audio data and the target image data into the third prediction model includes: First feature data is obtained from the target audio data, and second feature data is obtained from the target image data; Based on the first feature data and the second feature data, a correlation matrix between speech and image is generated; Based on the correlation matrix, feature selection is performed on the second feature data to obtain the third feature data; The first feature data and the third feature data are concatenated to form the fourth feature data, and the fourth feature data is input into the third prediction model.
8. The method according to claim 6, characterized in that, The method also includes the step of training a third prediction model: The image teacher model is trained using image data from image data and audio / video data, and the first max pooling output of the trained image teacher model is determined. The image teacher model predicts the confidence level of the wake word corresponding to the image data based on the image data; The audio teacher model is trained using audio data and audio data from the audio and video data, and the second max pooling output result of the trained audio teacher model is determined. The audio teacher model predicts the confidence level of the wake word corresponding to the audio data based on the audio data; The audio and video data are used to train the audio and video student model, and the third max pooling output result of the trained audio and video student model is determined. The audio-visual student model outputs the target wake word confidence scores of the target image data and the target audio data, and determines whether the voice interaction function of the device needs to be activated based on the target wake word confidence scores and the target wake word threshold. Based on the third max pooling output, the relative entropy between the first max pooling output and the third max pooling output, and the relative entropy between the second max pooling output and the third max pooling output, determine the total loss of the trained audio-visual student model. The parameters of the trained audio-visual student model are adjusted based on the total loss to obtain the third prediction model.
9. A voice-activated wake-up device, characterized in that, The device includes: The acquisition module is used to acquire target audio data and target image data collected from the user, wherein the target audio data and target image data are used to determine whether the user performs a voice wake-up operation; A threshold determination module is used to determine the target data quality of the target audio data and the target image data, and to determine the target wake word threshold corresponding to the target data quality; The confidence determination module is used to determine the confidence of the target wake word in the target audio data and the target image data. The wake-up module is used to perform voice wake-up when the confidence level of the target wake-up word is greater than the target wake-up word threshold. The threshold determination module is used to determine the target pronunciation quality score of the target audio data and the target image data. Image quality score; generate target data quality based on target pronunciation quality score and target image quality score; The threshold determination module is used to obtain weight information set for image quality and pronunciation quality, as well as a preset initial wake-up word threshold; and to calculate the target wake-up word threshold based on the weight information, target data quality, and initial wake-up word threshold.
10. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the voice wake-up method as described in any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the voice wake-up method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Voice awakening method and device
CN106782536A
Method and device for rousing voice recognition function
CN107871506A