Method, device, equipment and storage medium for voice keyword detection

By calculating the speech acoustic features and pan-semantic text space vectors and inputting keyword classification models for prediction, the problem of low detection accuracy of speech keywords in the prior art is solved, and a more efficient keyword detection effect is achieved.

CN115273815BActive Publication Date: 2025-05-06PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210906376.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-29
Publication Date
2025-05-06
Estimated Expiration
2042-07-29

AI Technical Summary

Technical Problem

The existing pronunciation keyword detection technology lacks semantic information modeling in pronunciation, resulting in large amounts of calculation and low keyword retrieval accuracy.

Method used

By calculating the speech acoustic features and the pre-calculated pan-semantic text space vectors, acoustic semantic context feature vectors are generated, and a keyword classification model is input for prediction, to realize keyword detection.

Benefits of technology

It integrates the characteristics of text and pronunciation, reduces the amount of extra calculation and improves the accuracy and effect of keyword detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115273815B_ABST
    Figure CN115273815B_ABST
Patent Text Reader

Abstract

The present invention provides a method, device, equipment and storage medium for voice keyword detection, which relates to the field of voice recognition technology; the method comprises: processing the voice data to be processed to obtain voice acoustic features; inputting the voice acoustic features into a preset voice coding network model to obtain a voice acoustic feature vector; taking out a pan-semantic text space vector under a preset storage path; performing attention calculation on the pan-semantic text space vector and the voice acoustic feature vector to obtain an acoustic semantic context feature vector; inputting the acoustic semantic context feature vector into a preset keyword classification model to obtain a predicted keyword. The method, device and storage medium for voice keyword detection of the embodiments of the present invention can improve the effect of keyword detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to, but are not limited to, the field of speech recognition technology, and in particular to a method, apparatus, device and storage medium for speech keyword detection. Background Art

[0002] Speech keyword detection mainly completes the process of retrieving predefined keywords in continuous speech streams. Traditional keyword retrieval methods include blank filling models, sample matching, and text retrieval based on large-scale speech recognition, but their defects are that they are mainly based on high-level feature sequence matching of acoustic features or text-level string matching based on large-scale speech recognition, and lack semantic information modeling in speech. In recent years, with the development of deep learning technology, scholars have proposed a variety of keyword retrieval systems that integrate acoustic features and keyword text features. However, whether it is based on training the fusion model of acoustic features and linguistic features, or based on similarity calculation and judgment of the two, they are all cyclic calculation and matching of keyword lists. The model has a large amount of calculation and due to the extraction of a single speech feature for keywords, it will also limit the diversity of command word expressions. The accuracy of keyword retrieval is low. Therefore, in related technologies, the effect of speech keyword detection is poor. Summary of the invention

[0003] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0004] The embodiments of the present invention provide a method, apparatus, device and storage medium for voice keyword detection, which can improve the effect of keyword detection.

[0005] In a first aspect, an embodiment of the present invention provides a method for detecting a speech keyword, comprising:

[0006] Processing the speech data to be processed to obtain speech acoustic features;

[0007] Inputting the speech acoustic feature into a preset speech coding network model to obtain a speech acoustic feature vector;

[0008] Retrieving the pan-semantic text space vector from a preset storage path;

[0009] Performing attention calculation on the pan-semantic text space vector and the speech acoustic feature vector to obtain an acoustic semantic context feature vector;

[0010] The acoustic semantic context feature vector is input into a preset keyword classification model to obtain predicted keywords.

[0011] According to some embodiments of the first aspect of the present invention, the pan-semantic text space vector is obtained by concatenating multiple generalized semantic feature vectors; and performing attention calculation on the pan-semantic text space vector and the speech acoustic feature vector to obtain the acoustic semantic context feature vector includes:

[0012] Inputting the pan-semantic text space vector and the speech acoustic feature vector into a preset attention model to perform attention calculation, and obtaining a plurality of weighted distribution data corresponding to the plurality of the generalized semantic feature vectors one by one;

[0013] The acoustic semantic context feature vector is obtained by combining a plurality of the weighted distribution data.

[0014] According to some embodiments of the first aspect of the present invention, the keyword classification model includes a feedforward neural network layer and a normalized network layer; the step of inputting the acoustic semantic context feature vector into a preset keyword classification model to obtain a predicted keyword includes:

[0015] Inputting the plurality of weighted distribution data included in the acoustic semantic context feature vector into the feedforward neural network layer to obtain probability update data;

[0016] Performing classification prediction on the probability update data through the normalized network layer to obtain a plurality of classification probabilities corresponding to a plurality of preset keywords;

[0017] A preset keyword corresponding to the largest classification probability is selected from the multiple classification probabilities as the keyword of the voice data.

[0018] According to some embodiments of the first aspect of the present invention, the processing the speech data to be processed to obtain speech acoustic features includes: extracting basic acoustic features from the speech data to obtain basic acoustic features of speech;

[0019] Correspondingly, the step of inputting the speech acoustic feature into a preset speech coding network model to obtain a speech acoustic feature vector includes:

[0020] The basic acoustic features of the speech are input into the speech coding network model to perform high-dimensional feature extraction to obtain the speech acoustic feature vector.

[0021] According to some embodiments of the first aspect of the present invention, the pan-semantic text space vector is calculated by the following steps:

[0022] Obtain a preset language representation model, a keyword sample sequence set, and a negative sample sequence set;

[0023] Extracting features from the keyword sample sequence set using the language representation model to obtain a plurality of keyword generalization feature vectors;

[0024] Extracting features from the negative sample sequence set using the language representation model to obtain at least one non-keyword feature vector;

[0025] The pan-semantic text space vector is obtained by concatenating a plurality of the keyword generalization feature vectors and at least one of the non-keyword feature vectors.

[0026] According to some embodiments of the first aspect of the present invention, extracting features from the keyword sample sequence set by using the language representation model to obtain a plurality of keyword generalization feature vectors includes:

[0027] Inputting each keyword sample sequence in the keyword sample sequence set into the language representation model respectively;

[0028] Extracting features from each generalized sample in the generalized sample set corresponding to the keywords of the keyword sample sequence by using the language representation model to obtain a generalized semantic feature set;

[0029] The generalized semantic feature set is averaged through the language representation model to obtain a keyword generalized feature vector corresponding to each keyword sample sequence.

[0030] According to some embodiments of the first aspect of the present invention, extracting features from the negative sample sequence set by using the language representation model to obtain at least one non-keyword feature vector includes:

[0031] Inputting each negative sample sequence in the negative sample sequence set into the language representation model respectively;

[0032] Randomly extracting non-keywords from the negative sample sequence using the language representation model to obtain a plurality of non-keyword data;

[0033] The language representation model is used to extract features from the plurality of non-keyword data and average the features to obtain the non-keyword feature vector.

[0034] In a second aspect, an embodiment of the present invention further provides a device for detecting speech keywords, including:

[0035] A preprocessing module, used for processing the speech data to be processed to obtain speech acoustic features;

[0036] An acoustic feature extraction module, used for inputting the speech acoustic feature into a preset speech coding network model to obtain a speech acoustic feature vector;

[0037] An acquisition module is used to retrieve the pan-semantic text space vector under a preset storage path;

[0038] An attention calculation module, used for performing attention calculation on the pan-semantic text space vector and the speech acoustic feature vector to obtain an acoustic semantic context feature vector;

[0039] The classification module is used to input the acoustic semantic context feature vector into a preset keyword classification model to obtain predicted keywords.

[0040] In a third aspect, an embodiment of the present invention further provides an electronic device, comprising: at least one processor, and a memory communicatively connected to the at least one processor; wherein the memory stores instructions, and the instructions are executed by the at least one processor so that when the at least one processor executes the instructions, the method for voice keyword detection as described in any one of the first aspects is implemented.

[0041] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the method for voice keyword detection described in any one of the first aspects.

[0042] The above-mentioned embodiment of the present invention has at least the following beneficial effects: by performing attention calculation on the extracted speech acoustic features and the pan-semantic text space vector to obtain the correlation between the two and predicting the keywords through the keyword classification model, the entire keyword search process integrates the features of both text and speech and combines the generalized pan-semantic text space vector obtained in advance, thereby reducing the additional amount of calculation in the keyword prediction process and improving the accuracy of keyword detection. Therefore, compared with the prior art, the embodiment of the present invention can improve the effect of keyword search.

[0043] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention and do not constitute a limitation on the technical solution of the present invention.

[0045] Figure 1 is a flow chart of a method for voice keyword detection according to an embodiment of the present invention;

[0046] Figure 2It is a structural diagram of an embodiment of a method for applying speech keyword detection in an embodiment of the present invention;

[0047] Figure 3 Schematic diagram of the attention mechanism principle of the method for voice keyword detection in an embodiment of the present invention;

[0048] Figure 4 It is a schematic diagram of a keyword classification model processing flow in a method for voice keyword detection in an embodiment of the present invention;

[0049] Figure 5 It is a schematic diagram of a process for obtaining a pan-semantic text space vector in a method for detecting speech keywords in an embodiment of the present invention;

[0050] Figure 6 is a schematic diagram of a module of a device corresponding to the method for voice keyword detection according to an embodiment of the present invention;

[0051] Figure 7 It is a hardware schematic diagram of a device corresponding to the method for voice keyword detection in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0054] In addition, the described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the present disclosure.

[0055] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0056] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.

[0057] The following is an explanation of some terms used in the present invention.

[0058] FNN: The full name is Feedforward Neutral Network, also known as feedforward neural network. In the process of calculating the output value, the input value propagates from the input layer unit to the output layer layer by layer, passes through the hidden layer and finally reaches the output layer to obtain the output. The units in the first layer of the feedforward network are connected to all the units in the second layer, and the second layer is connected to the units in the previous layer. There is no connection between the units in the same layer.

[0059] Fbank is the basic acoustic feature in the speech field. Its full name is Filter Bank. It is the output feature of speech after Mel filtering.

[0060] BERT, the full name of which is Bidirectional Encoder Representations from Transformer, is a pre-trained language representation model based on Transformer's bidirectional encoder representation. It emphasizes that it no longer uses the traditional unidirectional language model or the shallow concatenation of two unidirectional language models for pre-training as in the past, but instead uses a new masked language model (MLM) to generate deep bidirectional language representations.

[0061] The Softmax function is a probability function. The Max function means that if a>b, then a is taken and b is impossible to be taken. Softmax calculates the probability of each element being taken. If the probability of a being taken is greater than b, then a is often taken, but it is also possible to be taken. It is equivalent to pulling out all the elements to make a score, then normalizing them, and then sorting them.

[0062] Speech keyword detection mainly completes the process of retrieving predefined keywords in a continuous speech stream. Traditional keyword retrieval methods include blank filling models, sample matching, and text retrieval based on large-scale speech recognition, etc. However, their defects are that they are mainly based on high-level feature sequence matching of acoustic features or text-level string matching based on large-scale speech recognition, and lack semantic information modeling in speech. In recent years, with the development of deep learning technology, scholars have proposed a variety of keyword retrieval systems that integrate acoustic features and keyword text features. However, whether it is based on training a fusion model of acoustic features and linguistic features, or based on similarity calculation and judgment of the two, it is a cyclic calculation and matching of the keyword list. The model has a large amount of calculation and due to the extraction of a single speech feature for keywords, it will also limit the diversity of command word expressions. The accuracy of keyword retrieval is low. Therefore, in the related art, the effect of speech keyword detection is poor. Based on this, the embodiment of the present invention provides a method, device, equipment and storage medium for speech keyword detection, which can improve the effect of keyword detection.

[0063] First, refer to Figure 1 As shown, the method for voice keyword detection according to an embodiment of the present invention includes:

[0064] Step S100: Process the speech data to be processed to obtain speech acoustic features.

[0065] Step S200: input the speech acoustic features into a preset speech coding network model to obtain a speech acoustic feature vector.

[0066] Step S300: Retrieve the pan-semantic text space vector from a preset storage path.

[0067] It should be noted that the pan-semantic text space vector is calculated in advance and is composed of multiple generalized generalized semantic feature vectors, wherein each of the multiple generalized semantic feature vectors corresponds to a preset keyword. The generalized semantic feature vector is used to represent the generalized text features of multiple expressions of the keyword. It is obtained by generalizing the corresponding keyword and extracting features from the generalized multiple sentences. Among the multiple generalized semantic feature vectors, at least one generalized semantic feature vector represents the generalized text features of non-keywords. Therefore, the pan-semantic text space vector can enrich the semantics of multiple keywords.

[0068] It should be noted that for each keyword in the keyword list, multiple text languages ​​with different expressions can be obtained by exhaustive enumeration, and the language representation model is used to extract features from the exhaustively obtained text languages ​​to obtain a generalized semantic feature vector corresponding to the keyword. In other embodiments, multiple expressions of preset keywords can be periodically counted to obtain generalized data corresponding to the keywords.

[0069] It should be noted that the pan-semantic text space vector can be obtained by extracting features from the text language corresponding to the list of matching keywords using an existing language representation model, such as the BERT model. The calculation of the pan-semantic text space vector is independent of the keyword detection process.

[0070] Step S400: Perform attention calculation on the pan-semantic text space vector and the speech acoustic feature vector to obtain an acoustic semantic context feature vector.

[0071] It should be noted that the acoustic semantic context feature vector is used to characterize the probability distribution state of the speech acoustic feature vector in the pan-semantic text space vector; through attention calculation, the probability distribution of the speech acoustic feature vector relative to each generalized semantic feature vector included in the pan-semantic text space vector can be obtained, and then the probability distribution is multiplied to the corresponding generalized semantic feature vector to obtain the weighted distribution state of the current speech acoustic feature vector in the pan-semantic text space vector.

[0072] Step S500: input the acoustic semantic context feature vector into a preset keyword classification model to obtain predicted keywords.

[0073] It should be noted that the keyword classification model is used to perform probability prediction and normalization processing on the acoustic semantic context feature vector to obtain the probability distribution of the acoustic semantic context feature vector relative to multiple preset keywords, and then the predicted keyword can be determined based on the probability distribution. In some embodiments, in this step, the keyword with the largest probability value is selected as the predicted keyword. In other embodiments, in this step, the keyword corresponding to the N probabilities with the largest probability is selected as the predicted keyword. Preferably, in an embodiment of the present invention, the one with the largest probability value is selected as the predicted keyword output. When there are multiple maximum probability values, all of them are output as predicted keywords.

[0074] It should be noted that the keyword classification model is used to classify multiple preset keywords, and its specific model structure is not limited in this step. For example, the keyword classification model is obtained by combining at least one forward network layer and a normalization layer.

[0075] Therefore, by performing attention calculation on the extracted speech acoustic features and the pan-semantic text space vector to obtain the correlation between the two and predicting the keywords through the keyword classification model, the entire keyword search process integrates the features of both text and speech and combines the generalized pan-semantic text space vector obtained in advance, thereby reducing the additional amount of calculation in the keyword prediction process and improving the accuracy of keyword detection. Therefore, compared with the prior art, the embodiment of the present invention can improve the effect of keyword search.

[0076] It should be noted that the pan-semantic text space vector enriches the diversity of each keyword text expression. Therefore, when the pan-semantic text space vector and speech acoustic features are used for attention calculation, the accuracy of keyword prediction can be improved.

[0077] For example, refer to Figure 2 As shown, the speech data is preprocessed, and Fbank features are extracted by frame division and windowing; the Fbank features are input into the acoustic encoder to obtain the speech acoustic feature vector; the pre-stored pan-semantic text space vector and the speech acoustic feature vector are calculated through the attention mechanism to obtain the acoustic semantic context feature vector; and the context feature vector is input into the keyword classification model, and finally the most likely keyword is predicted and output.

[0078] It can be understood that the pan-semantic text space vector is obtained by concatenating multiple generalized semantic feature vectors; the pan-semantic text space vector and the speech acoustic feature vector are subjected to attention calculation to obtain an acoustic semantic context feature vector, including: inputting the pan-semantic text space vector and the speech acoustic feature vector into a preset attention model to perform attention calculation to obtain multiple weighted distribution data corresponding one-to-one to multiple generalized semantic feature vectors; and combining multiple weighted distribution data to obtain an acoustic semantic context feature vector.

[0079] It should be noted that, taking the attention model as an example using a multi-layer transformer structure, refer to Figure 3 The following is a schematic diagram of the attention mechanism principle. In the attention model, the constituent elements in Source are imagined as a series of<Key,Value> The data pairs are composed of a certain element Query in Target. By calculating the similarity or correlation between Query and each Key, the weight coefficient of the Value corresponding to each Key is obtained, and then the Value is weighted and summed to obtain the final Attention value. Each acoustic feature in the speech acoustic feature vector is used as Query, and each generalized semantic feature vector in the pan-semantic text space vector is used as Key, and the attention weight is applied to each generalized semantic feature vector. Then the following formula corresponding to Query, Key and Value can be obtained:

[0080]

[0081]

[0082]

[0083] Among them, H enc is the high-dimensional representation of the speech acoustic feature vector, H pre is a high-dimensional representation of pan-semantic features, Qenc , K pre 、V pre They correspond to the Query, Key and Value vectors in the attention mechanism respectively. They correspond to the mapping matrices of Query, Key and Value vectors in the attention mechanism respectively.

[0084] At this time, the weighted distribution data where d k Represents the feature vector dimension.

[0085] At this time, the acoustic semantic context feature vector can be expressed as Z = Att(C), which is used to represent the probability distribution of each generalized semantic feature vector obtained after each acoustic feature vector performs attention calculation on all generalized semantic feature vectors, and then multiplies it to each generalized semantic feature vector to obtain the weighted distribution of the current speech acoustic feature vector in the semantic space.

[0086] Understandably, referring to Figure 4 As shown, the keyword classification model includes a feedforward neural network layer and a normalized network layer; in the above step S500, the acoustic semantic context feature vector is input into the preset keyword classification model to obtain the predicted keywords, including:

[0087] Step S510: Input multiple weighted distribution data included in the acoustic semantic context feature vector into the forward neural network layer to obtain probability update data.

[0088] It should be noted that the number of layers of the feedforward neural network layer can be set to one or more layers, and those skilled in the art can specifically set it as needed. Preferably, in the embodiment of the present invention, the feedforward neural network layer is set to have two layers.

[0089] Step S520: classify and predict the probability update data through the normalized network layer to obtain multiple classification probabilities corresponding to multiple preset keywords.

[0090] It should be noted that there are multiple classification targets in the output connection of the forward neural network layer, each preset keyword corresponds to a classification target, and the normalized network layer is used to calculate the probability that the statistical probability update data falls on each classification target. In some embodiments, the multiple classification targets include N+1 classification targets, where N represents the number of keywords of the preset keywords, and 1 represents the number of non-keyword targets. In other embodiments, multiple classification targets each correspond to a preset keyword. Preferably, N+1 classification targets are used in the embodiment of the present invention, and the classification probability calculation is performed by introducing non-keyword targets, which can further improve the accuracy of classification.

[0091] It should be noted that the normalized network layer is composed of the Softmax function, where the loss function of the Softmax function is the cross entropy.

[0092] Step S530: Select a preset keyword corresponding to the maximum classification probability from the multiple classification probabilities as the keyword of the voice data.

[0093] It should be noted that the accuracy is higher if the keyword with the largest classification probability value is used as the keyword of the voice data. For example, assuming that there are 6 preset keywords, the classification probabilities corresponding to the 6 keywords calculated by the Softmax function are 60%, 20%, 10%, 7%, 2% and 1% respectively, among which the classification probability value of 60% is the largest among the 6, so the keyword corresponding to 60% is used as the keyword of the voice data.

[0094] For example, refer to Figure 3 As shown, Z = Att (C) is input into the keyword classification model, referring to Figure 2 As shown, the keyword classification model includes two layers of forward neural network layers FNN, the output of FNN is connected to multiple classification targets, and the FNN output is predicted by the Softmax function to obtain the probability prediction Pos=softmax(FNN(Z)) corresponding to each weighted distribution data classification target, where FNN(Z) represents the output of FNN. At this time, the keyword corresponding to the voice data can be determined according to the value of Pos=softmax(FNN(Z)).

[0095] It is understandable that, before obtaining the speech acoustic feature vector, the method further includes: extracting basic acoustic features from the speech data to obtain basic acoustic features of the speech.

[0096] Correspondingly, in step S100, the speech data to be processed is processed to obtain speech acoustic features, including: extracting basic acoustic features of the speech data to obtain basic speech acoustic features; correspondingly, in step S200, the speech acoustic features are input into a preset speech coding network model to obtain a speech acoustic feature vector, including: inputting the basic speech acoustic features into the speech coding network model to perform high-dimensional feature extraction to obtain a speech acoustic feature vector.

[0097] It should be noted that the basic acoustic features are extracted as Fbank features. The speech data contains multiple speech frames, so the Fbank features can be extracted by framing and windowing the speech data.

[0098] For example, taking the neural network structure of the speech coding network model using N layers of conformer as an example, for the input speech basic acoustic feature (i.e., Fbank) X = {x t}, the speech coding network model converts X = {xt}Convert to high-dimensional acoustic features Where t is the voice frame index in the voice data, then H enc =f enc (X), where f enc Represents the encoder network (i.e., N-layer conformer neural network).

[0099] Understandably, referring to Figure 5 As shown in Figure 2, the pan-semantic text space vector is calculated through the following steps:

[0100] Step S610: Obtain a preset language representation model, a keyword sample sequence set, and a negative sample sequence set.

[0101] It should be noted that the language representation model is pre-trained, such as the BERT model. The keyword sample sequence set is a set of samples used for keyword extraction. The negative sample sequence set is a set of samples used for non-keyword target extraction. The keyword sample sequence set can be a set of one or more samples; the negative sample sequence set can also be a set of one or more samples.

[0102] Step S620: extract features from the keyword sample sequence set using a language representation model to obtain a plurality of keyword generalization feature vectors.

[0103] It should be noted that the keyword generalization feature vector is a feature vector obtained by extracting features from samples after generalization of the corresponding keyword. For example, the keyword extracted from one of the samples in the keyword sample sequence set is "turn on the light", and the keyword is generalized to obtain generalized texts such as "turn on the light", "turn on the light" and "turn on the light". Multiple keyword generalization feature vectors are vectors obtained by extracting features from the corresponding "turn on the light", "turn on the light" and "turn on the light".

[0104] Step S630: extract features from the negative sample sequence set using a language representation model to obtain at least one non-keyword feature vector.

[0105] It should be noted that multiple non-keyword feature vectors can be set to enrich the negative sample semantic space to better distinguish keyword sequences from non-keyword sequences.

[0106] Step S640: concatenate multiple keyword generalization feature vectors and at least one non-keyword feature vector to obtain a pan-semantic text space vector.

[0107] It should be noted that the keyword generalization feature vector extracted from the keyword sample sequence set and the non-keyword feature vector extracted from the negative sample sequence set are both generalization semantic feature vectors.

[0108] It should be noted that the spliced ​​pan-semantic text space vector will be stored in a preset storage path, such as a database. In practical applications, feature extraction can be performed regularly through steps S510 to S540 to update the pan-semantic text space vector, so that the generalized semantic features used for calculation can be updated in real time when the voice data is used for keyword prediction. And steps S510 to S540 are execution steps independent of the keyword prediction process and can be processed asynchronously with the keyword prediction process. Therefore, while reducing the amount of calculation in the keyword prediction process, the accuracy of keyword detection can be further improved.

[0109] For example, the pre-trained BERT model is used as the language representation model. During the entire training process, the BERT encoder parameters are fixed and no training optimization is performed. i ={y j}, where i is the keyword and negative sample sequence index, and j is the word index in each sample sequence. The semantic feature representation of each input text sequence after BERT model alignment is: It means taking the output of the first word cls as the corresponding high-dimensional semantic feature vector. The high-dimensional semantic feature vector is expressed as follows: Finally, all high-dimensional semantic feature vectors are concatenated to obtain the pan-semantic text space vector

[0110] It is understandable that step S620, extracting features from the keyword sample sequence set through the language representation model to obtain multiple keyword generalization feature vectors, includes: inputting each keyword sample sequence in the keyword sample sequence set into the language representation model respectively; extracting features from each generalization sample in the generalization sample set corresponding to the keyword of the keyword sample sequence through the language representation model to obtain a generalization semantic feature set; and averaging the generalization semantic feature set through the language representation model to obtain a keyword generalization feature vector corresponding to each keyword sample sequence.

[0111] It should be noted that after inputting the keyword sample sequence into the language representation model, the corresponding keywords will be extracted. At this time, generalization processing can be performed based on the keyword or the corresponding generalization sample set can be directly obtained, and then the feature vectors corresponding to multiple generalization samples for each keyword can be obtained. Among them, each generalization sample corresponds to a feature dimension; the average processing means adding the average of all output feature vectors dimension by dimension to obtain the keyword generalization feature vector. At this time, each keyword generalization feature vector represents the characteristics of multiple expressions of a keyword. Therefore, the semantics of keywords can be enriched, making the expression of keywords diverse, more in line with actual application scenarios, and improving the accuracy of keyword detection.

[0112] It is understandable that step S630, extracting features from the negative sample sequence set through a language representation model to obtain at least one non-keyword feature vector, includes: inputting each negative sample sequence in the negative sample sequence set into the language representation model respectively; performing non-keyword random sampling on the negative sample sequence through the language representation model to obtain multiple non-keyword data; performing feature extraction on the multiple non-keyword data through the language representation model and averaging to obtain a non-keyword feature vector.

[0113] It should be noted that the more non-keyword data there is, the richer the negative sample semantic space is and the higher the accuracy of keyword detection is.

[0114] It should be noted that random extraction can make semantic expression random. In this way, non-keyword feature vectors are more referenceable. Therefore, the method of obtaining non-keyword data through random extraction can further improve the diversity of pan-semantic text space vectors, so as to improve the detection accuracy in the keyword detection process.

[0115] It should be noted that the negative sample sequence may be one group or multiple groups, which is not limited in the embodiment of the present invention. Preferably, multiple groups of negative sample sequences are selected in the embodiment of the present invention to extract non-keyword feature vectors.

[0116] Below, refer to Figures 1 to 5 The process of keyword detection in the embodiment of the present invention is described with a specific embodiment, and the specific detection is as follows:

[0117] Reference Figures 1 to 4 As shown in the figure, the speech data to be processed is preprocessed, and the Fbank features are extracted by frame windowing to obtain the basic acoustic features of the speech; the basic acoustic features of the speech are input into the speech coding network model to obtain the speech acoustic feature vector; and the pan-semantic text space vector output by the preset pre-trained language feature model is obtained, and then the pan-semantic text space vector and the speech feature vector are input into the attention mechanism model to obtain multiple weighted distribution data, and the multiple weighted distribution data are used as the input parameters of the two-layer FNN for probability adjustment to obtain the probability update data; through the normalization layer (corresponding to Figure 2 The probability update data corresponds to the classification probability of N+1 preset target classifications, and the target classification corresponding to the maximum classification probability is output as a keyword. Figure 5 As shown, the language feature model extracts features from the input keyword text sequence to obtain a keyword generalization feature vector, and simultaneously selects a certain amount of negative sample text sequences as input parameters corresponding to the non-keyword feature vector. Finally, all the obtained keyword generalization feature vectors and non-keyword feature vectors are concatenated to obtain a pan-semantic space vector representation.

[0118] The method of the present invention can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0119] Second, refer to Figure 6 As shown, the device for voice keyword detection according to an embodiment of the present invention includes:

[0120] A preprocessing module 100 is used to process the speech data to be processed to obtain speech acoustic features;

[0121] The acoustic feature extraction module 200 is used to input the speech acoustic feature into a preset speech coding network model to obtain a speech acoustic feature vector;

[0122] The acquisition module 300 is used to retrieve the pan-semantic text space vector in a preset storage path;

[0123] The attention calculation module 400 is used to perform attention calculation on the pan-semantic text space vector and the speech acoustic feature vector to obtain an acoustic semantic context feature vector;

[0124] The classification module 500 is used to input the acoustic semantic context feature vector into a preset keyword classification model to obtain predicted keywords.

[0125] It should be noted that the preprocessing module 100, the acoustic feature extraction module 200, the acquisition module 300, the attention calculation module 400 and the classification module 500 are all modules that need to be called in the keyword detection process. In some embodiments, the device for voice keyword detection also includes a pan-semantic text space vector extraction module, which is used to extract the pan-semantic text space vector that represents the keyword features. The pan-semantic text space vector extraction module is independent of the keyword prediction process. After the pan-semantic text space vector extraction module extracts the pan-semantic text space vector, it will store the pan-semantic text space vector in a preset storage path, decoupling the acquisition of the pan-semantic text space vector and the keyword detection process, so that in the keyword detection process, there is no need to calculate the pan-semantic text space vector corresponding to the keyword, thereby reducing the amount of calculation in the keyword detection process. And the pan-semantic text space vector represents a collection of multiple semantic features in the keyword list to be matched, so the accuracy of keyword matching can be improved.

[0126] It should be noted that, in some embodiments, the classification module 500 includes a feed-forward neural network layer and a normalization layer. The acoustic semantic context feature vector is processed by the feed-forward neural network layer, and the output of the feed-forward neural network layer is used as an input parameter of the normalization layer to further obtain the classification probability of the acoustic semantic context feature vector relative to a plurality of preset keywords, and the keyword corresponding to the maximum classification probability is used as the predicted keyword of the speech data.

[0127] An embodiment of the present invention further provides an electronic device, including:

[0128] at least one processor, and

[0129] a memory communicatively connected to at least one processor; wherein,

[0130] The memory stores instructions, and the instructions are executed by at least one processor, so that when the at least one processor executes the instructions, the method for voice keyword detection as described in the above embodiment of the present invention is implemented.

[0131] It should be noted that the electronic device is applied to any one of the voice keyword detection methods of the first aspect, and therefore has all the beneficial effects of the voice keyword detection method of the first aspect.

[0132] Combine the following Figure 7 The hardware structure of the computer device is described in detail. The electronic device includes: a processor 710 , a memory 720 , an input / output interface 730 , a communication interface 740 and a bus 750 .

[0133] The processor 710 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present disclosure;

[0134] The memory 720 can be implemented in the form of ROM (Read Only Memory), static storage device, dynamic storage device or RAM (Random Access Memory). The memory 720 can store operating systems and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program codes are stored in the memory 720, and the processor 710 calls and executes the training method of the model in the embodiment of the present disclosure;

[0135] Input / output interface 730, used to implement information input and output;

[0136] The communication interface 740 is used to realize the communication interaction between the device and other devices, and the communication can be realized by wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and the bus 750 is used to transmit information between the various components of the device (such as the processor 710, the memory 720, the input / output interface 730 and the communication interface 740);

[0137] The processor 710 , the memory 720 , the input / output interface 730 , and the communication interface 740 are connected to each other in communication within the device via the bus 750 .

[0138] It should be noted that, in some embodiments, the processor 710 executes steps S100 to S400 of the method for voice keyword detection of the first aspect; in other embodiments, the processor 710 executes steps S100 to S500 and steps S510 to S530 of the method for voice keyword detection; in other embodiments, the processor 710 executes steps S100 to S500, steps S510 to S530, and steps S610 to S640 of the method for voice keyword detection. In other embodiments, the processor 710 executes all steps of the method for voice keyword detection.

[0139] An embodiment of the present invention further provides a storage medium, which is a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the above-mentioned voice keyword detection method.

[0140] It should be noted that the storage medium can execute any one of the methods for voice keyword detection in the first aspect. Therefore, when any device of the computer class execution instruction of the storage medium runs, the corresponding device has all the beneficial effects of the method for voice keyword detection in the first aspect.

[0141] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0142] It will be appreciated by those skilled in the art that all or some of the steps and systems in the disclosed method above may be implemented as software, firmware, hardware and appropriate combinations thereof. Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor or a microprocessor, or may be implemented as hardware, or may be implemented as an integrated circuit, such as an application specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or a non-transitory medium) and a communication medium (or a temporary medium). As known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that may be used to store desired information and may be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0143] The embodiments described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art can appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0144] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.

[0145] The terms "including" and "having" and any variations thereof in the description of the present invention are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.

[0146] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0147] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above-mentioned implementation mode. Technical personnel familiar with the field can also make various equivalent deformations or substitutions without violating the spirit of the present invention. These equivalent deformations or substitutions are all included in the scope defined by the claims of the present invention.

Claims

1. A method for detecting speech keywords, characterized in that: The method comprises: Processing the speech data to be processed to obtain speech acoustic features; Inputting the speech acoustic feature into a preset speech coding network model to obtain a speech acoustic feature vector; Retrieving the pan-semantic text space vector from a preset storage path; Performing attention calculation on the pan-semantic text space vector and the speech acoustic feature vector to obtain an acoustic semantic context feature vector; Inputting the acoustic semantic context feature vector into a preset keyword classification model to obtain predicted keywords; The pan-semantic text space vector is obtained by concatenating a plurality of generalized semantic feature vectors; and the pan-semantic text space vector and the speech acoustic feature vector are subjected to attention calculation to obtain an acoustic semantic context feature vector, including: Inputting the pan-semantic text space vector and the speech acoustic feature vector into a preset attention model to perform attention calculation, and obtaining a plurality of weighted distribution data corresponding to the plurality of the generalized semantic feature vectors one by one; Combining a plurality of the weighted distribution data to obtain the acoustic semantic context feature vector; The pan-semantic text space vector is calculated by the following steps: Obtain a preset language representation model, a keyword sample sequence set, and a negative sample sequence set; Extracting features from the keyword sample sequence set using the language representation model to obtain a plurality of keyword generalization feature vectors; Extracting features from the negative sample sequence set using the language representation model to obtain at least one non-keyword feature vector; The pan-semantic text space vector is obtained by concatenating a plurality of the keyword generalization feature vectors and at least one of the non-keyword feature vectors.

2. The method for voice keyword detection according to claim 1, characterized in that: The keyword classification model includes a feedforward neural network layer and a normalized network layer; the acoustic semantic context feature vector is input into a preset keyword classification model to obtain predicted keywords, including: Inputting the plurality of weighted distribution data included in the acoustic semantic context feature vector into the feedforward neural network layer to obtain probability update data; Performing classification prediction on the probability update data through the normalized network layer to obtain a plurality of classification probabilities corresponding to a plurality of preset keywords; A preset keyword corresponding to the largest classification probability is selected from the multiple classification probabilities as the keyword of the voice data.

3. The method for voice keyword detection according to claim 1, characterized in that: The processing of the speech data to be processed to obtain speech acoustic features includes: Extracting basic acoustic features from the speech data to obtain basic acoustic features of the speech; Correspondingly, the step of inputting the speech acoustic feature into a preset speech coding network model to obtain a speech acoustic feature vector includes: The basic acoustic features of the speech are input into the speech coding network model to perform high-dimensional feature extraction to obtain the speech acoustic feature vector.

4. The method for voice keyword detection according to claim 1, characterized in that: The extracting features of the keyword sample sequence set by the language representation model to obtain a plurality of keyword generalization feature vectors includes: Inputting each keyword sample sequence in the keyword sample sequence set into the language representation model respectively; Extracting features from each generalized sample in the generalized sample set corresponding to the keywords of the keyword sample sequence by using the language representation model to obtain a generalized semantic feature set; The generalized semantic feature set is averaged through the language representation model to obtain a keyword generalized feature vector corresponding to each keyword sample sequence.

5. The method for voice keyword detection according to claim 1, characterized in that: The extracting features of the negative sample sequence set by using the language representation model to obtain at least one non-keyword feature vector includes: Inputting each negative sample sequence in the negative sample sequence set into the language representation model respectively; Randomly extracting non-keywords from the negative sample sequence using the language representation model to obtain a plurality of non-keyword data; The language representation model is used to extract features from the plurality of non-keyword data and average the features to obtain the non-keyword feature vector.

6. A device for detecting speech keywords, characterized in that: include: A preprocessing module is used to process the speech data to be processed to obtain speech acoustic features; An acoustic feature extraction module, used for inputting the speech acoustic feature into a preset speech coding network model to obtain a speech acoustic feature vector; An acquisition module is used to retrieve the pan-semantic text space vector under a preset storage path; An attention calculation module, used for performing attention calculation on the pan-semantic text space vector and the speech acoustic feature vector to obtain an acoustic semantic context feature vector; A classification module, used for inputting the acoustic semantic context feature vector into a preset keyword classification model to obtain predicted keywords; The pan-semantic text space vector is obtained by concatenating a plurality of generalized semantic feature vectors; and the pan-semantic text space vector and the speech acoustic feature vector are subjected to attention calculation to obtain an acoustic semantic context feature vector, including: Inputting the pan-semantic text space vector and the speech acoustic feature vector into a preset attention model to perform attention calculation, and obtaining a plurality of weighted distribution data corresponding to the plurality of the generalized semantic feature vectors one by one; Combining a plurality of the weighted distribution data to obtain the acoustic semantic context feature vector; The pan-semantic text space vector is calculated by the following steps: Obtain a preset language representation model, a keyword sample sequence set, and a negative sample sequence set; Extracting features from the keyword sample sequence set using the language representation model to obtain a plurality of keyword generalization feature vectors; Extracting features from the negative sample sequence set using the language representation model to obtain at least one non-keyword feature vector; The pan-semantic text space vector is obtained by concatenating a plurality of the keyword generalization feature vectors and at least one of the non-keyword feature vectors.

7. An electronic device, characterized in that: include: at least one processor, and a memory communicatively connected to at least one processor; wherein, The memory stores instructions, and the instructions are executed by at least one processor, so that when the at least one processor executes the instructions, the method for voice keyword detection as described in any one of claims 1 to 5 is implemented.

8. A computer-readable storage medium, characterized in that: Computer executable instructions are stored, and the computer executable instructions are used to execute at least the method for voice keyword detection as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Voice instruction recognition method and device, electronic equipment and storage medium

    CN110827816A

  • Information processing method and device and storage medium

    CN111583919A