Data processing method, device, equipment, storage medium and program product

By adopting a multilingual detection method based on factorization delay neural network in keyword detection, combining multi-task learning and singular value decomposition layer, the problem of multilingual keyword detection in the existing technology is solved, and efficient and intelligent multilingual keyword detection is achieved.

CN114333790BActive Publication Date: 2025-05-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111472195.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-05-09
Estimated Expiration
2041-12-03

AI Technical Summary

Technical Problem

The prior art is difficult to realize multilingual keyword detection, and traditional methods are done in monolingual and do not support multilingual keyword detection tasks, resulting in low detection efficiency and poor applicability.

Method used

The multilingual keyword detection method based on factorization delay neural network (TDNN-F) is adopted to realize keyword detection of multilingual speech data through multitask learning and weighted finite state converter (WFST) decoding. At the same time, a singular value decomposition layer (SVD) is introduced to reduce the training parameters of the acoustic model and improve the inference speed.

Benefits of technology

The automation and intelligence of multilingual keyword detection has been realized, the detection efficiency and accuracy have been improved, the applicability has been improved, and the correct output of keywords with the same accuracy rate has been increased by 1.2%-3.6%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333790B_ABST
    Figure CN114333790B_ABST
Patent Text Reader

Abstract

The present application proposes a data processing method, device, equipment, storage medium and program product, wherein the method includes: obtaining speech data to be recognized, and obtaining a target syllable sequence of a reference keyword; calling a keyword detection model to process the speech data to be recognized, determining a syllable sequence to be detected of the speech data to be recognized, and determining a keyword detection result of the speech data to be recognized based on the syllable sequence to be detected and the target syllable sequence; wherein the keyword detection model is trained using a training data set, and the training data set includes sample speech data of one or more language categories; the keyword detection model can perform keyword detection on speech data of any language category in one or more language categories. The present application can be applied to various scenarios such as cloud technology, artificial intelligence, smart platforms, and in-vehicle Internet, and can realize keyword detection in multiple languages, realize automation and intelligence of keyword detection, and improve the efficiency of keyword detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, specifically to the field of artificial intelligence technology, and specifically to a data processing method, a data processing device, a computer equipment, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the continuous development and application of computer technology, more and more scenarios require the use of data processing technology. For example, data processing technology is used to detect keywords in voice data to wake up smart devices and detect the frequency of keyword occurrence. However, how to achieve keyword detection is currently a hot research topic. Summary of the invention

[0003] The present application provides a data processing method, a data processing device, a computer device, a computer-readable storage medium, and a computer program product, which can realize keyword detection in multiple languages, realize automation and intelligence of keyword detection, and improve the efficiency of keyword detection.

[0004] The present application provides a data processing method, the method comprising: obtaining speech data to be recognized, and obtaining a target syllable sequence of a reference keyword;

[0005] Calling a keyword detection model to process the voice data to be recognized, determining a syllable sequence to be detected of the voice data to be recognized, and determining a keyword detection result of the voice data to be recognized according to the syllable sequence to be detected and the target syllable sequence;

[0006] Among them, the above-mentioned keyword detection model is trained using a training data set, and the above-mentioned training data set includes sample speech data of one or more language categories; the above-mentioned keyword detection model can perform keyword detection on speech data of any language category in the above-mentioned one or more language categories.

[0007] The present application provides a data processing device, the device comprising:

[0008] An acquisition module, used to acquire speech data to be recognized and a target syllable sequence of a reference keyword;

[0009] A processing module, used for calling a keyword detection model to process the above-mentioned voice data to be recognized, determining a syllable sequence to be detected of the above-mentioned voice data to be recognized, and determining a keyword detection result of the above-mentioned voice data to be recognized according to the above-mentioned syllable sequence to be detected and the above-mentioned target syllable sequence;

[0010] Among them, the above-mentioned keyword detection model is trained using a training data set, and the above-mentioned training data set includes sample speech data of one or more language categories; the above-mentioned keyword detection model can perform keyword detection on speech data of any language category in the above-mentioned one or more language categories.

[0011] The present application provides a computer device, comprising: a memory and a processor, wherein a data processing program is stored in the memory, and when the data processing program is executed by the processor, the steps of the data processing method are implemented.

[0012] The present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the program instructions are executed by a processor to perform the steps of the above-mentioned data processing method.

[0013] The present application provides a computer program product, which includes a computer program or computer instructions. The computer program or computer instructions are executed by a processor to implement the data processing method as described above.

[0014] The present application first obtains the target syllable sequence of the speech data to be recognized and the reference keyword, then calls the keyword detection model to process the speech data to be recognized, determines the syllable sequence to be detected of the speech data to be recognized, and then determines the keyword detection result of the speech data to be recognized according to the syllable sequence to be detected and the target syllable sequence, which can realize the automation and intelligence of keyword detection and improve the efficiency of keyword detection; the keyword detection model proposed in the present application can be trained with sample speech data of multiple language categories, and can perform keyword detection on speech data of multiple language categories, which improves the applicability of keyword detection and further improves the intelligence of keyword detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0016] Figure 1 is a schematic diagram of the architecture of a data processing system provided by an exemplary embodiment of the present application;

[0017] Figure 2 It is a flowchart of a data processing method provided by an exemplary embodiment of the present application;

[0018] Figure 3 is a flowchart of a keyword detection provided by an exemplary embodiment of the present application;

[0019] Figure 4 is a flowchart of another data processing method provided by an exemplary embodiment of the present application;

[0020] Figure 5 is a schematic diagram of a factorized time-delay neural network provided by an exemplary embodiment of the present application;

[0021] Figure 6 It is a schematic diagram of the structure and flow of a syllable recognition network provided by an exemplary embodiment of the present application;

[0022] Figure 7 is a flowchart of another keyword detection provided by an exemplary embodiment of the present application;

[0023] Figure 8 is a schematic block diagram of a data processing device provided by an exemplary embodiment of the present application;

[0024] Fig. 9 It is a schematic block diagram of a computer device provided by another exemplary embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0026] The embodiment of the present application provides a data processing method to realize automation and intelligence of keyword detection and improve the efficiency of keyword detection. The data processing method provided in the embodiment of the present application can be implemented by one or more technologies in artificial intelligence technology.

[0027] Artificial Intelligence (AI) is a theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that the machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level technology and software-level technology. Basic artificial intelligence technologies generally include technologies such as sensors, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technology mainly includes natural language processing technology, computer vision technology, machine learning / deep learning and other major directions. The solution provided in the embodiment of the present application involves technologies such as natural language processing and machine learning under artificial intelligence technology, and natural language processing technology and machine learning technology will be described below.

[0028] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language used by people in daily life, so it is closely related to the study of linguistics. Natural language processing technology generally includes technologies such as speech recognition, text processing, semantic understanding, machine translation, robot question and answer, knowledge graph, etc. This application mainly relates to speech recognition technology in natural language processing technology. Specifically, the terminal device obtains the voice data to be detected and converts the voice data into the corresponding syllable sequence form (that is, the syllable sequence to be detected and the target syllable sequence in this application) through speech recognition technology. Subsequently, keyword detection operations can be performed based on the syllable sequence to be detected and the target syllable sequence.

[0029] Machine Learning (ML) is a multi-disciplinary cross-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. Machine learning specifically studies how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning and other technologies. This application mainly relates to artificial neural networks in machine learning technology. Specifically, the terminal device trains the initialized neural network through a training data set to obtain a keyword detection model, and then uses the keyword detection model to process the speech data to be recognized to obtain a syllable sequence to be detected. Keyword detection is performed based on the syllable sequence to be detected and the target syllable sequence of the keyword, which improves the detection efficiency and intelligence of the keyword.

[0030] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, Internet of Vehicles, automatic driving, smart transportation, etc. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0031] Spoken Term Detection is a subfield of speech recognition, which aims to detect all occurrences of a specified word in a speech signal. Existing keyword detection technologies are generally based on keyword / filler detection methods. The main idea of ​​this method is to use a keyword model to output keyword detection results, and use a filler model to absorb non-keyword speech. In order to build a complete keyword / filler-based speech keyword detection system, the traditional method uses phonemes as modeling units, then uses various deep neural networks as acoustic models, and then uses loss functions to guide model training.

[0032] The main disadvantage of keyword detection methods based on keywords / fillers is that the acoustic model parameters for keyword detection are large and large, resulting in a slow inference speed. In addition, traditional keyword detection methods based on keywords / fillers are all done on a single language and do not support multi-language keyword detection tasks.

[0033] The present invention proposes a multilingual keyword detection method based on a factorized time delay neural network (TDNN-F). The method first mixes the corpus of each language in proportion, adopts a multi-task learning method to train an acoustic model (AM), and then adopts a weighted finite state translator (WFST) method to decode and output keyword results. In the multi-task learning process, we use syllables as output units, and each language has a separate output. The multi-task learning method can treat each language as a separate branch, and the encoder part of each branch is weight-shared, but the last normalization layer (Softmax layer) is independent. Therefore, it can not only learn the common information between each language, but also map the encoder output to the corresponding language through the Softmax layer. This method can be effectively applied to multilingual keyword detection tasks, while previous keyword detection methods are generally done on a single language and do not support multilingual keyword detection tasks.

[0034] In addition, this application proposes to add a singular value decomposition layer (SVD) to TDNN, which effectively reduces the training parameters of the acoustic model and significantly speeds up the inference speed while the effect is not much different. Experiments show that the effect of this method exceeds the cascade method of language recognition and monolingual systems, and the correct reporting of keywords at the same accuracy rate increases by 1.2%-3.6%.

[0035] The keyword detection method proposed in this application can be applied to the fields of smart device wake-up interaction and audio and video file voice keyword detection. In order to cope with the limited computing resources of smart devices and the processing requirements of massive audio and video files, we usually need the keyword detection speed to be fast enough, and then on this basis, we hope that the detection results are accurate enough, and finally we hope that the detection method has more functions and greater flexibility.

[0036] This application can be applied to various scenarios such as cloud technology, artificial intelligence, smart platform, and in-vehicle Internet. In the field of cloud technology, this application can store the voice data to be recognized, the syllable sequence to be detected, the reference keyword, the target syllable sequence, and the keyword detection results on the cloud server, which is convenient for data management and reuse. When the keyword detection results and other data are needed, they can be directly obtained from the cloud server; in the field of artificial intelligence, the keyword detection method can be used to wake up the smart device, making the smart device more humane, and developing more intelligent application services based on the keyword detection technology; in the field of smart platform, after obtaining the user's consent, the keyword detection method can be used to count and analyze the voice data input by the user, determine different types of user classification labels for different keywords, and make relevant recommendations based on user labels to ensure a good user experience; in the field of in-vehicle Internet, the keyword detection method can be used to wake up the in-vehicle device, or the voice data of the passengers in the car can be collected in real time, and intelligent interaction can be performed according to the detected keywords. For example, when the keywords "tired, sleepy, bored" are detected, the in-vehicle device can perform music playback operations, and when the keywords "smoking, hot, stuffy" are detected, the in-vehicle device can open the window.

[0037] This application will be specifically described by the following examples:

[0038] See also Figure 1 , Figure 1 FIG. 1 is a schematic diagram of the architecture of a data processing system provided by an exemplary embodiment of the present application. Figure 1 As shown, the data processing system may specifically include a terminal device 101 and a server 102, and the terminal device 101 and the server 102 are connected via a network, such as a wireless network connection. Based on the data processing method proposed in the present application, the terminal device 101 may collect the speech data to be recognized, and the terminal device 101 may perform feature extraction, syllable recognition, keyword matching, and other processing, and in the processing process, the collected speech data to be recognized, the keyword detection result, and the intermediate data in the processing process may be sent to the server 102, so as to facilitate the subsequent data management of the server 102; the server 102 may also perform feature extraction, syllable recognition, keyword matching, and other processing. When the server 102 executes, the terminal device 101 may collect the speech data to be recognized, and send the speech data to be recognized to the server 102 for the above processing to obtain the keyword detection result, and the server 102 returns the keyword detection result to the terminal device 101, and then performs subsequent operations.

[0039] In an embodiment of the present application, the terminal device 101 can obtain voice data to be recognized and reference keywords; the terminal device 101 can send the voice data to be recognized and the reference keywords to the server 102; the server 102 determines the syllable sequence to be detected of the voice data to be recognized based on the voice data to be recognized, and determines the target syllable sequence of the reference keyword based on the reference keyword; the server 102 calls the keyword detection model to match the syllable sequence to be detected and the target syllable sequence of the voice data to be recognized, obtains the keyword detection result of the voice data to be recognized, and sends the keyword detection result to the terminal device 101; the terminal device 101 displays the keyword detection result on the user interface based on the received keyword detection result.

[0040] The terminal device 101 is also referred to as a terminal, user equipment (UE), access terminal, user unit, mobile device, user terminal, wireless communication device, user agent or user device. The terminal device may be a smart home appliance, a handheld device with wireless communication function (such as a smart phone, a tablet computer), a computing device (such as a personal computer (PC), a vehicle terminal, an intelligent voice interaction device, a wearable device or other smart device, but is not limited thereto.

[0041] Server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), as well as big data and artificial intelligence platforms.

[0042] It is understandable that the system architecture diagram described in the embodiment of the present application is to more clearly illustrate the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided by the embodiment of the present application. Figure 1 In addition to the three devices shown in FIG. , more than three devices may be included; similarly, the server 102 includes Figure 1 In addition to the one server shown in the figure, it can also be composed of multiple servers (that is, a server cluster), and the server 102 can also be a local server established on the terminal device 101. Those skilled in the art will know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.

[0043] See also Figure 2 , Figure 2is a flow chart of a data processing method provided by an exemplary embodiment of the present application, in which the method is applied to Figure 1 Taking the terminal device in (referring to the above-mentioned terminal device 101, which will be described below) as an example, the method may include the following steps:

[0044] S201, obtaining speech data to be recognized, and obtaining a target syllable sequence of a reference keyword.

[0045] In the embodiment of the present application, the speech data to be recognized is the target speech that needs to be detected by keyword, the reference keyword is the keyword to be detected that needs to be detected by keyword, and the target syllable sequence of the reference keyword is obtained to obtain the original data for keyword detection, and the subsequent keyword detection methods are all based on the original data. The embodiment of the present application is to determine whether the target speech includes the reference keyword.

[0046] In one embodiment, the voice data to be recognized in the present application may be voice data obtained by the terminal device through its configured sound pickup device (such as a microphone), or it may be voice data obtained by the terminal device from other devices in the network through the Internet, or it may be voice data obtained directly from a local storage device (such as a USB flash drive).

[0047] In one embodiment, the speech data to be recognized can be speech data in an official language (such as Chinese and English) used in any country, or speech data in a local language (such as Cantonese, Fujian, and Sichuan) used in any region in any country.

[0048] In one embodiment, the reference keywords of the present application are pre-set, and it is necessary to perform keyword detection on the reference keywords in the speech data to be recognized to determine whether the reference keywords exist in the speech data to be recognized. The reference keywords can be keywords of the official language used by any country (such as Chinese, English, Russian), or keywords of the local language used in any region of any country (such as Cantonese, Fujian, Sichuan); in addition, the reference keywords can be one or more.

[0049] In one embodiment, the target syllable sequence of the reference keyword may be obtained by:

[0050] (1) Obtain reference keywords and obtain the target language category of the speech data to be recognized.

[0051] Among them, the reference keywords can be in different languages, and the voice data to be recognized can also be in different languages. Therefore, it is necessary to perform language matching processing on the reference keywords and the voice data to be recognized, so that keyword detection can be performed in subsequent steps based on the syllable sequence of the reference keywords and the syllable sequence of the voice data to be recognized.

[0052] (2) Acquire a syllable sequence of the reference keyword that matches the target language category, and determine the syllable sequence of the reference keyword that matches the target language category as a target syllable sequence.

[0053] In one embodiment, the language category of the speech data to be recognized (that is, the target language category) can be obtained first, and then a reference keyword can be obtained, and the reference keyword can be converted into speech to obtain a reference keyword syllable sequence (that is, a target syllable sequence) corresponding to the target language category.

[0054] In one embodiment, a target syllable sequence of a reference keyword can be generated by a pronunciation dictionary model. The pronunciation dictionary model can include multiple branch networks, and can perform text conversion operations on reference keywords in different languages ​​to generate a syllable sequence in a target language. Exemplarily, when the language category of the speech data to be recognized is language A, and the language category of the reference keyword is language B, the reference keyword of language B can be converted into a syllable sequence of language A according to the pronunciation dictionary model. It should be noted that the pronunciation dictionary model can also first translate the reference keyword of language B into language A, and then perform a syllable recognition operation on the translation result to obtain a target syllable sequence.

[0055] S202, calling a keyword detection model to process the speech data to be recognized, determining a syllable sequence to be detected of the speech data to be recognized, and determining a keyword detection result of the speech data to be recognized according to the syllable sequence to be detected and the target syllable sequence.

[0056] In an embodiment of the present application, the keyword detection model can obtain the syllable sequence of the speech data to be recognized (that is, the syllable sequence to be detected) through feature extraction, acoustic modeling and other operations. When the syllable sequence to be detected and the target syllable sequence are obtained, the segment with the highest similarity to the target syllable sequence in the syllable sequence to be detected can be detected based on the idea of ​​similarity judgment. When the similarity of the segment reaches a preset condition, it can be judged that the target syllable sequence exists in the syllable sequence to be detected (that is, it can be judged that the reference keyword exists in the speech data to be recognized), and the keyword detection result can be obtained.

[0057] In one embodiment, if the reference keyword contains multiple sub-keywords, that is, the target syllable sequence includes multiple sub-syllable sequences, the keyword detection model can perform keyword detection on the multiple sub-syllable sequences included in the target syllable sequence respectively, and output the sub-keywords that meet the preset conditions in the multiple sub-syllable sequences as the keyword detection results.

[0058] In one embodiment, the speech data to be recognized may also include a syllable time sequence corresponding to the syllable sequence to be detected. Therefore, after determining the sub-keyword that meets the preset conditions, the appearance time of the sub-keyword that meets the preset conditions can be determined based on the sub-keyword that meets the preset conditions and the syllable time sequence corresponding to the syllable sequence to be detected; then the sub-keyword that meets the preset conditions and the appearance time are output.

[0059] In one embodiment, the keyword detection model is trained using a training data set, which includes sample speech data in one or more language categories; the keyword detection model can perform keyword detection on speech data in any language category in the one or more language categories.

[0060] The keyword detection model may include a feature extraction network and a syllable recognition network. The keyword detection model performs keyword detection through a to-be-detected syllable sequence and a target syllable sequence, which is mainly implemented based on the syllable recognition network in the keyword detection model. The method of training the keyword detection model may include the following steps:

[0061] (1) Obtain a training data set, where the training data set includes one or more training data subsets, each of which includes sample speech data of any one of one or more language categories and a reference syllable sequence corresponding to the sample speech data.

[0062] When the training data set includes one training data subset (that is, the training data set includes sample speech data of only one language category), the syllable recognition network of the keyword detection model can be trained based on the sample speech data of this language category; when the training data set includes multiple training data subsets (that is, the training data set includes sample speech data of multiple languages ​​category), the corpus of each language can be mixed in proportion, for example, the training corpus of each language can be mixed with reference to the language distribution ratio of online speech data, so that the model can perform better in actual business.

[0063] It should be noted that during the keyword detection model training process, the weight of each language has a great impact on the bias of the final model. In order to make the final model not biased towards any one language, the language weights can be evenly divided. This method has a better training effect when the amount of training data is equal, but in actual business, the training corpus of each language is quite different. Therefore, the language weights need to be adjusted according to the number of languages, actual training data and business needs.

[0064] (2) Perform feature extraction processing on the sample speech data in one or more training data subsets to obtain sample speech features of the sample speech data.

[0065] (3) Using the obtained sample speech features and the corresponding reference syllable sequence, the initial syllable recognition network is trained to obtain a trained syllable recognition network.

[0066] In an embodiment of the present application, the sample speech features of the sample speech data are input into the initial syllable recognition network to obtain a predicted syllable sequence, and then based on the predicted syllable sequence and the reference syllable sequence, the network parameters of the initial syllable recognition network are adjusted in combination with a loss function. When the predicted syllable sequence of subsequent sample speech features meets preset conditions, it can be determined that the training of the initial syllable recognition network is completed, and a trained syllable recognition network is obtained.

[0067] In one embodiment, the model training method can be: using the loss function for back propagation to update the weights of each node in the initialized syllable recognition network, and when the loss function meets the preset conditions, determining the final weights of each node in the initialized model, wherein the loss function can use one or more of a 0-1 loss function, an absolute value loss function, a mean square error loss function, a logarithmic loss function, and an exponential loss function.

[0068] (4) Generate a trained keyword detection model based on the trained syllable recognition network.

[0069] After the syllable recognition network training is completed, a trained keyword detection model can be generated based on the feature extraction network, the syllable recognition network, etc. Among them, the feature extraction network can also be trained using the training data set to achieve more accurate feature extraction.

[0070] In one embodiment, the keyword detection model can be trained by using the same training data set to jointly train the feature extraction network and the syllable recognition network. The method of jointly training the feature extraction network and the syllable recognition network can be:

[0071] The feature extraction network and the syllable recognition network each include an optimizer, and each optimizer corresponds to the network parameters of the network. By fusing the two network parameters, joint training can be performed to obtain the trained feature extraction network and syllable recognition network. It should be noted that joint training is a method to improve the prediction accuracy of the model, and the prediction accuracy of the model is also affected by many factors, such as the quality of the training samples. Therefore, in the use of the method provided in this application, the appropriate training method should be selected in combination with the actual situation and the prediction results to achieve a higher model prediction accuracy.

[0072] Exemplarily, the keyword detection model includes a feature extraction network, a syllable recognition network, and a keyword matching network; the syllable recognition network includes one or more recognition sub-networks, each recognition sub-network is used to perform syllable recognition on speech data of a specified language category, and the specified language category is included in one or more language categories.

[0073] Each recognition subnetwork included in the syllable recognition network can perform syllable recognition on speech data of a specified language category, specifically including the following two situations: one is that the syllable recognition network includes one recognition subnetwork, and the second is that the syllable recognition network includes multiple recognition subnetworks.

[0074] In the first case, when the syllable recognition network includes a recognition subnetwork, the syllable recognition network can perform syllable recognition on speech data of a language category corresponding to the syllable recognition network to obtain a syllable sequence corresponding to the speech data (that is, a syllable sequence to be detected).

[0075] In the second case, when the syllable recognition network includes multiple recognition subnetworks, the network parameters of any two recognition subnetworks in the multiple recognition subnetworks match. In other words, the weights of the encoding layer of the recognition subnetworks corresponding to one or more language categories are shared. Weight sharing of neural networks refers to the application of information learned from a local area to other parts of the data to be processed, so as to reduce the amount of calculation and increase the speed of model training. It should be noted that the weights of any two recognition subnetworks in the multiple recognition subnetworks can be shared, or the weights of all recognition subnetworks in the multiple recognition subnetworks can be shared.

[0076] See also Figure 3 , Figure 3The present invention is a flowchart of keyword detection. First, the speech data to be recognized (for example, hello, world) is input into a feature extraction network to obtain the speech features of the speech data to be recognized; then the speech features of the speech data to be recognized are input into a syllable recognition network to obtain the syllable sequence to be detected of the speech data to be recognized; a vocabulary generation operation is performed on the reference keyword using a pronunciation dictionary (for example, the syllable sequence of "hello" is 'ni3 hao3', and the syllable sequence of "world" is 'shi4 jie4'. Among them, '3' or '4' represents the tone corresponding to the syllable, respectively), to obtain the target syllable sequence corresponding to the reference keyword; and then the keyword matching network is used to match the target syllable sequence with the syllable sequence to be detected to obtain an output result (that is, the keyword matching result).

[0077] Among them, the feature extraction network converts continuous speech information into discrete vector representations through digital signal processing algorithms. These vectors can effectively represent relevant speech features and are beneficial to subsequent speech tasks; the syllable recognition network mainly trains the model based on the given feature sequence and state sequence, where the state output probability of the model can be modeled by graph neural network (Graph Neural Network, GNN), deep neural network (Deep Neural Networks, DNN), convolutional neural network (Convolutional Neural Networks, CNN), recurrent neural network (Recurrent Neural Network, RNN), time delay neural network (Time Delay Neural Network, TDNN), etc., and the state jump probability of the model can be modeled by hidden Markov model (Hidden Markov Model, HMM), etc.; the word list generation operation is to generate the syllable sequence corresponding to the keyword through the dictionary; the keyword matching network mainly searches all possible state spaces through the trained model, and then finds the most likely state sequence of the input speech feature to maximize the output probability. The keyword detection model can detect the keywords in the speech and the start and end positions of the keywords by cascading the above networks and methods.

[0078] For example, the present application can use a multi-task learning method to train each language as a separate branch, with the encoder part of each branch sharing weights, and the final Softmax layer being independent. Therefore, through the above method, common information between languages ​​can be learned, and the encoder output can be mapped to the corresponding language through the Softmax layer.

[0079] For example, after obtaining the weight of the model as a floating point number, the relevant clustering method is used to divide the objects into different categories according to the similarity of the size. Each category only needs to save the weight of a cluster center and the corresponding cluster index (Index), which can greatly reduce the amount of data. When the weight is updated, back propagation is performed to calculate the gradient of each weight, accumulate the gradients of the same category of the previous clustering, and update the weight in combination with the learning rate (LR).

[0080] The present application first obtains the target syllable sequence of the speech data to be recognized and the reference keyword, then calls the keyword detection model to process the speech data to be recognized, determines the syllable sequence to be detected of the speech data to be recognized, and then determines the keyword detection result of the speech data to be recognized according to the syllable sequence to be detected and the target syllable sequence, which can realize the automation and intelligence of keyword detection and improve the efficiency of keyword detection; the keyword detection model proposed in the present application can be trained with sample speech data of multiple language categories, can perform keyword detection on speech data of multiple language categories, and can match the reference keyword with the language category of the speech data to be recognized, which improves the applicability of keyword detection and further improves the intelligence of keyword detection; in the training of the keyword detection model, weight sharing is used to improve the model training speed, and the training data set is mixed according to the language distribution ratio of the online speech data, which improves the detection accuracy of the keyword detection model.

[0081] See also Figure 4 , Figure 4 is a flow chart of a data processing method provided by another exemplary embodiment of the present application, in which the method is applied to Figure 1 Taking the terminal device in as an example, the method may include the following steps:

[0082] S401, obtaining speech data to be recognized, and obtaining a target syllable sequence of a reference keyword.

[0083] The specific implementation of step S401 refers to the relevant description of step S201 in the above embodiment, which will not be repeated here.

[0084] S402 , calling the feature extraction network of the keyword detection model to process the speech data to be recognized, and obtaining speech features of the speech data to be recognized.

[0085] In one embodiment, the above step S402 may be implemented by the following method:

[0086] (1) Call the feature extraction network to perform frame processing on the speech data to be recognized, and obtain the time domain features of the speech data to be recognized.

[0087] In the embodiment of the present application, since the duration of sound that can be heard by human ears is at least 10ms, the digital signal needs to be divided into multiple audible blocks, that is, frames. For example, the speech data to be recognized with a length of 10s can be divided into 25ms as a frame to obtain 400 speech frames, and these 400 speech frames are used as the time domain features of the speech data to be recognized.

[0088] (2) Call the feature extraction network to perform windowing and frequency domain transformation on the time domain features to obtain frequency domain features.

[0089] Frequency domain features are windowed based on time domain features, and then frequency domain transform processing is performed (typical methods include Fourier transform processing). Windowing means using a window function to process each frame, eliminating samples at both ends of a frame, so as to generate a periodic signal. If frequency domain transform processing is performed directly without windowing, spectral leakage will occur.

[0090] (3) Using frequency domain features as speech features of the speech data to be recognized.

[0091] After the frequency domain features of the speech data to be recognized are obtained, the frequency domain features can be used as speech features of the speech data to be recognized.

[0092] S403, calling the syllable recognition network of the keyword detection model to process the speech features to obtain a syllable sequence to be detected of the speech data to be recognized.

[0093] In the embodiment of the present application, the purpose of obtaining the syllable sequence to be detected of the speech data to be recognized is to match the syllable sequence to be detected of the speech data to be recognized with the target syllable sequence of the reference keyword to obtain a keyword detection result. The keyword detection model includes a syllable recognition network, and the function of the syllable recognition network is to generate a syllable sequence corresponding to the input speech data.

[0094] In one embodiment, the above step S403 may be implemented by the following method:

[0095] (1) Call the syllable recognition network of the keyword detection model to process the speech features and determine the target language category of the speech data to be recognized based on the speech features.

[0096] In an embodiment of the present application, the syllable recognition network of the keyword detection model will first determine the language category corresponding to the speech feature, so as to use the recognition sub-network corresponding to the language category to process the speech feature and improve the processing accuracy.

[0097] In one embodiment, the syllable recognition network includes a language recognition network, and the keyword detection model calls the language recognition network to process the speech data to be recognized to obtain the target language category of the speech data to be recognized. The language recognition network can use CV syllable division method, linear prediction residual and other methods to perform language recognition.

[0098] In one embodiment, the language recognition network is independent of the syllable recognition network, and is placed before the syllable recognition network and after the feature extraction network. The language recognition network first obtains the speech features output by the feature extraction network, processes the speech features, obtains the target language category of the speech data to be recognized, and then sends the speech features and the target language category of the speech data to be recognized to the syllable recognition network for processing.

[0099] (2) Call the recognition subnetwork corresponding to the target language category in the syllable recognition network to process the speech features and obtain the syllable distribution probability of the speech features.

[0100] According to the syllable characteristics of each language, the syllable categories corresponding to the language category can be pre-set. For example, there are 1000 syllable types in language A. Then, given a speech in the target language category, each frame of the speech can obtain the distribution probability of 1000 syllable types (that is, the syllable distribution probability) through the recognition subnetwork corresponding to the target language category.

[0101] (3) Determine the syllable sequence to be detected of the speech data to be recognized based on the syllable distribution probability.

[0102] In one embodiment, after calling the recognition subnetwork corresponding to the target language category in the syllable recognition network to process each frame of the speech features of the speech data to be recognized, the syllable distribution probability corresponding to each frame of speech is obtained, and the syllable corresponding to the maximum probability is used as the target syllable; when the syllable distribution probabilities of all speech features are determined, that is, multiple target syllables are obtained, these multiple target syllables are combined as a syllable sequence to be detected for the speech data to be recognized.

[0103] In one embodiment, the syllable recognition network is generated based on a time-delay neural network and a singular value decomposition network.

[0104] In the embodiment of the present application, the time-delay neural network (ie, TDNN) is used to solve the problem that the traditional method of speech recognition, the Hidden Markov Model (HMM), cannot adapt to the dynamic time domain changes in the speech signal. TDNN has fewer structural parameters, and speech recognition does not require the phonetic symbols to be aligned with the audio on the timeline in advance. Experiments have shown that TDNN performs better than HMM. The two obvious features of TDNN are dynamic adaptation to time domain feature changes and fewer parameters. The input layer of the traditional deep neural network is connected to the hidden layer one by one. TDNN has a change here, that is, the characteristics of the hidden layer are not only related to the input at the current moment, but also to the input at the future moment. Therefore, TDNN has the ability to express the relationship between speech features in time, and can facilitate the learning of the TDNN network through the weight sharing method.

[0105] Singular value decomposition (SVD) is a method for decomposing a matrix, but unlike the eigenvalue decomposition method, eigenvalue decomposition is only applicable to square matrices, but SVD does not require the matrix to be decomposed to be a square matrix and can be applied to the decomposition of any matrix. Using SVD, a smaller data set can be used to represent the original data set, thereby removing noise and redundant information, greatly reducing the amount of data and improving the operation speed.

[0106] The present invention uses syllables as modeling units, constructs an SVD network based on the SVD method, combines TDNN and uses a factorized time-delay neural network (i.e., TDNN-F) as a syllable recognition network. The main function of the syllable recognition network is to convert speech input into acoustically represented syllable output. Specifically, for each speech frame, the probability distribution of the syllable to which it belongs is calculated.

[0107] See also Figure 5 yes, Figure 5 TDNN-F network is a multi-layer network, each layer has a strong abstract ability for speech features, and it has a wider context vision, can capture a wider range of context information, and has a stronger modeling ability for speech time sequence dependency information. In addition to the common layer structure (such as Figure 5 A SVD layer can be inserted between some TDNN layers (such as Figure 5 For example, the TDNN-F model has N layers, and each TDNN layer and SVD layer has M and K nodes respectively, where K is much smaller than M. Then the number of model parameters D1 of TDNN is:

[0108] D1=N*M*M

[0109] The number of model parameters D2 of TDNN-F is:

[0110] D2=N*(M+M)*K

[0111] From the above two model parameter calculation formulas, we can see that the number of model parameters D1 of TDNN is much larger than the number of model parameters D2 of TDNN-F. Compared with the TDNN network, the TDNN-F network can effectively reduce the training parameters of the model and speed up the reasoning speed of the model.

[0112] It should be noted that a network model with larger parameters can achieve better results, but at the same time it will slow down the reasoning speed of the model. Therefore, in order to achieve a better balance between speed and effect, the network model usually needs to be parameter optimized. In the TDNN-F network proposed in this application, the main parameters that affect the network effect are the number of network layers, the number of nodes per layer, the number of SVD nodes, and the context field width.

[0113] According to the TDNN-F network proposed above, a syllable recognition network based on the TDNN-F network can be constructed. Figure 6 yes, Figure 6 This is a schematic diagram of the structure and flow of a syllable recognition network proposed in this application. The syllable recognition network is based on the TDNN-F network, and uses the Chain Model method (a method of sequence identification training) and the multi-task learning method to treat each language category as an independent branch. The specific process of the syllable recognition network is: the speech features of the speech data to be recognized are used as the input of the syllable recognition network, and the speech features of different language categories are processed separately through the Chain Model network, and then the syllable sequences corresponding to multiple language categories (such as language 1 and language N) are output (that is, the syllable sequences to be detected).

[0114] Among them, the Chain Model network is a multi-layer network structure ( Figure 6 The network layer where the black solid circle is located is the input layer, the layer where the hollow circle is located is the middle layer, and the layer where the slash circle is located is the output layer). The multi-layer network structure includes multi-layer TDNN-F networks. Each TDNN-F network consists of the i-th layer in the Chain Model network (i refers to any layer in the middle layer), the i-1 layer, the Dropout layer (neuron drop layer, used to set the probability of the specified neural network unit to zero when the neural network is used), the TDNN-F layer, and multiple channels; the multiple channels include direct channels from the i-1 layer to the i layer (that is, skip layer connections), and also include ordinary channels from the i-1 layer to the i layer through the TDNN-F layer and the Dropout layer; the model training is made more stable by skip layer connections and the Dropout method;

[0115] In one embodiment, the syllable recognition network based on the TDNN-F network constructed in the present application can use subsampling layers and shared weight methods to reduce the amount of calculation, and can also use the frame-level cross-entropy loss function and add the LF-MMI loss function on this basis to guide the error propagation of the TDNN-F model; in the network output stage, the syllable recognition network proposed in the present application will perform keyword matching operations on the branches of each language category. When the matching results of the keywords meet the preset conditions, the keywords that meet the conditions and the predicted probability and start and end positions of the keywords are output.

[0116] According to the syllable recognition network based on the TDNN-F network, this application proposes a multilingual speech keyword detection method based on TDNN-F, and the specific implementation method is as follows:

[0117] See also Figure 7 , Figure 7 This is a flow chart of the multi-language speech keyword detection method based on TDNN-F proposed in the present application. First, the speech data to be recognized (for example: hello, world) is input into the feature extraction network to obtain the speech features of the speech data to be recognized; then the speech features of the speech data to be recognized are input into the syllable recognition network to obtain the syllable sequence to be detected of the speech data to be recognized. The syllable recognition network is based on TDNN-F and multi-task learning, and can process the speech features of speech data to be recognized of different language types separately to obtain the syllable sequences to be detected corresponding to multiple languages ​​(for example, language 1 and language N). Then, the state jump of the model is realized by using a method based on hidden Markov (HMM) and Lattic-Free Maximum Mutual Information (LF-MMI). The multi-language speech keyword detection method based on TDNN-F proposed in the present application is based on HMM. HMM describes a Markov process with implicit unknown parameters. The Markov process contains visible states (for example, syllables in speech data, such as Figure 7 S1, S2, etc.) and implicit states (e.g., the contextual relationship between syllables in speech data, such as Figure 7 The method can make full use of the syllable sequence of the speech data to be detected and the implicit states between the syllable sequences, thereby improving the utilization rate of the data.

[0118] In the embodiments of the present application, a multilingual pronunciation dictionary can be used to perform a vocabulary generation operation on a reference keyword (i.e., the multilingual keyword in the figure). For example, the syllable sequences of the keyword "你好" include 'ni3 hao3','nei5 hou2', etc., and the syllable sequences of "世界" include'shi4 jie4','sai3 jaai3', etc., where '2', '3', or '4' respectively represent the tones corresponding to the syllables. After obtaining the target syllable sequences of multiple language types corresponding to the reference keyword; then, a keyword matching network is used to match the target syllable sequences and the syllable sequences to be detected to obtain an output result (i.e., the keyword matching result).

[0119] S404. Invoke the keyword matching network of the keyword detection model to process the syllable sequences to be detected and the target syllable sequences, and obtain the keyword detection result of the speech data to be recognized.

[0120] In one embodiment, the syllable sequences to be detected may include one or more syllable elements arranged in a first order. The first order may be the order of the syllables in the speech data to be recognized; the target syllable sequences include one or more syllable elements arranged in a second order. The second order may be the order of the syllables in the keyword, or may be sorted after adjustment according to the components of the keyword. For example, when the keyword is "你好,世界", it can be sorted according to the syllable order of the keyword, that is, a plurality of syllable elements with "你好世界" as the sorting reference are obtained; it can also be adjusted according to the components of the keyword (the keyword includes two parts, "你好" and "世界"), and the syllable order of "世界你好" is obtained, and a plurality of syllable elements with "世界你好" as the sorting reference are obtained.

[0121] In one embodiment, the above step S404 can be implemented by the following method:

[0122] (1). Invoke the keyword matching network of the keyword detection model to process the syllable sequences to be detected and the target syllable sequences, and detect whether there is a matching syllable element in the target syllable sequence that matches the first syllable element, where the first syllable element is any syllable element in the syllable sequence to be detected.

[0123] In the embodiments of the present application, the first syllable element can be taken from any syllable element in the syllable sequence to be detected. The above step can be understood as: invoking the keyword matching network of the keyword detection model, and respectively matching the target syllable sequence with each syllable element in the syllable sequence to be detected (in each matching process, the selected syllable element in the syllable sequence to be detected can be used as the first syllable element), and then detecting whether there is a matching syllable element in the target syllable sequence that matches the first syllable element.

[0124] In one embodiment, a method for determining whether there is a matching syllable element in a target syllable sequence that matches the first syllable element may be: calling a keyword matching network to match each syllable element in the target syllable sequence with the first syllable element respectively, obtaining multiple matching probabilities corresponding to each syllable element in the target syllable sequence, obtaining the maximum matching probability among the multiple matching probabilities, and if the maximum matching probability is greater than the syllable matching probability threshold, determining that the syllable element corresponding to the maximum matching probability is a matching syllable element, that is, determining that there is a matching syllable element in the target syllable sequence that matches the first syllable element.

[0125] (2) If there is a matching syllable element that matches the first syllable element, then detect whether there is a matching syllable element that matches the second syllable element in the target syllable sequence, where the second syllable element is the syllable element that is arranged after the first syllable element in the syllable sequence to be detected.

[0126] In one embodiment, after determining that there is a matching syllable element that matches the first syllable element in the target syllable sequence, the syllable element following the first syllable element can be used as the second syllable element, and according to the method provided in the above step (1), the keyword matching network is called to match the target syllable sequence and the second syllable element.

[0127] (3) If there is a matching syllable element that matches the second syllable element, a keyword detection result of the speech data to be recognized is determined based on the first syllable element and the second syllable element.

[0128] In one embodiment, the reference keyword has two syllable elements, and syllable matching is performed based on the methods of steps (1) and (2). If the target syllable sequence has a matching syllable element that matches the first syllable element, and the target syllable sequence has a matching syllable element that matches the second syllable element (in this case, the second syllable element contains one syllable element), it is determined that the target syllable sequence exists in the speech data to be recognized.

[0129] In one embodiment, the number of syllable elements of the reference keyword is K (K is greater than 2), then the second syllable element may include K-1 syllable elements (wherein the K-1 syllable elements may be sorted according to the second order), after determining that there is a matching syllable element matching the first syllable element in the target syllable sequence according to the method provided in step (1), keyword matching is performed on the first of the K-1 syllable elements according to the method in the above step (1), and when the match is successful, keyword judgment is performed on subsequent syllable elements in the K-1 syllable elements in turn. If the K-1 syllable elements can match the syllable elements in the syllable sequence to be detected, it is determined that the target syllable sequence exists in the speech data to be recognized, and the keyword detection result of the speech data to be recognized is output.

[0130] In one embodiment, outputting the keyword detection result of the speech data to be recognized may specifically include outputting the target keyword judged to exist in the speech data to be recognized, the start and end time of the target keyword, and the predicted probability of the target keyword.

[0131] In one embodiment, the predicted probability of the target keyword refers to the fusion of the matching probabilities of all syllable elements in the target keyword. For example, the target keyword contains K (for example, K=3) syllable elements, the syllable matching probability threshold is V (for example, V=0.9), and the matching probabilities of the K syllable elements contained in the target keyword are all greater than the syllable matching probability threshold (for example, the matching probabilities of the three syllable elements are 0.96, 0.98, and 0.94, respectively). The average value of the matching probabilities of the three syllable elements (for example, 0.96) can be used as the predicted probability of the target keyword, or the product of the matching probabilities of the three syllable elements (for example, 0.88) can be used as the predicted probability of the target keyword.

[0132] The keyword matching network may include multiple matching strategies. The matching probabilities obtained by the multiple matching strategies are used to determine the results of keywords, thereby improving the accuracy of keyword detection. The matching strategies may include methods such as dynamic programming, longest sequence first, and exhaustive optimal path. The matching strategies may be implemented based on one or more of the above methods. When keyword matching is performed based on multiple matching strategies, the specific steps are as follows:

[0133] (1) Call multiple matching strategies in the keyword matching network to process the syllable sequence to be detected (including the first syllable element and the second syllable element) and the target syllable sequence respectively to obtain multiple posterior probability sets, each of which includes multiple prediction probabilities of the reference keyword corresponding to one of the multiple matching strategies.

[0134] Exemplarily, the keyword matching network includes Z (for example, Z = 3) matching strategies, and the Z matching strategies are used to process the syllable sequence to be detected and the target syllable sequence to obtain Z posterior probability sets (for example, each posterior probability set includes 3 posterior probabilities, and the posterior probability sets are [0.94, 0.96, 0.98], [0.88, 0.94, 0.96] and [0.90, 0.94, 0.92] respectively).

[0135] (2) Determine the maximum prediction probability of multiple prediction probabilities in each posterior probability set as the strategy probability, and determine the prediction probability of the reference keyword based on the multiple strategy probabilities.

[0136] Exemplarily, determine the maximum prediction probability among multiple prediction probabilities based on the set of posterior probabilities (for example, the maximum prediction probability corresponding to the first matching strategy is 0.98, the maximum prediction probability corresponding to the second matching strategy is 0.96, and the maximum prediction probability corresponding to the third matching strategy is 0.94). Then, determine the prediction probability of the reference keyword based on the maximum prediction probability among multiple prediction probabilities (for example, the three maximum prediction probabilities corresponding to the three matching strategies are 0.98, 0.96, and 0.94 respectively, and the average of the three maximum prediction probabilities is calculated to obtain the prediction probability of the reference keyword as 0.96).

[0137] (3) Determine the keyword detection result of the to-be-recognized speech data based on the prediction probability of the reference keyword.

[0138] Exemplarily, the syllable matching probability threshold is V (for example, V = 0.9), and the prediction probability of the reference keyword is 0.96, which is greater than the syllable matching probability threshold. Then, determine that the reference keyword is the target keyword, and then output the keyword detection result.

[0139] During the keyword detection process, if the to-be-detected speech data is not accurate (for example, the user's pronunciation is not standard, so the to-be-detected syllable sequence recognized by the to-be-detected speech data is relatively fuzzy), a correction network can be added to the keyword detection model proposed in this application (it can be before the feature extraction network or after the syllable recognition network). The to-be-detected speech data or the to-be-detected syllable sequence is corrected based on the built-in grammar rules or syllable pronunciation rules based on multiple language categories to obtain a more accurate to-be-detected syllable sequence, which can improve the accuracy of subsequent key matching detection.

[0140] In addition, when the method proposed in this application generates the target syllable sequence based on the reference keyword, it also generates a fuzzy syllable sequence. The fuzzy syllable sequence is a non-standard form of the target syllable sequence (for example, the reference keyword is "world", the target syllable sequence of this keyword can be "shi4 jie4", and the fuzzy syllable sequence of this keyword can be "si4jie4"). By generating the fuzzy syllable sequence and performing keyword detection based on the target syllable sequence, the fuzzy syllable sequence, and the to-be-detected syllable sequence, the non-standard keyword speech data in the to-be-detected speech data can be detected, improving the intelligence and applicability of the method. This method can improve the user experience in some specific environments (such as keyword detection of speech data of the elderly, children, and people with non-standard pronunciation).

[0141] When performing keyword matching based on the fuzzy syllable sequence, the target syllable sequence, and the to-be-detected syllable sequence, it can be achieved in the following way:

[0142] The keyword matching network of the keyword detection model is called to process the syllable sequence to be detected, the target syllable sequence and the ambiguous syllable sequence to obtain the keyword detection result of the speech data to be recognized. According to the method provided in the above step S404, the first prediction probability of the reference keyword (that is, the target syllable sequence) and the second prediction probability of the ambiguous syllable sequence can be obtained respectively, and the keyword detection result of the speech data to be recognized is jointly determined according to the first prediction probability and the second prediction probability.

[0143] Exemplarily, the first prediction probability of the target syllable sequence is 0.88, the second prediction probability of the fuzzy syllable sequence is 0.92, and the first syllable matching probability threshold is 0.9. By analyzing that the first prediction probability of the target syllable sequence is less than the syllable matching probability threshold (this indicates that there is no standardized syllable sequence based on the keyword in the voice data to be detected), but the second prediction probability of the fuzzy syllable sequence is greater than the syllable matching probability threshold (this indicates that there is a non-standard syllable sequence based on the keyword in the voice data to be detected), it is determined that there is a non-standard form of the keyword in the voice data to be detected, and further judgment is performed. For example, a second syllable matching probability threshold P (for example, P is 0.85) can be set. If the first prediction probability of the target syllable sequence is less than the first syllable matching probability threshold, the first prediction probability of the target syllable sequence is greater than the second syllable matching probability threshold, and the second prediction probability of the fuzzy syllable sequence is greater than the syllable matching probability threshold, it is determined that the keyword exists in the voice data to be recognized.

[0144] This application tests the proposed keyword detection method based on TDNN-F. First, two target language categories (including language 1 and language 2) are selected, and 27.5 hours of corpus are selected in each language category as test data. In order to test the performance of the keyword detection model based on TDNN-F, a set of language identification + monolingual keyword detection cascade is also built as a baseline for comparison.

[0145] See the table below:

[0146] method Accuracy Recall Harmonic mean Language recognition + monolingual keyword detection cascade 91.6% 64.9% 76.0% Keyword detection based on TDNN-F 91.7% 68.5% 78.4% Improvement effect +0.1% +3.6% +2.45%

[0147] The table lists the effects of keyword detection in language 1. In language 1, the accuracy of the language identification + monolingual keyword detection cascade method is 91.6%, the recall rate is 64.9%, and the harmonic mean is 76.0%; the accuracy of the keyword detection method based on TDNN-F is 91.7%, the recall rate is 68.5%, and the harmonic mean is 78.4%. It is calculated that the keyword detection method based on TDNN-F has an accuracy improvement of 0.1%, a recall improvement of 3.6%, and a harmonic mean improvement of 2.45% compared with the language identification + monolingual keyword detection cascade method.

[0148] From the results, we can see that compared with the language recognition + monolingual keyword detection cascade method, the keyword detection method based on TDNN-F proposed in this application has a 3.6% higher coverage rate while maintaining the same accuracy. The results show that the keyword detection system based on TDNN-F can effectively increase the number of keyword recalls and has better results.

[0149] See the table below:

[0150] method Accuracy Recall Harmonic mean Language recognition + monolingual keyword detection cascade 84.2% 78.5% 81.1% Keyword detection based on TDNN-F 84.4% 79.4% 81.82% Improvement effect +0.2% +1.2% +0.73%

[0151] The table lists the effects of keyword detection in language 2. In language 2, the accuracy of the language identification + monolingual keyword detection cascade method is 84.2%, the recall rate is 78.5%, and the harmonic mean is 81.1%; the accuracy of the keyword detection method based on TDNN-F is 84.4%, the recall rate is 79.4%, and the harmonic mean is 81.82%. It is calculated that the keyword detection method based on TDNN-F has an accuracy improvement of 0.2%, a recall improvement of 1.2%, and a harmonic mean improvement of 0.73% compared with the language identification + monolingual keyword detection cascade method.

[0152] From the results, we can see that compared with the language recognition + monolingual keyword detection cascade method, the keyword detection method based on TDNN-F proposed in this application has a 1.2% higher coverage rate while maintaining basically the same accuracy. This structure once again verifies the effectiveness of the multilingual keyword detection method based on TDNN-F.

[0153] The syllable recognition network in the keyword detection model proposed in the present application includes a language recognition network, which first performs language recognition on the speech to be detected, and then calls the recognition sub-network corresponding to the target language category in the syllable recognition network to process the speech features, and determines the syllable sequence to be detected of the speech data to be recognized, so that the keyword detection model can perform targeted processing on speech data of different languages, thereby improving the processing efficiency; the present application selects syllables as modeling units, and uses singular value decomposition method and time-delay neural network to construct a keyword detection model based on factorization time-delay neural network, thereby improving the accuracy and efficiency of keyword detection, and uses skip-layer connection and neuron discarding methods to further improve the detection efficiency and model training stability; the present application also sorts the syllable sequence of the reference keyword and generates a fuzzy syllable sequence corresponding to the reference keyword, thereby enabling the keyword detection model to detect different syllable arrangement orders and non-standardized keywords, thereby improving the applicability and intelligence; by presetting multiple matching strategies in the keyword matching network, and determining the keyword detection result based on the keyword prediction probability obtained by the multiple matching strategies, the accuracy of keyword detection is further improved.

[0154] See also Figure 8 , Figure 8 : is a schematic block diagram of a data processing device provided in an embodiment of the present application. The data processing device may specifically include:

[0155] An acquisition module 801 is used to acquire speech data to be recognized and to acquire a target syllable sequence of a reference keyword;

[0156] The processing module 802 is used to call the keyword detection model to process the speech data to be recognized, determine the syllable sequence to be detected of the speech data to be recognized, and determine the keyword detection result of the speech data to be recognized according to the syllable sequence to be detected and the target syllable sequence;

[0157] The keyword detection model is trained using a training data set, which includes sample speech data of one or more language categories; the keyword detection model can perform keyword detection on speech data of any language category in the one or more language categories.

[0158] In one embodiment, the keyword detection model includes a feature extraction network, a syllable recognition network, and a keyword matching network; the syllable recognition network includes one or more recognition subnetworks, each recognition subnetwork is used to perform syllable recognition on speech data of a specified language category, and the specified language category is included in the one or more language categories; when the syllable recognition network includes multiple recognition subnetworks, network parameters of any two of the multiple recognition subnetworks are matched.

[0159] Optionally, the acquisition module 801, when used to acquire a target syllable sequence of a reference keyword, is specifically used to:

[0160] Acquire reference keywords, and acquire the target language category of the speech data to be recognized;

[0161] A syllable sequence of the reference keyword that matches the target language category is obtained, and the syllable sequence of the reference keyword that matches the target language category is determined as a target syllable sequence.

[0162] Optionally, the processing module 802, when used to call the keyword detection model to process the voice data to be recognized, determine the syllable sequence to be detected of the voice data to be recognized, and determine the keyword detection result of the voice data to be recognized according to the syllable sequence to be detected and the target syllable sequence, is specifically used to:

[0163] Calling a feature extraction network of a keyword detection model to process the speech data to be recognized, and obtaining speech features of the speech data to be recognized;

[0164] Calling the syllable recognition network of the keyword detection model to process the speech features to obtain a syllable sequence to be detected of the speech data to be recognized;

[0165] The keyword matching network of the keyword detection model is called to process the syllable sequence to be detected and the target syllable sequence to obtain a keyword detection result of the speech data to be recognized.

[0166] Optionally, the processing module 802, when the syllable recognition network for calling the keyword detection model processes the speech feature to obtain the syllable sequence to be detected of the speech data to be recognized, is specifically used to:

[0167] Calling the syllable recognition network of the keyword detection model to process the speech features, and determining the target language category of the speech data to be recognized according to the speech features;

[0168] Calling the recognition subnetwork corresponding to the target language category in the syllable recognition network to process the speech feature to obtain the syllable distribution probability of the speech feature;

[0169] The syllable sequence to be detected of the speech data to be recognized is determined according to the syllable distribution probability.

[0170] Optionally, the syllable sequence to be detected includes one or more syllable elements arranged in a first order, and the target syllable sequence includes one or more syllable elements arranged in a second order. The processing module 802, when the keyword matching network used to call the keyword detection model processes the syllable sequence to be detected and the target syllable sequence to obtain the keyword detection result of the speech data to be recognized, is specifically used to:

[0171] Calling the keyword matching network of the keyword detection model to process the syllable sequence to be detected and the target syllable sequence, and detecting whether there is a matching syllable element matching the first syllable element in the target syllable sequence, where the first syllable element is any syllable element in the syllable sequence to be detected;

[0172] If there is a matching syllable element that matches the first syllable element, then detecting whether there is a matching syllable element that matches the second syllable element in the target syllable sequence, where the second syllable element is a syllable element that is arranged one position after the first syllable element in the syllable sequence to be detected;

[0173] If there is a matching syllable element that matches the second syllable element, a keyword detection result of the speech data to be recognized is determined according to the first syllable element and the second syllable element.

[0174] In one embodiment, the syllable recognition network is generated based on a time-delay neural network and a singular value decomposition network.

[0175] Optionally, the processing module 802 is further configured to:

[0176] Acquire the training data set, wherein the training data set includes one or more training data subsets, each training data subset includes sample speech data of any one of the one or more language categories and a reference syllable sequence corresponding to the sample speech data;

[0177] Performing feature extraction processing on the sample speech data in the one or more training data subsets to obtain sample speech features of the sample speech data;

[0178] Using the obtained sample speech features and the corresponding reference syllable sequence, an initial syllable recognition network is trained to obtain a trained syllable recognition network;

[0179] Generate a trained keyword detection model based on the trained syllable recognition network.

[0180] It should be noted that the functions of each functional module of the image processing device in the embodiment of the present application can be specifically implemented according to the method in the above method embodiment. The specific implementation process can refer to the relevant description of the above method embodiment, which will not be repeated here.

[0181] See also Fig. 9 , FIG. is a schematic block diagram of a computer device provided in an embodiment of the present application. As shown in the figure, the computer device in the embodiment of the present application may include: a processor 901, a storage device 902, and a network interface 903. The processor 901, the storage device 902, and the network interface 903 may exchange data.

[0182] The above-mentioned storage device 902 may include a volatile memory (volatile memory), such as a random-access memory (RAM); the storage device 902 may also include a non-volatile memory (non-volatile memory), such as a flash memory, a solid-state drive (SSD), etc.; the above-mentioned storage device 902 may also include a combination of the above-mentioned types of memory.

[0183] The processor 901 may be a central processing unit (CPU). In one embodiment, the processor 901 may also be a graphics processing unit (GPU). The processor 901 may also be a combination of a CPU and a GPU. In one embodiment, the storage device 902 is used to store program instructions, and the processor 901 may call the program instructions to perform the following operations:

[0184] Acquire speech data to be recognized, and acquire a target syllable sequence of a reference keyword;

[0185] Calling a keyword detection model to process the speech data to be recognized, determining a syllable sequence to be detected of the speech data to be recognized, and determining a keyword detection result of the speech data to be recognized according to the syllable sequence to be detected and the target syllable sequence;

[0186] The keyword detection model is trained using a training data set, which includes sample speech data of one or more language categories; the keyword detection model can perform keyword detection on speech data of any language category in the one or more language categories.

[0187] In one embodiment, the keyword detection model includes a feature extraction network, a syllable recognition network, and a keyword matching network; the syllable recognition network includes one or more recognition subnetworks, each recognition subnetwork is used to perform syllable recognition on speech data of a specified language category, and the specified language category is included in the one or more language categories; when the syllable recognition network includes multiple recognition subnetworks, network parameters of any two of the multiple recognition subnetworks are matched.

[0188] Optionally, the processor 901, when used to obtain a target syllable sequence of a reference keyword, is specifically configured to:

[0189] Acquire reference keywords, and acquire the target language category of the speech data to be recognized;

[0190] A syllable sequence of the reference keyword that matches the target language category is obtained, and the syllable sequence of the reference keyword that matches the target language category is determined as a target syllable sequence.

[0191] Optionally, the processor 901, when used to call a keyword detection model to process the voice data to be recognized, determine a syllable sequence to be detected of the voice data to be recognized, and determine a keyword detection result of the voice data to be recognized according to the syllable sequence to be detected and the target syllable sequence, is specifically used to:

[0192] Calling a feature extraction network of a keyword detection model to process the speech data to be recognized, and obtaining speech features of the speech data to be recognized;

[0193] Calling the syllable recognition network of the keyword detection model to process the speech features to obtain a syllable sequence to be detected of the speech data to be recognized;

[0194] The keyword matching network of the keyword detection model is called to process the syllable sequence to be detected and the target syllable sequence to obtain a keyword detection result of the speech data to be recognized.

[0195] Optionally, the processor 901, when being used to call the syllable recognition network of the keyword detection model to process the speech feature to obtain the syllable sequence to be detected of the speech data to be recognized, is specifically used to:

[0196] Calling the syllable recognition network of the keyword detection model to process the speech features, and determining the target language category of the speech data to be recognized according to the speech features;

[0197] Calling the recognition subnetwork corresponding to the target language category in the syllable recognition network to process the speech feature to obtain the syllable distribution probability of the speech feature;

[0198] The syllable sequence to be detected of the speech data to be recognized is determined according to the syllable distribution probability.

[0199] Optionally, the syllable sequence to be detected includes one or more syllable elements arranged in a first order, and the target syllable sequence includes one or more syllable elements arranged in a second order, and the processor 901, when the keyword matching network used to call the keyword detection model processes the syllable sequence to be detected and the target syllable sequence to obtain the keyword detection result of the speech data to be recognized, is specifically used to:

[0200] Calling the keyword matching network of the keyword detection model to process the syllable sequence to be detected and the target syllable sequence, and detecting whether there is a matching syllable element matching the first syllable element in the target syllable sequence, where the first syllable element is any syllable element in the syllable sequence to be detected;

[0201] If there is a matching syllable element that matches the first syllable element, then detecting whether there is a matching syllable element that matches the second syllable element in the target syllable sequence, where the second syllable element is a syllable element that is arranged one position after the first syllable element in the syllable sequence to be detected;

[0202] If there is a matching syllable element that matches the second syllable element, a keyword detection result of the speech data to be recognized is determined according to the first syllable element and the second syllable element.

[0203] In one embodiment, the syllable recognition network is generated based on a time-delay neural network and a singular value decomposition network.

[0204] Optionally, the processor 901 is further configured to:

[0205] Acquire the training data set, wherein the training data set includes one or more training data subsets, each training data subset includes sample speech data of any one of the one or more language categories and a reference syllable sequence corresponding to the sample speech data;

[0206] Performing feature extraction processing on the sample speech data in the one or more training data subsets to obtain sample speech features of the sample speech data;

[0207] Using the obtained sample speech features and the corresponding reference syllable sequence, an initial syllable recognition network is trained to obtain a trained syllable recognition network;

[0208] Generate a trained keyword detection model based on the trained syllable recognition network.

[0209] In a specific implementation, the processor 901, the storage device 902, and the network interface 903 described in the embodiments of the present application can execute the embodiments of the present application. Figure 2 or Figure 4 The implementation method described in the relevant embodiments of the data processing method provided can also be implemented in the embodiments of the present application Figure 8 The implementation methods described in the relevant embodiments of the provided data processing device will not be repeated here.

[0210] In the several embodiments provided in the present application, it should be understood that the disclosed methods, devices and systems can be implemented in other ways. For example, the device embodiments described above are merely schematic; for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation; for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0211] In addition, it should be pointed out here that: the embodiment of the present application also provides a computer-readable storage medium, and the computer-readable storage medium stores a computer program executed by the image processing device mentioned above, and the computer program includes program instructions. When the processor executes the above program instructions, it can execute the above Figure 2 , Figure 4 The method in the corresponding embodiment, therefore, will not be repeated here. In addition, the description of the beneficial effects of the same method will not be repeated. For technical details not disclosed in the computer-readable storage medium embodiment involved in this application, please refer to the description of the method embodiment of this application. As an example, the program instructions can be deployed on a computer device, or executed on multiple computer devices located in one location, or, executed on multiple computer devices distributed in multiple locations and interconnected by a communication network, and multiple computer devices distributed in multiple locations and interconnected by a communication network can constitute a blockchain system.

[0212] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device can perform the above Figure 2 , Figure 4 The method in the corresponding embodiment will therefore not be described in detail here.

[0213] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing related hardware through a computer program, and the above-mentioned program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The above-mentioned storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0214] The above disclosure is only part of the embodiments of the present application, which certainly cannot be used to limit the scope of rights of the present application. Ordinary technicians in this field can understand that all or part of the processes of implementing the above embodiments and making equivalent changes according to the claims of this application are still within the scope of the invention.

Claims

1. A data processing method, characterized in that: The method comprises: Acquire speech data to be recognized, and acquire a target syllable sequence of a reference keyword; Calling a keyword detection model to process the speech data to be recognized, determining a syllable sequence to be detected of the speech data to be recognized, and determining a keyword detection result of the speech data to be recognized according to the syllable sequence to be detected and the target syllable sequence; Among them, the keyword detection model is trained using a training data set, and the training data set includes sample speech data of one or more language categories; the keyword detection model can perform keyword detection on speech data of any language category in the one or more language categories; the keyword detection model includes a feature extraction network, a syllable recognition network, and a keyword matching network; the syllable recognition network includes one or more recognition subnetworks, each recognition subnetwork is used to perform syllable recognition on speech data of a specified language category, and the specified language category is included in the one or more language categories; when the syllable recognition network includes multiple recognition subnetworks, the network parameters of any two recognition subnetworks in the multiple recognition subnetworks match.

2. The method according to claim 1, characterized in that The step of obtaining a target syllable sequence of a reference keyword includes: Acquire reference keywords, and acquire the target language category of the speech data to be recognized; A syllable sequence of the reference keyword that matches the target language category is obtained, and the syllable sequence of the reference keyword that matches the target language category is determined as a target syllable sequence.

3. The method according to claim 1, characterized in that The calling of the keyword detection model to process the speech data to be recognized, determining a syllable sequence to be detected of the speech data to be recognized, and determining a keyword detection result of the speech data to be recognized according to the syllable sequence to be detected and the target syllable sequence, comprises: Calling a feature extraction network of a keyword detection model to process the speech data to be recognized, and obtaining speech features of the speech data to be recognized; Calling the syllable recognition network of the keyword detection model to process the speech features to obtain a syllable sequence to be detected of the speech data to be recognized; The keyword matching network of the keyword detection model is called to process the syllable sequence to be detected and the target syllable sequence to obtain a keyword detection result of the speech data to be recognized.

4. The method according to claim 3, characterized in that The calling of the syllable recognition network of the keyword detection model to process the speech features to obtain the syllable sequence to be detected of the speech data to be recognized includes: Calling the syllable recognition network of the keyword detection model to process the speech features, and determining the target language category of the speech data to be recognized according to the speech features; Calling the recognition subnetwork corresponding to the target language category in the syllable recognition network to process the speech feature to obtain the syllable distribution probability of the speech feature; The syllable sequence to be detected of the speech data to be recognized is determined according to the syllable distribution probability.

5. The method according to claim 3, characterized in that: The to-be-detected syllable sequence includes one or more syllable elements arranged in a first order, and the target syllable sequence includes one or more syllable elements arranged in a second order; the keyword matching network that calls the keyword detection model processes the to-be-detected syllable sequence and the target syllable sequence to obtain a keyword detection result of the to-be-recognized speech data, including: Calling the keyword matching network of the keyword detection model to process the syllable sequence to be detected and the target syllable sequence, and detecting whether there is a matching syllable element matching the first syllable element in the target syllable sequence, where the first syllable element is any syllable element in the syllable sequence to be detected; If there is a matching syllable element that matches the first syllable element, then detecting whether there is a matching syllable element that matches the second syllable element in the target syllable sequence, where the second syllable element is a syllable element that is arranged one position after the first syllable element in the syllable sequence to be detected; If there is a matching syllable element that matches the second syllable element, a keyword detection result of the speech data to be recognized is determined according to the first syllable element and the second syllable element.

6. The method according to claim 1, characterized in that The syllable recognition network is generated based on a time-delay neural network and a singular value decomposition network.

7. The method according to claim 1, characterized in that The method further comprises: Acquire the training data set, wherein the training data set includes one or more training data subsets, each training data subset includes sample speech data of any one of the one or more language categories and a reference syllable sequence corresponding to the sample speech data; Performing feature extraction processing on the sample speech data in the one or more training data subsets to obtain sample speech features of the sample speech data; Using the obtained sample speech features and the corresponding reference syllable sequence, an initial syllable recognition network is trained to obtain a trained syllable recognition network; Generate a trained keyword detection model based on the trained syllable recognition network.

8. A data processing device, characterized in that: The device comprises: An acquisition module, used to acquire speech data to be recognized and a target syllable sequence of a reference keyword; A processing module, used for invoking a keyword detection model to process the speech data to be recognized, determining a syllable sequence to be detected of the speech data to be recognized, and determining a keyword detection result of the speech data to be recognized according to the syllable sequence to be detected and the target syllable sequence; Among them, the keyword detection model is trained using a training data set, and the training data set includes sample speech data of one or more language categories; the keyword detection model can perform keyword detection on speech data of any language category in the one or more language categories; the keyword detection model includes a feature extraction network, a syllable recognition network, and a keyword matching network; the syllable recognition network includes one or more recognition subnetworks, each recognition subnetwork is used to perform syllable recognition on speech data of a specified language category, and the specified language category is included in the one or more language categories; when the syllable recognition network includes multiple recognition subnetworks, the network parameters of any two recognition subnetworks in the multiple recognition subnetworks match.

9. A computer device, characterized in that: include: A memory and a processor, wherein a data processing program is stored in the memory, and when the data processing program is executed by the processor, it is used to implement the data processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium device, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the program instructions are executed by a processor to implement the data processing method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The computer program product comprises a computer program or a computer instruction. When the computer program or the computer instruction is executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Wake-up word recognition method and device, electronic device, and computer-readable storage medium

    CN109065044A

  • Multilingual voice keyword detection method, multilingual voice keyword model generation method and electronic equipment

    CN112185346A