Keyword detection method, device, equipment and storage medium

By constructing a first index structure based on the end-to-end speech recognition model, the problem of low keyword detection recall in the prior art is solved, and a more comprehensive keyword detection effect is achieved.

CN113823266BActive Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110832861.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-22
Publication Date
2025-05-27
Estimated Expiration
2041-07-22

AI Technical Summary

Technical Problem

When using end-to-end speech recognition model for keyword detection, it is difficult to comprehensively detect keywords in the speech signal, resulting in a low recall rate.

Method used

By acoustic features of the speech signal, the end-to-end speech recognition model is called for recognition, and the first speech recognition results of each frame of speech segment are obtained, the first index structure is constructed based on the results, and then keyword detection is performed.

Benefits of technology

This method can detect keywords present in the voice signal in a more comprehensive manner, improve the recall rate of keyword detection results, and improve the detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113823266B_ABST
    Figure CN113823266B_ABST
Patent Text Reader

Abstract

The present application discloses a keyword detection method, apparatus, device and storage medium, belonging to the field of artificial intelligence technology, and relating to the automatic speech recognition technology and speech keyword detection technology in the field of artificial intelligence technology. The method includes: obtaining a speech signal; calling an end-to-end speech recognition model to recognize the acoustic features of the speech signal, and obtaining first speech recognition results respectively corresponding to each frame of speech segments in the speech signal; based on the first speech recognition results respectively corresponding to each frame of speech segments, obtaining a first index structure, where the first index structure is an index structure corresponding to the speech signal for keyword detection; and obtaining a keyword detection result of the speech signal based on the first index structure. Such a process can more comprehensively detect the keywords existing in the speech signal, thereby improving the recall rate of the obtained keyword detection result, and the effect of keyword detection is better.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to a keyword detection method, apparatus, device, and storage medium. Background Art

[0002] With the development of artificial intelligence (AI) technology, the application of automatic speech recognition (ASR) technology is becoming more and more widespread. Keyword spotting (KWS) is an important application of automatic speech recognition technology, which is used to detect whether certain keywords exist in a speech signal, so as to wake up an electronic device or control the operation of the electronic device according to the detected keywords.

[0003] In the process of keyword detection for a speech signal, it is necessary to first call a speech recognition model to recognize the acoustic features of the speech signal to obtain a recognition result, and then implement keyword detection according to the recognition result. The existing speech recognition models include two types: one is a traditional hybrid speech recognition model including an acoustic model and a language model, and the other is an end-to-end speech recognition model trained based on the connectionist temporal classification (CTC) algorithm. Among them, the end-to-end speech recognition model has gradually become the mainstream speech recognition model.

[0004] In the related art, in the case of calling an end-to-end speech recognition model to recognize acoustic features, based on the obtained recognition result, the most matching speech unit corresponding to each frame of speech segment is determined, and then the optimal speech unit sequence composed of the most matching speech units corresponding to each frame of speech segment is matched with the candidate keyword in text to obtain a keyword detection result. In this way, the keyword detection result is determined according to the optimal speech unit sequence, and the optimal speech unit sequence is difficult to comprehensively reflect the obtained recognition result, resulting in a poor keyword detection effect, making it difficult to comprehensively detect the keywords existing in the speech signal, and the recall rate of the keyword detection result is low. Summary of the Invention

[0005] Embodiments of the present application provide a keyword detection method, apparatus, device, and storage medium, which can be used to comprehensively detect the keywords existing in a speech signal and improve the recall rate of the obtained keyword detection result. The technical solution is as follows:

[0006] On the one hand, embodiments of the present application provide a keyword detection method, and the method includes:

[0007] Obtain a speech signal;

[0008] Call an end-to-end speech recognition model to recognize the acoustic features of the speech signal, and obtain first speech recognition results respectively corresponding to each frame of speech segment in the speech signal. The first speech recognition result corresponding to one frame of speech segment includes the matching probabilities of the one frame of speech segment with each candidate speech unit respectively, and the candidate speech units are speech units set for keyword detection;

[0009] Based on the first speech recognition results respectively corresponding to each frame of speech segment, obtain a first index structure, where the first index structure is an index structure corresponding to the speech signal for keyword detection;

[0010] Based on the first index structure, obtain the keyword detection result of the speech signal.

[0011] On the other hand, a keyword detection device is provided, and the device includes:

[0012] A first acquisition unit, configured to acquire a speech signal;

[0013] An identification unit, configured to call an end-to-end speech recognition model to recognize the acoustic features of the speech signal, and obtain first speech recognition results respectively corresponding to each frame of speech segment in the speech signal. The first speech recognition result corresponding to one frame of speech segment includes the matching probabilities of the one frame of speech segment with each candidate speech unit respectively, and the candidate speech units are speech units set for keyword detection;

[0014] A second acquisition unit, configured to obtain a first index structure based on the first speech recognition results respectively corresponding to each frame of speech segment, where the first index structure is an index structure corresponding to the speech signal for keyword detection;

[0015] A third acquisition unit, configured to obtain the keyword detection result of the speech signal based on the first index structure.

[0016] In a possible implementation manner, the second acquisition unit includes:

[0017] An acquisition subunit, configured to obtain second speech recognition results respectively corresponding to each frame of speech segment based on the first speech recognition results respectively corresponding to each frame of speech segment. The second speech recognition result corresponding to one frame of speech segment includes the matching probability of the one frame of speech segment with the target speech unit corresponding to the one frame of speech segment, and the target speech unit corresponding to the one frame of speech segment is the candidate speech unit that meets the selection condition among the candidate speech units;

[0018] A conversion subunit, configured to convert the second speech recognition results respectively corresponding to the respective frame speech segments into a second index structure, where the second index structure is composed of nodes, edges between the nodes, and information on the edges, and the information on the edges includes an input label, an output label, and a probability weight;

[0019] A transformation subunit, configured to transform the second index structure to obtain the first index structure, where the speech unit sequence corresponding to the first index structure includes a subsequence that satisfies a reference condition of the speech unit sequence corresponding to the second index structure.

[0020] In a possible implementation manner, that the one frame of speech segment satisfies a selection condition means that the matching probability with the one frame of speech segment is not lower than a probability threshold; the obtaining subunit is configured to, for a first speech segment, use, as the target speech unit corresponding to the first speech segment, the candidate speech units among the respective candidate speech units whose matching probability with the first speech segment is not lower than the probability threshold, where the first speech segment is any one of the respective frame speech segments; extract, from the first speech recognition result corresponding to the first speech segment, the matching probability between the first speech segment and the target speech unit corresponding to the first speech segment, and use the matching probability between the first speech segment and the target speech unit corresponding to the first speech segment to form the second speech recognition result corresponding to the first speech segment.

[0021] In a possible implementation manner, the nodes in the second index structure correspond one-to-one to the reference time points in the speech signal; the transformation subunit is configured to add a virtual start node and a virtual end node to the second index structure; construct an edge pointing to a first associated node between the virtual start node and the first associated node, and construct an edge pointing to the virtual end node between a second associated node and the virtual end node, where the first associated node is a node in the second index structure that satisfies a first condition, and the second associated node is a node in the second index structure that satisfies a second condition; add a time weight to the edge whose pointed node is a node in the second index structure to obtain the first index structure, and the time weight on one edge is used to indicate the reference time point corresponding to the node pointed to by the one edge.

[0022] In a possible implementation, the nodes in the second index structure correspond one-to-one with the reference time points in the speech signal, and any two adjacent reference time points in the speech signal are used to identify a frame of speech segment; the conversion subunit is configured to use the target speech unit corresponding to the second speech segment as the input label and the output label on the edge between the first node and the second node, where the first node and the second node are any two adjacent nodes in the second index structure, and the second speech segment is the speech segment identified by the two reference time points corresponding to the first node and the second node; based on the matching probability between the second speech segment and the target speech unit corresponding to the second speech segment, determine the probability weight on the edge between the first node and the second node.

[0023] In a possible implementation, the conversion subunit is further configured to calculate the negative logarithm of the matching probability between the second speech segment and the target speech unit corresponding to the second speech segment, and use the negative logarithm as the probability weight on the edge between the first node and the second node.

[0024] In a possible implementation, the third obtaining unit is configured to convert the candidate keyword into a third index structure, where the third index structure is an index structure of the same form as the first index structure; perform a combination operation on the first index structure and the third index structure to obtain a fourth index structure, where the speech unit sequence corresponding to the fourth index structure hits the candidate keyword; based on the fourth index structure, obtain the keyword detection result of the speech signal.

[0025] In a possible implementation, the keyword detection result of the speech signal includes the hit information of the speech signal for the candidate keyword, and the third obtaining unit is further configured to obtain the weight information of the speech unit sequence corresponding to the fourth index structure; based on the weight information, obtain the hit information of the speech signal for the candidate keyword.

[0026] In a possible implementation, the third obtaining unit is configured to traverse the first index structure based on the candidate keyword to obtain the keyword detection result of the speech signal.

[0027] On the other hand, a computer device is provided, where the computer device includes a processor and a memory, and at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to enable the computer device to implement the keyword detection method described in any one of the above.

[0028] On the other hand, a computer-readable storage medium is also provided. At least one computer program is stored in the computer-readable storage medium. The at least one computer program is loaded and executed by a processor to enable a computer to implement the keyword detection method described in any one of the above.

[0029] On the other hand, a computer program product or a computer program is also provided. The computer program product or the computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the keyword detection method described in any one of the above.

[0030] The technical solutions provided in the embodiments of the present application at least bring the following beneficial effects:

[0031] In the embodiments of the present application, in the case of calling an end-to-end speech recognition model to recognize acoustic features, a first index structure is directly obtained according to the recognized result (that is, the first speech recognition result corresponding to each frame of speech segment). The first index structure can more comprehensively reflect the recognized result. Keyword detection based on the first index structure can more comprehensively detect the keywords existing in the speech signal, thereby improving the recall rate of the obtained keyword detection result, and the effect of keyword detection is better. Description of the Drawings

[0032] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0033] Figure 1 is a schematic diagram of a WFST provided by an embodiment of the present application;

[0034] Figure 2 is a schematic diagram of an implementation environment of a keyword detection method provided by an embodiment of the present application;

[0035] Figure 3 is a flowchart of a keyword detection method provided by an embodiment of the present application;

[0036] Figure 4 is a flowchart of a process for obtaining a first index structure based on the first speech recognition result corresponding to each frame of speech segment;

[0037] Figure 5It is a schematic diagram of a second index structure provided by an embodiment of the present application;

[0038] Figure 6 It is a schematic diagram of a first index structure provided by an embodiment of the present application;

[0039] Figure 7 It is a schematic diagram of another first index structure provided by an embodiment of the present application;

[0040] Figure 8 It is a schematic diagram of a third index structure provided by an embodiment of the present application;

[0041] Figure 9 It is a schematic diagram of a fourth index structure provided by an embodiment of the present application;

[0042] Figure 10 It is a schematic diagram of another fourth index structure provided by an embodiment of the present application;

[0043] Figure 11 It is a schematic diagram of a keyword detection process provided by an embodiment of the present application;

[0044] Figure 12 It is a schematic diagram of a keyword detection device provided by an embodiment of the present application;

[0045] Figure 13 It is a schematic diagram of the structure of a second acquisition unit provided by an embodiment of the present application;

[0046] Figure 14 It is a schematic diagram of the structure of a server provided by an embodiment of the present application;

[0047] Figure 15 It is a schematic diagram of the structure of a terminal provided by an embodiment of the present application. Detailed implementation manners

[0048] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0049] To facilitate understanding of the technical process of the embodiments of the present application, the following explains the nouns involved in the embodiments of the present application.

[0050] Weighted Finite State Transducer (WFST): WFST is a member of the Finite Automaton (FA) family. FA consists of five elements: (A, Q, E, I, F). Among them, Q is the set of states, representing the nodes in the graph; A is the set of Labels, representing the symbols on the edges; E is the set of transition functions. Two state nodes, the edge between them, and the label and weight on the edge constitute a Transition; I is the initial state, represented by a thicker circle in the graph, which is the starting point of the search; F is the final state, represented by a double-ring circle in the graph, which is the ending point of the search. The WFST structure is often used in speech recognition. For ease of understanding, combined with Figure 1 the following example to introduce WFST.

[0051] A typical WFST contains a starting node (i.e., the node identified by the number 0 in Figure 1 ) and a terminating node (i.e., the node identified by the number 3 in Figure 1 ). The nodes are connected by directed edges. Each edge is composed of (starting node, terminating node, input label: output label / weight). For example, Figure 1 there is an edge pointing to the node identified by the number 2 between the node identified by the number 1 and the node identified by the number 2 in

[0052] The input label on this edge is b, the output label is u, and the weight is 0.34.

[0052] Factor Automata (FA): Given two strings u and v, if u = xvy (x / y is any string), then v is called a factor (substring) of u. The factor automaton F(u) is the smallest deterministic finite-state acceptor (FSA) that recognizes all factors of u, and it can be considered to contain all substrings of the string u. F(A) can be extended and defined as the factor automaton that recognizes all strings (sub-paths) in the WFST A. The Factor Transducer (FT) mentioned in the embodiments of this application is a specific form of FA.

[0053] TFT (Timed Factor Transducer): A factor transducer with time information.

[0054] FTFT (Fast Timed Factor Transducer): A fast factor transducer with time information. FTFT can be regarded as a special WFST.

[0055] In an exemplary embodiment, the keyword detection method provided by the embodiments of the present application can be applied to the field of artificial intelligence technology. Next, the artificial intelligence technology will be introduced.

[0056] Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.

[0057] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation. The keyword detection method provided by the embodiments of the present application involves speech technology and machine learning and other technologies in artificial intelligence technology, and is specifically described through Figure 3 the embodiments shown.

[0058] The key technologies of speech technology include automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling the computer to listen, see, speak, and feel is the future development direction of human-computer interaction, and speech has become one of the most promising human-computer interaction methods in the future.

[0059] Machine learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0060] With the research and progress of artificial intelligence technology, artificial intelligence technology is being studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, vehicle networking, autonomous driving, intelligent transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0061] In an exemplary embodiment, the keyword detection method provided in the embodiments of the present application is implemented in a blockchain system. That is to say, the computer device used to execute the keyword detection method provided in the embodiments of the present application is a node device in the blockchain system. Based on this, the voice signal, the keyword detection result of the voice signal, etc. involved in the keyword detection method provided in the embodiments of the present application can be stored on the blockchain in the blockchain system for each node device in the blockchain system to use, thereby ensuring the security and reliability of the data.

[0062] Figure 2 A schematic diagram of the implementation environment of the keyword detection method provided in the embodiments of the present application is shown. The implementation environment may include: a terminal 11 and a server 12.

[0063] The keyword detection method provided in the embodiments of the present application may be executed by the terminal 11, or may be executed by the server 12, or may be jointly executed by the terminal 11 and the server 12. The embodiments of the present application do not limit this. For the case where the keyword detection method provided in the embodiments of the present application is jointly executed by the terminal 11 and the server 12, the server 12 undertakes the main computing work, and the terminal 11 undertakes the secondary computing work; or, the server 12 undertakes the secondary computing work, and the terminal 11 undertakes the main computing work; or, a distributed computing architecture is adopted between the server 12 and the terminal 11 for collaborative computing.

[0064] Taking the keyword detection method executed by the terminal 11 as an example, the terminal 11 obtains a voice signal by collecting the voice emitted by the object, then extracts the acoustic features of the voice signal, and obtains the first speech recognition result corresponding to each frame of the voice segment in the voice signal by recognizing the acoustic features of the voice signal; furthermore, based on the first speech recognition result corresponding to each frame of the voice segment, an index structure for keyword detection corresponding to the voice signal is obtained, that is, the first index structure; furthermore, keyword detection is automatically implemented based on the first index structure to obtain the keyword detection result of the voice signal. After obtaining the keyword detection result of the voice signal, the terminal 11 can determine whether there are certain keywords in the voice signal. If it is determined that there are certain keywords, it can further wake up certain clients installed in the terminal 11 or control certain clients installed in the terminal 11 to perform certain operations according to the keywords detected from the voice signal.

[0065] In a possible implementation, the terminal 11 can be any electronic product that can perform human-computer interaction with the user in one or more ways such as a keyboard, a touchpad, a touch screen, a remote control, voice interaction, or a handwriting device. For example, a PC (Personal Computer), a mobile phone, a smart phone, a PDA (Personal Digital Assistant), a wearable device, a PPC (Pocket PC), a tablet computer, a smart car machine, a smart TV, a smart speaker, etc. The server 12 can be a single server, a server cluster composed of multiple servers, or a cloud computing service center. The terminal 11 and the server 12 establish a communication connection through a wired or wireless network.

[0066] Those skilled in the art should understand that the above-mentioned terminal 11 and server 12 are only examples. Other existing or future possible terminals or servers that can be applied to this application should also be included within the protection scope of this application and are hereby incorporated by reference.

[0067] Based on the above Figure 2 shown implementation environment, the embodiments of this application provide a keyword detection method. This keyword detection method is executed by a computer device, which can be the terminal 11 or the server 12. The embodiments of this application do not limit this. As Figure 3 shown, the keyword detection method provided by the embodiments of this application includes the following steps 301 to step 304.

[0068] In step 301, a voice signal is obtained.

[0069] Keyword detection is performed on the voice signal. Therefore, before keyword detection, it is necessary to first obtain the voice signal, and then perform keyword detection on the voice signal. The embodiments of this application do not limit the manner of obtaining the voice signal. Exemplarily, the voice signal is collected through the microphone of a computer device (such as a mobile phone or a laptop); or, the voice signal is downloaded from a network database using a wired or wireless communication method; or, the voice signal is obtained by accessing a local database. The embodiments of this application do not limit the duration of the voice signal, which can be a voice signal with a longer duration or a voice signal with a shorter duration. Exemplarily, the duration of the voice signal is 10 seconds.

[0070] Keyword detection technology is widely used in fields such as intelligent device wake-up interaction and voice keyword detection of audio and video files. In order to cope with the limited computing resources of computer devices and the processing requirements of a large number of audio and video files, it is usually desired that the detection speed can be as fast as possible, and on this basis, it is further desired that the detection result is as accurate as possible.

[0071] Taking a computer device as an example of a terminal, a specific application of keyword detection for a voice signal is illustrated. Exemplarily, the voice signal is collected by the terminal, for example, it can be collected by an instant messaging software installed on the terminal during use. By performing keyword detection on the voice signal, it can be determined whether there is malicious voice content in the voice signal. If there is malicious voice content, it can be fed back to the application server corresponding to the instant messaging software for processing, such as punishing the terminal, performing silencing processing on the malicious voice content, etc. Also, by performing keyword detection on the voice signal, it can be determined whether there is a wake-up keyword in the voice signal. If there is, the terminal is woken up and the terminal is controlled to execute corresponding voice commands, such as reporting the weather forecast, playing multimedia, replying to messages, making a phone call, etc.

[0072] In an exemplary embodiment, after obtaining the voice signal, the voice signal is framed to obtain each frame of voice segment in the voice signal. Exemplarily, each frame of voice segment has its corresponding start and end positions and indexes, etc. Exemplarily, the start and end positions corresponding to a frame of voice segment are determined by two reference time points in the voice signal, and these two reference time points are respectively referred to as the start time point and the end time point of this frame of voice segment. Exemplarily, for two adjacent frames of voice segments in the voice signal, the end time point of the previous frame of voice segment is the same as the start time point of the next frame of voice segment.

[0073] Exemplarily, each time point used for framing is recorded as a reference time point in the voice signal, and a frame of voice segment is identified by two adjacent reference time points in the voice signal. That is to say, any two adjacent reference time points in the voice signal are used to identify a frame of voice segment.

[0074] The duration of a frame of voice segment in the embodiments of the present application is not limited, which is related to the way of framing the voice signal. Exemplarily, the way of framing the voice signal is: dividing a 1-second voice signal into 100 equal frames, then the duration of a frame of voice segment is 10 milliseconds.

[0075] In an exemplary embodiment, after obtaining the voice signal, the acoustic features of the voice signal are extracted to facilitate subsequent direct recognition of the acoustic features of the voice signal. Exemplarily, the acoustic features of the voice signal are extracted by analyzing each frame of voice segment in the voice signal.

[0076] The acoustic features of a speech signal are used to characterize the speech signal from an acoustic perspective. The process of extracting acoustic features is a process of converting the speech signal in the time domain into a more compact signal representation form such as a frequency domain signal. The embodiments of the present application do not limit the manner of extracting acoustic features, and any acoustic feature extraction technology can be used to extract the acoustic features of the speech signal. Exemplarily, one or more features of MFCC (Mel Frequency Cepstrum Coefficient), LPC (Linear Predictive Coefficient), LPCC (Linear Predictive Cepstrum Coefficient), PLPC (Perceptual Linear Predictive Coefficient), or FBank (Filter Bank) of the speech signal are extracted, and the extracted features are used as the acoustic features of the speech signal. Exemplarily, a deep learning network is used to extract the acoustic features of the speech signal.

[0077] In step 302, the end-to-end speech recognition model is called to recognize the acoustic features of the speech signal, and the first speech recognition results corresponding to each frame of speech segment in the speech signal are obtained. The first speech recognition result corresponding to one frame of speech segment includes the matching probabilities of one frame of speech segment with each candidate speech unit respectively.

[0078] After obtaining the acoustic features of the speech signal, the end-to-end speech recognition model is called to recognize the acoustic features of the speech signal, and the first speech recognition results corresponding to each frame of speech segment in the speech signal are obtained. Exemplarily, the end-to-end speech recognition model is trained based on the CTC algorithm. The first speech recognition result corresponding to one frame of speech segment includes the matching probabilities of this frame of speech segment with each candidate speech unit respectively, and the candidate speech units are speech units set for keyword detection.

[0079] Exemplarily, the sum of the matching probabilities of one frame of speech segment with each candidate speech unit is 1. The candidate speech units are speech units preset for keyword detection. Exemplarily, the candidate speech units are the basic units relied on in the process of keyword detection. The embodiments of the present application do not limit the type of candidate speech units, and can be flexibly set according to requirements. Exemplarily, the type of candidate speech units is syllables, single characters, or words, etc. The embodiments of the present application take the type of candidate speech units as syllables (such as jia1, wei1, ge4, etc.) as an example for illustration.

[0080] For each frame of speech segment, all candidate speech units are the same. Exemplarily, in addition to normal speech units, all candidate speech units also include special speech units. Exemplarily, in an application scenario where an end-to-end speech recognition model trained based on the CTC algorithm is invoked, the candidate speech units include a special speech unit, namely, the blank character. The embodiments of the present application do not limit the total number of all candidate speech units, which can be flexibly adjusted according to actual situations. For example, if the candidate speech units include N (N is an integer not less than 1) normal speech units and one special speech unit, the total number of all candidate speech units is (N + 1).

[0081] The first speech recognition result corresponding to a frame of speech segment includes the matching probabilities of this frame of speech segment with all candidate speech units respectively. That is to say, according to the first speech recognition result corresponding to a frame of speech segment, it can be known what the matching probabilities of this frame of speech segment with all candidate speech units are respectively. The greater the matching probability of a frame of speech segment with a candidate speech unit, the more likely it is that this frame of speech segment corresponds to this speech unit. It should be noted that a frame of speech segment corresponding to a speech unit means that this frame of speech segment is obtained by pronouncing this speech unit.

[0082] Exemplarily, if the total number of all candidate speech units is (N + 1), the first speech result corresponding to a frame of speech segment includes (N + 1) probabilities, and these (N + 1) probabilities refer to the matching probabilities of this frame of speech segment with these (N + 1) candidate speech units respectively. Assuming that the speech signal includes T (T is an integer not less than 1) frames of speech segments, the first speech recognition results respectively corresponding to each frame of speech segment can form a probability matrix of T * (N + 1) dimensions.

[0083] The first speech recognition results respectively corresponding to each frame of speech segment in the speech signal are obtained by invoking the end-to-end speech recognition model to recognize the acoustic features of the speech signal. That is to say, by inputting the acoustic features of the speech signal into the end-to-end speech recognition model, the first speech recognition results respectively corresponding to each frame of speech segment in the speech signal output by the end-to-end speech recognition model can be directly obtained. Exemplarily, the end-to-end speech recognition model is trained based on the Connectionist Temporal Classification (CTC) algorithm.

[0084] As an end-to-end training criterion, the CTC algorithm can quickly build a state-of-the-art end-to-end speech recognition model. Compared with the traditional NN-HMM (Neural Network Hidden Markov Model) Hybrid speech recognition model, the end-to-end speech recognition model trained based on the CTC algorithm can achieve obvious superiority in both recognition effect and inference speed, and has increasingly become a mainstream speech recognition model in speech recognition systems.

[0085] The CTC algorithm finds the alignment relationship between all sequences X(x1,..., xm) and Y(y1,..., yn) through a special blank symbol, where m is not less than n. When processing using the model trained based on the CTC algorithm, for the case where the number of normal speech units is N, the number of candidate speech units related to the true output probability is (N + 1). When recognizing the acoustic features of a speech signal, the output of the end-to-end speech recognition model can obtain a probability matrix of T*(N + 1), where T is the number of frames of speech segments in the speech signal.

[0086] In step 303, based on the first speech recognition results corresponding to each frame of speech segment, a first index structure is obtained, and the first index structure is an index structure corresponding to the speech signal for keyword detection.

[0087] After obtaining the first speech recognition results corresponding to each frame of speech segment, an index structure corresponding to the speech signal for keyword detection is further obtained, that is, the first index structure is obtained, and the first index structure can represent the speech signal. Exemplarily, the first index structure is an index structure in the form of WFST.

[0088] In a possible implementation manner, refer to Figure 4 , the process of obtaining the first index structure based on the first speech recognition results corresponding to each frame of speech segment includes the following steps 3031 to 3033:

[0089] Step 3031: Based on the first speech recognition results corresponding to each frame of speech segment, obtain the second speech recognition results corresponding to each frame of speech segment. The second speech recognition result corresponding to a frame of speech segment includes the matching probability between a frame of speech segment and the target speech unit corresponding to the frame of speech segment, and the target speech unit corresponding to a frame of speech segment is the candidate speech unit that satisfies the selection condition with the frame of speech segment among all candidate speech units.

[0090] The second speech recognition result corresponding to each frame of speech segment is obtained based on the first speech recognition result corresponding to its respective frame of speech segment. The principle of obtaining the second speech recognition results corresponding to different frames of speech segments is the same. In the embodiments of the present application, any one frame of speech segment (referred to as the first speech segment) in each frame of speech segment is taken as an example for illustration.

[0091] Since the second speech recognition result corresponding to the first speech segment includes the matching probability between the first speech segment and the target speech unit corresponding to the first speech segment, in the process of obtaining the second speech recognition result corresponding to the first speech segment, it is necessary to first determine the target speech unit corresponding to the first speech segment. The target speech unit corresponding to the first speech segment is the candidate speech unit that satisfies the selection condition with the first speech segment among each candidate speech unit. That is to say, which one or which of the candidate speech units among each candidate speech unit is the target speech unit corresponding to the first speech segment is related to the selection condition.

[0092] Exemplarily, it is considered that all candidate speech units satisfy the selection condition with the first speech segment. In this case, the target speech unit corresponding to the first speech segment is all candidate speech units. Based on this, the method for obtaining the second speech recognition result corresponding to the first speech segment based on the first speech recognition result corresponding to the first speech segment is: directly using the first speech recognition result corresponding to the first speech segment as the second speech recognition result corresponding to the first speech segment. In this way, the second speech recognition result is the same as the first speech recognition result, and the subsequent obtained index structure is all constructed based on the first speech recognition result obtained by recognizing the acoustic features of the speech signal, which is beneficial to ensuring the comprehensiveness and reliability of the finally obtained keyword detection result.

[0093] Exemplarily, satisfying the selection condition with a frame of speech segment means that the matching probability with the frame of speech segment is not lower than the probability threshold. In this case, the process of obtaining the second speech recognition result corresponding to the first speech segment based on the first speech recognition result corresponding to the first speech segment is: taking the candidate speech units whose matching probability with the first speech segment is not lower than the probability threshold among each candidate speech unit as the target speech units corresponding to the first speech segment; extracting the matching probability between the first speech segment and the target speech unit corresponding to the first speech segment from the first speech recognition result corresponding to the first speech segment, and forming the second speech recognition result corresponding to the first speech segment by the matching probability between the first speech segment and the target speech unit corresponding to the first speech segment. The probability threshold is set according to experience or flexibly adjusted according to the application scenario. The embodiments of the present application do not limit this, for example, the probability threshold is 50%.

[0094] In this way, candidate speech units with a relatively low matching probability with the first speech segment can be eliminated, so that the efficiency of obtaining the index structure based on the second speech recognition result is relatively high, and the obtained index structure can implement keyword detection more quickly.

[0095] According to the method of obtaining the second speech recognition result corresponding to the first speech segment, the second speech recognition results corresponding to each frame of speech segment can be obtained, and then step 3032 is executed.

[0096] Step 3032: Convert the second speech recognition results corresponding to each frame of speech segment into a second index structure. The second index structure is composed of nodes, edges between the nodes, and information on the edges. The information on the edges includes an input label, an output label, and a probability weight.

[0097] The second index structure is determined based on the second speech recognition results corresponding to all speech segments. The second index structure is composed of nodes, edges between the nodes, and information on the edges. That is to say, the process of converting the second speech recognition results corresponding to each frame of speech segment into the second index structure is realized by determining the nodes in the second index structure according to the second speech recognition results corresponding to each frame of speech segment, constructing edges between the nodes, and determining the information on the edges. Exemplarily, the second index structure is an index structure in the form of WFST, and the second index structure is a structured representation of the second speech recognition results corresponding to each frame of speech segment.

[0098] The nodes in the second index structure correspond one-to-one with the reference time points in the speech signal. That is to say, a node in the second index structure corresponds to a reference time point in the speech signal. Therefore, by determining the number of reference time points in the speech signal, the number of nodes in the second index structure can be determined. The order of arrangement of the nodes in the second index structure is the same as the order of the reference time points corresponding to the nodes. Based on this, the nodes in the second index structure can be determined. Exemplarily, the reference time points in the speech signal mentioned in the embodiments of the present application refer to the time points used to implement frame division in the speech signal.

[0099] After determining the nodes in the second index structure, it is necessary to further construct edges between the nodes in the second index structure. Exemplarily, the nodes in the second index structure correspond one-to-one with the reference time points in the speech signal, and any two adjacent reference time points in the speech signal are used to identify a frame of speech segment. In this case, the method of constructing edges between the nodes in the second index structure is as follows: If the speech segments identified by the two reference time points corresponding to two adjacent nodes correspond to M (M is an integer not less than 1) target speech units, then M edges pointing from the previous node to the next node among the two adjacent nodes are constructed between the two adjacent nodes.

[0100] After constructing the edges, it is necessary to further determine the information on the edges. The information on the edges includes input labels, output labels, and probability weights. The information on the edges between any two adjacent nodes in the second index structure is determined based on the second speech recognition results corresponding to the speech segments identified by the two reference time points corresponding to the two adjacent nodes.

[0101] Exemplarily, the second index structure includes a first node and a second node, and the first node and the second node are any two adjacent nodes in the second index structure. Taking the first node and the second node as an example, the process of determining the information on the edge between the first node and the second node is described. Exemplarily, the speech segments identified by the two reference time points corresponding to the first node and the second node are referred to as the second speech segments, and the information on the edge between the first node and the second node is determined based on the second speech recognition results corresponding to the second speech segments.

[0102] The second speech recognition results corresponding to the second speech segments include the matching probabilities between the second speech segments and the target speech units corresponding to the second speech segments. Based on the second speech recognition results corresponding to the second speech segments, the process of determining the information on the edge between the first node and the second node is as follows: taking the target speech units corresponding to the second speech segments as the input labels and output labels on the edge between the first node and the second node in the second index structure; determining the probability weight on the edge between the first node and the second node based on the matching probabilities between the second speech segments and the target speech units corresponding to the second speech segments.

[0103] It should be noted that for the case where the number of target speech units corresponding to the second speech segments is multiple, the edges between the first node and the second node correspond one-to-one with the target speech units corresponding to the second speech segments. The input label and output label on an edge are the target speech unit corresponding to the edge, and the probability weight on an edge is determined based on the matching probability between the second speech segment and the target speech unit corresponding to the edge.

[0104] Taking the target speech units corresponding to the second speech segments as an example, the manner of determining the probability weight on the edge between the first node and the second node based on the matching probabilities between the second speech segments and the target speech units corresponding to the second speech segments can be flexibly selected according to requirements, and the embodiments of the present application do not limit this. In an exemplary embodiment, the matching probability between the second speech segment and the target speech unit corresponding to the second speech segment is directly used as the probability weight on the edge between the first node and the second node.

[0105] In an exemplary embodiment, calculate the negative logarithm of the matching probability between the second speech segment and the target speech unit corresponding to the second speech segment, and use the negative logarithm as the probability weight on the edge between the first node and the second node. For example, assume that the matching probability between the second speech segment and a target speech unit corresponding to the second speech segment is a, then use -loga as the probability weight on the edge corresponding to this target speech unit between the first node and the second node.

[0106] According to the manner of determining the information on the edge between the first node and the second node, the information on each edge between every two adjacent nodes can be determined. After determining the nodes, constructing edges between any two adjacent nodes, and determining the information on each edge, a second index structure is obtained.

[0107] Since the CTC algorithm is a discriminative training criterion, the distribution of the matching probabilities between the output speech segments and the candidate speech units is usually relatively sharp, that is, the matching probabilities are usually concentrated on only a few candidate speech units. Therefore, under the condition that a frame of speech segment meets the selection condition, which means that the matching probability with this frame of speech segment is not lower than the probability threshold, the target speech unit selected for each frame of speech segment is a candidate speech unit with a matching probability higher than the probability threshold output by the end-to-end speech recognition model trained based on the CTC algorithm. In this way, for the case where the number of candidate speech units is (N + 1), the target speech unit corresponding to each frame of speech segment is less than (N + 1), so that the second index structure can be constructed quickly, reducing the size of the second index structure and the subsequent detection complexity. Exemplarily, the process of constructing the second index structure can be achieved by traversing the second speech recognition results corresponding to each frame of speech segment once.

[0108] Exemplarily, the second index structure is as Figure 5 shown. Figure 5 In the second index structure shown, there are 6 nodes, which are respectively identified by the numbers 0 - 5. This indicates that there are 6 reference time points in the speech signal, and one node corresponds to one reference time point. These 6 reference time points can identify 5 frames of speech segments. That is to say, the two reference time points corresponding to the two nodes identified by the numbers 0 and 1 identify the first frame of speech segment, the two reference time points corresponding to the two nodes identified by the numbers 1 and 2 identify the second frame of speech segment, the two reference time points corresponding to the two nodes identified by the numbers 2 and 3 identify the third frame of speech segment, the two reference time points corresponding to the two nodes identified by the numbers 3 and 4 identify the fourth frame of speech segment, and the two reference time points corresponding to the two nodes identified by the numbers 4 and 5 identify the fifth frame of speech segment.

[0109] Furthermore, according to Figure 5It can be seen that the target speech units corresponding to the first-frame speech segment are respectively <unk>Harmony jia1, based on the first-frame voice segment and the target voice unit <unk>The probability weight determined by the matching probability is 4.3871, and the probability weight determined by the matching probability between the first-frame speech segment and the target speech unit jia1 is 0.013275. Similarly, the target speech units corresponding to the second-frame speech segments are respectively <blank>, wei1 and ge4, based on the second-frame speech segment and the target speech unit <blank>The probability weight determined based on the matching probability is 0.073067. The probability weight determined based on the matching probability between the second speech segment and the target speech unit "wei1" is 4.5336. The probability weight determined based on the matching probability between the second speech segment and the target speech unit "ge4" is 2.9268.

[0110] The target speech unit corresponding to the third speech segment is "wei1", and the probability weight determined based on the matching probability between the third speech segment and the target speech unit "wei1" is 0.00055655. The target speech units corresponding to the fourth speech segment are respectively <blank>Sum to 1, based on the 4th frame of the speech segment and the target speech unit <blank>The probability weight determined based on the matching probability of the 4th frame of the speech segment and the target speech unit wei1 is 0.034106. The probability weight determined based on the matching probability of the 5th frame of the speech segment and the target speech unit xin4 is 3.4059. The target speech unit corresponding to the 5th frame of the speech segment is xin4, and the probability weight determined based on the matching probability of the 5th frame of the speech segment and the target speech unit xin4 is 0.00024697.

[0111] Step 3033: Transform the second index structure to obtain the first index structure. The speech unit sequence corresponding to the first index structure includes the subsequence of the speech unit sequence corresponding to the second index structure that satisfies the reference condition.

[0112] After obtaining the second index structure, transform the second index structure to obtain the index structure for keyword detection corresponding to the speech signal, that is, the first index structure. The speech unit sequence corresponding to the first index structure includes the subsequence of the speech unit sequence corresponding to the second index structure that satisfies the reference condition. The speech unit sequence corresponding to the index structure refers to the sequence of speech units on the complete path from the first node to the last node in the index structure. A speech unit sequence corresponds to a complete path from the first node to the last node, and a complete path is composed of multiple edges.

[0113] Exemplarily, for the first index structure and the second index structure, if the input label and the output label on an edge are the same, then the speech unit sequence corresponding to a complete path can refer to the sequence of input labels on each edge in a complete path, or can refer to the sequence of output labels on each edge in a complete path.

[0114] For example, taking the second index structure as Figure 5 shown as an example, from Figure 5 the first node (i.e., the node marked with the number 0) to the last node (i.e., the node marked with the number 5) in the second index structure shown, there are a total of 12 complete paths. Therefore, Figure 5 the second index structure shown corresponds to 12 speech unit sequences. Exemplarily, Figure 5 if the probability weight on the edge in the second index structure shown is the negative logarithm of the matching probability, then the smaller the probability weight, the better. The matching probability corresponding to a speech unit sequence can be obtained based on the probability weights of all the edges in the complete path corresponding to the speech unit sequence. Since the probability weight on the edge is the negative logarithm of the matching probability, the speech unit sequence with the maximum matching probability is "jia1 <blank>wei1 <blank>"xin4", wherein <blank>Represents empty. Figure 5 The original speech corresponding to the second index structure shown is a short sentence with relatively clear pronunciation and accurate recognition. Starting from Figure 5 it can be seen that the candidate keyword "jia1#wei1#xin4" to be detected exists in this second index structure.

[0115] A subsequence of a speech unit sequence that meets the reference conditions refers to a sequence of speech units on a sub-path that meets the reference conditions in the complete path corresponding to a speech unit sequence. Which sub-path or sub-paths are sub-paths that meet the reference conditions is set according to experience or flexibly adjusted according to the application scenario, and the embodiments of the present application do not limit this.

[0116] Exemplarily, a sub-path that meets the reference conditions refers to a path formed by at least one edge in the complete path. That is to say, in a sub-path that meets the reference conditions, the smallest sub-path is a path formed by one edge in the complete path, and the largest sub-path is the complete path. Based on this, a subsequence that meets the reference conditions refers to a sequence formed by at least one speech unit that makes up a speech unit sequence. That is to say, in a subsequence that meets the reference conditions, the smallest subsequence is one speech unit that makes up a speech unit sequence, and the largest subsequence is the entire speech unit sequence. In this case, the speech unit sequence corresponding to the first index structure includes all non-empty subsequences of the speech unit sequence corresponding to the second index structure.

[0117] The speech unit sequence corresponding to the first index structure includes the subsequence of the speech unit sequence corresponding to the second index structure that meets the reference conditions, so as to be able to quickly implement keyword detection based on the first index structure.

[0118] Exemplarily, if keyword detection requires not only detecting whether a certain keyword or certain keywords exist in the speech signal, but also locating the specific position of the detected keyword in the speech signal, then the first index structure needs to carry time information to use the time information to determine the position of the detected keyword in the speech signal; if keyword detection only needs to detect whether a certain keyword or certain keywords exist in the speech signal and does not need to locate the specific position of the detected keyword in the speech signal, then the first index structure does not need to carry time information.

[0119] In a possible implementation manner, for the case where the first index structure needs to carry time information, the process of transforming the second index structure to obtain the first index structure includes the following steps 1 to 3:

[0120] Step 1: Add a virtual start node and a virtual end node to the second index structure.

[0121] The virtual start node is added before the first node in the second index structure, and the virtual end node is added after the last node in the second index structure. That is to say, there are two more nodes in the first index structure than in the second index structure.

[0122] Exemplarily, the original nodes in the second index structure are identified with original identifiers in the second index structure. Since the nodes in the second index structure correspond one-to-one with the reference time points and the nodes correspond one-to-one with the original identifiers, there is a corresponding relationship between the original identifiers and the reference time points. It should be noted that the original nodes in the second index structure can be identified with the original identifiers or new identifiers in the first index structure. This application embodiment does not limit this. For the case of being identified with new identifiers, it is necessary to record the corresponding relationship between the new identifiers and the original identifiers, so as to determine the original identifier corresponding to the new identifier according to the corresponding relationship between the new identifier and the original identifier, and further determine the corresponding reference time point according to the original identifier.

[0123] Step 2: Construct an edge pointing to the first associated node between the virtual start node and the first associated node, and construct an edge pointing to the virtual end node between the second associated node and the virtual end node. The first associated node is a node in the second index structure that meets the first condition, and the second associated node is a node in the second index structure that meets the second condition.

[0124] The first associated node refers to the node that needs to be associated with the virtual start node to ensure that the speech unit sequence corresponding to the first index structure includes the subsequence that meets the reference condition of the speech unit sequence corresponding to the second index structure; the second associated node refers to the node that needs to be associated with the virtual end node to ensure that the speech unit sequence corresponding to the first index structure includes the subsequence that meets the reference condition of the speech unit sequence corresponding to the second index structure.

[0125] The first associated node is a node in the second index structure that meets the first condition, and the second associated node is a node in the second index structure that meets the second condition. The setting of the first condition and the second condition is related to the setting of the reference condition. Exemplarily, for the case where the subsequence that meets the reference condition refers to a sequence composed of at least one speech unit of the speech unit sequence, to ensure that the speech unit sequence corresponding to the first index structure includes all non-empty subsequences of the speech unit sequence corresponding to the second index structure, the node that meets the first condition is all other nodes in the second index structure except the last node, and the node that meets the second condition is all other nodes in the second index structure except the first node.

[0126] That is to say, edges are constructed between the virtual start node and all other nodes in the second index structure except the last node, and edges are constructed between all other nodes in the second index structure except the first node and the end node. It should be noted that the edges in the index structure are directed edges. The edge constructed between the virtual start node and the first associated node points to the first associated node, and the edge constructed between the second associated node and the virtual end node points to the virtual end node.

[0127] In an exemplary embodiment, after constructing new edges, it is necessary to supplement the information on the newly constructed edges. Exemplarily, the way to supplement the information on the newly constructed edges is as follows: Use the specified speech unit as the input label and output label on the newly constructed edge; Set the probability weight to be empty or to a reference label. The specified speech unit is set according to experience or flexibly adjusted according to the application scenario. The embodiments of the present application do not limit this, and exemplarily, the specified speech unit is <eps>The reference labels are set according to experience or flexibly adjusted according to the way of obtaining the probability weights on the existing edges. Exemplarily, for the case where the probability weight is the negative logarithm of the matching probability, the reference label can be 0; for the case where the probability weight is the matching probability, the reference label can be 1. Exemplarily, the reference label can also be a special symbol, which is used to indicate that when obtaining the matching probability corresponding to the speech unit sequence subsequently, the probability weight on this edge does not need to be concerned about.

[0128] In an exemplary embodiment, the index structure obtained through the above steps 1 and 2 can be regarded as an FT (Factor Transducer).

[0129] Step 3: Add a time weight to the edge whose pointed node is a node in the second index structure to obtain the first index structure. The time weight on an edge is used to indicate the reference time point corresponding to the node pointed to by the edge.

[0130] In the index structure obtained through steps 1 and 2, there are edges pointing to nodes in the second index structure, and there are also edges pointing to other nodes (such as, a virtual termination node). Since a node in the second index structure corresponds to a reference time point in the speech signal, in order to enable the obtained first index structure to not only obtain the probability information of the existence of keywords in the speech signal, but also obtain the position information of the existing keywords in the speech signal, a time weight is added to the edge whose pointed node is a node in the second index structure. After adding the time weight, the first index structure is obtained. The time weight on an edge is used to indicate the reference time node corresponding to the node pointed to by the edge.

[0131] Exemplarily, the form of the time weight on the edge in the embodiments of the present application is not limited, as long as it can be used to indicate the reference time node corresponding to the node pointed to by the edge. Exemplarily, directly use the reference time point corresponding to the node pointed to by the edge as the time weight on the edge.

[0132] Exemplarily, use the identifier of the node pointed to by the edge in the second index structure as the time weight on the edge. There is a corresponding relationship between the identifier of the node in the second index structure and the reference time point corresponding to the node, so that the corresponding reference time point can be queried according to the time weight in the corresponding relationship between the identifier of the node in the second index structure and the reference time point corresponding to the node.

[0133] Exemplarily, the identifier of the node pointed by the edge in the index structure obtained through Step 1 and Step 2 is used as the time weight on the edge. There is a corresponding relationship between the identifier of the node in the index structure obtained through Step 1 and Step 2 and the identifier of the node in the second index structure. The identifier of the node in the second index structure has a corresponding relationship with the reference time point corresponding to the node. Therefore, according to the time weight, in the corresponding relationship between the identifier of the node in the index structure obtained through Step 1 and Step 2 and the identifier of the node in the second index structure, the identifier of the node corresponding to the time weight in the second index structure can be queried; then, according to the identifier of the node in the second index structure, in the corresponding relationship between the identifier of the node in the second index structure and the reference time point corresponding to the node, the corresponding reference time point can be queried.

[0134] In an exemplary embodiment, for the case where the probability weight on the edge constructed in Step 2 is set to null, while adding the time weight to the edge whose pointed node is a node in the second index structure, if there is no probability weight on the edge, the probability weight needs to be added to avoid misinterpreting the time weight as the probability weight during the subsequent keyword detection process. Exemplarily, for the case where the probability weight on the edge with an existing probability weight is the negative logarithm of the matching probability, the probability weight added to the edge without a probability weight is 0; for the case where the probability weight on the edge with an existing probability weight is the matching probability, the probability weight added to the edge without a probability weight is 1.

[0135] Exemplarily, the time weight is in the form of a string, and the probability weight is in the form of a floating-point value.

[0136] Exemplarily, since the time weight is added in Step 3, the index structure carries time information. Therefore, the first index structure obtained based on Steps 1 to 3 can be called TFT (Timed Factor Transducer, factor transducer with time information).

[0137] Exemplarily, for the case where the target speech unit corresponding to a frame of speech segment is a candidate speech unit whose matching probability with a frame of speech unit is not lower than the probability threshold, since it can accelerate the efficiency of constructing the first index structure, the first index structure obtained based on Steps 1 to 3 can also be called FTFT (Fast Timed Factor Transducer, fast factor transducer with time information). The construction of FTFT directly acts on the second speech recognition result, using the two-tuple weight of (probability weight, time weight) as the co-occurrence weight. In order to obtain the accurate position of the keyword in the speech signal, FTFT encodes the start time of the sub-path in the second index structure on the weight from the FTFT start node to each sub-path.

[0138] Exemplarily, taking the second index structure as Figure 5 Taking the example shown, the first index structure obtained based on the above steps 1 to 3 is as Figure 6 shown. Figure 5 The nodes identified by the numbers 0 to 5 in Figure 6 become the nodes identified by the numbers 2 to 7 in Figure 6 An additional virtual start node (the node identified by the number 0) and a virtual end node (the node identified by the number 1) are added in . Edges are constructed between the virtual start node and the nodes identified by the numbers 2 to 6 (i.e., the first associated nodes), and also between the nodes identified by the numbers 3 to 7 (i.e., the second associated nodes) and the virtual end node. That is to say, the virtual start node can be connected to any one of the entity nodes except the entity node identified by the number 7, and all entity nodes except the entity node identified by the number 2 can jump to the common virtual end node.

[0139] In Figure 6 the time weight added to the edge whose pointed node is a node in the second index structure is the digital identifier of the pointed node in the first index structure. On the edge between the virtual start node and the first associated nodes, in addition to having input labels and output labels <eps>In addition to the time weight, it also has a probability weight of 0. On the edge between the second associated node and the virtual termination node, there are only input labels and output labels <eps>, without probability weight and time weight.

[0140] In a possible implementation, for the case where the first index structure does not need to carry time information, the process of transforming the second index structure to obtain the first index structure is as follows: adding a virtual start node and a virtual end node to the second index structure; constructing an edge pointing to the first associated node between the virtual start node and the first associated node, and constructing an edge pointing to the virtual end node between the second associated node and the virtual end node to obtain the first index structure. The implementation of this process can refer to steps 1 and 2 in the case where the first index structure needs to carry time information as described above, which will not be elaborated here. Exemplarily, taking the second index structure as shown in Figure 5 as an example, the first index structure obtained in the case where the first index structure does not need to carry time information is as shown in Figure 7 shown.

[0141] In step 304, based on the first index structure, obtain the keyword detection result of the speech signal.

[0142] Since the first index structure is the index structure corresponding to the speech signal for keyword detection, after obtaining the first index structure, the keyword detection result of the speech signal can be obtained based on the first index structure. According to the keyword detection result of the speech signal, information such as whether a certain keyword exists in the speech signal, the probability of the existence of a certain keyword, and the position of the existing certain keyword in the speech signal can be known. Exemplarily, the position of a certain keyword in the speech signal is determined based on the reference time point in the speech signal.

[0143] In a possible implementation, the methods for obtaining the keyword detection result of the speech signal based on the first index structure include but are not limited to the following two:

[0144] Method 1: Convert the candidate keyword into a third index structure, where the third index structure and the first index structure are index structures of the same form; perform a combination operation on the first index structure and the third index structure to obtain a fourth index structure; based on the fourth index structure, obtain the keyword detection result of the speech signal.

[0145] A candidate keyword refers to a keyword for which it is necessary to detect whether it exists in the speech signal. The candidate keyword is set in advance and is related to the actual application scenario. Exemplarily, when the actual application scenario is that the user wakes up the navigation client in the computer device with a speech signal containing the wake-up word "open navigation", the candidate keyword is "open navigation". Exemplarily, when the actual application scenario is to use a speech signal containing the specified word "take a photo" to command the camera client in the computer device to take a photo, the candidate keyword is "take a photo".

[0146] The number of candidate keywords can be one or more, and the embodiments of the present application do not limit this. In an exemplary embodiment, for the case where the number of candidate keywords is multiple, all candidate keywords are converted into a third index structure to improve the efficiency of obtaining the fourth index structure through a combination operation. The third index structure and the first index structure are index structures of the same form. Exemplarily, for the case where the first index structure is an index structure in the WFST form, the third index structure is an index structure in the WFST form.

[0147] The manner of converting candidate keywords into the third index structure is related to the type of candidate speech units. Exemplarily, first, the candidate keywords are represented as keyword speech sequences that match the type of candidate speech units; then, the third index structure is constructed according to the possible jump relationships between the speech units in the keyword speech sequences. Each speech unit sequence corresponding to the constructed third index structure is a speech unit sequence corresponding to a candidate keyword. Exemplarily, for the case where the number of candidate keywords is multiple, any speech unit sequence corresponding to the third index structure is a speech unit sequence of a certain candidate keyword. The speech unit sequence of a candidate keyword refers to the speech unit sequence of the keyword speech sequence that can be obtained by refinement to get the candidate keyword.

[0148] Exemplarily, taking the candidate keyword "add WeChat" as an example, for the case where the type of candidate speech unit is syllable, the candidate keyword is represented as the keyword speech sequence "jia1#wei1#xin4", and then the third index structure as shown in Figure 8 is constructed according to the possible jump relationships between the speech units in the keyword speech sequence.

[0149] In Figure 8 the third index structure shown, the nodes are virtual nodes determined according to the possible jump relationships between the speech units in the keyword speech sequence. The information on the edges between the nodes includes input labels and output labels, and the default probability weight is empty. To ensure that each speech unit sequence corresponding to the third index structure is a speech unit sequence corresponding to a candidate keyword, the keyword speech sequence is used as the output label on the edge between the first node and the second node. Exemplarily, the keyword speech sequence can also be used as the output label on the edge between the penultimate node and the last node.

[0150] After obtaining the third index structure, a Compose operation is performed on the first index structure and the third index structure to obtain a fourth index structure. The Compose operation is a standard operation in WFST. Compose(A,B) means taking the output of index structure A as the input of index structure B. Traverse A from the start node to the end node, and its output is used as the input of B to jump on the WFST graph of B starting from the start node until reaching the end node of B. The input and output results that successfully jump to the end node of B are recorded and combined as a new WFST result for output.

[0151] The speech unit sequence corresponding to the fourth index structure obtained after performing the Compose operation on the first index structure and the third index structure hits the candidate keyword. It should be noted that for the case where the number of candidate keywords is multiple, the speech unit sequence corresponding to the fourth index structure hitting the candidate keyword means that each speech unit sequence corresponding to the fourth index structure hits a certain candidate keyword. Exemplarily, the speech unit sequence hitting the candidate keyword means that the keyword speech sequence that can obtain the candidate keyword after the speech unit sequence is streamlined.

[0152] Exemplarily, for Figure 6 the first index structure shown and Figure 8 the third index structure shown, after performing the Compose operation, a fourth index structure as shown in Figure 9 can be obtained. For Figure 7 the first index structure shown and Figure 8 the third index structure shown, after performing the Compose operation, a fourth index structure as shown in Figure 10 can be obtained. Figure 9 In the fourth index structure shown, in addition to the input label, output label, and probability weight, the edge also includes a time weight. Figure 10 In the fourth index structure shown, the edge does not include a time weight.

[0153] Exemplarily, in Figure 9 the fourth index structure shown, the common time weight sequence corresponding to each complete path (i.e., 2_3_4_5_6_7) is used as the time weight on the edge between the first node and the second node to make the fourth index structure more streamlined. It should be noted that since the nodes in the third index structure are virtual nodes, the nodes in the fourth index structure are also virtual nodes.

[0154] After obtaining the fourth index structure, based on the fourth index structure, obtain the key detection result of the speech signal. Exemplarily, the keyword detection result of the speech signal includes the hit information of the speech signal for the candidate keyword. The hit information of the speech signal for the candidate keyword includes at least the probability that the candidate keyword exists in the speech signal. Exemplarily, the hit information of the speech signal corresponding to the candidate keyword further includes the position where the candidate keyword exists in the speech signal.

[0155] In a possible implementation manner, the process of obtaining the keyword detection result of the speech signal based on the fourth index structure includes: obtaining the weight information of the speech unit sequence corresponding to the fourth index structure; based on the weight information, obtaining the hit information of the speech signal for the candidate keyword.

[0156] A speech unit sequence corresponding to the fourth index structure refers to a sequence composed of speech units represented by input labels on each edge in a complete path from the first node to the last node in the fourth index structure. The manner of obtaining the weight information of the speech unit sequence corresponding to the fourth index structure is related to whether there is a time weight on the edge of the first index structure.

[0157] Exemplarily, if there is a time weight on the edge of the first index structure, the weight information of the speech unit sequence corresponding to the fourth index structure includes probability weight information and time weight information; if there is no time weight on the edge of the first index structure, the weight information of the speech unit sequence corresponding to the fourth index structure only includes probability weight information.

[0158] Taking the case where there is a time weight on the edge of the first index structure as an example, the process of obtaining the weight information of the speech unit sequence corresponding to the fourth index structure is realized by obtaining probability weight information and time weight information.

[0159] The manner of obtaining probability weight information is related to the manner of obtaining the probability weight on the edge. If the probability weight on the edge is the matching probability, multiply the probability weights on each edge in the complete path corresponding to the speech unit sequence to obtain the probability weight information corresponding to the speech unit sequence. This probability weight information can be considered as the probability that the candidate keyword hit by the speech unit sequence exists in the speech signal.

[0160] If the probability weight on the edge is the negative logarithm of the matching probability, add the probability weights on each edge in the complete path corresponding to the speech unit sequence to obtain the probability weight information corresponding to the speech unit sequence. In this case, the probability weight information corresponding to the speech unit sequence is the negative logarithm of the probability that the candidate keyword hit by the speech unit sequence exists in the speech signal. By performing a mathematical transformation on the probability weight information corresponding to the speech unit sequence, the probability that the candidate keyword hit by the speech unit sequence exists in the speech signal can be obtained.

[0161] The way to obtain the time weight information is as follows: concatenate the time weights on each edge in the complete path corresponding to the speech unit sequence to obtain the time weight information. Based on this time weight information, the position of the candidate keyword hit by the speech unit sequence in the speech signal can be determined. Exemplarily, according to Figure 9 the fourth index structure shown, it can be determined that the time weight information corresponding to the speech unit sequence "jia1 wei1 wei1 wei1 xin4" is 2_3_4_5_6_7. Based on this time weight information, the position of the candidate keyword "add WeChat" hit by this speech unit sequence in the speech signal can be determined as the position between the reference time point corresponding to the node identified by the number 2 in the first index structure and the reference time point corresponding to the node identified by the number 7.

[0162] After obtaining the weight information of the speech unit sequence corresponding to the fourth index structure, based on the weight information of the speech unit sequence corresponding to the fourth index structure, obtain the hit information of the speech signal for the candidate keyword. Exemplarily, the hit information of the speech signal for the candidate keyword is used to indicate the probability that the speech signal contains the candidate keyword and the position where the candidate keyword exists. The probability that the speech signal contains the candidate keyword is determined according to the probability weight information corresponding to the speech unit sequence, and the position where the speech signal contains the candidate keyword is determined according to the time weight information corresponding to the speech unit sequence.

[0163] Method 2: Based on the candidate keyword, traverse the first index structure to obtain the keyword detection result of the speech signal.

[0164] In this Method 2, on the basis of the known candidate keyword, by directly traversing the first index structure, the keyword detection result of the speech signal is obtained. Exemplarily, the way to traverse the first index structure based on the candidate keyword is as follows: start from the first node in the first index structure, and sequentially jump to the last node along each path. Each time jumping to the last node along a path, the speech unit sequence corresponding to that path and the weight information of that speech unit sequence are obtained. If the speech unit sequence corresponding to that path hits a certain candidate keyword, it means that the candidate keyword exists in the speech signal, and the specific information (such as probability, position, etc.) of the existence of the candidate keyword in the speech signal is determined according to the weight information of the speech unit sequence. After traversing all the paths, the keyword detection result of the speech signal can be obtained.

[0165] The embodiment of the present application proposes a fast framework for keyword detection adaptable to any keyword, which can have good detection performance even in a small neural network model, for an end-to-end speech recognition model trained based on the CTC algorithm. According to the first speech recognition result output by calling the end-to-end speech recognition model trained based on the CTC algorithm, this method retains some candidate speech units (i.e., target speech units) with relatively high matching probabilities for each frame of speech segment, and then converts the matching probabilities (i.e., the second speech recognition result) of the speech segment and the retained candidate speech units into an index structure in the WFST form. The candidate speech keywords are also converted into the corresponding index structure in the WFST form. Through the standard Compose operation of the WFST structure, an index structure in the WFST form that can obtain the keyword detection result can be obtained.

[0166] Since keyword detection in speech usually requires locating the position where a specific keyword appears, the embodiment of the present application proposes an FTFT method, which saves time information as a string weight and combines it with the probability weight into a co-occurrence weight (Pair Weight), enabling the method proposed in the embodiment of the present application to quickly and accurately obtain the probability and position of the existence of a certain keyword in the speech signal.

[0167] Exemplarily, the keyword detection process is as Figure 11 shown. Obtain a speech signal and extract the acoustic features of the speech signal; call the end-to-end speech recognition model trained based on the CTC algorithm to recognize the acoustic features of the speech signal, and obtain the first speech recognition result corresponding to each frame of speech segment. Based on the first speech recognition result corresponding to each frame of speech segment, obtain a first index structure; convert the candidate keywords into a third index structure. Perform a combination operation on the first index structure and the third index structure to obtain a fourth index structure, and based on the fourth index structure, obtain the keyword detection result of the speech signal.

[0168] The FTFT index structure provided by the embodiments of the present application can be directly constructed according to the speech recognition results output by an end-to-end speech recognition model trained based on the CTC algorithm. The construction complexity is very low and it does not affect keyword detection. The method provided by the embodiments of the present application is implemented based on an end-to-end speech recognition model trained based on the CTC algorithm and the FTFT index structure. Using an end-to-end speech recognition model trained based on the CTC algorithm can ensure the optimal acoustic modeling ability under the same amount of training data and the same number of parameters; the proposed FTFT index structure can be quickly constructed by effectively using the fact that the probabilities output by the end-to-end speech recognition model trained based on the CTC algorithm are relatively sharp, achieving the goal of "fast and accurate" speech keyword detection. Experiments prove that when the underlying acoustic model is small, the overall performance is still very excellent, and the overall performance loss is very small. In particular, there are obvious improvements in the recall rate of keywords and the coverage rate of keywords. This is beneficial to ensuring the speech keyword detection effect in resource-constrained or resource-efficient scenarios and can significantly reduce the hardware cost.

[0169] In the embodiments of the present application, when calling an end-to-end speech recognition model to recognize acoustic features, a first index structure is directly obtained according to the recognition results (that is, the first speech recognition results corresponding to each frame of speech segment). The first index structure can comprehensively reflect the recognition results obtained. Keyword detection based on the first index structure can comprehensively detect the keywords existing in the speech signal, thereby improving the recall rate of the obtained keyword detection results, and the keyword detection effect is good.

[0170] See Figure 12 , the embodiments of the present application provide a keyword detection device, and the device includes:

[0171] A first acquisition unit 1201, configured to acquire a speech signal;

[0172] An identification unit 1202, configured to call an end-to-end speech recognition model to recognize the acoustic features of the speech signal, and obtain first speech recognition results corresponding to each frame of speech segment in the speech signal. The first speech recognition result corresponding to one frame of speech segment includes the matching probabilities of one frame of speech segment with each candidate speech unit, and the candidate speech unit is a speech unit set for implementing keyword detection;

[0173] A second acquisition unit 1203, configured to obtain a first index structure based on the first speech recognition results corresponding to each frame of speech segment. The first index structure is an index structure corresponding to the speech signal for performing keyword detection;

[0174] A third acquisition unit 1204, configured to obtain a keyword detection result of the speech signal based on the first index structure.

[0175] In a possible implementation, refer to Figure 13 , the second acquisition unit 1203 includes:

[0176] An acquisition subunit 12031, configured to obtain a second speech recognition result corresponding to each frame of speech segment based on the first speech recognition result corresponding to each frame of speech segment. The second speech recognition result corresponding to a frame of speech segment includes the matching probability between a frame of speech segment and a target speech unit corresponding to the frame of speech segment. The target speech unit corresponding to a frame of speech segment is a candidate speech unit that satisfies the selection condition with the frame of speech segment among all candidate speech units;

[0177] A conversion subunit 12032, configured to convert the second speech recognition result corresponding to each frame of speech segment into a second index structure. The second index structure is composed of nodes, edges between the nodes, and information on the edges. The information on the edges includes an input label, an output label, and a probability weight;

[0178] A transformation subunit 12033, configured to transform the second index structure to obtain a first index structure. The speech unit sequence corresponding to the first index structure includes a subsequence that satisfies the reference condition of the speech unit sequence corresponding to the second index structure.

[0179] In a possible implementation, that a frame of speech segment satisfies the selection condition means that the matching probability with the frame of speech segment is not lower than a probability threshold. The acquisition subunit 12031 is configured to, for the first speech segment, use, as the target speech unit corresponding to the first speech segment, a candidate speech unit among all candidate speech units whose matching probability with the first speech segment is not lower than the probability threshold. The first speech segment is any one of the frames of speech segments. From the first speech recognition result corresponding to the first speech segment, extract the matching probability between the first speech segment and the target speech unit corresponding to the first speech segment. The matching probability between the first speech segment and the target speech unit corresponding to the first speech segment constitutes the second speech recognition result corresponding to the first speech segment.

[0180] In a possible implementation, the nodes in the second index structure correspond one by one to the reference time points in the speech signal. The transformation subunit 12033 is configured to add a virtual start node and a virtual end node to the second index structure; construct an edge pointing to the first associated node between the virtual start node and the first associated node, and construct an edge pointing to the virtual end node between the second associated node and the virtual end node. The first associated node is a node in the second index structure that satisfies the first condition, and the second associated node is a node in the second index structure that satisfies the second condition; add a time weight to the edge whose pointed node is a node in the second index structure to obtain the first index structure. The time weight on an edge is used to indicate the reference time point corresponding to the node pointed to by the edge.

[0181] In a possible implementation manner, the nodes in the second index structure correspond one-to-one with the reference time points in the speech signal, and any two adjacent reference time points in the speech signal are used to identify a frame of speech segment; the conversion subunit 12032 is configured to use the target speech unit corresponding to the second speech segment as the input label and the output label on the edge between the first node and the second node, where the first node and the second node are any two adjacent nodes in the second index structure, and the second speech segment is the speech segment identified by the two reference time points corresponding to the first node and the second node; and determine the probability weight on the edge between the first node and the second node based on the matching probability between the second speech segment and the target speech unit corresponding to the second speech segment.

[0182] In a possible implementation manner, the conversion subunit 12032 is further configured to calculate the negative logarithm of the matching probability between the second speech segment and the target speech unit corresponding to the second speech segment, and use the negative logarithm as the probability weight on the edge between the first node and the second node.

[0183] In a possible implementation manner, the third obtaining unit 1204 is configured to convert the candidate keyword into a third index structure, where the third index structure and the first index structure are index structures of the same form; perform a combination operation on the first index structure and the third index structure to obtain a fourth index structure, and the speech unit sequence corresponding to the fourth index structure hits the candidate keyword; and obtain the keyword detection result of the speech signal based on the fourth index structure.

[0184] In a possible implementation manner, the keyword detection result of the speech signal includes the hit information of the speech signal for the candidate keyword, and the third obtaining unit 1204 is further configured to obtain the weight information of the speech unit sequence corresponding to the fourth index structure; and obtain the hit information of the speech signal for the candidate keyword based on the weight information.

[0185] In a possible implementation manner, the third obtaining unit 1204 is configured to traverse the first index structure based on the candidate keyword to obtain the keyword detection result of the speech signal.

[0186] In the embodiments of the present application, when the end-to-end speech recognition model is called to recognize the acoustic features, the first index structure is directly obtained according to the recognized result (that is, the first speech recognition result corresponding to each frame of speech segment). The first index structure can comprehensively reflect the recognized result. Performing keyword detection according to the first index structure can comprehensively detect the keywords existing in the speech signal, thereby improving the recall rate of the obtained keyword detection result, and the effect of keyword detection is good.

[0187] It should be noted that when the device provided in the above embodiments realizes its functions, only the division of the above-mentioned functional modules is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.

[0188] In an exemplary embodiment, a computer device is further provided. The computer device includes a processor and a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by one or more processors so that the computer device implements any of the above keyword detection methods. The computer device can be a server or a terminal, and the embodiments of the present application do not limit this. Next, the structures of the server and the terminal will be introduced respectively.

[0189] Figure 14 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server may vary greatly due to configuration or performance differences, and may include one or more processors (Central Processing Units, CPUs) 1401 and one or more memories 1402. Among them, at least one computer program is stored in the one or more memories 1402, and the at least one computer program is loaded and executed by the one or more processors 1401 so that the server implements the keyword detection methods provided by the above various method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The server may further include other components for implementing the functions of the device, which will not be elaborated here.

[0190] Figure 15 FIG. is a schematic structural diagram of a terminal provided by an embodiment of the present application. The terminal may be: a smart phone, a tablet computer, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, or a desktop computer. The terminal may also be referred to by other names such as a user device, a portable terminal, a laptop terminal, a desktop terminal, etc.

[0191] Generally, the terminal includes a processor 1501 and a memory 1502.

[0192] The processor 1501 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1501 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1501 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1501 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1501 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0193] The memory 1502 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1502 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1502 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1501 so that the terminal implements the keyword detection method provided in the method embodiments of the present application.

[0194] In some embodiments, the terminal may further optionally include: a peripheral device interface 1503 and at least one peripheral device. The processor 1501, the memory 1502, and the peripheral device interface 1503 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1503 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1504, a display screen 1505, a camera module 1506, an audio circuit 1507, and a power supply 1509.

[0195] The peripheral device interface 1503 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1501 and the memory 1502. In some embodiments, the processor 1501, the memory 1502, and the peripheral device interface 1503 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1501, the memory 1502, and the peripheral device interface 1503 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0196] The radio frequency circuit 1504 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1504 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1504 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1504 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1504 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1504 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0197] The display screen 1505 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1505 is a touch display screen, the display screen 1505 also has the ability to collect touch signals on or above the surface of the display screen 1505. The touch signals can be input as control signals to the processor 1501 for processing. At this time, the display screen 1505 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1505, which is disposed on the front panel of the terminal; in other embodiments, there may be at least two display screens 1505, which are respectively disposed on different surfaces of the terminal or are in a foldable design; in other embodiments, the display screen 1505 may be a flexible display screen, which is disposed on a curved surface or a folding surface of the terminal. Even, the display screen 1505 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1505 can be prepared from materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0198] The camera module 1506 is used to collect images or videos. Optionally, the camera module 1506 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to implement functions such as background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fused shooting functions. In some embodiments, the camera module 1506 may further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0199] The audio circuit 1507 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1501 for processing, or input to the radio frequency circuit 1504 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1501 or the radio frequency circuit 1504 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1507 may further include a headphone jack.

[0200] The power supply 1509 is used to supply power to each component in the terminal. The power supply 1509 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 1509 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0201] In some embodiments, the terminal further includes one or more sensors 1510. The one or more sensors 1510 include but are not limited to: an acceleration sensor 1511, a gyroscope sensor 1512, a pressure sensor 1513, an optical sensor 1515, and a proximity sensor 1516.

[0202] The acceleration sensor 1511 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal. For example, the acceleration sensor 1511 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1501 can control the display screen 1505 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1511. The acceleration sensor 1511 can also be used for collecting game or user movement data.

[0203] The gyroscope sensor 1512 can detect the body direction and rotation angle of the terminal. The gyroscope sensor 1512 can cooperate with the acceleration sensor 1511 to collect the 3D actions of the user on the terminal. According to the data collected by the gyroscope sensor 1512, the processor 1501 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.

[0204] The pressure sensor 1513 can be disposed on the side frame of the terminal and / or the lower layer of the display screen 1505. When the pressure sensor 1513 is disposed on the side frame of the terminal, it can detect the holding signal of the user on the terminal, and the processor 1501 performs left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor 1513. When the pressure sensor 1513 is disposed on the lower layer of the display screen 1505, the processor 1501 controls the operable controls on the UI interface according to the pressure operation of the user on the display screen 1505.

[0205] The optical sensor 1515 is used to collect the ambient light intensity. In one embodiment, the processor 1501 can control the display brightness of the display screen 1505 according to the ambient light intensity collected by the optical sensor 1515. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1505 is increased; when the ambient light intensity is low, the display brightness of the display screen 1505 is decreased. In another embodiment, the processor 1501 can also dynamically adjust the shooting parameters of the camera module 1506 according to the ambient light intensity collected by the optical sensor 1515.

[0206] The proximity sensor 1516, also known as the distance sensor, is usually disposed on the front panel of the terminal. The proximity sensor 1516 is used to collect the distance between the user and the front of the terminal. In one embodiment, when the proximity sensor 1516 detects that the distance between the user and the front of the terminal is gradually decreasing, the processor 1501 controls the display screen 1505 to switch from the lit state to the off state; when the proximity sensor 1516 detects that the distance between the user and the front of the terminal is gradually increasing, the processor 1501 controls the display screen 1505 to switch from the off state to the lit state.

[0207] Those skilled in the art can understand that Figure 15 the structures shown in

[0208] do not constitute a limitation on the terminal, and may include more or fewer components than shown in the figures, or combine certain components, or adopt different component arrangements.

[0209] In a possible implementation manner, the above computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0210] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions that are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes any one of the above keyword detection methods.

[0211] It should be noted that the terms "first", "second", etc. in the specification and claims of this application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. The embodiments described in the above exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0212] It should be understood that "a plurality of" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents that the associated objects before and after are in an "or" relationship.

[0213] The above are only exemplary embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.< / eps> < / eps> < / eps> < / blank> < / blank> < / blank> < / blank> < / blank> < / blank> < / blank> < / unk> < / unk>

Claims

1. A keyword detection method, characterized in that, the method includes: Obtain a speech signal; Call an end-to-end speech recognition model to recognize the acoustic features of the speech signal, and obtain first speech recognition results respectively corresponding to each frame of speech segment in the speech signal. The first speech recognition result corresponding to one frame of speech segment includes the matching probabilities of the one frame of speech segment with each candidate speech unit respectively, and the candidate speech units are speech units set for keyword detection; Based on the first speech recognition results respectively corresponding to each frame of speech segment, obtain second speech recognition results respectively corresponding to each frame of speech segment. The second speech recognition result corresponding to one frame of speech segment includes the matching probability of the one frame of speech segment with the target speech unit corresponding to the one frame of speech segment, and the target speech unit corresponding to the one frame of speech segment is the candidate speech unit that satisfies the selection condition with the one frame of speech segment among the various candidate speech units; Convert the second speech recognition results respectively corresponding to each frame of speech segment into a second index structure; wherein, the process of converting the second speech recognition results respectively corresponding to each frame of speech segment into a second index structure is realized by determining the nodes in the second index structure according to the second speech recognition results respectively corresponding to each frame of speech segment, constructing edges between the nodes in the second index structure, and determining the information on the edges; Transform the second index structure to obtain a first index structure. The speech unit sequence corresponding to the first index structure includes the subsequence that satisfies the reference condition of the speech unit sequence corresponding to the second index structure, and the first index structure is the index structure for keyword detection corresponding to the speech signal; Based on the first index structure, obtain the keyword detection result of the speech signal.

2. The method according to claim 1, characterized in that, the information on the edge includes an input label, an output label, and a probability weight.

3. The method according to claim 1, characterized in that, the one that satisfies the selection condition with the one frame of speech segment means that the matching probability with the one frame of speech segment is not lower than a probability threshold; the obtaining of the second speech recognition results respectively corresponding to each frame of speech segment based on the first speech recognition results respectively corresponding to each frame of speech segment includes: For the first speech segment, use the candidate speech units among the various candidate speech units whose matching probabilities with the first speech segment are not lower than the probability threshold as the target speech unit corresponding to the first speech segment, and the first speech segment is any one of the frames of speech segments; Extract the matching probability of the first speech segment with the target speech unit corresponding to the first speech segment from the first speech recognition result corresponding to the first speech segment, and the matching probability of the first speech segment with the target speech unit corresponding to the first speech segment constitutes the second speech recognition result corresponding to the first speech segment.

4. The method according to claim 1, characterized in that, The nodes in the second index structure correspond one-to-one with the reference time points in the speech signal; the transformation of the second index structure to obtain the first index structure includes: Adding a virtual start node and a virtual end node to the second index structure; Constructing an edge pointing to the first associated node between the virtual start node and the first associated node, and constructing an edge pointing to the virtual end node between the second associated node and the virtual end node, where the first associated node is a node in the second index structure that satisfies the first condition, and the second associated node is a node in the second index structure that satisfies the second condition; Adding a time weight to the edge whose pointed node is a node in the second index structure to obtain the first index structure, and the time weight on an edge is used to indicate the reference time point corresponding to the node pointed to by the edge.

5. The method according to claim 2, wherein, the nodes in the second index structure correspond one-to-one with the reference time points in the speech signal, and any two adjacent reference time points in the speech signal are used to identify a frame of speech segment; the conversion of the second speech recognition results corresponding to the respective frames of speech segments into the second index structure includes: Taking the target speech unit corresponding to the second speech segment as the input label and the output label on the edge between the first node and the second node, where the first node and the second node are any two adjacent nodes in the second index structure, and the second speech segment is the speech segment identified by the two reference time points corresponding to the first node and the second node; Determining the probability weight on the edge between the first node and the second node based on the matching probability between the second speech segment and the target speech unit corresponding to the second speech segment.

6. The method according to claim 5, wherein, the determining the probability weight on the edge between the first node and the second node based on the matching probability between the second speech segment and the target speech unit corresponding to the second speech segment includes: Calculating the negative logarithm of the matching probability between the second speech segment and the target speech unit corresponding to the second speech segment, and taking the negative logarithm as the probability weight on the edge between the first node and the second node.

7. The method according to any one of claims 1-6, wherein, the obtaining the keyword detection result of the speech signal based on the first index structure includes: Converting the candidate keyword into a third index structure, where the third index structure is an index structure of the same form as the first index structure; Performing a combination operation on the first index structure and the third index structure to obtain a fourth index structure, and the speech unit sequence corresponding to the fourth index structure hits the candidate keyword; Obtaining the keyword detection result of the speech signal based on the fourth index structure.

8. The method according to claim 7, wherein, The keyword detection result of the speech signal includes the hit information of the speech signal for the candidate keyword. Obtaining the keyword detection result of the speech signal based on the fourth index structure includes: Obtaining the weight information of the speech unit sequence corresponding to the fourth index structure; Based on the weight information, obtaining the hit information of the speech signal for the candidate keyword.

9. According to the method of any one of claims 1-6, characterized in that Obtaining the keyword detection result of the speech signal based on the first index structure includes: Traversing the first index structure based on the candidate keyword to obtain the keyword detection result of the speech signal.

10. A keyword detection device, characterized in that The device includes: A first acquisition unit for acquiring a speech signal; An identification unit for calling an end-to-end speech recognition model to identify the acoustic features of the speech signal, and obtaining a first speech recognition result corresponding to each frame of speech segment in the speech signal. The first speech recognition result corresponding to one frame of speech segment includes the matching probabilities of the one frame of speech segment with each candidate speech unit respectively. The candidate speech unit is a speech unit set for keyword detection; A second acquisition unit, including an acquisition subunit, a conversion subunit, and a transformation subunit; The acquisition subunit is configured to obtain a second speech recognition result corresponding to each frame of speech segment based on the first speech recognition result corresponding to each frame of speech segment. The second speech recognition result corresponding to one frame of speech segment includes the matching probability of the one frame of speech segment with the target speech unit corresponding to the one frame of speech segment. The target speech unit corresponding to one frame of speech segment is the candidate speech unit that satisfies the selection condition among the various candidate speech units for the one frame of speech segment; The conversion subunit is configured to convert the second speech recognition result corresponding to each frame of speech segment into a second index structure; wherein, the process of converting the second speech recognition result corresponding to each frame of speech segment into a second index structure is realized by determining the nodes in the second index structure according to the second speech recognition result corresponding to each frame of speech segment, constructing edges between the nodes in the second index structure, and determining the information on the edges; The transformation subunit is configured to transform the second index structure to obtain a first index structure. The speech unit sequence corresponding to the first index structure includes a subsequence that satisfies the reference condition of the speech unit sequence corresponding to the second index structure. The first index structure is the index structure corresponding to the speech signal for keyword detection; A third acquisition unit for obtaining the keyword detection result of the speech signal based on the first index structure.

11. According to the device of claim 10, characterized in that The information on the edge includes an input label, an output label, and a probability weight.

12. According to the device of claim 10, characterized in that Saying that the frame of speech segment meets the selection condition means that the matching probability with the frame of speech segment is not lower than the probability threshold; The obtaining subunit is configured to, for a first speech segment, use, as the target speech unit corresponding to the first speech segment, a candidate speech unit in each of the candidate speech units whose matching probability with the first speech segment is not lower than the probability threshold, where the first speech segment is any one of the frames of speech segments; extract, from a first speech recognition result corresponding to the first speech segment, a matching probability between the first speech segment and the target speech unit corresponding to the first speech segment, and use the matching probability between the first speech segment and the target speech unit corresponding to the first speech segment to form a second speech recognition result corresponding to the first speech segment.

13. The apparatus according to claim 10, wherein, nodes in the second index structure correspond one-to-one to reference time points in the speech signal; The transformation subunit is configured to add a virtual start node and a virtual end node to the second index structure; construct an edge pointing to a first associated node between the virtual start node and the first associated node, and construct an edge pointing to the virtual end node between a second associated node and the virtual end node, where the first associated node is a node in the second index structure that meets a first condition, and the second associated node is a node in the second index structure that meets a second condition; add a time weight to an edge whose pointed node is a node in the second index structure to obtain the first index structure, and a time weight on an edge is used to indicate a reference time point corresponding to the node pointed to by the edge.

14. The apparatus according to claim 11, wherein, nodes in the second index structure correspond one-to-one to reference time points in the speech signal, and any two adjacent reference time points in the speech signal are used to identify a frame of speech segment; The conversion subunit is configured to use a target speech unit corresponding to a second speech segment as an input label and an output label on an edge between a first node and a second node, where the first node and the second node are any two adjacent nodes in the second index structure, and the second speech segment is a speech segment identified by two reference time points corresponding to the first node and the second node; determine a probability weight on the edge between the first node and the second node based on a matching probability between the second speech segment and the target speech unit corresponding to the second speech segment.

15. The apparatus according to claim 14, wherein, The conversion subunit is configured to calculate a negative logarithm of a matching probability between the second speech segment and the target speech unit corresponding to the second speech segment, and use the negative logarithm as a probability weight on the edge between the first node and the second node.

16. The apparatus according to any one of claims 10-15, wherein, The third acquisition unit is configured to convert the candidate keyword into a third index structure, where the third index structure is in the same form as the first index structure; perform a combination operation on the first index structure and the third index structure to obtain a fourth index structure, and the speech unit sequence corresponding to the fourth index structure hits the candidate keyword; Based on the fourth index structure, obtain the keyword detection result of the speech signal.

17. The apparatus according to claim 16, wherein, The keyword detection result of the speech signal includes the hit information of the speech signal on the candidate keyword. The third acquisition unit is configured to obtain the weight information of the speech unit sequence corresponding to the fourth index structure; based on the weight information, obtain the hit information of the speech signal on the candidate keyword.

18. The apparatus according to any one of claims 10-15, wherein, The third acquisition unit is configured to traverse the first index structure based on the candidate keyword to obtain the keyword detection result of the speech signal.

19. A computer device, wherein, The computer device includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor so that the computer device implements the keyword detection method according to any one of claims 1 to 9.

20. A computer-readable storage medium, wherein, At least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor so that a computer implements the keyword detection method according to any one of claims 1 to 9.

21. A computer program product, wherein, The computer program product includes computer instructions. The computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computer device executes the keyword detection method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Voice recognition method, device and system and storage medium

    CN112133294A