Voice verification method and device, storage medium and electronic device
By using acoustic models trained by reading text and digital strings, the number strings read aloud by the target object are verified and the first and second probability of the recognition results are calculated, which solves the problem of low speech verification accuracy in the prior art and achieves higher verification accuracy.
Patent Information
- Application Number
- CN202010753151.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-07-30
AI Technical Summary
In the prior art, the accuracy of speech verification is low, resulting in the accuracy of verification target objects.
By obtaining the target speech generated by the target object reading the target number string and inputting it into the acoustic model trained using the reading text and the number string, the first and second probability of the recognition results are calculated, and the recognition results with high semantic understanding and matching the target number string are selected for verification.
It improves the accuracy of the speech verification process and ensures that the verification results of the target object are more reliable.
Smart Images

Figure CN111933117B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a voice verification method and device, a storage medium and an electronic device. Background Art
[0002] In the prior art, in many scenarios such as account login, money transfer, etc., it is necessary to verify the target object to verify whether the object performing the operation is a robot.
[0003] The prior art provides a means to verify the target object by acquiring the sound of the target object reading the target content, using a model to recognize the sound, and comparing whether the sound matches the target content.
[0004] However, in the above process, since the model has low accuracy in recognizing the sound of the read numbers, the accuracy of verifying the target object is further low.
[0005] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0006] The embodiments of the present invention provide a voice verification method and device, a storage medium and an electronic device to at least solve the technical problem of low voice verification accuracy.
[0007] According to one aspect of an embodiment of the present invention, a voice verification method is provided, comprising: obtaining a target voice generated by a target object reading a target digital string; inputting the target voice into an acoustic model to obtain multiple recognition results of the target voice and a first probability of each recognition result, wherein the acoustic model is a model for recognizing the target voice obtained by training using a first training sample and a second training sample, the first training sample is a sample obtained by reading a text, and the second training sample is a sample obtained by reading a digital string, and the first probability is used to indicate the possibility that the recognition result is the same as the target voice; calculating the second probability of each of the multiple recognition results, wherein the second probability is used to indicate the degree of semantic understanding of the recognition result; determining a target recognition result based on the first probability and the second probability, wherein the target recognition result is a recognition result for which the second probability is less than a predetermined threshold and the first probability is the largest; in the case where the target recognition result is the same as the target digital string, sending a first prompt message, wherein the first prompt message is used to prompt the target object to pass the verification corresponding to the target digital string.
[0008] According to another aspect of an embodiment of the present invention, a speech verification device is further provided, comprising: a first acquisition unit, configured to acquire a target speech generated by a target object reading a target digital string; an input unit, configured to input the target speech into an acoustic model, and obtain multiple recognition results of the target speech and a first probability of each recognition result, wherein the acoustic model is a model for recognizing the target speech obtained by training using a first training sample and a second training sample, the first training sample is a sample obtained by reading a text, and the second training sample is a sample obtained by reading a digital string, and the first probability is used to indicate the possibility that the recognition result is the same as the target speech; a calculation unit, configured to calculate a second probability of each of the multiple recognition results, wherein the second probability is used to indicate the semantic understanding degree of the recognition result; a determination unit, configured to determine a target recognition result according to the first probability and the second probability, wherein the target recognition result is a recognition result in which the second probability is less than a predetermined threshold and the first probability is the largest; a first sending unit, configured to send a first prompt message when the target recognition result is the same as the target digital string, wherein the first prompt message is used to prompt the target object to pass the verification corresponding to the target digital string.
[0009] As an optional example, the above-mentioned device also includes: a second sending unit, used to send a second prompt message after the above-mentioned target recognition result is determined according to the above-mentioned first probability and the above-mentioned second probability, if the above-mentioned target recognition result is different from the above-mentioned target digital string, wherein the above-mentioned second prompt message is used to prompt that the above-mentioned target object has not passed the verification corresponding to the above-mentioned target digital string.
[0010] As an optional example, the above-mentioned device also includes: a second acquisition unit, used to acquire the above-mentioned first training sample and the above-mentioned second training sample before inputting the above-mentioned target speech into the above-mentioned acoustic model; a first training unit, used to train the original model with the above-mentioned first training sample until the number of training times reaches a predetermined number of times or the accuracy of the above-mentioned original model reaches a first accuracy; a second training unit, used to train the above-mentioned original model trained with the above-mentioned first training sample with the above-mentioned second training sample, to obtain the above-mentioned acoustic model.
[0011] As an optional example, the calculation unit includes: an acquisition module, used to acquire the multiple recognition results; and a calculation module, used to calculate the second probability of each of the multiple recognition results using a target language model.
[0012] As an optional example, the acquisition unit further includes: a first training module, used to train the first language model using a third training sample to obtain a second language model before using the language model to calculate the second probability of each of the multiple recognition results, wherein the third training sample is a text sample; a second training module, used to train the first language model using a fourth training sample to obtain a third language model, wherein the fourth training sample is a digital string sample; and a merging module, used to merge the trained second language model and the third language model into the target language model.
[0013] As an optional example, the above-mentioned determination unit includes: a deletion module, used to delete the recognition results whose second probability is greater than or equal to the above-mentioned predetermined threshold among the above-mentioned multiple recognition results; a first determination module, used to determine the recognition result whose first probability is the largest among the remaining above-mentioned recognition results as the above-mentioned target recognition result.
[0014] As an optional example, the above-mentioned first acquisition unit includes: a display module, used to display the above-mentioned target digital string on a display interface; a prompt module, used to prompt the above-mentioned target object to read the above-mentioned target digital string; a recording module, used to start recording when the above-mentioned target object is prompted to read the above-mentioned target digital string, and end recording after recording for a first time length; a second determination module, used to determine the recorded recording as the above-mentioned target voice.
[0015] As an optional example, the above-mentioned device also includes: a receiving unit, used to receive a login request from the above-mentioned target object before obtaining the above-mentioned target voice generated by the above-mentioned target object reading the above-mentioned target number string, wherein the above-mentioned login request is used to request to log in to the target application; a display unit, used to display the above-mentioned target number string, prompting the above-mentioned target object to read the above-mentioned target number.
[0016] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above-mentioned voice verification method.
[0017] According to another aspect of an embodiment of the present invention, there is provided an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the voice verification method through the computer program.
[0018] In an embodiment of the present invention, a target speech generated by a target object reading a target digital string is obtained; the target speech is input into an acoustic model to obtain multiple recognition results of the target speech and a first probability of each recognition result, wherein the acoustic model is a model for recognizing the target speech obtained by training using a first training sample and a second training sample, the first training sample is a sample obtained by reading a text, and the second training sample is a sample obtained by reading a digital string, and the first probability is used to indicate the possibility that the recognition result is the same as the target speech; the second probability of each of the multiple recognition results is calculated, wherein the second probability is used to indicate the degree of semantic understanding of the recognition result; the target recognition result is determined based on the first probability and the second probability. The result is that the target recognition result is a recognition result in which the second probability is less than a predetermined threshold and the first probability is the largest; when the target recognition result is the same as the target digital string, a first prompt message is sent, wherein the first prompt message is used to prompt the target object to pass the verification method corresponding to the target digital string. In the above method, in the process of verifying the voice, an acoustic model trained successively by using samples obtained by reading text and samples obtained by reading digital strings is used to recognize the voice, and the first probability of the recognition result output by the acoustic model and the second probability of the obtained recognition result are used to screen the recognition result, so that a more accurate voice recognition result can be obtained, thereby improving the accuracy of the voice verification process, and thus solving the technical problem of low voice verification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0020] Figure 1 is a schematic diagram of an application environment of an optional voice verification method according to an embodiment of the present invention;
[0021] Figure 2 is a schematic diagram of an application environment of another optional voice verification method according to an embodiment of the present invention;
[0022] Figure 3 is a flow chart of an optional voice verification method according to an embodiment of the present invention;
[0023] Figure 4 is a schematic diagram of an interface of an optional voice verification method according to an embodiment of the present invention;
[0024] Figure 5 is a verification flow diagram of an optional voice verification method according to an embodiment of the present invention;
[0025] Figure 6 is a verification schematic diagram of an optional voice verification method according to an embodiment of the present invention;
[0026] Figure 7 is a verification schematic diagram of another optional voice verification method according to an embodiment of the present invention;
[0027] Figure 8 is a schematic structural diagram of an optional voice verification device according to an embodiment of the present invention;
[0028] Fig. 9 is a schematic structural diagram of another optional voice verification device according to an embodiment of the present invention;
[0029] Fig.10 is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0030] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0032] Artificial Intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.
[0033] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0034] The key technologies of speech technology include automatic speech recognition technology (ASR), text-to-speech technology (TTS) and voiceprint recognition technology. Enabling computers to listen, see, speak and feel is the future development direction of human-computer interaction, among which speech has become one of the most promising human-computer interaction methods in the future.
[0035] Natural language processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, that is, the language people use in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph and other technologies.
[0036] Machine Learning (ML) is a multi-disciplinary subject that involves probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and self-learning.
[0037] The solutions provided in the embodiments of the present application involve artificial intelligence speech technology, natural language processing technology, machine learning and other technologies, which are specifically explained through the following embodiments.
[0038] According to one aspect of an embodiment of the present invention, a voice verification method is provided. Optionally, as an optional implementation, the voice verification method can be but is not limited to being applied to: Figure 1 in the environment shown.
[0039] Figure 1 In the example, a user 102 and a user device 104 can perform human-computer interaction. The user device 104 includes a memory 106 for storing interaction data and a processor 108 for processing interaction data. The user device 104 can perform data interaction with a server 112 via a network 110. The server 112 includes a database 114 for storing interaction data and a processing engine 116 for processing interaction data. The user device 102 can obtain a target voice and send the target voice to the server 112, which verifies the target voice and returns a verification result.
[0040] As an optional implementation, the above voice verification method can be applied to, but is not limited to, Figure 2 in the environment shown.
[0041] Figure 2 In the example, a user 202 and a user device 204 can perform human-computer interaction. The user device 204 includes a memory 206 for storing interaction data and a processor 208 for processing interaction data. The user device 202 can obtain a target voice, verify the target voice, and return a verification result.
[0042] Optionally, the user device 104 or user device 204 in the present application can be a terminal such as a mobile phone, a tablet computer, a laptop computer, a PC, or other terminals with storage and computing capabilities. The user device 104 includes a memory 106 and a processor 108, and the user device 204 includes a memory 206 and a processor 208. The memory 106 and the memory 206 can store the computer program in the present application, and the computer program includes the acoustic model and the target language model in the present application. The processor 108 and the processor 208 can be, but not limited to, calling the computer program in the memory 106 and the memory 206 to execute the voice verification method in the present application. Specifically, the user device obtains the target voice generated by the target object reading the target digital string, and then the processor calls the acoustic model, inputs the target voice into the acoustic model, and obtains multiple recognition results and the first probability of each recognition result. Then, the processor calls the target language model to calculate the second probability of each recognition result, and finally, the target recognition result is determined according to the first probability and the second probability. Finally, the user device may compare the received target digital string with the target recognition result to verify whether the target object reads the target digital string correctly.
[0043] Optionally, the user equipment 104 and the user equipment 204 may include, but are not limited to, more components, such as a transmission device, a display device, a recording device, a connection device, etc. The transmission device may receive or send data via a network, the display may display specific content of the verification, etc., the recording device may be used to record the target voice, and the connection device is used to connect various components in the user equipment.
[0044] Optionally, as an optional implementation, as Figure 3 As shown, the above voice verification method includes:
[0045] S302, obtaining a target voice generated by a target object reading a target digital string;
[0046] S304, inputting the target speech into the acoustic model, obtaining multiple recognition results of the target speech and a first probability of each recognition result, wherein the acoustic model is a model for recognizing the target speech obtained by training using a first training sample and a second training sample, the first training sample is a sample obtained by reading a text, and the second training sample is a sample obtained by reading a digital string, and the first probability is used to indicate the possibility that the recognition result is the same as the target speech;
[0047] S306, calculating a second probability of each of the multiple recognition results, wherein the second probability is used to indicate a semantic understanding degree of the recognition result;
[0048] S308, determining a target recognition result according to the first probability and the second probability, wherein the target recognition result is a recognition result in which the second probability is less than a predetermined threshold and the first probability is the largest;
[0049] S310, when the target recognition result is the same as the target digital string, sending a first prompt message, wherein the first prompt message is used to prompt the target object to pass the verification corresponding to the target digital string.
[0050] Optionally, the second probability in the present application is a probability value, the size of which is used to indicate the degree of semantic understanding of the recognition result. The above-mentioned degree of semantic understanding indicates whether humans understand a sentence correctly or not. If the recognition result is very easy to understand from a human perspective without any misunderstanding, the calculated second probability of the recognition result is smaller. If the recognition result is understood from a human perspective and its meaning cannot be understood, the second probability is larger. The second probability of the recognition result is calculated in the present application, that is, the calculated probability is used as the possibility of whether the recognition result is a content that is easy to understand. The larger the second probability, the smaller the degree of semantic understanding, indicating that the recognition result is difficult to understand its meaning, and the smaller the second probability, the greater the degree of sentence understanding, indicating that the recognition result is easy to understand and there is no difficulty in understanding.
[0051] Optionally, the second probability in the present application may also indicate the possibility that the recognition result does not contain grammatical errors, or the second probability is used to indicate the number of grammatical errors in the recognition result.
[0052] Optionally, the above voice verification method can be applied, but not limited to, in a login process, or a transfer process, or a download process, or a file opening process, or a payment process.
[0053] Taking the login process as an example, when logging in, it is usually necessary to verify the current login user. During verification, the target digital string can be displayed on the interface and the current user can be prompted to read the target digital string. Record the target voice of the current user reading the target digital string, and then input the target voice into the acoustic model. Obtain the recognition result of the target voice recognized by the acoustic model. For example, the target digital string is "Login Authorization" and the user reads "Login Authorization", and the results that may be recognized are "Login Collection", "Login Authorization" and other results. Then, the target recognition result is determined from the recognition result. The target recognition result is the recognition result in which the second probability is less than the predetermined threshold and the first probability is the largest. Then compare the recognition result with the target digital string to see if they are the same. If they are the same, the verification is successful and the user can log in. If they are different, the verification fails and the user cannot log in.
[0054] Through the above method, in the process of verifying speech, the acoustic model trained successively by using samples obtained by reading text and samples obtained by reading digital strings is used to recognize speech, and the first probability of the recognition result output by the acoustic model and the second probability of the recognition result obtained are used to screen the recognition result, so that more accurate speech recognition results can be obtained, thereby improving the accuracy of the speech verification process.
[0055] Optionally, in the present application, the target string can be displayed on the terminal. When the current user needs to be verified, the target string can be displayed on the display interface first. The target string can be a string randomly extracted from the database. After the target string is displayed, the current user can be prompted to read the target string. Recording starts when the target object is prompted to read the target string, and ends after the first recording time. The first time length can be a preset time length, which can be set according to the length of the target string. For example, one character of the target string corresponds to one second, and the first time length is as many seconds as the number of characters in the target string. After stopping recording, the recorded recording is the target speech, and the target speech is recognized by the acoustic model.
[0056] like Figure 4 As shown, Figure 4 It is an optional verification interface. When logging in, you need to read a string of numbers. The client completes the verification in the background to determine whether to allow login.
[0057] It should be noted that when recording a target audio, a button can be set to allow the user to determine the start and end timings.
[0058] like Figure 5 As shown, Figure 5 It is the verification logic. After the target voice 502 is obtained, the target voice 502 can be verified. If the verification fails, the target voice verification is continued. The number of verifications can be required. If it exceeds the number, the login is not allowed within the predetermined time. If the verification is successful, the login is allowed.
[0059] Optionally, the acoustic model in the present application may be a model trained using the first training sample and the second training sample.
[0060] Before training the acoustic model, you first need to obtain the first training sample and the second training sample. The first training sample is a training sample obtained by reading the text content aloud. For example, the first training sample is obtained by reading the text content aloud in different dialects or Mandarin. The second training sample is obtained by reading the digital string aloud in different dialects or Mandarin. Then, the original model is first trained using the first training sample. After training to a certain extent, the second training sample is used to continue training to obtain the acoustic model. The acoustic model can output a variety of text contents after the target speech is input. Each text content is a possible content in the target speech, and each text content corresponds to a first probability.
[0061] Optionally, in the present application, after acquiring the target speech, a feature extraction model may be used to extract features in the target speech, and then the features may be input into an acoustic model for recognition.
[0062] After obtaining the plurality of recognition results and the plurality of first probabilities, the target language model may be used to calculate the second probability of each of the plurality of recognition results.
[0063] The above second probability can be understood as the probability that the identified recognition result is a relatively normal sentence. For example, "Have you eaten?" conforms to language logic and is a correct sentence, while "Have you eaten?" is likely to be a wrong sentence or a sentence that is misrecognized. Therefore, the second probability of each recognition result is calculated through the language model, and then the recognition results whose second probability exceeds the predetermined threshold are deleted, so that the unlikely recognition results in the recognition results can be filtered out, and the remaining recognition results are left.
[0064] After filtering the recognition results output by the acoustic model using the target language model, the recognition result with the highest first probability is determined as the target recognition result among the remaining results. The target recognition result here can be determined as the text content contained in the user's target speech. By comparing the target recognition result with the target digital string, the user can be verified.
[0065] Optionally, in the present application, after obtaining multiple recognition results of the acoustic model to obtain multiple first probabilities, the target language model can be used to calculate the second probability of each recognition result, and then the first probability and the second probability are weighted and summed to obtain the final total probability, and then the recognition result corresponding to the largest probability in the total probability is determined as the target recognition result. The weight can be positive or negative.
[0066] The target language model in this application may be a pre-trained language model.
[0067] In the process of determining the target language model, the first language model can be trained with the third training sample to obtain the second language model, wherein the third training sample is a text type sample, and the first language model can be trained with the fourth training sample to obtain the third language model, wherein the fourth training sample is a digital string sample; the digital string sample is a sample composed of numbers, such as 12345. The trained second language model and the third language model are merged into the target language model. The merging process can be performed by interpolation.
[0068] This application does not require too much annotated corpus and computing resources. The acoustic model and language model adaptation method can be used in any ASR-based voice digit string verification product, especially for voice scenarios that lack sufficient annotated corpus.
[0069] The present application is explained below in conjunction with a specific embodiment, for example, the present application is applied to the process of user login verification.
[0070] First, a target number string is displayed on the front-end interface, for example, 6913, and the user is prompted to read 6913. The system can display a recording bar to indicate that recording is in progress. When the user reads 6913, it will be recorded. When the recording is completed, the target voice is obtained. Through this application, it can be determined whether the target voice is 6913, thereby realizing the verification of the user. In this process, it is also possible that there may be no voice in the target voice obtained, for example, the user did not read it aloud.
[0071] The following are the specific technical contents.
[0072] The present application can be applied to an ASR system based on a Hidden Markov Model (HMM) and an end-to-end ASR system.
[0073] Figure 6 This is a schematic diagram of a robust speech digital string verification method based on HMM. In the adaptive learning process of the acoustic model, the network input of the present application is a 40-dimensional MFCC and a 100-dimensional i-vector feature. That is to say, the present application can extract the 40-dimensional MFCC and 100-dimensional i-vector features of the training sample through the feature extraction model, and then train the original model to obtain the acoustic model. The network structure of the acoustic model adopts the FTDNN structure. The FTDNN structure adopts semi-orthogonal low-rank matrix decomposition and subsampling technology, which can maintain the recognition performance while speeding up the training and decoding speed.
[0074] The training samples in this application include a first training sample and a second training sample. The first training sample is a general corpus, which is a sample obtained by reading a text, and the second training sample is a digital string corpus, which is a sample obtained by reading a digital string. An acoustic model based on FTDNN is trained in the general annotated corpus, and then it is used as an initialization model to continue training in the target corpus, thereby obtaining the acoustic model used in this application.
[0075] In the learning process of language model adaptation, the method first selects the first N commonly used words in the general text, where N is a positive integer, to form the third training sample. The third training sample is used to train the first model. The first model is a language model (second model) generated using the N-gram model. At the same time, the application also performs N-gram training on the digital string text to obtain the language model (third model) of the digital string text. Then, the method interpolates and merges the two language models of the second model and the third model according to a certain weight ratio to obtain the target language model. This language model adaptation method can effectively reduce the reporting of non-digital content and more accurately identify the content of the voice digital string. Finally, the application can reduce the size of the target language model by the method of model pruning.
[0076] In the HMM training and decoding process, the present application adopts the LF-MMI criterion, which can make the training and decoding speed of the entire HMM faster. For the acoustic features of a speech digital string, the present application can decode it through the trained HMM model to obtain the decoding sequence of the audio. Finally, the present application compares this decoding sequence with the given digital string to verify whether the speech digital string recognition is correct.
[0077] In addition, the acoustic model and language model adaptation method of the present application can also be combined with an end-to-end ASR method. Figure 7 It is an end-to-end robust speech digital string verification method, which uses the end-to-end model framework LAS. The entire LAS model is generally composed of two parts: an encoder and an attention-based decoder. The encoder uses a neural network to encode acoustic features, and the attention-based decoder first uses the attention mechanism to calculate the similarity between the decoded content at the current moment and the encoder output and generates a context vector at the corresponding moment, then decodes according to this context vector, and finally directly outputs the decoded sequence through a softmax layer.
[0078] In the encoder learning process, this application also uses transfer learning to do acoustic model adaptation. Specifically, this application first trains an encoder based on a PBLSTM network on a large amount of general corpus (first training sample), and then uses it as an initialization model to continue training in the target corpus (second training sample) to obtain the acoustic model in this application. Although the end-to-end ASR method can directly output a decoding sequence, this decoding sequence generally deviates in the actual task scenario. Therefore, this application also uses a target language model adaptive method for rescoring to generate a more accurate decoding sequence.
[0079] This application randomly selected 5000 audio samples for testing, and statistically analyzed the accuracy and real-time rate of different methods in voice digit string verification. From the results in Table 1, it can be seen that compared with the HMM-based ASR method, the HMM-based and end-to-end robust voice digit string verification methods of this application have significantly improved accuracy. At the same time, the real-time rate of these two methods in voice digit string verification is significantly reduced.
[0080] Table 1
[0081] method Accuracy Real-time rate HMM-based ASR 65.60% 0.0250 The present invention is based on HMM robust speech digital string verification 91.32% 0.0125 The present invention is based on end-to-end robust voice digital string verification 92.05% 0.0100
[0082] Through the present application and the above method, in the process of verifying speech, an acoustic model trained successively by using samples obtained by reading text and samples obtained by reading digital strings is used to recognize speech, and the first probability of the recognition result output by the acoustic model and the second probability of the obtained recognition result are used to screen the recognition result, thereby obtaining more accurate speech recognition results and improving the accuracy of the speech verification process.
[0083] As an optional implementation scheme, after determining the target recognition result according to the first probability and the second probability, the method further includes:
[0084] When the target recognition result is different from the target digital string, a second prompt message is sent, wherein the second prompt message is used to prompt that the target object has not passed the verification corresponding to the target digital string.
[0085] In the present application, after sending the first prompt information, the user may also be assigned corresponding permissions, that is, permissions to allow login. Alternatively, after sending the second prompt information, the user is prohibited from logging in.
[0086] Through this application and the above method, corresponding prompt information can be sent on the basis of completing the verification, and the user can be allowed to log in or prohibited to log in, thereby achieving the effect of improving the accuracy of verification of the target object.
[0087] As an optional implementation, before inputting the target speech into the acoustic model, the method further includes:
[0088] Obtaining a first training sample and a second training sample;
[0089] Using the first training sample to train the original model until the number of training times reaches a predetermined number or the accuracy of the original model reaches a first accuracy;
[0090] The original model trained with the first training sample is trained with the second training sample to obtain an acoustic model.
[0091] Optionally, the first training sample in the present application may be a general corpus, which includes speech contents of various styles and contents, and the second training sample may be a target corpus, which includes speech contents of digital strings.
[0092] The original model is first trained using the first training sample, and then the original model trained using the first training sample is trained using the second training sample to obtain an acoustic model, which can improve the accuracy of the model.
[0093] As an optional implementation, calculating the second probability of each recognition result in the plurality of recognition results includes:
[0094] Multiple recognition results are obtained;
[0095] A second probability of each of the plurality of recognition results is calculated using the target language model.
[0096] The target language model in the present application can be a target language model obtained by training a first model with different samples, obtaining a second model and a third model, and then merging the second model and the third model. The merged target language model has stronger recognition and identification capabilities.
[0097] As an optional implementation, before using the language model to calculate the second probability of each of the multiple recognition results, the method further includes:
[0098] Using a third training sample to train the first language model to obtain a second language model, wherein the third training sample is a text sample;
[0099] Using the fourth training sample to train the first language model to obtain a third language model, wherein the fourth training sample is a digital string sample;
[0100] The trained second language model and the third language model are merged into the target language model.
[0101] Through the above steps in the present application, a target language model with stronger recognition ability can be obtained, thereby further improving the verification accuracy of the target object.
[0102] As an optional implementation scheme, determining the target recognition result according to the first probability and the second probability includes:
[0103] Deleting, from among the multiple recognition results, recognition results whose second probability is greater than or equal to a predetermined threshold;
[0104] Among the remaining recognition results, the recognition result with the highest first probability is determined as the target recognition result.
[0105] That is to say, in this application, an acoustic model is used to recognize the target speech, multiple possibilities are obtained, and then the target speech model is used to screen out unlikely results, and the recognition result with the highest first probability among the remaining results is determined as the target recognition result, so that accurate recognition results can be obtained and the accuracy of verification can be improved.
[0106] As an optional implementation scheme, obtaining the target speech generated by the target object reading the target digital string includes:
[0107] Displaying the target digital string on the display interface;
[0108] Prompt the target subject to read the target number string;
[0109] Begin recording when the target subject is prompted to read the target digit string;
[0110] End the recording after recording the first duration;
[0111] The recorded audio is identified as the target voice.
[0112] Through the present application, the accuracy of verifying the target object is improved by displaying the target digital string on the display interface and recording the target voice, and verifying the target object after recognizing the target voice.
[0113] As an optional implementation, before obtaining the target speech generated by the target object reading the target digital string, the method further includes:
[0114] A login request from a target object is received, wherein the login request is used to request to log into a target application, display a target number string, and prompt the target object to read aloud the target number.
[0115] That is to say, the present application is applied in the login process, thereby improving the accuracy of verifying the target object in the login process.
[0116] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0117] According to another aspect of the embodiments of the present invention, a voice verification device for implementing the above-mentioned voice verification method is also provided. Figure 8 As shown, the device comprises:
[0118] The first acquisition unit 802 is used to acquire the target speech generated by the target object reading the target digital string;
[0119] An input unit 804 is used to input the target speech into the acoustic model to obtain multiple recognition results of the target speech and a first probability of each recognition result, wherein the acoustic model is a model for recognizing the target speech obtained by training using a first training sample and a second training sample, the first training sample is a sample obtained by reading a text, and the second training sample is a sample obtained by reading a digital string, and the first probability is used to indicate the possibility that the recognition result is the same as the target speech;
[0120] A calculation unit 806, configured to calculate a second probability of each of the plurality of recognition results, wherein the second probability is used to indicate a semantic understanding degree of the recognition result;
[0121] A determination unit 808, configured to determine a target recognition result according to the first probability and the second probability, wherein the target recognition result is a recognition result in which the second probability is less than a predetermined threshold and the first probability is the largest;
[0122] The first sending unit 810 is used to send a first prompt message when the target recognition result is the same as the target digital string, wherein the first prompt message is used to prompt the target object to pass the verification corresponding to the target digital string.
[0123] Optionally, the above-mentioned voice verification device can be applied to, but not limited to, a login process, a transfer process, a download process, a file opening process, or a payment process.
[0124] Taking the login process as an example, when logging in, it is usually necessary to verify the current login user. During verification, the target digital string can be displayed on the interface and the current user can be prompted to read the target digital string. Record the target voice of the current user reading the target digital string, and then input the target voice into the acoustic model. Obtain the recognition result of the target voice recognized by the acoustic model. For example, the target digital string is "Login Authorization" and the user reads "Login Authorization", and the results that may be recognized are "Login Collection", "Login Authorization" and other results. Then, the target recognition result is determined from the recognition result. The target recognition result is the recognition result in which the second probability is less than the predetermined threshold and the first probability is the largest. Then compare the recognition result with the target digital string to see if they are the same. If they are the same, the verification is successful and the user can log in. If they are different, the verification fails and the user cannot log in.
[0125] Through the above method, in the process of verifying speech, the acoustic model trained successively by using samples obtained by reading text and samples obtained by reading digital strings is used to recognize speech, and the first probability of the recognition result output by the acoustic model and the second probability of the recognition result obtained are used to screen the recognition result, so that more accurate speech recognition results can be obtained, thereby improving the accuracy of the speech verification process.
[0126] Optionally, in the present application, the target string can be displayed on the terminal. When the current user needs to be verified, the target string can be displayed on the display interface first. The target string can be a string randomly extracted from the database. After the target string is displayed, the current user can be prompted to read the target string. Recording starts when the target object is prompted to read the target string, and ends after the first recording time. The first time length can be a preset time length, which can be set according to the length of the target string. For example, one character of the target string corresponds to one second, and the first time length is as many seconds as the number of characters in the target string. After stopping recording, the recorded recording is the target speech, and the target speech is recognized by the acoustic model.
[0127] like Figure 4 As shown, Figure 4 It is an optional verification interface. When logging in, you need to read a string of numbers. The client completes the verification in the background to determine whether to allow login.
[0128] It should be noted that when recording a target audio, a button can be set to allow the user to determine the start and end timings.
[0129] like Figure 5 As shown, Figure 5It is the verification logic. After the target voice 502 is obtained, the target voice 502 can be verified. If the verification fails, the target voice verification is continued. The number of verifications can be required. If it exceeds the number, the login is not allowed within the predetermined time. If the verification is successful, the login is allowed.
[0130] Optionally, the acoustic model in the present application may be a model trained using the first training sample and the second training sample.
[0131] Before training the acoustic model, you first need to obtain the first training sample and the second training sample. The first training sample is a training sample obtained by reading the text content aloud. For example, the first training sample is obtained by reading the text content aloud in different dialects or Mandarin. The second training sample is obtained by reading the digital string aloud in different dialects or Mandarin. Then, the original model is first trained using the first training sample. After training to a certain extent, the second training sample is used to continue training to obtain the acoustic model. The acoustic model can output a variety of text contents after the target speech is input. Each text content is a possible content in the target speech, and each text content corresponds to a first probability.
[0132] Optionally, in the present application, after acquiring the target speech, a feature extraction model may be used to extract features in the target speech, and then the features may be input into an acoustic model for recognition.
[0133] After obtaining the plurality of recognition results and the plurality of first probabilities, the target language model may be used to calculate the second probability of each of the plurality of recognition results.
[0134] The above second probability can be understood as the probability that the identified recognition result is a relatively normal sentence. For example, "Have you eaten?" conforms to language logic and is a correct sentence, while "Have you eaten?" is likely to be a wrong sentence or a sentence that is misrecognized. Therefore, the second probability of each recognition result is calculated through the language model, and then the recognition results whose second probability exceeds the predetermined threshold are deleted, so that the unlikely recognition results in the recognition results can be filtered out, and the remaining recognition results are left.
[0135] After filtering the recognition results output by the acoustic model using the target language model, the recognition result with the highest first probability is determined as the target recognition result among the remaining results. The target recognition result here can be determined as the text content contained in the user's target speech. By comparing the target recognition result with the target digital string, the user can be verified.
[0136] The target language model in this application may be a pre-trained language model.
[0137] In the process of determining the target language model, the first language model can be trained with the third training sample to obtain the second language model, wherein the third training sample is a text type sample, and the first language model can be trained with the fourth training sample to obtain the third language model, wherein the fourth training sample is a digital string sample; the digital string sample is a sample composed of numbers, such as 12345. The trained second language model and the third language model are merged into the target language model. The merging process can be performed by interpolation.
[0138] This application does not require too much annotated corpus and computing resources. The acoustic model and language model adaptation method can be used in any ASR-based voice digit string verification product, especially for voice scenarios that lack sufficient annotated corpus.
[0139] Through the present application and the above-mentioned device, in the process of verifying speech, an acoustic model trained successively by using samples obtained by reading aloud text and samples obtained by reading aloud digital strings is used to recognize speech, and the first probability of the recognition result output by the acoustic model and the second probability of the obtained recognition result are used to screen the recognition result, thereby obtaining more accurate speech recognition results and improving the accuracy of the speech verification process.
[0140] As an optional implementation scheme, Fig. 9 As shown, the above device also includes:
[0141] The second sending unit 902 is used to send a second prompt message after determining the target recognition result according to the first probability and the second probability, if the target recognition result is different from the target digital string, wherein the second prompt message is used to prompt that the target object has not passed the verification corresponding to the target digital string.
[0142] In the present application, after sending the first prompt information, the user may also be assigned corresponding permissions, that is, permissions to allow login. Alternatively, after sending the second prompt information, the user is prohibited from logging in.
[0143] Through this application and the above method, corresponding prompt information can be sent on the basis of completing the verification, and the user can be allowed to log in or prohibited to log in, thereby achieving the effect of improving the accuracy of verification of the target object.
[0144] As an optional implementation scheme, the above device also includes:
[0145] A second acquisition unit, used for acquiring a first training sample and a second training sample before inputting the target speech into the acoustic model;
[0146] A first training unit, configured to train the original model using the first training sample until the number of training times reaches a predetermined number of times or the accuracy of the original model reaches a first accuracy;
[0147] The second training unit is used to use the second training sample to train the original model trained with the first training sample to obtain an acoustic model.
[0148] Optionally, the first training sample in the present application may be a general corpus, which includes speech contents of various styles and contents, and the second training sample may be a target corpus, which includes speech contents of digital strings.
[0149] The original model is first trained using the first training sample, and then the original model trained using the first training sample is trained using the second training sample to obtain an acoustic model, which can improve the accuracy of the model.
[0150] As an optional implementation scheme, the above-mentioned calculation unit includes:
[0151] An acquisition module, used to obtain multiple recognition results;
[0152] The calculation module is used to calculate a second probability of each recognition result in the plurality of recognition results using the target language model.
[0153] The target language model in the present application may be a target language model formed by merging the second model and the third model after training the first model with different samples to obtain the second model and the third model.
[0154] As an optional implementation scheme, the acquisition unit further includes:
[0155] A first training module is used to train the first language model using a third training sample to obtain a second language model before using the language model to calculate the second probability of each recognition result in the plurality of recognition results, wherein the third training sample is a text sample;
[0156] A second training module is used to train the first language model using a fourth training sample to obtain a third language model, wherein the fourth training sample is a digital string sample;
[0157] The merging module is used to merge the trained second language model and the third language model into a target language model.
[0158] Through the above steps in the present application, a target language model with stronger recognition ability can be obtained, thereby further improving the verification accuracy of the target object.
[0159] As an optional implementation scheme, the above-mentioned determination unit includes:
[0160] A deletion module, used to delete, from among the multiple recognition results, the recognition results whose second probability is greater than or equal to a predetermined threshold;
[0161] The first determination module is used to determine the recognition result with the highest first probability among the remaining recognition results as the target recognition result.
[0162] That is to say, in this application, an acoustic model is used to recognize the target speech, multiple possibilities are obtained, and then the target speech model is used to screen out unlikely results, and the recognition result with the highest first probability among the remaining results is determined as the target recognition result, so that accurate recognition results can be obtained and the accuracy of verification can be improved.
[0163] As an optional implementation scheme, the first acquisition unit includes:
[0164] A display module, used for displaying the target digital string on a display interface;
[0165] A prompting module, used for prompting the target object to read a target digital string;
[0166] A recording module, used to start recording when prompting the target subject to read a target digital string, and end recording after recording a first duration;
[0167] The second determination module is used to determine the recorded audio as the target speech.
[0168] Through the present application, the accuracy of verifying the target object is improved by displaying the target digital string on the display interface and recording the target voice, and verifying the target object after recognizing the target voice.
[0169] As an optional implementation scheme, the above device also includes:
[0170] A receiving unit, configured to receive a login request from the target object before acquiring the target voice generated by the target object reading aloud the target digital string, wherein the login request is used to request login to the target application;
[0171] The display unit is used to display the target number string and prompt the target object to read the target number aloud.
[0172] That is to say, the present application is applied in the login process, thereby improving the accuracy of verifying the target object in the login process.
[0173] According to another aspect of the embodiment of the present invention, an electronic device for implementing the above-mentioned voice verification method is also provided. Fig.10 As shown, the electronic device includes a memory 1002 and a processor 1004. The memory 1002 stores a computer program, and the processor 1004 is configured to execute the steps in any of the above method embodiments through the computer program.
[0174] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0175] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0176] Obtaining the target speech produced by the target object reading the target digital string;
[0177] Inputting the target speech into the acoustic model, obtaining multiple recognition results of the target speech and a first probability of each recognition result, wherein the acoustic model is a model for recognizing the target speech obtained by training using a first training sample and a second training sample, the first training sample is a sample obtained by reading a text, and the second training sample is a sample obtained by reading a digital string, and the first probability is used to indicate the possibility that the recognition result is the same as the target speech;
[0178] Calculating a second probability of each of the plurality of recognition results, wherein the second probability is used to indicate a semantic understanding degree of the recognition result;
[0179] Determine a target recognition result according to the first probability and the second probability, wherein the target recognition result is a recognition result in which the second probability is less than a predetermined threshold and the first probability is the largest;
[0180] When the target recognition result is the same as the target digital string, a first prompt message is sent, wherein the first prompt message is used to prompt the target object to pass the verification corresponding to the target digital string.
[0181] Alternatively, a person skilled in the art may understand that: Fig.10 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Fig.10 The structure of the electronic device is not limited. Fig.10 More or fewer components (such as network interfaces, etc.) as shown in, or with Fig.10 Different configurations shown.
[0182] Among them, the memory 1002 can be used to store software programs and modules, such as the program instructions / modules corresponding to the voice verification method and device in the embodiment of the present invention. The processor 1004 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002, that is, to implement the above-mentioned voice verification method. The memory 1002 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1002 may further include a memory remotely arranged relative to the processor 1004, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1002 can be specifically, but not limited to, used to store information such as acoustic models, target speech models, and multiple recognition results. As an example, such as Fig.10 As shown, the memory 1002 may include, but is not limited to, the first acquisition unit 802, input unit 804, calculation unit 806, determination unit 808, and first sending unit 810 in the voice verification device. In addition, it may also include, but is not limited to, other module units in the voice verification device, which will not be repeated in this example.
[0183] Optionally, the transmission device 1006 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one example, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 1006 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0184] In addition, the electronic device further includes: a display 1008 for displaying the verification result; and a connection bus 1010 for connecting various module components in the electronic device.
[0185] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0186] Optionally, in this embodiment, the storage medium may be configured to store a computer program for performing the following steps:
[0187] Obtaining the target speech produced by the target object reading the target digital string;
[0188] Inputting the target speech into the acoustic model, obtaining multiple recognition results of the target speech and a first probability of each recognition result, wherein the acoustic model is a model for recognizing the target speech obtained by training using a first training sample and a second training sample, the first training sample is a sample obtained by reading a text, and the second training sample is a sample obtained by reading a digital string, and the first probability is used to indicate the possibility that the recognition result is the same as the target speech;
[0189] Calculating a second probability of each of the plurality of recognition results, wherein the second probability is used to indicate a semantic understanding degree of the recognition result;
[0190] Determine a target recognition result according to the first probability and the second probability, wherein the target recognition result is a recognition result in which the second probability is less than a predetermined threshold and the first probability is the largest;
[0191] When the target recognition result is the same as the target digital string, a first prompt message is sent, wherein the first prompt message is used to prompt the target object to pass the verification corresponding to the target digital string.
[0192] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0193] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0194] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention.
[0195] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0196] In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0197] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0198] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0199] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A voice verification method, characterized in that: include: Obtaining the target speech produced by the target object reading the target digital string; Inputting the target speech into an acoustic model to obtain a plurality of recognition results of the target speech and a first probability of each recognition result, wherein the acoustic model is a model for recognizing the target speech obtained by training using a first training sample and a second training sample, the first training sample is a sample obtained by reading a text, the second training sample is a sample obtained by reading a digital string, and the first probability is used to indicate the possibility that the recognition result is the same as the target speech; Inputting the multiple recognition results into a target language model, and calculating a second probability of each of the multiple recognition results, wherein the target language model is obtained by combining a first language model and a second language model, the first language model is obtained by training an original language model using a text sample, and the second language model is obtained by training the original language model using a digital string sample, and the second probability is used to indicate a semantic understanding degree of the recognition result; Determine a target recognition result according to the first probability and the second probability, wherein the target recognition result is a recognition result in which the second probability is less than a predetermined threshold and the first probability is the largest; In the case where the target recognition result is identical to the target digital string, a first prompt message is sent, wherein the first prompt message is used to prompt the target object to pass the verification corresponding to the target digital string.
2. The method according to claim 1, characterized in that After determining the target recognition result according to the first probability and the second probability, the method further includes: In the case that the target recognition result is different from the target digital string, a second prompt information is sent, wherein the second prompt information is used to prompt that the target object has not passed the verification corresponding to the target digital string.
3. The method according to claim 1, characterized in that Before inputting the target speech into the acoustic model, the method further includes: Acquire the first training sample and the second training sample; Using the first training sample to train the original model until the number of training times reaches a predetermined number or the accuracy of the original model reaches a first accuracy; The original model trained with the first training sample is trained with the second training sample to obtain the acoustic model.
4. The method according to claim 1, characterized in that: Before inputting the plurality of recognition results into the target language model and calculating a second probability of each of the plurality of recognition results, the method further includes: Using a third training sample to train the original language model to obtain the first language model, wherein the third training sample is a text sample; Using a fourth training sample to train the original language model to obtain the second language model, wherein the fourth training sample is a digital string sample; The trained first language model and the second language model are merged into the target language model.
5. The method according to claim 1, characterized in that Determining the target recognition result according to the first probability and the second probability includes: Deleting, from the multiple recognition results, the recognition results whose second probability is greater than or equal to the predetermined threshold; Among the remaining recognition results, the recognition result with the highest first probability is determined as the target recognition result.
6. The method according to any one of claims 1 to 5, characterized in that: The step of obtaining the target speech generated by the target object reading the target digital string includes: Displaying the target digital string on a display interface; prompting the target object to read aloud the target digital string; Start recording when the target subject is prompted to read the target digital string; End the recording after recording the first duration; The recorded audio is determined as the target speech.
7. The method according to any one of claims 1 to 5, characterized in that: Before obtaining the target voice generated by the target object reading the target number string, the method also includes: receiving a login request from the target object, wherein the login request is used to request to log in to a target application, display the target number string, and prompt the target object to read the target number.
8. A voice verification device, characterized in that: include: A first acquisition unit is used to acquire a target voice generated by a target object reading a target digital string; an input unit, configured to input the target speech into an acoustic model, and obtain a plurality of recognition results of the target speech and a first probability of each recognition result, wherein the acoustic model is a model for recognizing the target speech obtained by training using a first training sample and a second training sample, the first training sample is a sample obtained by reading a text, and the second training sample is a sample obtained by reading a digital string, and the first probability is used to indicate the possibility that the recognition result is the same as the target speech; a calculation unit, configured to input the plurality of recognition results into a target language model, and calculate a second probability of each of the plurality of recognition results, wherein the target language model is obtained by combining a first language model and a second language model, the first language model is obtained by training an original language model using a text sample, and the second language model is obtained by training the original language model using a digital string sample, and the second probability is used to indicate a semantic understanding degree of the recognition result; a determination unit, configured to determine a target recognition result according to the first probability and the second probability, wherein the target recognition result is a recognition result in which the second probability is less than a predetermined threshold and the first probability is the largest; A first sending unit is used to send a first prompt message when the target recognition result is the same as the target digital string, wherein the first prompt message is used to prompt the target object to pass the verification corresponding to the target digital string.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method described in any one of claims 1 to 7 is implemented.
10. An electronic device comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the processor executes the computer program, the method described in any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Method and apparatus for constructing speech decoding network in digital speech recognition
CN105869624A
Recognizing the Numeric Language in Natural Spoken Dialogue
US20100049519A1