Voice data recognition method and device, storage medium and electronic device
By automatically processing speech data and performing voiceprint recognition in parallel, and controlling the voiceprint recognition process based on the number of characters, the problem of poor timeliness in voiceprint recognition is solved, thereby improving recognition efficiency and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-03-27
AI Technical Summary
In existing technologies, voiceprint recognition has poor timeliness in home appliance voice interaction scenarios, resulting in response delays and increased computing power consumption.
By performing automatic speech recognition (ASR) and voiceprint recognition in parallel, the voiceprint recognition process is controlled and optimized based on the number of currently recognized characters meeting the set conditions.
It shortens the recognition latency, reduces the waste of computing resources, and improves recognition efficiency and user experience.
Smart Images

Figure CN121747585A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart home technology, and more specifically, to a method, apparatus, storage medium, and electronic device for recognizing voice data. Background Technology
[0002] Currently, in home appliance voice interaction scenarios, traditional solutions typically rely on a voiceprint recognition module to identify the full duration of audio. For example, even if the voiceprint recognition module is sufficient to identify the user, it still needs to wait for the speech recognition or natural language processing of the audio to complete before it can retrieve the recognition results. Alternatively, the voiceprint recognition module may need to wait for speech recognition or natural language processing to complete before triggering voiceprint recognition. These solutions, due to the voiceprint recognition module idling, not only exacerbate response delays but also increase unnecessary computational consumption, resulting in poor timeliness of voiceprint recognition.
[0003] Therefore, no effective solution has yet been proposed to address the technical problem of poor timeliness in voiceprint recognition. Summary of the Invention
[0004] This application provides a method, apparatus, storage medium, and electronic device for recognizing voice data, in order to at least solve the technical problem of poor timeliness in voiceprint recognition in related technologies.
[0005] According to one embodiment of this application, a method for recognizing voice data is provided, comprising: performing automatic speech recognition (ASR) and voiceprint recognition in parallel on voice data from a terminal device; and determining the recognition process of the voiceprint recognition based on the recognition result of the current character when it is determined that the number of current characters recognized by the ASR meets a set condition.
[0006] In an exemplary embodiment, determining the recognition process of the voiceprint recognition based on the recognition result of the current character includes: obtaining the interaction words of the interaction object interacting with the terminal device from the recognition result; stopping ASR of the voice data when it is determined that the interaction words belong to preset words and the number of words in the interaction words is equal to or greater than a preset number; determining the first recognition process of the voiceprint recognition within the first recognition time corresponding to the ASR, wherein the preset words include at least one of the following: intention words for indicating interaction intention, personalized words of the interaction object, and the first recognition process includes one of the following: recognition completed, recognition not completed.
[0007] In an exemplary embodiment, determining the recognition process of the voiceprint recognition based on the recognition result of the current character includes: if it is determined that the interactive word belongs to a preset word and the number of words in the interactive word is less than a preset number, continuing to perform ASR on the voice data.
[0008] In an exemplary embodiment, determining the recognition process of the voiceprint recognition within the recognition time corresponding to the ASR includes: obtaining complete recognized text after the ASR of the voice data is fully completed, and determining a second recognition time used to complete the ASR, wherein the second recognition time is greater than the first recognition time; obtaining a second recognition process of the voiceprint recognition within the second recognition time, wherein the second recognition process includes one of the following: recognition completed, recognition not completed.
[0009] In an exemplary embodiment, the method further includes: when it is determined that both the first identification process and the second identification process have been completed, obtaining a first object identity obtained after the first identification process has been completed and a second object identity obtained after the second identification process has been completed; when it is determined that the first object identity and the second object identity are consistent, determining that the voiceprint recognition result is correct.
[0010] In one exemplary embodiment, the method further includes: acquiring the current recognized text obtained by performing ASR on the voice data; dividing the current recognized text into characters to obtain multiple groups of characters; deleting invalid characters from each group of characters to obtain updated multiple groups of characters; determining updated text based on the updated multiple groups of characters, and acquiring a first count of the characters in the updated text; setting the setting condition using the relationship between the first count and the preset number of characters corresponding to the device type of the terminal device, wherein, if the first count is determined to be equal to the preset number of characters corresponding to the device type, the count of the current characters is determined to satisfy the setting condition.
[0011] In one exemplary embodiment, the method further includes: when it is determined that the first quantity is greater than a preset number of characters corresponding to the device type, obtaining a quantity difference between the first quantity and the preset number of characters; when it is determined that the quantity difference is greater than a preset difference, updating the preset number of characters according to the first quantity.
[0012] In an exemplary embodiment, the method further includes: acquiring historical interaction data of all terminal devices during voice interaction; filtering out unintentional words and phrases without interaction intent and invalid words and phrases unrelated to voice services from the historical interaction data; parsing the unintentional words and phrases and the invalid words and phrases to obtain invalid characters, and establishing an invalid character library based on the invalid characters.
[0013] In one exemplary embodiment, performing automatic speech recognition (ASR) and voiceprint recognition in parallel on voice data from a terminal device includes: splitting the voice data into multiple voice segments; performing ASR on the voice data, while simultaneously calling at least one voiceprint recognition thread to perform voiceprint recognition on the multiple voice segments in parallel.
[0014] According to another aspect of the embodiments of this application, a voice data recognition device is also provided, comprising: an obtaining module, configured to perform automatic speech recognition (ASR) and voiceprint recognition in parallel on voice data from a terminal device; and a determining module, configured to determine the recognition process of the voiceprint recognition based on the recognition result of the current character when it is determined that the number of current characters recognized by the ASR meets a set condition.
[0015] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer-readable storage medium, and the computer program is configured to execute the above-described voice data recognition method when running.
[0016] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the above-described voice data recognition method through the computer program.
[0017] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the above-described method for recognizing voice data.
[0018] In this embodiment, automatic speech recognition (ASR) and voiceprint recognition are performed in parallel on voice data from a terminal device. When the number of characters recognized by ASR meets a set condition, the current stage of voiceprint recognition is determined based on the recognition result of the current character. By processing voice data in parallel, i.e., performing ASR and voiceprint recognition simultaneously, the accumulated latency in the traditional serial mode can be reduced. Furthermore, by controlling the voiceprint recognition process through set conditions—that is, when the number of recognized characters reaches a preset threshold, the decision to continue voiceprint recognition is based on the recognition result of the current character—this technical solution solves the problem of poor timeliness in voiceprint recognition in related technologies. It not only reduces unnecessary waste of computing resources but also accelerates the entire recognition process, improves the efficiency of voice data recognition, and enhances the user experience. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of the hardware environment for a voice data recognition method according to an embodiment of this application;
[0022] Figure 2 This is a flowchart of a voice data recognition method according to an embodiment of this application;
[0023] Figure 3 This is a schematic diagram illustrating the setting of the number of characters according to an embodiment of this application;
[0024] Figure 4 This is a schematic diagram of characters recognized by ASR streaming according to an embodiment of this application;
[0025] Figure 5 This is a schematic diagram (a) illustrating the principle of streaming dual-channel recognition according to an embodiment of this application.
[0026] Figure 6 This is a schematic diagram (II) illustrating the principle of streaming dual-channel recognition according to an embodiment of this application.
[0027] Figure 7 This is a schematic diagram illustrating the principle of a voice data recognition method according to an embodiment of this application;
[0028] Figure 8 This is a structural block diagram of a voice data recognition device according to an embodiment of this application. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] According to one aspect of the embodiments of this application, a method for recognizing voice data is provided. This voice data recognition method is widely used in whole-house intelligent digital control application scenarios such as smart homes, smart home ecosystems, and intelligence house ecosystems. Optionally, in this embodiment, the above-mentioned voice data recognition method can be applied to, for example... Figure 1 The hardware environment shown consists of terminal device 102 and server 104. For example... Figure 1 As shown, server 104 is connected to terminal device 102 via a network and can be used to provide services (such as application services) to the terminal or clients installed on the terminal. A database can be set up on the server or independently of the server to provide data storage services for server 104. Cloud computing and / or edge computing services can be configured on the server or independently of the server to provide data processing services for server 104.
[0032] The aforementioned network may include, but is not limited to, at least one of the following: wired network, wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: wide area network, metropolitan area network, local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: Wi-Fi (Wireless Fidelity), Bluetooth. The terminal device 102 may not be limited to PC, mobile phone, tablet computer, smart air conditioner, smart range hood, smart refrigerator, smart oven, smart stove, smart washing machine, smart water heater, smart washing equipment, smart dishwasher, smart projector, smart TV, smart clothes rack, smart curtains, smart audio-visual equipment, smart socket, smart speaker, smart speaker box, smart fresh air equipment, smart kitchen and bathroom equipment, smart bathroom equipment, smart robot vacuum cleaner, smart window cleaning robot, smart mopping robot, smart air purifier, smart steam oven, smart microwave oven, smart water heater, smart air purifier, smart water dispenser, smart door lock, etc.
[0033] This embodiment provides a method for recognizing voice data, applied to the aforementioned terminal device. Figure 2 This is a flowchart of a speech data recognition method according to an embodiment of this application, which includes the following steps:
[0034] Step S202: Perform automatic speech recognition (ASR) and voiceprint recognition in parallel on the voice data from the terminal device;
[0035] Step S204: If the number of current characters identified by the ASR meets the set conditions, determine the recognition process of the voiceprint recognition based on the recognition result of the current characters.
[0036] Through the above steps, Automatic Speech Recognition (ASR) and Voiceprint Recognition are performed in parallel on the voice data from the terminal device. When the number of characters recognized by ASR meets a set condition, the current stage of voiceprint recognition is determined based on the recognition result of the current character. By processing voice data in parallel, i.e., performing ASR and voiceprint recognition simultaneously, the accumulated latency in the traditional serial mode can be reduced. Furthermore, by controlling the voiceprint recognition process through set conditions—that is, when the number of recognized characters reaches a preset threshold, the decision to continue voiceprint recognition is based on the recognition result of the current character—this technical solution solves the technical problem of poor timeliness in voiceprint recognition in related technologies. It not only reduces unnecessary waste of computing resources but also accelerates the entire recognition process, improves the efficiency of voice data recognition, and enhances the user experience.
[0037] In an exemplary embodiment, the process of determining the recognition process of the voiceprint recognition based on the recognition result of the current character includes: obtaining the interaction words of the interaction object interacting with the terminal device from the recognition result; stopping ASR of the voice data when it is determined that the interaction words belong to preset words and the number of words of the interaction words is equal to or greater than a preset number; determining the first recognition process of the voiceprint recognition within a first recognition time corresponding to the ASR, wherein the preset words include at least one of the following: intent words used to indicate interaction intent, personalized words of the interaction object, and the first recognition process includes one of the following: recognition completed, recognition not completed. This embodiment further refines the control strategy of the voiceprint recognition process. When the recognized words belong to preset business-related words and reach a preset number, the subsequent processing of ASR can be stopped immediately. This means that the device can respond faster, reduce resource waste, and improve the timeliness and accuracy of voiceprint recognition in scenarios that require rapid identification of user identity, while also reducing the risk of misidentification.
[0038] In an exemplary embodiment, the scheme for determining the current stage of the voiceprint recognition process based on the recognition result of the current character includes: if it is determined that the interactive word belongs to a preset word and the number of words in the interactive word is less than a preset number, continuing to perform ASR on the speech data. In this embodiment, when the recognized words belong to the preset word but the number is insufficient, ASR processing continues until the condition is met, thereby increasing the probability of obtaining effective recognition results. Even in situations where the user speaks slowly or unclearly, it can avoid the possibility of missing important information due to premature termination of ASR, thus balancing efficiency and comprehensiveness of recognition and improving the accuracy of voiceprint recognition.
[0039] In one exemplary embodiment, a scheme for determining the voiceprint recognition progress within the recognition time corresponding to the ASR includes: obtaining complete recognized text after the ASR of the voice data is fully completed, and determining a second recognition time used to complete the ASR, wherein the second recognition time is greater than the first recognition time; obtaining a second recognition progress of the voiceprint recognition within the second recognition time, the second recognition progress including one of the following: recognition completed, recognition incomplete. This embodiment focuses on determining the voiceprint recognition progress after the ASR is fully completed. After determining that the ASR is completed and the second recognition time of the ASR is greater than the first recognition time during the ASR recognition process, the final state of the voiceprint recognition is confirmed based on the recognition progress within the second recognition time. This method improves the flexibility and accuracy of voiceprint recognition.
[0040] In an exemplary embodiment, if both the first and second identification processes are determined to be complete, the first object identity obtained after the first identification process and the second object identity obtained after the second identification process are completed can be acquired. If the first object identity and the second object identity are determined to be consistent, the voiceprint recognition result is determined to be correct. This embodiment effectively verifies the accuracy of voiceprint recognition by comparing the consistency of ASR and voiceprint recognition results at different time points. That is, if the identified object identities are consistent when both ASR and voiceprint recognition are complete, the voiceprint recognition is considered correct. This dual-check mechanism proposed in this embodiment improves the credibility of voiceprint recognition, enhances the accuracy of user identity verification, and provides a solid foundation for subsequent personalized services.
[0041] In an exemplary embodiment, further, obtain the current recognized text obtained by performing ASR on the speech data; perform character division on the current recognized text to obtain multiple groups of characters; respectively perform a deletion operation on the invalid characters in each group of characters in the multiple groups of characters to obtain updated multiple groups of characters; determine an updated text according to the updated multiple groups of characters, and obtain a first quantity for counting the characters in the updated text one by one; set the set condition by using the quantity relationship between the first quantity and the preset character quantity corresponding to the device type of the terminal device, wherein, when it is determined that the first quantity is equal to the preset character quantity corresponding to the device type, it is determined that the quantity of the current characters meets the set condition. This embodiment proposes a method for optimizing the ASR process. Through the processes of character division, invalid character filtering, and recognizing the character quantity of the updated text, a recognition strategy that dynamically adapts to different device types is realized. It can not only adjust the set condition according to the device type, but also improve the adaptability and robustness to different audio qualities, and achieve the balance between recognition speed and accuracy.
[0042] Optionally, as Figure 3 shown, different device types correspond to different character quantities. For example, an air conditioner corresponds to 3 characters, a speaker corresponds to 4 characters, and a refrigerator corresponds to 5 characters. If the device type is not specified, the default character quantity, that is, 4 characters, can be adopted. Among them, the result of ASR streaming recognition is, for example, Figure 4 shown. If it is set that an air conditioner corresponds to 3 characters, the situations where the recognized number of characters is less than 3 include, for example, recognizing "hit" or "turn on", and the situations where the recognized number of characters is greater than or equal to 3 include, for example, "turn on the air conditioner", etc.
[0043] In an exemplary embodiment, when it is determined that the first quantity is greater than the preset character quantity corresponding to the device type, obtain the quantity difference between the first quantity and the preset character quantity; when it is determined that the quantity difference is greater than the preset difference, update the preset character quantity according to the first quantity. This embodiment introduces an intelligent adjustment mechanism for the preset character quantity. When the recognized character quantity exceeds the preset quantity, the quantity difference will be calculated, and the preset character quantity will be updated when the difference is large. This adjustment strategy can self-optimize based on the actual usage situation, reduce the inadaptability that may be brought by the fixed threshold, and can improve the recognition processing speed and enhance the overall recognition performance when facing diverse user groups and speech habits.
[0044] In one exemplary embodiment, historical interaction data of all terminal devices during voice interaction can be acquired; meaningless phrases without interaction intent and invalid phrases unrelated to voice services can be filtered out from the historical interaction data; the meaningless phrases and invalid phrases can be parsed to obtain invalid characters, and an invalid character library can be established based on the invalid characters. This embodiment constructs an invalid character library. By analyzing historical interaction data, meaningless phrases without interaction intent and invalid phrases are filtered and parsed out, and invalid characters are extracted from them. In this way, meaningless characters can be excluded in advance during the recognition process, eliminating unnecessary repeated calculations and storage, thus improving the recognition efficiency and accuracy of ASR. In addition, the invalid character library can be dynamically updated, thereby ensuring the accuracy and stability of the invalid character library.
[0045] Optionally, the above meaningless words and phrases, for example Figure 3 The "ah ah ah" etc. shown.
[0046] In an exemplary embodiment, the process of performing automatic speech recognition (ASR) and voiceprint recognition in parallel on voice data from a terminal device includes: splitting the voice data into multiple voice segments; performing ASR on the voice data; and simultaneously invoking at least one voiceprint recognition thread to perform voiceprint recognition on the multiple voice segments in parallel. This embodiment proposes an efficient data processing mode by splitting the voice data into multiple segments and simultaneously performing parallel processing of ASR and voiceprint recognition. The parallel voiceprint recognition thread improves the concurrency capability of recognition, accelerates the recognition speed, and further improves recognition efficiency.
[0047] To better understand the process of the above-mentioned voice data recognition method, the flow of the above-mentioned voice data recognition method will be described below in conjunction with optional embodiments, but it is not intended to limit the technical solution of the embodiments of this application.
[0048] Optionally, in one embodiment, combined with Figure 5 The dual-stream channel processing shown illustrates the parallel execution of Automatic Speech Recognition (ASR) and voiceprint recognition. For example... Figure 5 As shown, after the audio is uploaded, a dual-thread system is initialized. The ASR thread continuously acquires audio and performs ASR recognition, while the voiceprint recognition thread continuously acquires audio and performs voiceprint recognition. Then, the next step is performed based on the recognition results. The details of the next step based on the recognition results are as follows: Figure 6As shown, when the number of characters recognized by ASR is greater than or equal to the threshold number of characters, and the database determines that the recognized characters contain valid words such as "turn on the air conditioner," a termination message is immediately sent to the voiceprint thread, while ASR recognition continues until the audio upload is complete. Otherwise, ASR continues to perform ASR recognition on newly uploaded audio. After receiving the termination message, the voiceprint recognition thread only performs feature extraction on the first N milliseconds of received audio, skips the remaining audio processing, and then ends the voiceprint recognition process. Figure 7 As shown, the database contains meaningless corpus. This corpus can be used to analyze the identified characters to determine whether they contain valid words, such as excluding invalid words like "ah ah ah". Furthermore, combined with... Figure 7 right Figure 6 The voiceprint recognition part is explained below. When the number of characters in the ASR streaming recognition reaches the threshold and contains valid words, the ASR thread immediately sends a termination symbol to the voiceprint recognition thread, and the voiceprint recognition thread ends the voiceprint recognition process. Since voiceprint recognition and ASR recognition do not interfere with each other, if the ASR recognition result reaches the threshold number of characters and contains valid words, and a termination message is sent to the voiceprint thread in advance, the voiceprint result can be returned partially before the ASR is completely finished, for use by subsequent business logic. For example, the ASR recognition result of the corpus is "Turn on the air conditioner and set it to 26 degrees," but the actual voiceprint recognition part is the preceding "Turn on the air conditioner" part. Through this embodiment, invalid audio processing can be terminated in advance based on the ASR recognition result, reducing the computational load of the voiceprint recognition thread and improving recognition efficiency.
[0049] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of this application.
[0050] Figure 8 This is a structural block diagram of a voice data recognition device according to an embodiment of this application; as shown below. Figure 8 As shown, it includes:
[0051] Module 82 is obtained, which is used to perform automatic speech recognition (ASR) and voiceprint recognition in parallel on the voice data from the terminal device;
[0052] The determination module 84 is used to determine the recognition process of the voiceprint recognition based on the recognition result of the current character when the number of current characters recognized by the ASR meets the set conditions.
[0053] The aforementioned device performs Automatic Speech Recognition (ASR) and Voiceprint Recognition in parallel on voice data from the terminal device. When the number of characters recognized by ASR meets a preset condition, the recognition process of the voiceprint is determined based on the recognition result of the current character. By processing voice data in parallel, i.e., performing ASR and voiceprint recognition simultaneously, the accumulated latency in the traditional serial mode can be reduced. Furthermore, by controlling the voiceprint recognition process through preset conditions—that is, when the number of recognized characters reaches a preset threshold, the recognition result of the current character determines whether voiceprint recognition should continue—this technical solution solves the problem of poor timeliness in voiceprint recognition in related technologies. It not only reduces unnecessary waste of computing resources but also accelerates the entire recognition process, improves the efficiency of voice data recognition, and enhances the user experience.
[0054] In an exemplary embodiment, the determining module is further configured to: obtain from the recognition result the interaction words of the interaction object interacting with the terminal device; stop performing ASR on the voice data when it is determined that the interaction words belong to preset words and the number of words of the interaction words is equal to or greater than a preset number; determine the first recognition process of the voiceprint recognition within the first recognition time corresponding to the ASR, wherein the preset words include at least one of the following: intention words for indicating interaction intention, personalized words of the interaction object, and the first recognition process includes one of the following: recognition completed, recognition not completed.
[0055] In an exemplary embodiment, the determining module is further configured to: continue performing ASR on the voice data if it is determined that the interactive word belongs to a preset word and the number of words in the interactive word is less than a preset number.
[0056] In an exemplary embodiment, the determining module is further configured to: obtain complete recognized text after the ASR of the voice data is fully completed, and determine a second recognition time used to complete the ASR, wherein the second recognition time is greater than the first recognition time; obtain a second recognition process of the voiceprint recognition within the second recognition time, wherein the second recognition process includes one of the following: recognition completed, recognition not completed.
[0057] In an exemplary embodiment, the determining module is further configured to: when it is determined that both the first identification process and the second identification process have been completed, obtain the first object identity obtained after the first identification process has been completed and the second object identity obtained after the second identification process has been completed; and when it is determined that the first object identity and the second object identity are consistent, determine that the result of the voiceprint recognition is correct.
[0058] In an exemplary embodiment, the apparatus is further configured to: acquire current recognized text obtained by performing ASR on the voice data; divide the current recognized text into characters to obtain multiple groups of characters; delete invalid characters in each of the multiple groups of characters to obtain updated multiple groups of characters; determine updated text based on the updated multiple groups of characters, and acquire a first count of characters in the updated text; set the setting condition using the relationship between the first count and the preset number of characters corresponding to the device type of the terminal device, wherein, if the first count is determined to be equal to the preset number of characters corresponding to the device type, the count of the current characters is determined to satisfy the setting condition.
[0059] In an exemplary embodiment, the apparatus is further configured to: when it is determined that the first quantity is greater than a preset number of characters corresponding to the device type, obtain a quantity difference between the first quantity and the preset number of characters; when it is determined that the quantity difference is greater than a preset difference, update the preset number of characters according to the first quantity.
[0060] In an exemplary embodiment, the apparatus is further configured to: acquire historical interaction data of all terminal devices during voice interaction; filter out unintentional words and phrases without interaction intent and invalid words and phrases unrelated to voice services from the historical interaction data; parse the unintentional words and phrases and the invalid words and phrases to obtain invalid characters, and establish an invalid character library based on the invalid characters.
[0061] In one exemplary embodiment, the obtaining module is further configured to: split the speech data to obtain multiple speech segments; perform ASR on the speech data, and simultaneously call at least one voiceprint recognition thread to perform voiceprint recognition on the multiple speech segments in parallel.
[0062] Embodiments of this application also provide a storage medium including a stored program, wherein the program executes any of the methods described above when it is run.
[0063] Optionally, in this embodiment, the storage medium may be configured to store program code for performing the following steps:
[0064] S1 performs automatic speech recognition (ASR) and voiceprint recognition in parallel on the voice data from the terminal device;
[0065] S2, if it is determined that the number of current characters identified by the ASR meets the set conditions, the recognition process of the voiceprint recognition is determined according to the recognition result of the current characters.
[0066] Embodiments of this application also provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0067] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0068] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0069] S1 performs automatic speech recognition (ASR) and voiceprint recognition in parallel on the voice data from the terminal device;
[0070] S2, if it is determined that the number of current characters identified by the ASR meets the set conditions, the recognition process of the voiceprint recognition is determined according to the recognition result of the current characters.
[0071] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0072] Optionally, embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0073] Optionally, embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0074] Optionally, embodiments of this application also provide a computer program that includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.
[0075] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated here.
[0076] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0077] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A method for recognizing voice data, characterized in that, include: Automatic speech recognition (ASR) and voiceprint recognition are performed in parallel on voice data from terminal devices. If the number of current characters identified by the ASR meets the set conditions, the recognition process of the voiceprint recognition is determined based on the recognition result of the current characters.
2. The method for recognizing voice data according to claim 1, characterized in that, Determining the current stage of the voiceprint recognition process based on the recognition result of the current character includes: The interaction words of the interaction object that interacts with the terminal device are obtained from the recognition results; If it is determined that the interactive word belongs to a preset word and the number of words in the interactive word is equal to or greater than the preset number, stop performing ASR on the voice data; The voiceprint recognition process is determined within the first recognition time corresponding to the ASR, wherein the preset words include at least one of the following: intent words used to indicate interaction intent, personalized words of the interaction object, and the first recognition process includes one of the following: recognition completed, recognition not completed.
3. The method for recognizing voice data according to claim 2, characterized in that, Determining the current stage of the voiceprint recognition process based on the recognition result of the current character includes: If it is determined that the interactive word belongs to a preset word and the number of words in the interactive word is less than the preset number, ASR will continue to be performed on the voice data.
4. The method for recognizing voice data according to claim 2, characterized in that, Determining the recognition process of the voiceprint recognition within the recognition time corresponding to the ASR includes: After the ASR of the voice data is fully completed, the complete recognized text is obtained, and the second recognition time used to complete the ASR is determined, wherein the second recognition time is greater than the first recognition time. The second recognition process of the voiceprint recognition within the second recognition time is obtained, and the second recognition process includes one of the following: recognition completed, recognition not completed.
5. The method for recognizing voice data according to claim 4, characterized in that, The method further includes: If it is determined that both the first identification process and the second identification process have been completed, the first object identity obtained after the first identification process is completed and the second object identity obtained after the second identification process is completed are obtained. If the identity of the first object is determined to be consistent with the identity of the second object, the result of the voiceprint recognition is determined to be correct.
6. The method for recognizing voice data according to claim 1, characterized in that, The method further includes: Obtain the current recognized text obtained by performing ASR on the speech data; The currently recognized text is divided into multiple groups of characters. The invalid characters in each of the multiple sets of characters are deleted to obtain the updated multiple sets of characters; The updated text is determined based on the updated sets of characters, and a first count is obtained by counting each character in the updated text. The setting condition is set by using the quantitative relationship between the first quantity and the preset number of characters corresponding to the device type of the terminal device, wherein, when it is determined that the first quantity is equal to the preset number of characters corresponding to the device type, it is determined that the current number of characters satisfies the setting condition.
7. The method for recognizing voice data according to claim 6, characterized in that, The method further includes: If it is determined that the first quantity is greater than the preset number of characters corresponding to the device type, the difference between the first quantity and the preset number of characters is obtained; If the quantity difference is determined to be greater than a preset difference, the preset character quantity is updated according to the first quantity.
8. The method for recognizing voice data according to claim 6, characterized in that, The method further includes: Acquire historical interaction data of all terminal devices during the voice interaction process; Filter out meaningless words and phrases that do not have an interactive intent and invalid words and phrases that are unrelated to voice services from the historical interaction data; The meaningless and invalid words and phrases are parsed to obtain invalid characters, and an invalid character library is established based on the invalid characters.
9. The method for recognizing voice data according to claim 1, characterized in that, Automatic speech recognition (ASR) and voiceprint recognition are performed in parallel on voice data from terminal devices, including: The voice data is split into multiple voice segments; The speech data is subjected to ASR, and at the same time, at least one voiceprint recognition thread is invoked to perform voiceprint recognition on the multiple speech segments in parallel.
10. A voice data recognition device, characterized in that, include: The module is used to perform automatic speech recognition (ASR) and voiceprint recognition in parallel on voice data from terminal devices. The determination module is used to determine the current recognition process of the voiceprint recognition based on the recognition result of the current character when the number of current characters recognized by the ASR meets the set conditions.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program, when executed, performs the method described in any one of claims 1 to 9.
12. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 9 through the computer program.