Speech recognition method, apparatus, device, and storage medium
By setting the recognition time within the entire audio track, the speech recognition function is automatically activated, enabling automated continuous recognition of the entire audio track. This solves the problems of low efficiency and insufficient accuracy in speech recognition testing in existing technologies, thereby improving recognition efficiency and accuracy.
Patent Information
- Application Number
- CN202210717484.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2042-06-23
AI Technical Summary
Existing technologies for speech recognition testing are inefficient, require complex manual operation, are subject to significant noise interference in the testing environment, make it difficult to achieve automated recognition of the entire audio track, and cannot avoid missed recognition and misoperation.
By determining the start time of the first query statement in the entire audio track and setting the recognition time in advance, the speech recognition function is automatically activated, enabling automated continuous recognition of all query statements in the entire audio track, including simplex and full-duplex speech recognition models, avoiding manual segmentation and wake-up word interference.
It improves the efficiency of speech recognition testing, reduces labor costs, ensures the accuracy and stability of recognition, reduces testing time and storage maintenance costs, and avoids missed recognition and misoperation.
Smart Images

Figure CN115240642B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to the fields of artificial intelligence, natural language processing, speech technology, cloud computing, and deep learning. Background Technology
[0002] With the rapid development of speech recognition technology, the effectiveness of speech recognition has become particularly important for smart terminals. Speech recognition performance not only needs to reflect the intelligence of hardware products, but also needs to focus on the user's actual experience. Only with faster and more accurate interaction can a product cater to the market and achieve market dominance. Summary of the Invention
[0003] This disclosure provides a method, apparatus, device, and storage medium for speech recognition.
[0004] According to one aspect of this disclosure, a speech recognition method is provided, comprising:
[0005] The first recognition time is determined based on the start time of the first query statement in the entire audio track; where the entire audio track includes N query statements;
[0006] When the playback time of the entire audio track reaches the first recognition time, perform speech recognition on N query statements respectively; and
[0007] Determine the test identification results for N query statements.
[0008] According to another aspect of this disclosure, a speech recognition apparatus is provided, comprising:
[0009] The first determining module is used to determine the first recognition time based on the start time of the first query statement in the whole audio track; wherein, the whole audio track includes N query statements;
[0010] The recognition module is used to perform speech recognition on N query statements when the playback time of the entire audio track reaches the first recognition time; and
[0011] The second determination module is used to determine the test recognition results of N query statements.
[0012] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0013] At least one processor; and
[0014] The memory is communicatively connected to the at least one processor; wherein,
[0015] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the methods of any embodiment of the present disclosure.
[0016] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform a method according to any embodiment of this disclosure.
[0017] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a method according to any embodiment of this disclosure.
[0018] According to the scheme disclosed herein, automated continuous speech recognition of each query statement in an entire audio track can be achieved, thereby improving the efficiency of speech recognition testing.
[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0020] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0021] Figure 1 This is a schematic flowchart of a speech recognition method according to an embodiment of the present disclosure;
[0022] Figure 2 This is a schematic diagram illustrating an application scenario of the speech recognition method according to embodiments of the present disclosure;
[0023] Figure 3 This is a schematic flowchart of a speech recognition method according to another embodiment of the present disclosure;
[0024] Figure 4 This is a schematic flowchart of a speech recognition method according to another embodiment of the present disclosure;
[0025] Figure 5 This is a schematic diagram of the structure of a speech recognition device according to an embodiment of the present disclosure;
[0026] Figure 6 This is a block diagram of an electronic device used to implement the speech recognition method of the embodiments of this disclosure. Detailed Implementation
[0027] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0028] According to embodiments of this disclosure, such as Figure 1 As shown, a speech recognition method is provided, which may include the following steps:
[0029] S101: Determine the first recognition time based on the start time of the first query statement in the entire audio track. The entire audio track contains N query statements.
[0030] S102: When the playback time of the entire audio track reaches the first recognition time, perform speech recognition on each of the N query statements.
[0031] S103: Determine the test identification results of N query statements.
[0032] Based on the above steps of the embodiments of this disclosure, it should be noted that:
[0033] A full-track audio file can be understood as a single, complete audio file. It contains multiple queries arranged sequentially. Each query can have a different duration. The semantic content of each query can belong to the same domain; for example, each query in the full-track audio file could be about the weather. Alternatively, the semantic content of each query can belong to different domains; for example, some queries might be about music, while others might be about vehicle navigation. The file format of the full-track audio file can be selected and adjusted as needed, and no specific limitations are imposed here. For example, the full-track audio file can be a multi-channel interleaved audio file in PCM or WAV format.
[0034] The first query statement can be understood as the first query statement played during the playback of the entire audio track. The start time of the first query statement can be understood as the playback time of the first audio frame corresponding to the first query statement in the audio stream of the entire audio track. For example, if the total playback time of the entire audio track is 5 minutes, and the first query statement starts playing from the 10th second of the entire audio track, then the 10th second is the start time of the first query statement.
[0035] The first recognition time can be understood as the moment when the speech recognition function is activated to perform speech recognition on the entire audio track. The first recognition time is earlier than the start time of the first query statement by K milliseconds (ms), where K is a positive integer. The specific value of K can be selected and adjusted as needed; for example, the first recognition time can be 100ms, 200ms, etc., before the start time of the first query statement.
[0036] When the playback time of the entire audio track reaches the first recognition time, speech recognition is performed on each of the N query statements. This can be understood as follows: when the audio data output from the entire audio track corresponds to the audio data at the first recognition time, the speech recognition function is automatically activated. The speech recognition function then performs speech recognition on each query statement in the entire audio track sequentially, starting from the first recognition time. Alternatively, it can be understood as follows: when the audio data output from the entire audio track corresponds to the audio data at the first recognition time, the speech recognition function is automatically activated. The speech recognition function enters the speech recognition state from the first recognition time, and then begins speech recognition on each query statement in the entire audio track sequentially when the audio data output from the audio stream corresponds to the start time of the first query statement.
[0037] The computational module implementing the speech recognition function in this embodiment can be a speech recognition module on any device equipped with speech recognition capabilities. For example, it can be a speech recognition module in a smart speaker, smart home product, in-vehicle smart cockpit, smart robot, or mobile terminal. The speech recognition module can recognize each query statement in the entire audio track using a pre-trained machine learning model or other existing speech recognition technologies.
[0038] The test results for query statement recognition can be understood as the semantics and content of each recognized query statement.
[0039] According to the solution of this disclosure, by determining the first recognition time of the first query statement, automated continuous speech recognition of each query statement in the audio stream of the entire audio track can be achieved, improving the efficiency of speech recognition testing. This effectively reduces the labor cost of speech recognition verification testing in the prior art, which involves recording speech sentence by sentence by pressing buttons, accelerates the start-up and testing time of automated speech recognition testing, ensures the stability of problem reproduction, and makes problem localization faster and more effective.
[0040] The speech recognition method of any embodiment of this disclosure can be applied to an automated verification system based on simplex and full-duplex speech recognition. It receives full-track audio by integrating the interface exposed by the speech SDK (Software Development Kit) to read audio files, and further tests the speech recognition performance of devices equipped with speech recognition capabilities. The interface of the speech SDK supports automated testing for evaluating the overall link and model performance of both simplex and full-duplex speech recognition, enhancing the reusability of test sets (containing multiple full-track audio tracks) and improving scalability. The device can be a smart speaker, smart home product, in-vehicle smart cockpit, smart robot, mobile terminal, etc. The performance test of the speech recognition function can be a test of the performance of simplex speech recognition or a test of the performance of full-duplex speech recognition. Simplex speech recognition can be understood as follows: before each query is sent, the device's speech recognition function needs to be activated by a wake word, button press, or facial recognition. That is, after each query is recognized, the device's speech recognition function is disabled by default and needs to be activated to receive and recognize queries again.
[0041] In one example, the speech recognition scheme of the above embodiments of this disclosure can be applied to speech recognition models capable of achieving full-duplex speech recognition. By determining the first recognition time based on the start time of the first query statement in the entire audio track, automated recognition of all query statements in the entire audio track can be achieved.
[0042] In one embodiment, the speech recognition method provided in this disclosure includes steps S101 to S103, and may further include the step:
[0043] Based on the start time of the second to Nth query statements in the whole audio track, determine the second recognition time corresponding to the second to Nth query statements respectively.
[0044] Based on the above steps of the embodiments of this disclosure, it should be noted that:
[0045] The second identification time for each query statement in the second to Nth query statements can be exactly the same, partially the same and partially different, or all different; no specific limitation is made here. It can be determined based on the blank time interval between two consecutive query statements.
[0046] The second recognition time can be understood as the moment when the speech recognition function is activated to perform speech recognition on any one of the 2nd to Nth query statements. The second recognition time is earlier than the start time L milliseconds of the 2nd to Nth query statements, where L is a positive integer. The specific value of L can be selected and adjusted as needed. For example, the second recognition time can be 50ms, 100ms, 200ms, etc., before the start time of the first query statement.
[0047] According to the scheme of this disclosure embodiment, the second recognition time of the second to Nth query statements can be determined, which can reserve a recognition time in advance for the speech recognition of each query statement, avoid the situation of missing recognition of one or more audio frames before the query statement, and achieve accurate recognition of each query statement.
[0048] In one example, endpoint detection technology can be used to determine the start time of each query in the entire audio track. Machine learning can be used to label the start time of each query in the entire audio track. The specific method for determining the start time of each query in the entire audio track can be selected and adjusted as needed, and is not specifically limited here.
[0049] In one example, the speech recognition scheme of the above embodiments of this disclosure can be applied to a speech recognition model capable of performing simplex speech recognition. By determining the second recognition time corresponding to each of the second to Nth query statements based on their start times in the entire audio track, automated recognition of all query statements in the entire audio track can be achieved.
[0050] The speech recognition method provided in the embodiments of this disclosure can be applied to, for example... Figure 2 Within the scene framework shown. Figure 2 In this document, 10 represents a terminal, 20 represents a server, and 30 represents a distributed computer system. The speech recognition method of this disclosure can be jointly executed by terminal 10, server 20, or distributed computer system 30, or by one or more of these three systems. Terminal 10 can be used to report / send a full-track audio track labeled with the start time of a query statement to server 20 or distributed computer system 30. After completing the speech recognition method of this disclosure, server 20 or distributed computer system 30 can feed back the test recognition results of N query statements (or, the speech recognition performance determined by comparing the test recognition results of N query statements with the corresponding labeled recognition results) to terminal 10. Terminal 10, server 20, or distributed computer system 30 can also independently complete the speech recognition method of this disclosure.
[0051] In one embodiment, the speech recognition method provided by this disclosure includes steps S101 to S103, wherein step S102: when the playback time of the entire audio track reaches the first recognition time, speech recognition is performed on N query statements respectively, which may further include:
[0052] S1021: When the playback time of the entire audio track reaches the first recognition time, perform speech recognition on the first query statement.
[0053] S1021: When the playback time of the entire audio track reaches the second recognition time corresponding to the second to Nth query statements respectively, perform speech recognition on the second to Nth query statements respectively.
[0054] Based on the above steps of the embodiments of this disclosure, it should be noted that:
[0055] When the playback time of the entire audio track reaches the second recognition time corresponding to the 2nd to Nth query statements, speech recognition is performed on each of the 2nd to Nth query statements. This can be understood as the speech recognition function automatically activating when the audio data output from the entire audio track corresponds to the audio data of the second recognition time of the 2nd to Nth query statements. The speech recognition function then performs speech recognition on each query statement starting from its second recognition time. Alternatively, it can be understood as the speech recognition function automatically activating when the audio data output from the entire audio track corresponds to the audio data of the second recognition time of the 2nd to Nth query statements. The speech recognition function enters speech recognition mode starting from the second recognition time, and then begins speech recognition on each query statement sequentially when the audio data output from the audio stream corresponds to the start time of each query statement.
[0056] According to the scheme of this disclosure embodiment, by determining the second recognition time of the 2nd to Nth query statements, automated continuous speech recognition of each query statement in the audio stream of the entire audio track can be achieved, improving the efficiency of speech recognition testing. This effectively reduces the labor cost of speech recognition verification testing in the prior art, which involves recording speech sentence by sentence by pressing buttons, speeds up the start-up and testing time of automated speech recognition testing, ensures the stability of problem reproduction, and makes problem localization faster and more effective. Simultaneously, determining the second recognition time of the 2nd to Nth query statements allows for a pre-allocated recognition time for each query statement, avoiding the possibility of missing recognition of one or more audio frames preceding the query statement, thus achieving accurate recognition of each query statement.
[0057] In one example, the speech recognition scheme of the above embodiments of this disclosure can be applied to a speech recognition model capable of performing simplex speech recognition. By determining the second recognition time corresponding to each of the second to Nth query statements based on their start times in the entire audio track, automated recognition of all query statements in the entire audio track can be achieved.
[0058] In one embodiment, the speech recognition method provided by this disclosure includes steps S101 to S103, wherein step S1021: when the playback time of the entire audio track reaches the second recognition time corresponding to the second to Nth query statements respectively, speech recognition is performed on the second to Nth query statements respectively, which may further include:
[0059] If the playback time of the entire audio track reaches the second recognition time corresponding to the Mth query statement, determine whether the (M-1)th query statement has been successfully recognized. The Mth query statement can be any one of the 2nd to Nth query statements.
[0060] If the speech recognition of the (M-1)th query statement is completed, then the speech recognition of the Mth query statement is performed.
[0061] Based on the above steps of the embodiments of this disclosure, it should be noted that:
[0062] M-1 query statements can be understood as being arranged in playback order, with the query statement preceding the Mth query statement.
[0063] According to the solution of this disclosure embodiment, it can be ensured that every query statement in the entire audio track can be recognized, avoiding the situation where a query statement is missed.
[0064] In one embodiment, the speech recognition method provided by this disclosure includes steps S101 to S103, wherein step S1021: when the playback time of the entire audio track reaches the second recognition time corresponding to the second to Nth query statements respectively, speech recognition is performed on the second to Nth query statements respectively, which may further include:
[0065] If the playback time of the entire audio track reaches the second recognition time corresponding to the Mth query statement, determine whether the (M-1)th query statement has been successfully recognized. The Mth query statement can be any one of the 2nd to Nth query statements.
[0066] If speech recognition for the (M-1)th query statement is not completed, stop performing speech recognition on the (M-1)th query statement.
[0067] Perform speech recognition on the Mth query statement.
[0068] Based on the above steps of the embodiments of this disclosure, it should be noted that:
[0069] M-1 query statements can be understood as being arranged in playback order, with the query statement preceding the Mth query statement.
[0070] According to the solution of this embodiment, by stopping the recognition of the (M-1)th query statement and continuing to recognize the Mth query statement, the overall speech recognition efficiency of the entire audio track can be improved, avoiding a situation where the speech recognition of the entire audio track cannot continue if the recognition of a certain query statement is abnormal. Priority is given to ensuring the completion of the overall speech recognition of the entire audio track.
[0071] In one example, stopping speech recognition for the (M-1)th query statement if it is determined that speech recognition has not been completed can include the following steps:
[0072] If it is determined that the speech recognition of the (M-1)th query statement has not been completed, a stop recognition command is issued to stop the speech recognition of the (M-1)th query statement.
[0073] If it is determined that the speech recognition of the M-1th query statement has not stopped, a cancellation command is issued to cancel the speech recognition of the M-1th query statement.
[0074] In one example, the speech recognition method provided in this disclosure includes steps S101 to S103, and further includes the step:
[0075] If it is determined that the speech recognition of the (M-1)th query statement has not been cancelled, an anomaly has been identified in the speech recognition, and the recognition of the entire audio track is stopped.
[0076] In one example, such as Figure 3 As shown, the speech recognition method of this disclosure can realize speech recognition of single-track audio. Specifically, it includes:
[0077] 1. Obtain a complete audio track. The audio file contains multiple voice queries. Mark the start and end times of each voice query in the audio file to accurately obtain the number of bytes at the corresponding time point of each voice query in the audio file. Summarize the start time information of the recognized voice in the audio file and send it along with the complete audio track into the encapsulated voice reading data module of the automated verification system based on simplex and full-duplex voice recognition.
[0078] 2. The automated verification system's verification module simulates time by splitting the audio stream into multiple packets of a fixed number of bytes for the recognition module to read. When a custom time period of 200ms before the start time of the voice query (the first recognition time) is read, the recognition module of the automated verification system is started in a simulated recognition manner. The recognition module begins working and automatically detects the end of the voice query using endpoint detection technology, thus determining the end of the current recognition.
[0079] 3. The recognition module continuously reads audio data. If it is simplex speech recognition, when it encounters a custom time period before the start time of the next speech query, the verification module first determines whether the previous recognition has ended. If the previous recognition has ended, it directly starts the next recognition.
[0080] 4. If the current recognition process is still in progress, pause reading the audio file and stop the previous round of recognition by calling the stop command.
[0081] 5. If the recognition process ends within a few seconds, continue reading and start the next recognition process.
[0082] 6. If the question "End the previous round of recognition" still appears after a few seconds, the speech recognition module's operation can be canceled by calling the "cancel" command to force an exit.
[0083] 7. Wait a few seconds. If the previous round of recognition is forcibly exited, continue reading the audio file, restart the recognition module and start the next recognition process to continue recognizing the next voice query of the abnormally recognized voice query. At the same time, the recognition of the next voice query requires the execution of steps 3 to 6.
[0084] 8. If the system does not force exit, the monitoring and alarm mechanism will synchronously identify module abnormalities.
[0085] 9. After recognizing each voice query in the entire audio track, exit the recognition process.
[0086] In one embodiment, the speech recognition method provided in this disclosure includes steps S101 to S103, and may further include the step:
[0087] The test recognition results of N query statements are compared with the corresponding labeled recognition results.
[0088] Based on the comparison results, the speech recognition performance is determined.
[0089] Based on the above steps of the embodiments of this disclosure, it should be noted that:
[0090] The test recognition results can be understood as the recognition results obtained by the speech recognition model that needs to be tested for performance when it performs speech recognition on each query statement of the entire audio track.
[0091] The labeling and recognition results can be understood as the correct recognition results of each query statement in the entire audio track, which are known in advance.
[0092] According to the scheme of this disclosure embodiment, by comparing the test recognition result of the query statement with the corresponding labeled recognition result, the accuracy of the result obtained by speech recognition can be accurately determined, thereby further evaluating whether the performance of the speech recognition model meets the requirements, so as to further adjust the algorithm or logic of the speech recognition model.
[0093] In one embodiment, the speech recognition method provided by this disclosure includes steps S101 to S103, and, before determining a first recognition time based on the start time of the first query statement in the entire audio track in step S101, the method may further include the step of:
[0094] Identify wake words and query statements in a speech dataset.
[0095] Use query statements to construct the entire audio track.
[0096] Based on the above steps of the embodiments of this disclosure, it should be noted that:
[0097] The voice data in the voice dataset can be voice data collected by triggering recognition listening state through button mode, screen click button mode, or face wake-up mode.
[0098] According to the scheme of this disclosure embodiment, since the entire audio track is constructed using only query statements, the needs of speech recognition testing can be better identified, and the speech recognition performance of specific query statements can be tested more directly, avoiding the interference of wake word statements on speech recognition performance testing.
[0099] like Figure 4 As shown, the speech recognition method of this disclosure can achieve speech recognition of full-duplex full-track audio. Specifically:
[0100] 1. Obtain a complete audio track. The audio file contains multiple voice queries. Mark the start and end times of each voice query in the audio file to accurately obtain the number of bytes at the corresponding time point of each voice query in the audio file. Summarize the start time information of the recognized voice in the audio file and send it along with the complete audio track into the encapsulated voice reading data module of the automated verification system based on simplex and full-duplex voice recognition.
[0101] 2. The verification module of the automated verification system simulates time by splitting the audio stream into multiple packets of a fixed number of bytes for the recognition module to read the audio file. When a custom time period of 200ms before the start time of the voice query (the first recognition time) is read, the recognition module of the automated verification system is started in a simulated recognition manner. The recognition module starts working and returns the start and end callbacks of the first voice query through endpoint detection technology. Recognition does not exit and continues to read the audio stream, starting 2 to n (2-n) recognitions until the last voice query is recognized. The full-duplex recognition ends and the exit callback is returned normally.
[0102] The solution disclosed herein also addresses the problem that existing technologies can only evaluate the performance of speech recognition models through actual manual testing or audio playback from devices. Existing technologies are susceptible to the influence of background noise in the testing environment and the instability of the sound source, leading to fluctuations in the final results and hindering data accuracy and problem localization. Most methods rely heavily on manual testing, with one person initiating recognition via device buttons, screen buttons, or facial recognition, while others manually call out the recognition query according to a defined test scenario at the corresponding location. Log screening is then used to determine the accuracy of words and sentences in the recognition results. This purely manual testing approach is prone to significant data fluctuations due to factors such as time, space, network, and human voice. Reproducing and localizing problems still requires numerous repeated tests by a large number of people, resulting in a long timeframe. In contrast, the solution of this disclosure can directly achieve automated recognition of entire audio files without manual operation.
[0103] The solution disclosed herein also solves the problem that existing technologies cannot automatically recognize each query statement in an audio stream of an entire audio track. Existing technologies require manual segmentation of the audio stream, saving each speech recognition statement as a separate audio file. When the speech recognition module is started, each audio file is processed once, then the process ends, and the next audio file is read and the process continues. This not only increases workload but also introduces errors. If 1000 sentences from an entire audio track are tested, each speech query takes an average of 7-8 seconds to segment, adding up to 2 hours to the overall time. The larger the test dataset, the more the time increases, and manual operation is prone to errors. The solution disclosed herein solves these problems, automatically recognizing each query statement in an entire audio track without any other operations. The speech recognition performance can be evaluated directly from a single audio track. The solution disclosed herein eliminates the need to split each query statement of the entire audio track into multiple audio files, enabling speech recognition of all query statements using only a single audio file. This effectively reduces the number of test audio files and lowers storage and maintenance costs.
[0104] Furthermore, the solution of this disclosure also solves the problem in the prior art that speech recognition needs to be activated by keywords before testing can be performed. Full-duplex speech recognition has traditionally used keyword activation to initiate recognition, determining the start of the process and reading the audio stream following the keyword. If the smart terminal uses button mode, screen click button mode, or face activation mode to activate the speech function, and does not support keyword activation, then for recorded multi-sentence audio files, direct reading for effect verification is not possible.
[0105] According to embodiments of this disclosure, such as Figure 5 As shown, a speech recognition device is provided, comprising:
[0106] The first determining module 510 is used to determine the first recognition time based on the start time of the first query statement in the entire audio track. The entire audio track includes N query statements.
[0107] The recognition module 520 is used to perform speech recognition on N query statements when the playback time of the entire audio track reaches the first recognition time.
[0108] The second determining module 530 is used to determine the test recognition results of N query statements.
[0109] In one embodiment, the speech recognition device further includes:
[0110] The third determining module is used to determine the second recognition time corresponding to the second to Nth query statements based on the start time of the second to Nth query statements in the whole audio track.
[0111] In one embodiment, the identification module 520 includes:
[0112] The first recognition submodule is used to perform speech recognition on the first query statement when the playback time of the entire audio track reaches the first recognition time.
[0113] The second recognition submodule is used to perform speech recognition on the second to Nth query statements when the playback time of the entire audio track reaches the second recognition time corresponding to the second to Nth query statements respectively.
[0114] In one implementation, the second recognition submodule is used to determine whether the (M-1)th query statement has completed speech recognition when the playback time of the entire audio track reaches the second recognition time corresponding to the Mth query statement. Furthermore, if it is determined that the (M-1)th query statement has completed speech recognition, speech recognition is performed on the Mth query statement. Here, the Mth query statement is any one of the 2nd to Nth query statements.
[0115] In one implementation, the second recognition submodule is used to determine whether the (M-1)th query statement has completed speech recognition when the playback time of the entire audio track reaches the second recognition time corresponding to the Mth query statement. If it is determined that the (M-1)th query statement has not completed speech recognition, speech recognition for the (M-1)th query statement is stopped. Then, speech recognition is performed on the Mth query statement. Here, the Mth query statement is any one of the 2nd to Nth query statements.
[0116] In one embodiment, the speech recognition device further includes:
[0117] The comparison module is used to compare the test recognition results of N query statements with the corresponding labeled recognition results.
[0118] The fourth determination module is used to determine the speech recognition performance based on the comparison results.
[0119] In one embodiment, the speech recognition device further includes:
[0120] The recognition module is used to identify wake words and query statements in the speech dataset.
[0121] The audio building module is used to construct a complete audio track using query statements.
[0122] The specific functions and examples of each module and submodule of the apparatus in this disclosure can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.
[0123] The acquisition, storage, and application of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0124] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0125] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0126] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded from storage unit 608 into random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0127] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0128] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the speech recognition method. For example, in some embodiments, the speech recognition method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the speech recognition method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform the speech recognition method by any other suitable means (e.g., by means of firmware).
[0129] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0130] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0131] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0132] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0133] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0134] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0135] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0136] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for voice recognition, comprising: determining a first recognition time according to a start time of a first query sentence in a whole track audio, wherein the whole track audio comprises N query sentences, the first recognition time is a time point for voice recognition of the whole track audio, and the first recognition time is K milliseconds earlier than the start time of the first query sentence, K being a positive integer; in a case where a playing time of the whole track audio reaches the first recognition time, starting voice recognition of the N query sentences in a time sequence from the first recognition time to determine a test recognition result of the N query sentences, wherein the test recognition result is a semantic and / or sentence content of each recognized query sentence. 2.The method of claim 1, further comprising: determining second recognition times corresponding to the second to Nth query sentences respectively according to start times of the second to Nth query sentences in the whole track audio, wherein the second recognition time is a time point for voice recognition of any one of the second to Nth query sentences, and the second recognition time is L milliseconds earlier than the start time of the second to Nth query sentences, L being a positive integer.
3. The method of claim 2, wherein, the voice recognition of the N query sentences respectively in the case where the playing time of the whole track audio reaches the first recognition time comprises: in the case where the playing time of the whole track audio reaches the first recognition time, voice recognition of the first query sentence; in the case where the playing time of the whole track audio reaches the second recognition time corresponding to the second to Nth query sentences respectively, starting voice recognition of the second to Nth query sentences in a time sequence from the second recognition time.
4. The method of claim 3, wherein, the voice recognition of the second to Nth query sentences respectively in the case where the playing time of the whole track audio reaches the second recognition time corresponding to the second to Nth query sentences respectively comprises: in the case where the playing time of the whole track audio reaches the second recognition time corresponding to the Mth query sentence, determining whether the M-1th query sentence is completed voice recognition, wherein the Mth query sentence is any one of the second to Nth query sentences; in a case where it is determined that the M-1th query sentence is completed voice recognition, voice recognition of the Mth query sentence.
5. The method of claim 3, wherein, the voice recognition of the second to Nth query sentences respectively in the case where the playing time of the whole track audio reaches the second recognition time corresponding to the second to Nth query sentences respectively comprises: in the case where the playing time of the whole track audio reaches the second recognition time corresponding to the Mth query sentence, determining whether the M-1th query sentence is completed voice recognition, wherein the Mth query sentence is any one of the second to Nth query sentences; in a case where it is determined that the M-1th query sentence is not completed voice recognition, stopping voice recognition of the M-1th query sentence; and voice recognition of the Mth query sentence.
6. The method according to any one of claims 1 to 5, wherein, The test recognition result is speech recognition of each query sentence in the whole track audio by the speech recognition model needing performance test, and a recognition result obtained by the speech recognition; The method further comprises: Comparing the test recognition result of the N query sentences with a corresponding labeled recognition result, wherein the labeled recognition result is a correct recognition result of each query sentence in the whole track audio known in advance; According to the comparison result, the speech recognition performance of the speech recognition model is determined.
7. The method of any one of claims 1 to 5, before the first recognition time determined according to the start time of the first query sentence in the whole track audio, further comprising: Recognizing the wake-up word sentence and the query sentence in the speech data set; Using the query sentence to construct the whole track audio.
8. An apparatus for speech recognition, comprising: A first determination module configured to determine a first recognition time according to a start time of a first query sentence in a whole track audio, wherein the whole track audio comprises N query sentences, wherein the first recognition time is a time for speech recognition of the whole track audio, and the first recognition time is K milliseconds earlier than the start time of the first query sentence, K being a positive integer; An identification module configured to, when a playing time of the whole track audio reaches the first recognition time, perform speech recognition on the N query sentences in sequence from the first recognition time, to determine a test recognition result of the N query sentences, wherein the test recognition result is a semantic meaning and / or a sentence content of each recognized query sentence.
9. The apparatus of claim 8, further comprising: A third determination module configured to determine second recognition times corresponding to the second to Nth query sentences respectively according to start times of the second to Nth query sentences in the whole track audio, wherein the second recognition time is a time for speech recognition of any one of the second to Nth query sentences, and the second recognition time is L milliseconds earlier than the start time of the second to Nth query sentences, L being a positive integer.
10. The apparatus of claim 9, wherein, The identification module comprises: A first identification submodule configured to, when the playing time of the whole track audio reaches the first recognition time, perform speech recognition on the first query sentence; A second identification submodule configured to, when the playing time of the whole track audio reaches the second recognition time corresponding to the second to Nth query sentences respectively, perform speech recognition on the second to Nth query sentences in sequence from the second recognition time.
11. The apparatus of claim 10, wherein, The second identification submodule is configured to, when the playing time of the whole track audio reaches the second recognition time corresponding to an Mth query sentence, determine whether an M-1th query sentence has completed speech recognition, and, when it is determined that the M-1th query sentence has completed speech recognition, perform speech recognition on the Mth query sentence, wherein the Mth query sentence is any one of the second to Nth query sentences.
12. The apparatus of claim 10, wherein, The second recognition submodule is used to determine whether the (M-1)th query statement has completed speech recognition when the playback time of the entire audio track reaches the second recognition time corresponding to the Mth query statement; if it is determined that the (M-1)th query statement has not completed speech recognition, stop continuing to perform speech recognition on the (M-1)th query statement; and perform speech recognition on the Mth query statement; wherein the Mth query statement is any one of the 2nd to Nth query statements.
13. The apparatus of any one of claims 8 to 12, wherein, The test recognition result is the recognition result obtained by the speech recognition model that needs to be tested for performance on each query statement of the whole track audio. The device further includes: The comparison module is used to compare the test recognition results of the N query statements with the corresponding annotation recognition results; wherein, the annotation recognition results are the correct recognition results of each query statement in the whole track audio that are known in advance; The fourth determining module is used to determine the speech recognition performance of the speech recognition model based on the comparison results.
14. The apparatus according to any one of claims 8 to 12, further comprising: The recognition module is used to identify wake words and query statements in the speech dataset; An audio building module is used to construct a complete audio track using the query statement.
15. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 7.
16. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 7.
17. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Man-machine interaction processing method and device and electronic equipment
CN111241245A
Voice response processing method and device based on artificial intelligence, equipment and medium
CN111429899A