Audio file identification method and apparatus
Through the asynchronous execution of thread pool mechanism, long audio files are processed in frames and segments, and the recognition results are sorted and spliced, solving the problems of information loss and response delay in long audio file recognition in the existing technology, and achieving efficient and accurate audio file recognition.
Patent Information
- Application Number
- PCT/CN2024/125555
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-10-17
- Publication Date
- 2025-06-26
AI Technical Summary
When existing speech recognition systems process long audio files, there are problems such as loss of information and reduced recognition accuracy, especially when the voice signal crosses the boundaries of the segment. Meanwhile, the sequential execution of VAD and ASR results in a higher response delay.
The first thread and the second thread pool are adopted for asynchronous execution. The first thread is responsible for framed and segmented processing of the audio files to be identified to generate audio clips; the second thread is responsible for identifying these clips and sorting the recognition results and splicing them. In this way, information loss is avoided and identification efficiency is improved.
It realizes efficient identification of long audio files, reduces information loss, improves recognition accuracy, and reduces response delay.
Smart Images

Figure CN2024125555_26062025_PF_FP_ABST
Abstract
Description
Audio file recognition method and device
[0001] Related applications
[0002] This application claims priority to Chinese patent application number 2023117904816, filed on December 22, 2023, entitled “Audio File Recognition Method and Device,” the entire text of which is hereby incorporated by reference. Technical Field
[0003] The present application relates to the field of computer technology, and in particular to a method and device for audio file recognition. Background Art
[0004] Recognizing long audio files has important applications in audio transcription. With the advancement of artificial intelligence (AI) technology, automatic speech recognition (ASR) systems typically use deep learning-based neural network models to transcribe audio files. Because existing speech recognition models have a large number of parameters and computational complexity, graphics card acceleration is required to ensure acceptable recognition speeds. However, due to limited GPU storage, speech recognition models running on GPUs are still unable to process very long audio files all at once. Existing speech recognition systems use two approaches to transcribe long audio files. One approach is streaming processing: using streaming speech recognition technology, continuous audio streams can be processed in real time without waiting for the entire audio file to be fully uploaded. This approach is suitable for scenarios with high real-time requirements. The second approach is segmentation: long audio files are divided into shorter segments, each of which can be adjusted to suit different usage scenarios. These shorter segments are easier to process, reducing memory and computing resource requirements. The recognized text can be merged or processed later. However, when dividing a long audio file into multiple shorter segments, the segment length must be determined. To determine the length of the segment, a fixed value can be pre-set according to the application scenario, and then the long audio can be rigidly cut into several fixed-length segments. Although this cutting method is relatively simple, it may cause information loss when dividing long speech into fixed-length segments. In particular, when an important part of the speech signal spans two segments, key information may be lost, making it difficult for the recognition system to understand the complete semantics and context, resulting in a decrease in speech recognition accuracy. At the same time, another way to solve the problem of long audio file recognition in related technologies is to combine VAD (Voice Activity Detection) with ASR, but the ASR processing is performed only after the VAD processing is completed. The VAD and ASR executed sequentially have a relationship in which the latter waits for the former. In the case of high concurrency and fewer resources, this will result in a higher overall response delay of the service.
[0005] To address the above-mentioned problems, no effective solutions have been proposed so far.
[0006] Summary of the Invention
[0007] According to a first aspect of an embodiment of the present application, a method for audio file recognition is provided, comprising: receiving an audio file to be recognized; using a first thread in a first thread pool to process audio frames in the audio file to be recognized in sequence according to the arrangement order of the audio frames in the audio file to be recognized, to obtain multiple audio clips; using a second thread in a second thread pool to recognize the multiple audio clips in sequence to obtain multiple recognition results, and sorting and splicing the multiple recognition results to obtain a final recognition result, wherein the first thread and the second thread are asynchronously executed threads.
[0008] Optionally, the first thread in the first thread pool is used to process the audio frames in the audio file to be identified in sequence according to the arrangement order of the audio frames in the audio file to be identified to obtain multiple audio clips, including: verifying the audio file to be identified, and if the verification passes, placing the audio file to be identified into a first input queue; using the first thread to perform frame processing on the audio file to be identified in the first input queue to obtain multiple audio frames, processing each audio frame in sequence, and determining the start and end time points of each audio clip in sequence; and segmenting the audio file to be identified in sequence according to the start and end time points of each audio clip to obtain the multiple audio clips.
[0009] Optionally, the second thread in the second thread pool is used to identify the multiple audio clips in sequence to obtain multiple recognition results, including: using the first thread to generate a target number and collecting the clip number of the current audio clip to be identified, the target number is used to mark the audio file to be identified; using the first thread to construct a structure to be identified based on the current audio clip to be identified, the clip number of the previous audio clip to be identified and the target number; using the second thread to identify the structure to be identified, obtain the recognition result corresponding to the current audio clip to be identified, and store the recognition result corresponding to the current audio clip to be identified in the clip recognition result queue corresponding to the target number.
[0010] Optionally, sorting the multiple recognition results and splicing them together to obtain a final recognition result includes: using the second thread to detect the received structure to be recognized; when an end signal is detected in the structure to be recognized, using the second thread to sort the multiple recognition results and splicing them together to obtain the final recognition result, wherein the end signal is sent to the second thread by the first thread after processing all audio frames in the audio file to be recognized, and the end signal is used to indicate that all audio frames in the audio file to be recognized have been processed.
[0011] Optionally, the second thread is used to sort the multiple recognition results and then splice them to obtain the final recognition result, including: constructing a target key-value pair in a preset global dictionary, the key of the target key-value pair is the target number, and the value of the target key-value pair includes: the number of recognized fragments, the total number of fragments and a fragment recognition result queue, wherein the number of recognized fragments is updated by the second thread after completing the recognition of each audio fragment, the total number of fragments is updated by the first thread after all audio frames in the audio file to be recognized are processed, and the fragment recognition result queue is updated by the second thread after completing the recognition of each audio fragment, and is used to store the multiple recognition results; the second thread is used to extract the number of recognized fragments and the total number of fragments from the target key-value pair according to the target number, and when the number of recognized fragments and the total number of fragments are equal, the multiple recognition results are taken out one by one from the fragment recognition result queue; the multiple recognition results are spliced together in the order of fragment numbers to obtain the final recognition result.
[0012] Optionally, the method further includes: when the number of identified fragments is less than the total number of fragments, using the second thread to extract the structure to be identified from the second input queue for identification to obtain an identification result, wherein the second input queue is used to store the structure to be identified constructed by the first thread.
[0013] Optionally, after obtaining the final recognition result, the method further includes: storing the final recognition result in a preset database; retrieving the final recognition result from the preset database in response to a query instruction from the client; and outputting the final recognition result to the client.
[0014] Optionally, the first thread pool includes multiple first threads, and each first thread processes only one audio file to be recognized.
[0015] Optionally, the first input queue, the fragment recognition result queue, and the second input queue are all concurrent queues.
[0016] According to the second aspect of the embodiment of the present application, an audio file recognition device is also provided, including: a receiving module for receiving an audio file to be recognized; a segmentation module for using a first thread in a first thread pool to process the audio frames in the audio file to be recognized in sequence according to the arrangement order of the audio frames in the audio file to be recognized, to obtain multiple audio clips; an identification module for using a second thread in a second thread pool to recognize the multiple audio clips in sequence to obtain multiple recognition results, and sorting the multiple recognition results and splicing them to obtain a final recognition result, wherein the first thread and the second thread are asynchronously executed threads.
[0017] According to a third aspect of an embodiment of the present application, a non-volatile storage medium is further provided. The non-volatile storage medium includes a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above-mentioned audio file recognition method.
[0018] According to a fourth aspect of an embodiment of the present application, a computer device is further provided, including a memory and a processor, wherein the processor is configured to run a program, wherein the above-mentioned audio file recognition method is executed when the program is run.
[0019] The details of one or more embodiments of the present application are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the present application will become apparent from the description, drawings, and claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0021] FIG1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing an audio file recognition method according to an embodiment of the present application;
[0022] FIG2 is a flow chart of an audio file recognition method according to an embodiment of the present application;
[0023] FIG3 is a flowchart of a first thread processing an audio file to be recognized according to an embodiment of the present application;
[0024] FIG4 is a flowchart of a second thread identifying a structure to be identified according to an embodiment of the present application;
[0025] FIG5 is a flow chart of a method for identifying an audio file to be identified according to an embodiment of the present application;
[0026] FIG6 is a schematic structural diagram of an audio file recognition device according to an embodiment of the present application. DETAILED DESCRIPTION
[0027] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0029] According to an embodiment of the present application, an embodiment of an audio file recognition method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0030] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal, a cloud server or a similar computing device. Figure 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing an audio file recognition method. As shown in Figure 1, the computer terminal 10 (or mobile device 10) may include one or more (102a, 102b, ..., 102n are used in the figure to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply and / or a camera. It will be understood by those skilled in the art that the structure shown in Figure 1 is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may also include more or fewer components than those shown in Figure 1, or have a configuration different from that shown in Figure 1.
[0031] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry". The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuitry may be a single independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10 (or mobile device). As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0032] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the audio file recognition method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned audio file recognition method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0033] The transmission module 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0034] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0035] According to an embodiment of the present application, an embodiment of an audio file recognition method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0036] FIG2 is a flow chart of an audio file recognition method according to an embodiment of the present application. As shown in FIG2 , the method includes the following steps:
[0037] Step S202, receiving an audio file to be recognized;
[0038] Step S204: using the first thread in the first thread pool to sequentially process the audio frames in the audio file to be recognized according to the arrangement order of the audio frames in the audio file to be recognized, to obtain a plurality of audio clips;
[0039] Step S206: Use the second thread in the second thread pool to sequentially recognize the multiple audio clips to obtain multiple recognition results, sort the multiple recognition results and then splice them to obtain a final recognition result, wherein the first thread and the second thread are asynchronously executed threads.
[0040] Through the above steps, it is possible to implement the method of receiving an audio file to be identified; using the first thread in the first thread pool to process the audio frames in the audio file to be identified in sequence according to the arrangement order of the audio frames in the audio file to be identified, and obtain multiple audio clips; using the second thread in the second thread pool to identify the multiple audio clips in sequence to obtain multiple recognition results, and then sorting the multiple recognition results and splicing them to obtain a final recognition result, wherein the first thread and the second thread are asynchronously executed threads. Through the asynchronous execution of the first thread and the second thread, the purpose of the first thread and the second thread not having to wait for each other is achieved, thereby achieving the technical effect of improving the efficiency of long audio file recognition, and further solving the technical problem in the related technology that ASR processing must be performed only after VAD processing is completed, resulting in low audio file recognition efficiency.
[0041] In step S202 , the audio file to be recognized may be a long audio file.
[0042] In step S204, the first thread in the first thread pool is used to execute the VAD (Voice Activity Detection) program. It should be noted that the voice endpoint detection technology (VAD) is used to determine the length of the segment. VAD technology can automatically determine the part of voice activity based on the characteristics of the voice signal, thereby adaptively cutting long audio. This ensures that only the part containing voice is extracted and processed, avoiding the loss of important information, thereby saving computing resources and processing time for voice recognition, and improving processing efficiency. At the same time, VAD technology can reduce the interference of non-voice parts, thereby reducing the error rate of the voice recognition system. This is crucial to improving the accuracy of voice recognition, especially in the presence of noise or background interference.
[0043] In step S206, the second thread in the second thread pool is used to execute ASR (Automatic Speech Recognition).
[0044] The above steps S202 to S206 are described in detail below through a specific embodiment.
[0045] In some embodiments, step S204 includes: verifying the audio file to be identified, and if the verification passes, placing the audio file to be identified into a first input queue; using the first thread to perform frame processing on the audio file to be identified in the first input queue to obtain multiple audio frames, processing each audio frame in turn, and determining the start and end time points of each audio segment in turn; segmenting the audio file to be identified in turn according to the start and end time points of each audio segment to obtain the multiple audio segments.
[0046] Specifically, the client concurrently sends HTTP (Hypertext Transfer Protocol) requests for identifying long audio files to the server. After receiving the requests, the server verifies and filters the request information, returns a corresponding error code for illegal requests, and places the verified legitimate request information into the VAD input queue (the first input queue). It is understandable that each first thread processes only one audio file to be identified at a time and can generate a target number (each audio file to be identified corresponds to a unique target number) to mark each audio file to be identified.
[0047] It should be noted that there are multiple first threads in the first thread pool, and the server puts the request information into the VAD input queue. Then the threads in the VAD thread pool compete to receive the request, and multiple first threads can synchronously process multiple audio files to be recognized. Each first thread corresponds one-to-one to each audio file to be recognized. After that, the VAD thread will communicate with the streaming VAD engine to send the long audio file and multiple short audio clips into the ASR input queue. In order to improve efficiency, the process of each thread in the VAD thread pool sending the audio clip into the ASR input queue (second input queue) is concurrent, which means that the ASR input queue is mixed with clips from multiple long audio files.
[0048] In some embodiments of the present application, the specific process of using the second thread in the second thread pool to sequentially identify the multiple audio clips to obtain multiple recognition results includes: using the first thread to generate a target number and collecting the clip number of the current audio clip to be identified, and the target number is used to mark the audio file to be identified; using the first thread to construct a structure to be identified based on the current audio clip to be identified, the clip number of the previous audio clip to be identified and the target number; using the second thread to identify the structure to be identified, obtain the recognition result corresponding to the current audio clip to be identified, and store the recognition result corresponding to the current audio clip to be identified in the clip recognition result queue corresponding to the target number.
[0049] Specifically, threads in the second thread pool also compete to extract audio clips from the second input queue, then communicate with the non-streaming ASR engine for recognition and maintain relevant information in the global dictionary. Because each thread in the ASR thread pool recognizes different audio clips concurrently, it is necessary to use the relevant information in the global dictionary to determine whether the audio file to be recognized is completely recognized.
[0050] In some embodiments of the present application, the specific steps of sorting the multiple recognition results and then splicing them to obtain the final recognition result are as follows: using the second thread to detect the received structure to be recognized; when an end signal is detected in the structure to be recognized, using the second thread to sort the multiple recognition results and then splicing them to obtain the final recognition result, wherein the end signal is sent to the second thread by the first thread after processing all audio frames in the audio file to be recognized, and the end signal is used to indicate that all audio frames in the audio file to be recognized have been processed.
[0051] In an optional manner, a target key-value pair is constructed in a preset global dictionary, the key of the target key-value pair is the target number, and the value of the target key-value pair includes: the number of identified fragments, the total number of fragments, and a fragment recognition result queue, wherein the number of identified fragments is updated by the second thread after completing the recognition of each audio fragment, the total number of fragments is updated by the first thread after all audio frames in the audio file to be recognized are processed, and the fragment recognition result queue is updated by the second thread after completing the recognition of each audio fragment, and is used to store the multiple recognition results; the second thread is used to extract the number of identified fragments and the total number of fragments from the target key-value pair according to the target number, and when the number of identified fragments and the total number of fragments are equal, the multiple recognition results are taken out one by one from the fragment recognition result queue; the multiple recognition results are spliced in the order of fragment numbers to obtain the final recognition result.
[0052] Figure 3 illustrates the first thread's processing flow for verified audio files to be recognized. As shown in Figure 3, each first thread generates a globally unique target number for the current request for record keeping and creates a key-value pair corresponding to the current request in the global dictionary, where the key is the target number and the value is the number of recognized segments, the total number of segments, and a segment recognition result queue. The first thread then frames the preprocessed audio and sends each frame sequentially to the streaming endpoint detection engine. The streaming endpoint detection engine returns the start and end time points of each audio segment to the first thread. Whenever the first thread obtains the start and end time points of a segment, it extracts an audio segment from the long audio file based on these start and end time points. For each segment, the first thread enters a structure consisting of the target number, segment number, and audio segment to be recognized into the second input queue. After the first thread completes processing all audio frames, it updates the total number of segments in the global dictionary based on the task number and enters a structure consisting of the target number and a completion signal into the second input queue.
[0053] Figure 4 illustrates the second thread's recognition process for a structure to be recognized. As shown in Figure 4, a thread in the second thread pool retrieves the structure from the second input queue (ASR input queue) and parses it. Using the task number in the structure, the second thread obtains the corresponding number of recognized segments, total number of segments, and segment recognition result queue from the global dictionary. If the structure contains a segment number and an audio segment, the second thread sends the audio segment to the non-streaming ASR engine. The non-streaming ASR engine processes the received audio segment through feature extraction, CTC (Connectionist Temporal Classification, a neural network model used for predictive sequence labeling tasks), and attention rescoring, and asynchronously returns the recognition result to the second thread. The second thread then updates the number of recognized segments and sends a structure consisting of the recognition result and segment number to the segment recognition result queue. It then begins checking whether the number of recognized segments is equal to the total number of segments. If the structure contains an end signal, the second thread immediately begins checking whether the number of recognized segments is equal to the total number of segments. If the number of recognized fragments is less than the total number of fragments, the second thread extracts the next structure from the second input queue and starts processing; if they are equal, it means that the long audio file corresponding to the current task number has been recognized, and the second thread takes out the fragment recognition results one by one from the fragment recognition result queue and reorders them according to the fragment number to obtain the final complete recognition result.
[0054] It should be noted that all queues in the embodiments of the present application are concurrent queues.
[0055] In some embodiments of the present application, when the number of identified fragments is less than the total number of fragments, the second thread is used to extract the structure to be identified from the second input queue for identification to obtain an identification result, wherein the second input queue is used to store the structure to be identified constructed by the first thread.
[0056] After obtaining the final recognition result, the final recognition result is stored in a preset database; in response to a query instruction from the client, the final recognition result is retrieved from the preset database; and the final recognition result is output to the client.
[0057] Specifically, the second thread determines whether the number of identified fragments is equal to the total number of fragments. If the number of identified fragments is less than the total number of fragments, the structure to be identified is extracted from the second input queue for processing; if they are equal, it means that the long audio file corresponding to the current task number has been identified. The second thread takes out the fragment recognition results one by one from the fragment recognition result queue, and reorders them according to the fragment number to obtain the final complete recognition result, and stores it in the database, waiting for the client to send a query request.
[0058] Figure 5 illustrates another audio file recognition process. As shown in Figure 5, after the client concurrently sends requests to the server, the server places the request information into the VAD input queue. Threads in the VAD thread pool then compete to retrieve the request. Each VAD thread generates a globally unique task number for the current request, which it uses to record the request and maintains the relevant information in a global dictionary to ensure data synchronization when multiple tasks are running simultaneously. Furthermore, the globally unique task number means that each VAD thread is responsible for only one request at a time—that is, one long audio file. The VAD thread then communicates with the streaming VAD engine to send the long audio file, along with multiple short audio clips, into the ASR input queue. To improve efficiency, each thread in the VAD thread pool concurrently sends audio clips to the ASR input queue, which means the ASR input queue contains a mixture of clips from multiple long audio files. Simultaneously, threads in the ASR thread pool also compete to retrieve audio clips from the ASR input queue, then communicate with the non-streaming ASR engine for recognition and maintain the relevant information in the global dictionary. Since each thread in the ASR thread pool also recognizes different audio segments concurrently, it is necessary to use the relevant information in the global dictionary to determine whether a long audio is completely recognized.
[0059] The audio file recognition method provided in the embodiment of the present application is also applied to an audio file recognition device provided in the embodiment of the present application, as shown in Figure 6, including: a receiving module 60, used to receive an audio file to be recognized; a segmentation module 62, used to use a first thread in a first thread pool to process the audio frames in the audio file to be recognized in sequence according to the arrangement order of the audio frames in the audio file to be recognized, to obtain multiple audio clips; an identification module 64, used to use a second thread in a second thread pool to recognize the multiple audio clips in sequence to obtain multiple recognition results, and sort the multiple recognition results and then splice them to obtain a final recognition result, wherein the first thread and the second thread are asynchronously executed threads.
[0060] The segmentation module 62 includes: a segmentation submodule, which is used to verify the audio file to be identified and, if the verification is passed, put the audio file to be identified into a first input queue; use the first thread to perform frame processing on the audio file to be identified in the first input queue to obtain multiple audio frames, process each audio frame in turn, and determine the start and end time points of each audio segment in turn; and segment the audio file to be identified in turn according to the start and end time points of each audio segment to obtain the multiple audio segments.
[0061] The segmentation submodule includes: an identification unit, which is used to use the first thread to generate a target number and collect the segment number of the current audio segment to be identified, where the target number is used to mark the audio file to be identified; use the first thread to construct a structure to be identified based on the current audio segment to be identified, the segment number of the previous audio segment to be identified, and the target number; use the second thread to identify the structure to be identified, obtain a recognition result corresponding to the current audio segment to be identified, and store the recognition result corresponding to the current audio segment to be identified in a segment recognition result queue corresponding to the target number.
[0062] The recognition module 64 includes: a detection submodule, which is used to use the second thread to detect the received structure to be recognized; when an end signal is detected in the structure to be recognized, the second thread is used to sort the multiple recognition results and then splice them to obtain the final recognition result, wherein the end signal is sent to the second thread by the first thread after processing all audio frames in the audio file to be recognized, and the end signal is used to indicate that all audio frames in the audio file to be recognized have been processed.
[0063] The detection submodule includes: a splicing unit, which is used to construct a target key-value pair in a preset global dictionary, the key of the target key-value pair is the target number, and the value of the target key-value pair includes: the number of identified fragments, the total number of fragments and the fragment recognition result queue, wherein the number of identified fragments is updated by the second thread after completing the recognition of each audio fragment, the total number of fragments is updated by the first thread after all audio frames in the audio file to be recognized are processed, and the fragment recognition result queue is updated by the second thread after completing the recognition of each audio fragment, and is used to store the multiple recognition results; the second thread is used to extract the number of identified fragments and the total number of fragments from the target key-value pair according to the target number, and when the number of identified fragments and the total number of fragments are equal, the multiple recognition results are taken out one by one from the fragment recognition result queue; the multiple recognition results are spliced in the order of fragment numbers to obtain the final recognition result.
[0064] The splicing unit includes: a comparison subunit, which is used to use the second thread to extract the structure to be identified from the second input queue for identification when the number of identified fragments is less than the total number of fragments, to obtain an identification result, wherein the second input queue is used to store the structure to be identified constructed by the first thread.
[0065] The recognition module 64 includes: an output submodule, configured to store the final recognition result in a preset database; retrieve the final recognition result from the preset database in response to a query instruction from the client; and output the final recognition result to the client.
[0066] According to another aspect of an embodiment of the present application, a non-volatile storage medium is provided, including a stored program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the above-mentioned audio file recognition method.
[0067] According to another aspect of an embodiment of the present application, a computer device is further provided, including a memory and a processor, wherein the processor is configured to run a program, wherein the above-mentioned audio file recognition method is executed when the program is run.
[0068] The above-mentioned computer device executes the above-mentioned audio file recognition method, and receives the audio file to be recognized; uses the first thread in the first thread pool to process the audio frames in the audio file to be recognized in sequence according to the arrangement order of the audio frames in the audio file to be recognized, and obtains multiple audio clips; uses the second thread in the second thread pool to recognize the multiple audio clips in sequence to obtain multiple recognition results, and sorts and splices the multiple recognition results to obtain a final recognition result, wherein the first thread and the second thread are asynchronously executed threads. Through the asynchronous execution of the first thread and the second thread, the purpose of the first thread and the second thread not having to wait for each other is achieved, thereby achieving the technical effect of improving the efficiency of long audio file recognition, and further solving the technical problem in the related technology that ASR processing must be performed only after VAD processing is completed, resulting in low audio file recognition efficiency.
[0069] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0070] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0071] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0072] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.
[0073] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0074] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program code.
[0075] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for audio file recognition, comprising: Receiving an audio file to be recognized; Using a first thread in a first thread pool, sequentially processing the audio frames in the audio file to be recognized according to the arrangement order of the audio frames in the audio file to be recognized, to obtain a plurality of audio clips; The second thread in the second thread pool is used to recognize the multiple audio clips in sequence to obtain multiple recognition results, and the multiple recognition results are sorted and spliced to obtain a final recognition result, wherein the first thread and the second thread are asynchronously executed threads.
2. The method according to claim 1, wherein: The first thread in the first thread pool is used to process the audio frames in the audio file to be recognized in sequence according to the arrangement order of the audio frames in the audio file to be recognized, so as to obtain a plurality of audio clips, including: Verifying the audio file to be recognized, and if the verification passes, placing the audio file to be recognized into a first input queue; Using the first thread to perform frame processing on the audio file to be recognized in the first input queue to obtain multiple audio frames, processing each audio frame in turn, and determining the start and end time points of each audio segment in turn; The audio file to be recognized is segmented in sequence according to the start and end time points of each audio segment to obtain the multiple audio segments.
3. The method according to claim 2, wherein: Using the second thread in the second thread pool to sequentially recognize the multiple audio clips to obtain multiple recognition results, including: The first thread is used to generate a target number, and the segment number of the current audio segment to be identified is collected, wherein the target number is used to mark the audio file to be identified; Using the first thread to construct a structure to be recognized according to the current audio segment to be recognized, the segment number of the previous audio segment to be recognized, and the target number; The second thread is used to identify the structure to be identified, to obtain an identification result corresponding to the current audio segment to be identified, and the identification result corresponding to the current audio segment to be identified is stored in a segment identification result queue corresponding to the target number.
4. The method according to claim 1, wherein: The multiple recognition results are sorted and then concatenated to obtain a final recognition result, including: Using the second thread to detect the received structure to be identified; When an end signal is detected in the structure to be identified, the second thread is used to identify the plurality of The results are sorted and spliced to obtain the final recognition result, wherein the end signal is sent to the second thread by the first thread after processing all audio frames in the audio file to be recognized, and the end signal is used to indicate that all audio frames in the audio file to be recognized have been processed.
5. The method according to claim 4, wherein: The second thread is used to sort the multiple recognition results and then concatenate them to obtain the final recognition result, including: Constructing a target key-value pair in a preset global dictionary, wherein the key of the target key-value pair is the target number, and the value of the target key-value pair includes: the number of identified segments, the total number of segments, and a segment recognition result queue, wherein the number of identified segments is updated by the second thread after completing the recognition of each audio segment, the total number of segments is updated by the first thread after all audio frames in the audio file to be recognized are processed, and the segment recognition result queue is updated by the second thread after completing the recognition of each audio segment, and is used to store the multiple recognition results; extracting the number of identified fragments and the total number of fragments from the target key-value pair according to the target number by the second thread, and taking out the multiple recognition results one by one from the fragment recognition result queue when the number of identified fragments and the total number of fragments are equal; The multiple recognition results are concatenated in the order of the fragment numbers to obtain the final recognition result.
6. The method according to claim 5, further comprising: When the number of identified fragments is less than the total number of fragments, the second thread is used to extract the structure to be identified from the second input queue for identification to obtain an identification result, wherein the second input queue is used to store the structure to be identified constructed by the first thread.
7. The method according to claim 1, wherein: After obtaining the final recognition result, the method further includes: Storing the final recognition result in a preset database; Retrieving the final recognition result from the preset database in response to a query instruction from the client; The final recognition result is output to the client.
8. The method according to claim 1, wherein the first thread pool comprises a plurality of first threads, and each first thread processes only one audio file to be recognized.
9. The method according to claim 3, wherein the first input queue, the fragment identification result queue, and the second input queue are all concurrent queues.
10. An audio file recognition device, comprising: A receiving module, used for receiving an audio file to be recognized; A segmentation module, configured to use a first thread in a first thread pool to sequentially process the audio frames in the audio file to be identified according to the arrangement order of the audio frames in the audio file to be identified, so as to obtain a plurality of audio segments; The recognition module is used to use the second thread in the second thread pool to sequentially recognize the multiple audio clips to obtain multiple recognition results, sort the multiple recognition results and then splice them to obtain a final recognition result, wherein the first thread and the second thread are asynchronously executed threads.
11. A non-volatile storage medium, the non-volatile storage medium comprising a stored program, wherein: When the program is running, the device where the non-volatile storage medium is located is controlled to execute the audio file recognition method according to any one of claims 1 to 9.
12. A computer device comprising a memory and a processor, wherein the processor is used to run a program, wherein: When the program is run, the audio file recognition method according to any one of claims 1 to 9 is executed.
Citation Information
Patent Citations
Communication method and device in voice recognition scene
CN112992141A
Voice recognition method and device
CN113707152A
Speech recognition method and device, equipment and storage medium
CN115938397A
Audio file identification method and device
CN117765970A
Real-time transcription of conference calls
US20110112833A1
Cited By
Audio data processing method and system
CN121545546A