Speech Recognition Method, Speech Recognition Device and Speech Recognition System
By receiving and processing single-channel audio data sent by the sending end at the receiving end, the bandwidth waste and computing burden caused by multi-channel audio transmission in the prior art is solved, and more efficient audio data transmission is achieved.
Patent Information
- Application Number
- CN202211131469.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-09-16
AI Technical Summary
In the prior art, the sending end sends multi-channel audio to the receiver, resulting in wasted bandwidth resources and low transmission efficiency.
On the receiving end, the single-channel audio data sent by the receiving end is obtained by encapsulating multiple role audio data, and the role audio data has a role mark. After collecting multi-channel audio data, the sending end recognizes roles and adds role marks to form a single-channel audio data transmission.
By sending single-channel audio data, the bandwidth resources occupied during the transmission process are reduced, the computing burden on the receiver is reduced, and the transmission efficiency is improved.
Smart Images

Figure CN115512706B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of speech processing, and in particular, to a speech recognition method, a speech recognition device, a computer-readable storage medium, and a speech recognition system. Background Art
[0002] With the continuous development of technology, the applications of artificial intelligence (AI) technologies such as speech recognition and role separation are becoming more and more widespread. The prior art usually adopts the following methods for speech recognition and role separation: The first one: multi-channel audio is collected through relevant hardware devices (such as audio recording devices, clients), and then methods such as phase difference, amplitude difference, and fundamental frequency detection are used to perform role separation on the multi-channel audio, so as to perform speech recognition on the audio of one of the channels. The second one: relevant clustering algorithms are adopted to cluster the voice features in the multi-channel audio to achieve the purpose of role separation and speech recognition.
[0003] There are also the following problems in adopting the above-mentioned speech recognition and role separation: All the relevant audio processing steps of the first method above are executed by the receiving end. For the sending end, it only needs to send the collected multi-channel audio to the receiving end; while for the receiving end, the amount of calculation is large, that is, the calculation burden on the receiving end is large. And during the process of sending from the sending end to the receiving end, it is necessary to transmit multi-channel audio, which results in a waste of bandwidth resources and low transmission efficiency. The second method above has high requirements for the audio quality and the accuracy of the voiceprint recognition algorithm.
[0004] Therefore, there is an urgent need for a method that can reduce the bandwidth resources for transmitting multi-channel audio between the sending end and the receiving end. Summary of the Invention
[0005] The main purpose of the present application is to provide a speech recognition method, a speech recognition device, a computer-readable storage medium, and a speech recognition system to solve the problem of relatively wasteful bandwidth resources caused by the sending end sending multi-channel audio to the receiving end in the prior art.
[0006] According to one aspect of an embodiment of the present invention, a speech recognition method is provided. The speech recognition method is applied in a receiving end, and the speech recognition method includes: receiving single-channel audio data sent by a sending end, where the single-channel audio data is single-channel audio data encapsulated from multiple role audio data, the role audio data is multi-channel audio data with role tags, and the multi-channel audio data is audio data collected by the sending end; performing speech recognition processing on the single-channel audio data to obtain the speech recognition text information of each role.
[0007] Optionally, the role audio data is obtained by adding the role tag to the multi-channel audio data. The process of adding the role tag is as follows: the sending end performs role recognition on each piece of multi-channel audio data to obtain the role information corresponding to each piece of multi-channel audio data; then, according to the role information corresponding to each piece of multi-channel audio data, the corresponding role tag is added to the multi-channel audio data to obtain a plurality of pieces of role audio data.
[0008] Optionally, speech recognition processing is performed on the single-channel audio data to obtain the speech recognition text information of each role, including: determining the audio start and end time ranges and byte offsets corresponding to each role in the single-channel audio data according to the role tag and the audio bit rate, where the audio start and end time ranges are the timestamps of the audio data corresponding to each role; using a speech recognition algorithm to perform speech recognition on the single-channel audio data to obtain the target text information and the phrase time range, where the phrase time range is the timestamp of each phrase in the target text information; and obtaining the speech recognition text information corresponding to each role according to the audio start and end time ranges and the byte offsets of each role, as well as the target text information and the phrase time range.
[0009] Optionally, determining the audio start and end time ranges and byte offsets corresponding to each role in the single-channel audio data according to the role tag and the audio bit rate includes: performing segmentation processing on the single-channel audio data to obtain a plurality of pieces of role audio data; determining the corresponding audio start and end time ranges according to the role tags corresponding to each piece of role audio data; and determining the byte offsets corresponding to each role according to the audio start and end time ranges and the audio bit rate.
[0010] Optionally, obtaining the speech recognition text information corresponding to each role according to the audio start and end time ranges and the byte offsets of each role, as well as the target text information and the phrase time range, includes: starting from the phrase time range corresponding to the first phrase in the target text information, comparing each phrase time range with the audio start and end time ranges to obtain the preset text information of each role; and decoding the preset text information of each role according to the byte offset of each role to obtain the speech recognition text information of each role.
[0011] Optionally, starting from the phrase time range corresponding to the first phrase in the target text information, compare each phrase time range with the audio start and end time ranges to obtain the preset text information of each character, including: when there is an intersection between the target phrase time range and both audio start and end time ranges, determine the target time point, where the target time is the overlapping time point of the two audio start and end time ranges, and the target phrase time range is one of the multiple phrase time ranges; calculate the difference between the upper time point of the target phrase time range and the target time point to obtain a first difference, and calculate the difference between the lower time point of the target phrase time range and the target time point to obtain a second difference; based on the comparison result of the first difference and the second difference, determine the audio start and end time range to which the phrase corresponding to the target phrase time range belongs, to obtain the preset text information of each character.
[0012] Optionally, before determining the corresponding audio start and end time ranges according to the character markers corresponding to the audio data of each character, the speech recognition method further includes: using a sliding window algorithm to smooth the audio data of each character to obtain multiple smoothed audio data of each character.
[0013] According to another aspect of the embodiments of the present invention, there is also provided a speech recognition device, which is applied in a receiving end. The speech recognition device includes: a receiving unit, configured to receive single-channel audio data sent by a sending end, where the single-channel audio data is single-channel audio data encapsulated from multiple character audio data, the character audio data is multi-channel audio data with character markers, and the multi-channel audio data is audio data collected by the sending end; a processing unit, configured to perform speech recognition processing on the single-channel audio data to obtain the speech recognition text information of each character.
[0014] According to still another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and the program executes any one of the speech recognition methods.
[0015] According to yet another aspect of the embodiments of the present invention, there is also provided a speech recognition system, including: a receiving end, where the receiving end is configured to execute any one of the speech recognition methods; a sending end, communicatively connected to the receiving end.
[0016] In an embodiment of the present invention, the speech recognition method is applied to a receiving end. When the receiving end receives single-channel audio data sent by a sending end, it performs speech recognition processing on the received single-channel audio data to obtain speech recognition text information corresponding to each role. Among them, the single-channel audio data is single-channel audio data obtained by encapsulating multiple role audio data, and the role audio data has role tags. That is to say, after the sending end collects multi-channel audio data, it performs role recognition on the multi-channel audio data to identify the corresponding role (i.e., the corresponding speaker), and then adds role tags to the multi-channel audio data to obtain role audio data. In the speech recognition method of the present application, since the sending end sends single-channel audio data to the receiving end, this ensures that the bandwidth resources occupied during the transmission of the audio data are less, and it also ensures that the sending end can transmit the single-channel audio data to the receiving end relatively quickly. Since the receiving end does not need to perform role recognition, etc., and only needs to perform speech recognition on the single-channel audio data, this also ensures that the calculation amount of the receiving end is less, thus solving the problem of relatively wasting bandwidth resources caused by the sending end sending multi-channel audio to the receiving end in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0018] Figure 1 shows a flowchart of a speech recognition method according to an embodiment of the present application;
[0019] Figure 2 shows a schematic structural diagram of single-channel audio data according to an embodiment of the present application;
[0020] Figure 3 shows a schematic structural diagram of a speech recognition device according to an embodiment of the present application.
[0021] Among them, the above-mentioned accompanying drawings include the following reference numerals:
[0022] 100, multi-channel audio data; 101, role tag; 102, role audio data; 103, single-channel audio data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be encapsulated with each other. The following will refer to the accompanying drawings and combine with the embodiments to detail this application.
[0024] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of this application described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0026] As described in the background art, in the prior art, the sending end sends multi-channel audio to the receiving end, which results in a relatively waste of bandwidth resources. To solve the above problems, in a typical implementation manner of this application, a speech recognition method, a speech recognition device, a computer-readable storage medium, and a speech recognition system are provided.
[0027] According to an embodiment of this application, a speech recognition method is provided.
[0028] Figure 1 is a flowchart of the speech recognition method according to an embodiment of this application. This speech recognition method is applied in the receiving end, as Figure 1 shown, this speech recognition method includes the following steps:
[0029] Step S101, receiving the single-channel audio data sent by the sending end, where the single-channel audio data is the single-channel audio data encapsulated from multiple role audio data, the role audio data is the multi-channel audio data with role tags, and the multi-channel audio data is the audio data collected by the sending end;
[0030] Step S102, performing speech recognition processing on the single-channel audio data to obtain the speech recognition text information of each role.
[0031] The above voice recognition method is applied to the receiving end. When the receiving end receives the single-channel audio data sent by the sending end, it performs voice recognition processing on the received single-channel audio data to obtain the voice recognition text information corresponding to each role. Among them, the above single-channel audio data is a single-channel audio data obtained by encapsulating multiple role audio data, and the role audio data has role tags. That is to say, after the sending end collects multi-channel audio data, it performs role recognition on the multi-channel audio data to identify the corresponding role (i.e., the corresponding speaker), and then adds role tags to the multi-channel audio data to obtain role audio data. In the voice recognition method of the present application, since the sending end sends single-channel audio data to the receiving end, this ensures that the bandwidth resources occupied during the transmission of the audio data are less, and it also ensures that the sending end can transmit the single-channel audio data to the receiving end relatively quickly. Since the receiving end does not need to perform role recognition, etc., it only needs to perform voice recognition on the single-channel audio data, which also ensures that the calculation amount of the receiving end is less, thus solving the problem of relatively wasting bandwidth resources caused by the sending end sending multi-channel audio to the receiving end in the prior art.
[0032] In the actual application process, after the sending end (such as a microphone array, an audio recording device, etc.) collects multi-channel audio data, it performs role recognition on the collected multi-channel audio data through the microphone array, the audio recording device or related edge devices to obtain the role corresponding to the multi-channel audio data (i.e., the corresponding speaker). Then, role tags are added to the multi-channel audio data to obtain role audio data. The sending end encapsulates multiple role audio data to obtain a single-channel audio data. Finally, the single-channel audio data is sent to the receiving end. After the receiving end receives the single-channel audio data, it performs voice recognition processing on the single-channel audio data to obtain the voice recognition text information corresponding to each role. In the voice recognition method of the present application, since the sending end sends single-channel audio data to the receiving end instead of multi-channel audio data, this ensures that the bandwidth resources occupied during the transmission process are less. At the same time, in the voice recognition method of the present application, the process of role recognition is placed at the sending end. In this way, for the receiving end, it only needs to perform voice recognition on the single-channel audio data, ensuring that the calculation amount and the computing resources occupied during the voice recognition of the receiving end are less. For the receiving end, in addition to having role tags, the original multi-channel audio data in the received single-channel audio data is not modified, which can retain the authenticity of the multi-channel audio data. The receiving end only needs to perform voice recognition, which also ensures that the voice recognition effect and accuracy of the receiving end are better.
[0033] In a specific embodiment of the present application, the above multi-channel audio data may be audio data within a predetermined duration collected by the sending end. The above predetermined duration may be 1 to 200 ms. Of course, the above predetermined duration is not limited to 1 to 200 ms, and may also be other suitable durations. In the present application, the size of the above predetermined duration is not limited, and can be specifically adjusted according to the actual situation.
[0034] In another specific embodiment of the present application, the above role marker may be hexadecimal role information. Of course, it is not limited to using hexadecimal role information as the role marker. For example, octal role information may also be used as the role marker, etc. That is, the present application does not limit the actual representation form of the role marker, as long as the receiving end can identify the corresponding role (speaker) in each role audio data subsequently.
[0035] In the actual application process, the above role marker can be added to the lower byte part of the multi-channel audio data. For example, Figure 2 As shown in the single-channel audio data 103, the single-channel audio data 103 is obtained by encapsulating multiple role audio data 102. Among them, the role audio data 102 is obtained by encapsulating the multi-channel audio data 100 and the role marker 101. Of course, in the actual application process, the role marker 101 is not limited to being located after the multi-channel audio data 100, that is, the lower byte part of the role audio data. The role marker 101 can also be before the multi-channel audio data 100, that is, the upper byte part of the role audio data. In the present application, the specific position of the above role marker 101 in the multi-channel audio data is not limited, as long as each multi-channel audio data can be marked with a role.
[0036] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0037] In one embodiment of the present application, the above-mentioned role audio data is obtained by adding the above-mentioned role tags to the above-mentioned multi-channel audio data. The process of adding the above-mentioned role tags is as follows: the sending end performs role recognition on each of the above-mentioned multi-channel audio data to obtain the role information corresponding to each of the above-mentioned multi-channel audio data; then, according to the above-mentioned role information corresponding to each of the above-mentioned multi-channel audio data, the corresponding above-mentioned role tags are added to the above-mentioned multi-channel audio data to obtain a plurality of the above-mentioned role audio data. In this embodiment, the sending end performs role recognition on the collected multi-channel audio data, adds the corresponding role tags to obtain a plurality of role audio data, and then the sending end encapsulates the plurality of role audio data to obtain single-channel audio data and sends it to the receiving end, further ensuring that the single-channel audio data transmitted between the sending end and the receiving end occupies less bandwidth and further reducing the transmission latency.
[0038] In the actual application process, the sending end can use methods such as phase difference, amplitude difference, or fundamental frequency detection to perform role recognition on the multi-channel audio data. Of course, it is not limited to the above-listed methods such as phase difference, amplitude difference, or fundamental frequency detection to perform role recognition on the multi-channel audio data, and it can use any feasible method in the prior art to perform role recognition on the multi-channel audio data.
[0039] In order to further obtain the speech recognition text information more accurately, in another embodiment of the present application, the above-mentioned single-channel audio data is subjected to speech recognition processing to obtain the speech recognition text information of each role, including: determining, according to the above-mentioned role tags and audio bit rate, the audio start and end time ranges and byte offsets corresponding to each of the above-mentioned roles in the above-mentioned single-channel audio data, where the above-mentioned audio start and end time ranges are the timestamps of the audio data corresponding to each of the above-mentioned roles; using a speech recognition algorithm to perform speech recognition on the above-mentioned single-channel audio data to obtain target text information and phrase time ranges, where the above-mentioned phrase time ranges are the timestamps of each phrase in the above-mentioned target text information; and obtaining the above-mentioned speech recognition text information corresponding to each of the above-mentioned roles according to the above-mentioned audio start and end time ranges and the above-mentioned byte offsets of each of the above-mentioned roles, as well as the above-mentioned target text information and the above-mentioned phrase time ranges.
[0040] Specifically, in the above-mentioned embodiment, the above-mentioned phrase can also be referred to as a short phrase, and a phrase is a certain encapsulation relationship composed of two or more words. For example, "open", "close", "yes", "very good", etc. That is, the division of the phrases in the present application is based on Chinese pinyin. Of course, in the actual application process, the above-mentioned phrase time range can also be the time range corresponding to one character.
[0041] In another embodiment of the present application, according to the above-mentioned role tags and audio bit rate, the audio start and end time ranges and byte offsets corresponding to each of the above-mentioned roles in the above-mentioned single-channel audio data are determined, including: performing segmentation processing on the above-mentioned single-channel audio data to obtain a plurality of the above-mentioned role audio data; determining the corresponding above-mentioned audio start and end time ranges according to the above-mentioned role tags corresponding to each of the above-mentioned role audio data; and determining the byte offsets corresponding to each of the above-mentioned roles according to each of the above-mentioned audio start and end time ranges and the above-mentioned audio bit rate. In this embodiment, the receiving end performs segmentation processing on the received single-channel audio data to obtain a plurality of role audio data, and then determines the audio start and end time ranges corresponding to each role according to the role tags corresponding to each role audio data, so as to ensure that the obtained audio start and end time ranges are relatively accurate. Then, according to the audio start and end time ranges and the audio bit rate, the byte offsets corresponding to each role are determined, so as to ensure that the obtained byte offsets are relatively accurate, and further ensure that the subsequent obtained speech recognition text information is relatively accurate.
[0042] In the actual application process, the above-mentioned single-channel audio data can be segmented according to the byte size to obtain a plurality of the above-mentioned role audio data. Of course, the above-mentioned single-channel audio data can also be segmented according to a predetermined duration. In the present application, the method for segmenting the above-mentioned single-channel audio data is not limited, and any feasible method in the prior art can be used to segment the above-mentioned single-channel audio data, as long as the audio data of each role is obtained.
[0043] In a specific embodiment of the present application, the sending end encapsulates the multi-channel audio data of 16k 16-bit uncompressed at every 20 ms, and transmits the audio data of 200 ms at a time. That is to say, the duration of the single-channel audio data received by the receiving end is 200 ms. After receiving the single-channel audio data, the receiving end can segment it according to a predetermined duration of 20 ms to obtain 10 role audio data. Starting from the first role audio data to the tenth role audio data, the audio start and end time ranges corresponding to each role audio data are 0 ms to 20 ms, 20 ms to 40 ms, 40 ms to 60 ms, 60 ms to 80 ms, 80 ms to 100 ms, 100 ms to 120 ms, 120 ms to 140 ms, 140 to 160 ms, 160 ms to 180 ms, and 180 ms to 200 ms respectively. Of course, in the actual application process, the above-mentioned consecutive several audio start and end time ranges can correspond to the same role.
[0044] In yet another embodiment of the present application, according to the above-mentioned audio start and end time ranges and the above-mentioned byte offsets of each of the above-mentioned roles, as well as the above-mentioned target text information and the above-mentioned phrase time ranges, the above-mentioned speech recognition text information corresponding to each of the above-mentioned roles is obtained, including: starting from the above-mentioned phrase time range corresponding to the first above-mentioned phrase in the above-mentioned target text information, comparing each of the above-mentioned phrase time ranges with the above-mentioned audio start and end time ranges to obtain the preset text information of each of the above-mentioned roles; decoding the preset text information of each of the above-mentioned roles according to the above-mentioned byte offsets of each of the above-mentioned roles to obtain the above-mentioned speech recognition text information corresponding to each of the above-mentioned roles, which further ensures the accuracy of the obtained speech recognition text information.
[0045] In a specific embodiment of the present application, before decoding the preset text information of each of the above-mentioned roles according to the above-mentioned byte offsets of each of the above-mentioned roles to obtain the above-mentioned speech recognition text information corresponding to each of the above-mentioned roles, the above-mentioned speech recognition method further includes: combining multiple consecutive preset text information corresponding to the same role to obtain the combined preset text information.
[0046] Specifically, in the actual application process, the above-mentioned role and the above-mentioned speech recognition text information are in one-to-one correspondence, that is, when there are three roles (speakers), the finally obtained speech recognition text information is also three. For the receiving end, although the received is single-channel audio data, after speech recognition, the receiving end will restore the single-channel audio data to three-channel audio data.
[0047] Specifically, in the above-mentioned embodiment, starting from the phrase time range corresponding to the first phrase in the target text information, the specific process of comparing each phrase time range with the audio start and end time ranges to obtain the preset text information of each role is as follows: comparing the phrase time range corresponding to the phrase with each audio start and end time range. If the phrase time range is within a certain audio start and end time range, then the phrase corresponding to the phrase time range is the phrase within the audio start and end time range. Then, combining the phrases within the same audio start and end time range to obtain multiple preset text information. In addition, since punctuation marks are the text expression products after speech recognition, have no direct correspondence with the audio, and do not occupy the time in the audio data, the punctuation marks can be merged upwards, that is, the punctuation marks are merged with the phrase before the punctuation mark, which can ensure that the punctuation marks will not be separated to the beginning of the sentence of the preset text information of a certain role.
[0048] In the actual application process, when the speaker (i.e., the role) is making a speech, there will be a situation where the pronunciation of the last phrase is elongated. In this case, there is a conflict between the phrase recognized by speech recognition and the role information marked at the sending end. That is, when the time range corresponding to this phrase straddles two roles at the same time, it will be misjudged as belonging to two adjacent roles. Therefore, in order to further ensure that the obtained preset text information is relatively accurate, in an embodiment of the present application, starting from the phrase time range corresponding to the first phrase in the above target text information, each of the above phrase time ranges is compared with the above audio start and end time ranges to obtain the preset text information of each of the above roles, including: when the target phrase time range has an intersection with both of the above audio start and end time ranges, determining a target time point, where the above target time is the overlapping time point of the two above audio start and end time ranges, and the above target phrase time range is one of the above multiple phrase time ranges; calculating the difference between the upper time point of the above target phrase time range and the above target time point to obtain a first difference, and calculating the difference between the lower time point of the above target phrase time range and the above target time point to obtain a second difference; according to the comparison result of the above first difference and the second difference, determining the above audio start and end time range to which the phrase corresponding to the above target phrase time range belongs, so as to obtain the above preset text information of each of the above roles.
[0049] In order to enable those skilled in the art to clearly and clearly understand the above technical solution, the following will be explained with a specific embodiment. When the first audio start and end time range is 0ms to 20ms, the second audio start and end time range is 20ms to 40ms, and the target phrase time range is 18ms to 21ms, that is to say, this target phrase time range includes the demarcation point of the two audio start and end time ranges. In this case, it is necessary to determine the target time point, that is, the overlapping time point of the first audio start and end time range and the second audio start and end time range, which is 20ms. Then, calculate the difference between 18ms (i.e., the upper time point of the target phrase time range) and 20ms (i.e., the target time point) to obtain a first difference (2ms), and calculate the difference between 21ms (i.e., the lower time point of the target phrase time range) and 20ms (i.e., the target time point) to obtain a second difference (1ms). Since 2ms is greater than 1ms, the target phrase corresponding to the target phrase time range belongs to the first audio start and end time range, which further ensures that the obtained preset text information is relatively accurate.
[0050] In another embodiment of the present application, before determining the corresponding audio start and end time ranges according to the above-mentioned role tags corresponding to each role audio data, the speech recognition method further includes: using a sliding window algorithm to smooth each of the role audio data to obtain a plurality of smoothed role audio data, which can avoid frequent changes of roles in a short period of time.
[0051] An embodiment of the present application further provides a speech recognition device. It should be noted that the speech recognition device of the embodiment of the present application can be used to execute the speech recognition method provided by the embodiment of the present application. The speech recognition device provided by the embodiment of the present application is introduced below.
[0052] Figure 3 It is a schematic structural diagram of a speech recognition device according to an embodiment of the present application. The speech recognition device is applied in a receiving end, such as Figure 3 shown, the speech recognition device includes:
[0053] A receiving unit 10, configured to receive single-channel audio data sent by a sending end, where the single-channel audio data is single-channel audio data encapsulated from a plurality of role audio data, the role audio data is multi-channel audio data with role tags, and the multi-channel audio data is audio data collected by the sending end;
[0054] A processing unit 20, configured to perform speech recognition processing on the single-channel audio data to obtain speech recognition text information of each role.
[0055] The above voice recognition device is applied in the receiving end. The voice recognition device includes a receiving unit and a processing unit. Among them, the receiving unit is used to receive the single-channel audio data sent by the sending end, and the processing unit is used to perform voice recognition processing on the received single-channel audio data to obtain the voice recognition text information corresponding to each role. Among them, the above single-channel audio data is a single-channel audio data obtained by encapsulating multiple role audio data, and the role audio data has role tags. That is to say, after the sending end collects multi-channel audio data, it performs role recognition on the multi-channel audio data to identify the corresponding role (i.e., the corresponding speaker), and then adds role tags to the multi-channel audio data to obtain role audio data. In the voice recognition device of the present application, since the sending end sends single-channel audio data to the receiving end, this ensures that the bandwidth resources occupied during the transmission of the audio data are less, and it also ensures that the sending end can transmit the single-channel audio data to the receiving end relatively quickly. Since the receiving end does not need to perform role recognition, etc., and only needs to perform voice recognition on the single-channel audio data, this also ensures that the calculation amount of the receiving end is less, thus solving the problem of relatively wasting bandwidth resources caused by the sending end sending multi-channel audio to the receiving end in the prior art.
[0056] In the actual application process, after the sending end (such as a microphone array, an audio recording device, etc.) collects multi-channel audio data, it can perform role recognition on the collected multi-channel audio data through the microphone array, the audio recording device or related edge devices to obtain the role corresponding to the multi-channel audio data (i.e., the corresponding speaker). Then, role tags are added to the multi-channel audio data to obtain role audio data. The sending end encapsulates multiple role audio data to obtain a single-channel audio data. Finally, the single-channel audio data is sent to the receiving end. After the receiving end receives the single-channel audio data, it performs voice recognition processing on the single-channel audio data to obtain the voice recognition text information corresponding to each role. In the voice recognition method of the present application, since the sending end sends single-channel audio data to the receiving end instead of multi-channel audio data, this ensures that the bandwidth resources occupied during the transmission process are less. At the same time, the voice recognition method of the present application places the role recognition process at the sending end. In this way, for the receiving end, it only needs to perform voice recognition on the single-channel audio data, ensuring that the calculation amount and the computing resources occupied during the voice recognition of the receiving end are less. For the receiving end, in addition to having role tags in the received single-channel audio data, the original multi-channel audio data is not modified, so the authenticity of the multi-channel audio data can be retained. The receiving end only needs to perform voice recognition, which also ensures that the voice recognition effect and accuracy of the receiving end are better.
[0057] In a specific embodiment of the present application, the above multi-channel audio data may be audio data within a predetermined duration collected by the sending end. The above predetermined duration may be 10 to 200 ms. Of course, the above predetermined duration is not limited to 1 to 200 ms, and may also be other appropriate durations. In the present application, the size of the above predetermined duration is not limited, and can be specifically adjusted according to the actual situation.
[0058] In another specific embodiment of the present application, the above role marker may be role information in hexadecimal. Of course, it is not limited to using role information in hexadecimal as the role marker. For example, role information in octal may also be used as the role marker, etc. That is, the present application does not limit the actual representation form of the role marker, as long as the receiving end can identify the corresponding role (speaker) in each role audio data subsequently.
[0059] In the actual application process, the above role marker may be added to the lower byte part of the multi-channel audio data. For example, Figure 2 as shown in the single-channel audio data 103, which is encapsulated from multiple role audio data 102. Among them, the role audio data 102 is encapsulated from the multi-channel audio data 100 and the role marker 101. Of course, in the actual application process, the role marker 101 is not limited to being located after the multi-channel audio data 100, that is, the lower byte part of the role audio data. The role marker 101 may also be before the multi-channel audio data 100, that is, the upper byte part of the role audio data. In the present application, the specific position of the above role marker 101 in the multi-channel audio data is not limited, as long as the role marking of each multi-channel audio data can be realized.
[0060] In an embodiment of the present application, the above role audio data is obtained by adding the above role marker to the above multi-channel audio data. The process of adding the above role marker is as follows: the sending end performs role recognition on each of the above multi-channel audio data to obtain the role information corresponding to each of the above multi-channel audio data; then, according to the above role information corresponding to each of the above multi-channel audio data, the corresponding above role marker is added to the above multi-channel audio data to obtain multiple above role audio data. In this embodiment, the sending end performs role recognition on the collected multi-channel audio data, adds the corresponding role marker to obtain multiple role audio data, and then the sending end encapsulates the multiple role audio data to obtain single-channel audio data and sends it to the receiving end, further ensuring that the single-channel audio data transmitted between the sending end and the receiving end occupies less bandwidth and further reducing the transmission latency.
[0061] In the actual application process, the sending end can use methods such as phase difference, amplitude difference, or fundamental frequency detection to perform role recognition on multi-channel audio data. Of course, it is not limited to the methods of phase difference, amplitude difference, or fundamental frequency detection listed above for role recognition of multi-channel audio data. Any feasible method in the prior art can be used to perform role recognition on multi-channel audio data.
[0062] In another embodiment of the present application, in order to more accurately obtain the speech recognition text information, the above processing unit includes a first determination module, a recognition module, and a second determination module. Among them, the first determination module is used to determine, according to the above role mark and audio bit rate, the audio start and end time ranges and byte offsets corresponding to each of the above roles in the above single-channel audio data. The above audio start and end time ranges are the timestamps of the audio data corresponding to each of the above roles; the recognition module is used to perform speech recognition on the above single-channel audio data using a speech recognition algorithm to obtain target text information and phrase time ranges. The above phrase time ranges are the timestamps of each phrase in the above target text information; the second determination module is used to obtain the above speech recognition text information corresponding to each of the above roles according to the above audio start and end time ranges and the above byte offsets of each of the above roles, as well as the above target text information and the above phrase time ranges.
[0063] Specifically, in the above embodiment, the above phrase can also be referred to as a short phrase. A phrase is a certain encapsulation relationship composed of two or more words. For example, "open", "close", "yes", "very good", etc. That is, the division of the phrases in the present application is based on Chinese pinyin. Of course, in the actual application process, the above phrase time range can also be the time range corresponding to one character.
[0064] In another embodiment of the present application, the first determination module includes a segmentation sub-module, a first determination sub-module, and a second determination sub-module. Among them, the segmentation sub-module is used to perform segmentation processing on the above single-channel audio data to obtain multiple above role audio data; the first determination sub-module is used to determine the corresponding above audio start and end time ranges according to the above role marks corresponding to each of the above role audio data; the second determination sub-module is used to determine the byte offsets corresponding to each of the above roles according to the above audio start and end time ranges and the above audio bit rate. In this embodiment, the receiving end performs segmentation processing on the received single-channel audio data to obtain multiple role audio data, and then determines the audio start and end time ranges corresponding to each role according to the role marks corresponding to each role audio data, so as to ensure that the obtained audio start and end time ranges are relatively accurate. Then, according to the audio start and end time ranges and the audio bit rate, the byte offsets corresponding to each role are determined, so as to ensure that the obtained byte offsets are relatively accurate, and further ensure that the subsequent obtained speech recognition text information is relatively accurate.
[0065] In the actual application process, the above single-channel audio data can be segmented according to the byte size to obtain multiple pieces of the above character audio data. Of course, the above single-channel audio data can also be segmented according to a predetermined duration. In this application, the method for segmenting the above single-channel audio data is not limited, and any feasible method in the prior art can be used to segment the above single-channel audio data, as long as the audio data of each character is obtained.
[0066] In a specific embodiment of this application, the sending end encapsulates the multi-channel audio data of 16k 16-bit uncompressed every 20 ms, and transmits the audio data of 200 ms at a time. That is to say, the duration of the single-channel audio data received by the receiving end is 200 ms. After receiving the single-channel audio data, the receiving end can segment it according to a predetermined duration of 20 ms to obtain 10 pieces of character audio data. Starting from the first piece of character audio data to the tenth piece of character audio data, the audio start and end time ranges corresponding to each piece of character audio data are 0 ms to 20 ms, 20 ms to 40 ms, 40 ms to 60 ms, 60 ms to 80 ms, 80 ms to 100 ms, 100 ms to 120 ms, 120 ms to 140 ms, 140 to 160 ms, 160 ms to 180 ms, and 180 ms to 200 ms respectively. Of course, in the actual application process, the above-mentioned consecutive several audio start and end time ranges can correspond to the same character.
[0067] In another embodiment of this application, the above second determination module includes a comparison sub-module and a third determination sub-module. Among them, the above comparison sub-module is used to start from the above phrase time range corresponding to the first above phrase in the above target text information, compare each above phrase time range with the above audio start and end time ranges to obtain the preset text information of each above character; the above third determination sub-module is used to decode the preset text information of each above character according to the above byte offset of each above character to obtain the above speech recognition text information of each above character, which further ensures that the obtained speech recognition text information is relatively accurate.
[0068] In a specific embodiment of this application, before decoding the preset text information of each above character according to the above byte offset of each above character to obtain the above speech recognition text information of each above character, the above speech recognition method further includes: merging multiple consecutive preset text information corresponding to the same character to obtain the merged preset text information.
[0069] Specifically, in the actual application process, the above-mentioned roles and the above-mentioned speech recognition text information are in one-to-one correspondence. That is, when there are three roles (speakers), the finally obtained speech recognition text information is also three. For the receiving end, although the received is single-channel audio data, after speech recognition, the receiving end will restore the single-channel audio data to three-channel audio data.
[0070] Specifically, in the above-mentioned embodiments, starting from the phrase time range corresponding to the first phrase in the target text information, the specific process of comparing each phrase time range with the audio start and end time ranges to obtain the preset text information of each role is as follows: Compare the phrase time range corresponding to the phrase with each audio start and end time range. If the phrase time range is within a certain audio start and end time range, then the phrase corresponding to the phrase time range is the phrase within the audio start and end time range. Then combine the phrases within the same audio start and end time range to obtain multiple preset text information. In addition, since punctuation marks are the text expression products after speech recognition, have no direct correspondence with the audio, and do not occupy time in the audio data, the punctuation marks can be merged upwards, that is, the punctuation marks are merged with the phrase before the punctuation mark, so as to ensure that the punctuation marks will not be assigned to the beginning of the sentence of the preset text information of a certain role.
[0071] In the actual application process, when the speaker (i.e., the role) is speaking, there will be a situation where the pronunciation of the last phrase is elongated. In this case, there is a conflict between the phrases recognized by speech recognition and the role information marked by the sending end, that is, when the time range corresponding to this phrase straddles two roles at the same time, it will be misjudged as belonging to two adjacent roles. Therefore, in order to further ensure that the obtained preset text information is relatively accurate, in an embodiment of the present application, the above-mentioned comparison sub-module includes a fourth determination sub-module, a calculation sub-module, and a fifth determination sub-module. Among them, the above-mentioned fourth determination sub-module is used to determine the target time point when the target phrase time range has an intersection with both of the above-mentioned audio start and end time ranges. The above-mentioned target time is the overlapping time point of the two above-mentioned audio start and end time ranges, and the above-mentioned target phrase time range is one of the multiple above-mentioned phrase time ranges; the above-mentioned calculation sub-module is used to calculate the difference between the upper limit time point of the above-mentioned target phrase time range and the above-mentioned target time point to obtain a first difference, and calculate the difference between the lower limit time point of the above-mentioned target phrase time range and the above-mentioned target time point to obtain a second difference; the above-mentioned fifth determination sub-module is used to determine the audio start and end time range to which the phrase corresponding to the above-mentioned target phrase time range belongs according to the comparison result of the above-mentioned first difference and the second difference, so as to obtain the above-mentioned preset text information of each of the above-mentioned roles.
[0072] In order to enable those skilled in the art to clearly and comprehensively understand the above technical solutions, the following will explain with a specific embodiment. In the first audio start and end time range of 0 ms to 20 ms, the second audio start and end time range of 20 ms to 40 ms, and the target phrase time range of 18 ms to 21 ms. That is to say, the target phrase time range includes the demarcation point of the two audio start and end time ranges. In this case, it is necessary to determine the target time point, that is, the overlapping time point of the first audio start and end time range and the second audio start and end time range, which is 20 ms. Then, calculate the difference between 18 ms (i.e., the upper limit time point of the target phrase time range) and 20 ms (i.e., the target time point) to obtain the first difference (2 ms), and calculate the difference between 21 ms (i.e., the lower limit time point of the target phrase time range) and 20 ms (i.e., the target time point) to obtain the second difference (1 ms). Since 2 ms is greater than 1 ms, the target phrase corresponding to the target phrase time range belongs to the first audio start and end time range, which further ensures the accuracy of the obtained preset text information.
[0073] In another embodiment of the present application, the above voice recognition device further includes a smoothing processing unit, which is used to perform smoothing processing on each of the above role audio data by using a sliding window algorithm before determining the corresponding above audio start and end time range according to the above role markers corresponding to each of the above role audio data, so as to obtain multiple above role audio data after smoothing processing, which can avoid frequent changes of roles in a short period of time.
[0074] The above voice recognition device includes a processor and a memory. The above receiving unit and processing unit and the like are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions.
[0075] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem of waste of bandwidth resources caused by the sending end sending multi-channel audio to the receiving end in the prior art can be solved.
[0076] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory includes at least one storage chip.
[0077] The embodiment of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the above voice recognition method is implemented.
[0078] An embodiment of the present invention provides a processor, which is used to run a program. When the program runs, the above-mentioned speech recognition method is executed.
[0079] In a typical embodiment of the present application, a speech recognition system is further provided. The speech recognition system includes a receiving end and a sending end. The receiving end is communicatively connected to the sending end, and the receiving end is used to execute any one of the above-mentioned speech recognition methods.
[0080] The above-mentioned speech recognition system includes a receiving end and a sending end. The receiving end is used to execute any one of the above-mentioned speech recognition methods. In the above-mentioned speech recognition method, when the receiving end receives the single-channel audio data sent by the sending end, it performs speech recognition processing on the received single-channel audio data to obtain the speech recognition text information corresponding to each role. Among them, the above-mentioned single-channel audio data is a single-channel audio data obtained by encapsulating multiple role audio data, and the role audio data has a role mark. That is to say, after the sending end collects the multi-channel audio data, it performs role recognition on the multi-channel audio data to identify the corresponding role (i.e., the corresponding speaker), and then adds a role mark to the multi-channel audio data to obtain the role audio data. In the speech recognition method of the present application, since the sending end sends single-channel audio data to the receiving end, this ensures that the bandwidth resources occupied during the transmission of the audio data are less, and it also ensures that the sending end can transmit the single-channel audio data to the receiving end relatively quickly. Since the receiving end does not need to perform role recognition, etc., and only needs to perform speech recognition on the single-channel audio data, this also ensures that the computational amount of the receiving end is less, thus solving the problem of relatively wasteful bandwidth resources caused by the sending end sending multi-channel audio to the receiving end in the prior art.
[0081] An embodiment of the present invention provides a device. The device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it realizes at least the following steps:
[0082] Step S101: Receive the single-channel audio data sent by the sending end. Among them, the above-mentioned single-channel audio data is a single-channel audio data obtained by encapsulating multiple role audio data, the above-mentioned role audio data is multi-channel audio data with a role mark, and the above-mentioned multi-channel audio data is the audio data collected by the above-mentioned sending end;
[0083] Step S102: Perform speech recognition processing on the above-mentioned single-channel audio data to obtain the speech recognition text information of each role.
[0084] The device in this article can be a server, a PC, a PAD, a mobile phone, etc.
[0085] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps:
[0086] Step S101: Receive the single-channel audio data sent by the sending end. Among them, the single-channel audio data is the single-channel audio data encapsulated from multiple role audio data, the role audio data is the multi-channel audio data with role tags, and the multi-channel audio data is the audio data collected by the sending end;
[0087] Step S102: Perform speech recognition processing on the single-channel audio data to obtain the speech recognition text information of each role.
[0088] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0089] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the above division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0090] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0091] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0092] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above method in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0093] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0094] 1) The speech recognition method of the present application is applied in the receiving end. When the receiving end receives the single-channel audio data sent by the sending end, it performs speech recognition processing on the received single-channel audio data to obtain the speech recognition text information corresponding to each role. Among them, the above single-channel audio data is a single-channel audio data obtained by encapsulating multiple role audio data, and the role audio data has a role mark. That is to say, after the sending end collects multi-channel audio data, it performs role recognition on the multi-channel audio data to identify the corresponding role (i.e., the corresponding speaker), and then adds a role mark to the multi-channel audio data to obtain the role audio data. In the speech recognition method of the present application, since the sending end sends single-channel audio data to the receiving end, this ensures that the bandwidth resources occupied during the transmission of the audio data are less, and it also ensures that the sending end can transmit the single-channel audio data to the receiving end relatively quickly. Since the receiving end does not need to perform role recognition, etc., and only needs to perform speech recognition on the single-channel audio data, this also ensures that the calculation amount of the receiving end is less, thus solving the problem of relatively wasteful bandwidth resources caused by the sending end sending multi-channel audio to the receiving end in the prior art.
[0095] 2) The voice recognition device of the present application is applied in the receiving end, and the voice recognition device includes a receiving unit and a processing unit. Among them, the receiving unit is used to receive the single-channel audio data sent by the sending end, and the processing unit is used to perform voice recognition processing on the received single-channel audio data to obtain the voice recognition text information corresponding to each role. Among them, the above single-channel audio data is a single-channel audio data obtained by encapsulating multiple role audio data, and the role audio data has role tags. That is to say, after the sending end collects multi-channel audio data, it performs role recognition on the multi-channel audio data to identify the corresponding role (i.e., the corresponding speaker), and then adds role tags to the multi-channel audio data to obtain role audio data. In the voice recognition device of the present application, since the sending end sends single-channel audio data to the receiving end, this ensures that the bandwidth resources occupied during the transmission of audio data are less, and it also ensures that the sending end can transmit the single-channel audio data to the receiving end relatively quickly. Since the receiving end does not need to perform role recognition, etc., and only needs to perform voice recognition on the single-channel audio data, this also ensures that the calculation amount of the receiving end is less, thus solving the problem of relatively wasting bandwidth resources caused by the sending end sending multi-channel audio to the receiving end in the prior art.
[0096] 3) The voice recognition system of the present application includes a receiving end and a sending end. Among them, the receiving end is used to execute any one of the above voice recognition methods. In the above voice recognition method, when the receiving end receives the single-channel audio data sent by the sending end, it performs voice recognition processing on the received single-channel audio data to obtain the voice recognition text information corresponding to each role. Among them, the above single-channel audio data is a single-channel audio data obtained by encapsulating multiple role audio data, and the role audio data has role tags. That is to say, after the sending end collects multi-channel audio data, it performs role recognition on the multi-channel audio data to identify the corresponding role (i.e., the corresponding speaker), and then adds role tags to the multi-channel audio data to obtain role audio data. In the voice recognition method of the present application, since the sending end sends single-channel audio data to the receiving end, this ensures that the bandwidth resources occupied during the transmission of audio data are less, and it also ensures that the sending end can transmit the single-channel audio data to the receiving end relatively quickly. Since the receiving end does not need to perform role recognition, etc., and only needs to perform voice recognition on the single-channel audio data, this also ensures that the calculation amount of the receiving end is less, thus solving the problem of relatively wasting bandwidth resources caused by the sending end sending multi-channel audio to the receiving end in the prior art.
[0097] The above are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A speech recognition method, characterized in that, the speech recognition method is applied in a receiving end, and the speech recognition method includes: receiving single-channel audio data sent by a sending end, where the single-channel audio data is single-channel audio data encapsulated from multiple character audio data, and the character audio data is multi-channel audio data with character markers, and the multi-channel audio data is audio data collected by the sending end; determining, according to the character markers and the audio bit rate, in the single-channel audio data, the audio start and end time ranges and byte offsets corresponding to each character, where the audio start and end time ranges are the timestamps of the audio data corresponding to each character; using a speech recognition algorithm to perform speech recognition on the single-channel audio data to obtain target text information and phrase time ranges, where the phrase time ranges are the timestamps of each phrase in the target text information; in the case where the target phrase time range has an intersection with both of the two audio start and end time ranges, determining a target time point, where the target time is the overlapping time point of the two audio start and end time ranges, and the target phrase time range is one of the multiple phrase time ranges; calculating the difference between the upper time point of the target phrase time range and the target time point to obtain a first difference, and calculating the difference between the lower time point of the target phrase time range and the target time point to obtain a second difference; according to the comparison result of the first difference and the second difference, determining the audio start and end time range to which the phrase corresponding to the target phrase time range belongs, to obtain the preset text information of each character; decoding the preset text information of each character according to the byte offset of each character to obtain the speech recognition text information of each character.
2. The speech recognition method according to claim 1, characterized in that, the character audio data is obtained by adding the character markers to the multi-channel audio data, and the process of adding the character markers is: the sending end performs character recognition on each multi-channel audio data to obtain the character information corresponding to each multi-channel audio data; then, according to the character information corresponding to each multi-channel audio data, adding the corresponding character markers to the multi-channel audio data to obtain multiple character audio data.
3. The speech recognition method according to claim 1, characterized in that, determining, according to the character markers and the audio bit rate, in the single-channel audio data, the audio start and end time ranges and byte offsets corresponding to each character includes: performing segmentation processing on the single-channel audio data to obtain multiple character audio data; determining the corresponding audio start and end time ranges according to the character markers corresponding to each character audio data; determining the byte offsets corresponding to each character according to each audio start and end time range and the audio bit rate.
4. The speech recognition method according to claim 3, characterized in that, Before determining the corresponding audio start and end time ranges according to the role tags corresponding to the respective role audio data, the speech recognition method further includes: Using a sliding window algorithm to smooth each of the role audio data to obtain multiple smoothed role audio data.
5. A speech recognition device Characterized in that The speech recognition device is applied in a receiving end, and the speech recognition device includes: A receiving unit, configured to receive single-channel audio data sent by a sending end, where the single-channel audio data is single-channel audio data encapsulated from multiple role audio data, the role audio data is multi-channel audio data with role tags, and the multi-channel audio data is audio data collected by the sending end; A processing unit, configured to determine, according to the role tags and the audio bit rate, in the single-channel audio data, the audio start and end time ranges and byte offsets corresponding to each role, where the audio start and end time ranges are time stamps of the audio data corresponding to each role; use a speech recognition algorithm to perform speech recognition on the single-channel audio data to obtain target text information and phrase time ranges, where the phrase time ranges are time stamps of each phrase in the target text information; in a case where the target phrase time range has an intersection with both of the two audio start and end time ranges, determine a target time point, where the target time is the overlapping time point of the two audio start and end time ranges, and the target phrase time range is one of the multiple phrase time ranges; calculate a difference between an upper time point of the target phrase time range and the target time point to obtain a first difference, and calculate a difference between a lower time point of the target phrase time range and the target time point to obtain a second difference; determine, according to a comparison result of the first difference and the second difference, the audio start and end time range to which the phrase corresponding to the target phrase time range belongs, to obtain the preset text information of each role; and decode the preset text information of each role according to the byte offset of each role to obtain the speech recognition text information of each role.
6. A computer-readable storage medium Characterized in that The computer-readable storage medium includes a stored program, where the program executes the speech recognition method according to any one of claims 1 to 4.
7. A speech recognition system Characterized in that Includes: A receiving end, where the receiving end is configured to execute the speech recognition method according to any one of claims 1 to 4; A sending end, communicatively connected to the receiving end.
Citation Information
Patent Citations
Audio transmission method and device and computer readable storage medium
CN112346700A
Voice signal processing method, device and system and computer readable storage medium
CN113012700A
Method and device for converting single-channel audio into text, electronic equipment and storage medium
CN114495941A
Audio annotation method, annotation device and annotation system
CN115662402A