Audio recognition method, apparatus, device, and storage medium

By calculating and fusing the correlation of multiple audio signals in the time and channel dimensions, the problem of inaccurate audio recognition when multiple speakers speak simultaneously is solved, thus improving recognition accuracy.

CN115641835BActive Publication Date: 2026-03-27ALIBABA DAMO (HANGZHOU) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-30
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

When multiple speakers are speaking at the same time, existing automatic speech recognition technology struggles to accurately identify multiple audio signals, resulting in inaccurate recognition results.

Method used

An audio recognition method is adopted, which acquires multiple raw audio signals, calculates the correlation in the time and channel dimensions using the first attention mechanism, outputs multiple target audio signals, and then fuses them to predict text information.

Benefits of technology

It improves the accuracy of audio signal recognition when multiple speakers are speaking at the same time by making full use of correlation in the time and channel dimensions to obtain finer-grained channel correlation, while taking into account the contextual information of each frame of audio signal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115641835B_ABST
    Figure CN115641835B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio recognition method, device, equipment and storage medium. The present disclosure obtains multiple original audio signals, and processes the multiple original audio signals according to the context information of each frame of original audio signal in the time dimension, the first correlation between multiple frames of original audio signals with the same time, and the second correlation between the context information of each of the multiple frames of original audio signals with the same time and each frame of original audio signal with the same time, to obtain multiple target audio signals, so that the correlation between the multiple original audio signals can be calculated in the time dimension and the channel dimension at the same time. Since the correlation between the multiple original audio signals in the time dimension and the channel dimension is fully utilized, not only the correlation between the channels with finer granularity can be obtained, but also the context information of each frame of original audio signal is considered, so that when multiple speakers speak at the same time, the recognition accuracy of the multiple original audio signals is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of information technology, and in particular, to an audio recognition method and device, equipment and a storage medium. BACKGROUND

[0002] The current automatic speech recognition (ASR) is widely used, for example, in a conference scenario, the audio of a speaker can be converted into text by ASR.

[0003] However, when multiple speakers speak at the same time, the recognition result of ASR is not accurate enough. SUMMARY

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an audio recognition method, device, equipment and storage medium to improve the recognition accuracy of multiple original audio signals.

[0005] In a first aspect, an embodiment of the present disclosure provides an audio recognition method, comprising:

[0006] Obtaining multiple original audio signals, each original audio signal comprising multiple frames of original audio signals;

[0007] Taking the multiple original audio signals as inputs of a first attention mechanism, the first attention mechanism is used to output multiple target audio signals according to the context information of each frame of original audio signal in each original audio signal in the time dimension, the first correlation between multiple frames of original audio signals with the same time in the multiple original audio signals, and the second correlation between the context information of each of the multiple frames of original audio signals with the same time and each frame of original audio signal in the multiple frames of original audio signals with the same time;

[0008] Fusing the multiple target audio signals to obtain a single-channel fusion result;

[0009] According to the multiple target audio signals, the single-channel fusion result and first text information that has been recognized from the multiple original audio signals, predicting second text information subsequent to the first text information.

[0010] In a second aspect, an embodiment of the present disclosure provides an audio recognition model, the audio recognition model comprising an encoder, a convolution module and a decoder.

[0011] The encoder comprises a first attention mechanism, an input of the first attention mechanism is a plurality of original audio signals, each original audio signal comprises a plurality of frames of original audio signals, the first attention mechanism is configured to output a plurality of target audio signals according to context information of each frame of original audio signal in each original audio signal in a time dimension, first association between a plurality of frames of original audio signals with the same time in the plurality of original audio signals, and second association of respective context information of the plurality of frames of original audio signals with the same time with each frame of original audio signal in the plurality of frames of original audio signals with the same time;

[0012] The convolution module is configured to fuse the plurality of target audio signals to obtain a single-channel fusion result.

[0013] The decoder is configured to predict second text information subsequent to first text information that has been recognized from the plurality of original audio signals according to the plurality of target audio signals, the single-channel fusion result, and the first text information.

[0014] In a third aspect, an audio recognition apparatus is provided, comprising:

[0015] The obtaining module is configured to obtain a plurality of original audio signals, each original audio signal comprising a plurality of frames of original audio signals.

[0016] The input module is configured to input the plurality of original audio signals as an input of a first attention mechanism, the first attention mechanism being configured to output a plurality of target audio signals according to context information of each frame of original audio signal in each original audio signal in a time dimension, first association between a plurality of frames of original audio signals with the same time in the plurality of original audio signals, and second association of respective context information of the plurality of frames of original audio signals with the same time with each frame of original audio signal in the plurality of frames of original audio signals with the same time.

[0017] The fusion module is configured to fuse the plurality of target audio signals to obtain a single-channel fusion result.

[0018] The prediction module is configured to predict second text information subsequent to first text information that has been recognized from the plurality of original audio signals according to the plurality of target audio signals, the single-channel fusion result, and the first text information.

[0019] In a fourth aspect, an electronic device is provided, comprising:

[0020] a memory;

[0021] a processor; and

[0022] a computer program;

[0023] The computer program is stored in the memory and is configured to be executed by the processor to implement the method according to the first aspect.

[0024] In a fifth aspect, the embodiments of the present disclosure provide a computer readable storage medium, having stored thereon a computer program, the computer program being executed by a processor to implement the method according to the first aspect.

[0025] The audio recognition method, device, equipment and storage medium provided by the embodiments of the present disclosure can obtain multiple original audio signals, and process the multiple original audio signals according to the context information of each frame of original audio signal in each original audio signal in the time dimension, the first correlation between multiple frames of original audio signals with the same time in the multiple original audio signals, and the second correlation between the context information of each frame of original audio signal in the multiple frames of original audio signals with the same time and each frame of original audio signal in the multiple frames of original audio signals with the same time, to obtain multiple target audio signals, so as to calculate the correlation between the multiple original audio signals in the time dimension and the channel dimension. Further, the multiple target audio signals are fused to obtain a single-channel fusion result, and the second text information subsequent to the first text information is predicted according to the multiple target audio signals, the single-channel fusion result and the first text information that has been recognized from the multiple original audio signals. Since the correlation between the multiple original audio signals in the time dimension and the channel dimension is fully utilized, the embodiments can not only obtain more fine-grained correlation between channels, but also consider the context information of each frame of original audio signal in the time dimension, so as to improve the recognition accuracy of the multiple original audio signals when multiple speakers speak at the same time. BRIEF DESCRIPTION OF DRAWINGS

[0026] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the description, further serve to explain the principles of the present disclosure.

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained from these drawings without creative labor.

[0028] Figure 1 The principle diagram of the single-channel attention mechanism provided by the embodiments of the present disclosure;

[0029] Figure 2 The principle diagram of the frame-level cross-channel attention mechanism provided by the embodiments of the present disclosure;

[0030] Figure 3A schematic diagram of averaging provided for an embodiment of the present disclosure;

[0031] Figure 4 A schematic diagram of a channel-level cross-channel attention mechanism provided for an embodiment of the present disclosure;

[0032] Figure 5 A flowchart of an audio recognition method provided for an embodiment of the present disclosure;

[0033] Figure 6 A schematic diagram of an application scenario provided for another embodiment of the present disclosure;

[0034] Figure 7 A schematic diagram of context information provided for another embodiment of the present disclosure;

[0035] Figure 8 A schematic diagram of a first attention mechanism provided for another embodiment of the present disclosure;

[0036] Figure 9 A flowchart of an audio recognition method provided for another embodiment of the present disclosure;

[0037] Figure 10 A schematic diagram of an audio recognition model provided for another embodiment of the present disclosure;

[0038] Figure 11 A schematic diagram of a multi-layer convolution module provided for an embodiment of the present disclosure;

[0039] Figure 12 A flowchart of an audio recognition method provided for another embodiment of the present disclosure;

[0040] Figure 13 A flowchart of a training method of an audio recognition model provided for another embodiment of the present disclosure;

[0041] Figure 14 A structural schematic diagram of an audio recognition apparatus provided for an embodiment of the present disclosure;

[0042] Figure 15 A structural schematic diagram of an electronic device embodiment provided for an embodiment of the present disclosure. DETAILED DESCRIPTION

[0043] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the schemes of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0044] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0045] Typically, audio can be converted into text using Automatic Speech Recognition (ASR) technology. For example... Figure 1 The diagram shown illustrates the principle of single-channel attention. Figure 1 A single-channel audio X can be processed to obtain processed audio H. Further, ASR (Automatic Signal Processing) can be performed on the processed audio H to obtain text. The calculation process within the single-channel attention mechanism is shown in formula (1) below:

[0046]

[0047] Where X represents the input of a single-channel audio signal, i.e., a single-channel attention mechanism, which may include a first linear layer, a second linear layer, and a third linear layer; Q represents the calculation result of X after passing through the first linear layer; K represents the calculation result of X after passing through the second linear layer; and V represents the calculation result of X after passing through the third linear layer. W q and b q W represents the parameters of the first linear layer. k and b k W represents the parameters of the second linear layer. v and b v ... express Figure 1 In this context, A and H represent the output of the single-channel attention mechanism. T represents the number of frames of the audio signal included in a single-channel audio X. For example, if one frame of audio signal corresponds to one moment, then T can represent the time length of the single-channel audio X, which includes T moments. D represents the high-dimensional representation of each frame of audio signal. For example, D can be a 256-dimensional representation vector. Figure 1 In this context, A represents the correlation between the audio signals of the T frames. This correlation can be, for example, the contribution factor. Figure 1 In this matrix, A is a T*T matrix. The first row of the matrix represents the contribution of each of the T frames of audio signal to the first frame of audio signal, the second row represents the contribution of each of the T frames of audio signal to the second frame of audio signal, and so on. Each contribution can be a value in the range of 0-1.

[0048] However, in a meeting setting, there are usually multiple people speaking, therefore Figure 1The single-channel attention mechanism shown will not be applicable to the conference scenario.

[0049] However, in the conference scenario, when multiple speakers speak at the same time, the recognition result of the ASR is not accurate enough. For example Figure 2 The principle diagram of the frame-level cross-channel attention mechanism is shown, which is processed by Figure 2 The multi-channel audio can be processed to obtain the processed multi-channel audio H, and further, the processed multi-channel audio H is subjected to ASR to obtain the text.

[0050] The calculation process inside the frame-level cross-channel attention mechanism is shown in the following formula (2):

[0051]

[0052] Wherein, C represents the total number of channels. represents the average of the multi-channel audio . In formula (2), A represents Figure 2 in A. The frame-level cross-channel attention mechanism includes a first linear layer, a second linear layer and a third linear layer, Q represents the calculation result of X after the first linear layer, K represents the calculation result of X after the second linear layer, and V represents the calculation result of X after the third linear layer. W q and b q respectively represent the parameters of the first linear layer, W k and b k respectively represent the parameters of the second linear layer, and W v and b v respectively represent the parameters of the third linear layer. H represents the output of the frame-level cross-channel attention mechanism. Specifically, the multi-channel audio can be represented as X0represents the audio on the first channel, and X C-1 represents the audio on the Cth channel. is represented as the following formula (21), any element in A is represented as the following formula (22), 0≤c≤C-1.

[0053]

[0054]

[0055] wherein, X n represents Any element in the equation. Assume that in this embodiment, the total number of channels C is 4, and each channel has 5 frames of audio signal, i.e., T = 5. For example... Figure 3 As shown, 300 represents the above-mentioned... 301 indicates as described above like Figure 3 As shown, 11-15 represent 5 frames of audio signal on channel 1, and each frame can be represented by a 256-dimensional representation vector. 21-25 represent 5 frames of audio signal on channel 2, 31-35 represent 5 frames of audio signal on channel 3, and 41-45 represent 5 frames of audio signal on channel 4. That is to say, as... Figure 3 The numbers 11-15 shown correspond to X0 as described above, 21-25 correspond to X1, 31-35 correspond to X2, and 41-45 correspond to X. C-1 51-55 correspond to the above as described above. 61-65 correspond to 71-75 correspond to 81-85 correspond to Because of formula (22) express any element in, therefore It can represent Any one of them, assuming yes It includes elements 51, 52, 53, 54, and 55, where element 51 is the average value of audio signals 21, 31, and 41; element 52 is the average value of audio signals 22, 32, and 42; element 53 is the average value of audio signals 23, 33, and 43; element 54 is the average value of audio signals 24, 34, and 44; and element 55 is the average value of audio signals 25, 35, and 45.

[0056] in addition, Figure 2 In this context, 'A' represents the correlation, or contribution, between the multiple frames of audio signal in that channel before averaging and the corresponding multiple frames of audio signal in that channel after averaging. For example, for... Figure 3 Channel 1, as shown, calculates the contributions of audio signals 11-15 to element 51, 52, 53, 54, and 55 respectively, thus obtaining T*T = 25 contributions.

[0057] Because frame-level cross-channel attention mechanisms can perform multi-channel audio processing in the time dimension. Directly averaging obtains information across channels, but direct averaging destroys some unique information on each channel. When the number of channels is small, there is a large loss in performance. This leads to inaccurate recognition results of ASR when multiple speakers speak at the same time.

[0058] In addition, Figure 4 The principle diagram of the channel-level cross-channel attention mechanism is shown, and the principle diagram of the channel-level cross-channel attention mechanism is shown. Figure 4 The multi-channel audio can also be processed to obtain the processed multi-channel audio H, and further, the processed multi-channel audio H is subjected to ASR to obtain text.

[0059] The calculation process inside the channel-level cross-channel attention mechanism is shown in the following formula (3):

[0060]

[0061] Among them, Q, K, V, W k and b k , W v and b v , W q and b q have similar meanings as described above. In formula (3), A in Figure 3 , A in Figure 3 indicates the contribution degree between channels for each of the T time instants. The contribution degree between channels includes the contribution degree of each channel to the first channel, the contribution degree of each channel to the second channel, and so on.

[0062] Since the channel-level cross-channel attention mechanism does not consider the temporal correlation of the channels, but only considers the correlation between a few channels at the current time instant, the recognition result of ASR when multiple speakers speak at the same time is not accurate enough.

[0063] To solve the above problems, the present embodiment provides an audio recognition method, which will be introduced below in combination with specific embodiments.

[0064] Figure 5 The audio recognition method flowchart provided by the present embodiment. The method can be executed by an audio recognition device, which can be realized by software and / or hardware. The device can be configured in an electronic device, such as a server or a terminal, wherein the terminal specifically includes a mobile phone, a computer, or a tablet computer, etc. In addition, the audio recognition method described in the present embodiment can be applied to the application scenario as shown in Figure 6 The application scenario as shown in​Figure 6 As shown, this application scenario includes an audio acquisition array 61 and a computing device 62. The audio acquisition array 61 can include multiple audio acquisition devices, each acquiring one audio signal and corresponding to one channel. Specifically, the audio acquisition array 61 is a microphone array, which includes multiple microphones, each acquiring one audio signal, with one microphone corresponding to one channel. It is understood that the multiple audio acquisition devices in the audio acquisition array 61 can be integrated into one device, or the multiple audio acquisition devices can be independent devices distributed in different locations in the conference room. Since there are usually multiple speakers in a conference scenario, each audio signal as described above includes the voices of multiple speakers. The computing device 62 can be a terminal or server with computing and processing functions. The computing device 62 can use the audio recognition method provided in this embodiment to recognize the multiple audio signals acquired by the audio acquisition array 61. It is understood that if the audio acquisition array 61 has computing and processing functions, then the audio acquisition array 61 can also recognize multiple audio signals. The following is in conjunction with... Figure 6 This method will be described in detail, such as Figure 5 As shown, the specific steps of this method are as follows:

[0065] S501. Acquire multiple raw audio signals, each raw audio signal including multiple frames of raw audio signal.

[0066] like Figure 6 As shown, assuming that each audio acquisition device in the audio acquisition array 61 acquires one audio signal, which is denoted as one raw audio signal, and assuming that the audio acquisition array 61 includes four audio acquisition devices, then the audio acquisition array 61 can acquire four raw audio signals. Furthermore, the audio acquisition array 61 can send these four raw audio signals to the computing device 62, thereby enabling the computing device 62 to acquire these four raw audio signals. Specifically, each raw audio signal includes multiple frames of raw audio signal.

[0067] S502, the multiple original audio signals are used as input to a first attention mechanism. The first attention mechanism is used to output multiple target audio signals based on the context information of each frame of the original audio signal in the time dimension, the first correlation between the multiple frames of the original audio signals with the same time in the multiple original audio signals, and the second correlation between the context information of each frame of the original audio signals with the same time and each frame of the original audio signals.

[0068] Specifically, the computing device 62 can use the four original audio signals as input to a first attention mechanism, which can simultaneously calculate the correlation between the four original audio signals in both the time and channel dimensions. Specifically, the computing device 62 can determine the context information of each frame of the original audio signal in the time dimension within each original audio signal, and further calculate the first correlation between multiple frames of original audio signals with the same time frame among the four original audio signals, and the second correlation between the context information of each of the multiple frames of original audio signals with the context information of each frame of original audio signals with the same time frame.

[0069] like Figure 7 As shown, 11-15 represent 5 frames of original audio signal in channel 1 (the first channel of original audio signal), 21-25 represent 5 frames of original audio signal in channel 2 (the second channel of original audio signal), 31-35 represent 5 frames of original audio signal in channel 3 (the third channel of original audio signal), and 41-45 represent 5 frames of original audio signal in channel 4 (the fourth channel of original audio signal). Taking the first frame of original audio signal 11 in the first channel of original audio signal as an example, the 0th frame of original audio signal 10 on channel 1 and the 2nd frame of original audio signal 12 on channel 1 are the context information of original audio signal 11 in the time dimension. The context information of other original audio signals in the time dimension is similar, and will not be described in detail here. For example, the first frame of the original audio signal 11 in the first channel, the first frame of the original audio signal 21 in the second channel, the first frame of the original audio signal 31 in the third channel, and the first frame of the original audio signal 41 in the fourth channel are multiple frames of original audio signals with the same time. The first correlation between these multiple frames of original audio signals with the same time may include the contribution of the original audio signals 11-41 to the original audio signal 11, the contribution of the original audio signals 11-41 to the original audio signal 21, the contribution of the original audio signals 11-41 to the original audio signal 31, and the contribution of the original audio signals 11-41 to the original audio signal 41. The second correlation, as described above, includes the contribution of the context information of each of the original audio signals 11-41 to the original audio signal 11, the contribution of the context information of each of the original audio signals 11-41 to the original audio signal 21, the contribution of the context information of each of the original audio signals 11-41 to the original audio signal 31, and the contribution of the context information of each of the original audio signals 11-41 to the original audio signal 41.

[0070] Optionally, the temporal context information of each frame of the original audio signal includes at least one frame of historical audio signal and at least one frame of future audio signal. Optionally, the number of frames in the at least one frame of historical audio signal and the number of frames in the at least one frame of future audio signal are the same. The number of frames is determined by the difference between the time when the first audio acquisition device in the audio acquisition array receives the audio signal and the time when the second audio acquisition device receives the audio signal. The first audio acquisition device is the earliest audio acquisition device in the audio acquisition array to receive the audio signal, and the second audio acquisition device is the latest audio acquisition device in the audio acquisition array to receive the audio signal.

[0071] like Figure 7 As shown, original audio signal 10 is a historical audio signal of original audio signal 11, and original audio signal 12 is a future audio signal of original audio signal 11. It is understandable that... Figure 7 The context information of each frame of the original audio signal shown includes only one frame of historical audio signal and one frame of future audio signal in the time dimension. In other embodiments, the context information of each frame of the original audio signal may include multiple frames of historical audio signal and multiple frames of future audio signal. For example, for a certain frame of original audio signal, the context information of that frame of original audio signal in the time dimension includes F frames of historical audio signal and F frames of future audio signal, that is, the number of frames of historical audio signal and the number of frames of future audio signal are the same. In other words, F is the number of frames that look back at the past or look forward to the future at each moment. The length of F is a trade-off between performance and efficiency.

[0072] It can be understood that, since the audio acquisition array includes a plurality of audio acquisition devices, and the plurality of audio acquisition devices are different in position and / or angle relative to the speaker, when a certain speaker is speaking, some of the plurality of audio acquisition devices can receive the audio signal first, and some of the plurality of audio acquisition devices can receive the audio signal last. Assuming that the audio acquisition device that receives the audio signal first is recorded as a first audio acquisition device, the audio acquisition device that receives the audio signal last is recorded as a second audio acquisition device, and the difference between the time at which the first audio acquisition device receives the audio signal and the time at which the second audio acquisition device receives the audio signal is recorded as a delay, for example, when the audio acquisition device is a microphone, the delay can be recorded as a microphone delay. The size of F can be determined by the delay. It can be understood that, in addition to calculating the delay between the first audio acquisition device and the second audio acquisition device as described above, the embodiment can also calculate the delay between any two microphones in the microphone array as described above. In this case, using a direction of arrival (DOA) estimation method, the position and direction of the speaker can be determined according to the delay between any two microphones in the microphone array.

[0073] Further, the computing device 62 can process the plurality of original audio signals to obtain a plurality of target audio signals according to the context information of each frame of the original audio signal in the time dimension in each original audio signal and the first correlation and the second correlation as described above, for example, one original audio signal corresponds to one target audio signal.

[0074] S503, fuse the plurality of target audio signals to obtain a single-channel fusion result.

[0075] For example, the computing device 62 can fuse the plurality of target audio signals to obtain a single-channel fusion result.

[0076] S504, according to the plurality of target audio signals, the single-channel fusion result, and the first text information that has been identified from the plurality of original audio signals, predict second text information subsequent to the first text information.

[0077] Specifically, the computing device 62 can recognize the text information corresponding to the multi-channel original audio signal in a word-by-word manner. Assuming that the text information currently recognized by the computing device 62 is denoted as first text information, the computing device 62 can predict second text information subsequent to the first text information according to the multi-channel target audio signal, the single-channel fusion result, and the first text information. For example, in the first prediction, the computing device 62 can predict the first word of the content spoken by the speaker in the multi-channel original audio signal according to the multi-channel target audio signal, the single-channel fusion result, and a preset start symbol. In the second prediction, the second word is predicted according to the multi-channel target audio signal, the single-channel fusion result, the preset start symbol, and the first word. In the third prediction, the third word is predicted according to the multi-channel target audio signal, the single-channel fusion result, the preset start symbol, the first word, and the second word. In this way, the prediction is continued until a preset end symbol is predicted. For example, when the first text information is the first word, the second text information is the second word. When the first text information is the first word and the second word, the second text information is the third word.

[0078] The embodiment of the present disclosure can obtain the multi-channel original audio signal, and process the multi-channel original audio signal according to the context information of each frame of original audio signal in each original audio signal in the time dimension, the first correlation between the multiple frames of original audio signal with the same time in the multi-channel original audio signal, and the second correlation between the context information of each frame of original audio signal and the multiple frames of original audio signal with the same time in the multi-channel original audio signal, to obtain the multi-channel target audio signal. Thus, the correlation between the multi-channel original audio signal can be calculated in the time dimension and the channel dimension. Further, the single-channel fusion result is obtained by fusing the multi-channel target audio signal, and the second text information subsequent to the first text information is predicted according to the multi-channel target audio signal, the single-channel fusion result, and the first text information. Since the correlation between the multi-channel original audio signal in the time dimension and the channel dimension is fully utilized, the embodiment can obtain the correlation between the channels with finer granularity, and the context information of each frame of original audio signal in the time dimension is also considered. Thus, the recognition accuracy of the multi-channel original audio signal is improved when multiple speakers speak at the same time.

[0079] Figure 8 The principle diagram of the first attention mechanism is shown as above, and the multi-channel target audio signal is obtained by Figure 8 The multi-channel original audio signal can be processed The processed multi-channel target audio signal H is obtained. Further, ASR is performed on the multi-channel target audio signal H to obtain the text. In this embodiment, the first attention mechanism can also be called a multi-frame cross-channel attention mechanism. The calculation process inside the multi-frame cross-channel attention mechanism is shown in the following formula (4):

[0080]

[0081] Among them, Q, K, V, W k and b k W v and b v W q and b q The meaning is similar to that described above. In formula (4) express Figure 8 A in the middle. This indicates that multiple raw audio signals will be... The temporal context information of each frame of the original audio signal in each of the multiple original audio signals. The result after splicing. F represents the number of historical audio signal frames included in the temporal context information of each original audio signal, or F represents the number of future audio signal frames included in the temporal context information of each original audio signal. When F=1, Figure 8 The historical audio signal shown corresponds to Figure 7 The historical audio signal shown is Figure 8 The future audio signal shown corresponds to Figure 7 The audio signal shown is for the future. Figure 8 shown Corresponding to Figure 7 The multiple original audio signals shown. Assuming that in formula (4)... This can be expressed as the following formula (41):

[0082]

[0083] The meaning of T is the same as that described above, such as... Figure 7 As shown, T = 5. Corresponding to Figure 7 The 12 elements in the first row shown Corresponding to Figure 7 The 12 elements in row t shown are... Corresponding to Figure 7 The 12 elements in row 5 shown. This can be expressed as the following formula (42):

[0084]

[0085] Since Figure 7 The 12 elements in the t-th row shown by the formula (1) are composed of 4 frames of historical audio signals, 4 frames of audio signals corresponding to the t-th moment, and 4 frames of future audio signals, therefore, Corresponding to Figure 7 The 4 frames of historical audio signals in the t-th row shown by the formula (1), Corresponding to Figure 7 The 4 frames of audio signals corresponding to the t-th moment shown by the formula (1), Corresponding to Figure 7 The 4 frames of future audio signals in the t-th row shown by the formula (1).

[0086] In addition, as shown by the formula (2), since the 4 frames of historical audio signals corresponding to the 0-th frame and the 4 frames of future audio signals corresponding to the 6-th frame are not included in the multi-channel original audio signals, in some embodiments, the 256-dimensional representation vector of each frame of historical audio signal corresponding to the 0-th frame can be recorded as all 0, and the 256-dimensional representation vector of each frame of future audio signal corresponding to the 6-th frame can be recorded as all 0. It can be understood that, Figure 8 The number of grids corresponding to each character shown by the formula (2) is used to illustrate the relationship between the characters, since the embodiment adopts the formula (1) to explain the formula (2), therefore, Figure 7 The principles shown by the formula (1) and the principles shown by the formula (2) are consistent, only Figure 8 The number of grids corresponding to a certain character shown by the formula (2) and the specific value of the character shown by the formula (2) may be slightly different, for example, Figure 7 T=5 shown by the formula (2), while Figure 8 T in the formula (1) corresponds to 4 grids, but Figure 8 The principles shown by the formula (1) and the principles shown by the formula (2) are consistent, only Figure 7 The number of grids corresponding to a certain character shown by the formula (2) and the specific value of the character shown by the formula (2) may be slightly different, for example, Figure 7 T=5 shown by the formula (2), while Figure 8 T in the formula (1) corresponds to 4 grids, but Figure 7 The principles shown by the formula (1) and the principles shown by the formula (2) are consistent, only Figure 8 The number of grids corresponding to a certain character shown by the formula (2) and the specific value of the character shown by the formula (2) may be slightly different, for example, Figure 8 T=5 shown by the formula (2), while Figure 7 T in the formula (1) corresponds to 4 grids, but Figure 8 The principles shown by the formula (1) and the principles shown by the formula (2) are consistent, only Figure 8 The number of grids corresponding to a certain character shown by the formula (2) and the specific value of the character shown by the formula (2) may be slightly different, for example, Figure 9 T=5 shown by the formula (2), while

[0087] Figure 8An audio recognition method flowchart is provided for another embodiment of the present disclosure. In this embodiment, the first attention mechanism includes a first linear layer, a second linear layer, and a third linear layer, and outputs a plurality of target audio signals according to the context information of each frame of the original audio signal in the time dimension of each original audio signal, the first correlation between a plurality of frames of original audio signals with the same time in the plurality of original audio signals, and the second correlation of each frame of original audio signal in the plurality of frames of original audio signals with the same time with the respective context information of the plurality of frames of original audio signals with the same time, including the following steps:

[0088] S901, processing the plurality of original audio signals through the first linear layer to obtain a first intermediate result.

[0089] As shown in Figure 8 , it is assumed that the first attention mechanism includes a first linear layer, a second linear layer, and a third linear layer. The plurality of original audio signals are processed by the first linear layer to obtain a first intermediate result Q.

[0090] S902, processing the context information of each frame of the original audio signal in the time dimension of each original audio signal and the plurality of original audio signals through the second linear layer to obtain a second intermediate result.

[0091] As shown in Figure 8 , the historical audio signal of each frame of the original audio signal in the time dimension of each original audio signal in the plurality of original audio signals , the plurality of original audio signals , and the future audio signal of each frame of the original audio signal in the time dimension of each original audio signal in the plurality of original audio signals constitute are processed by the second linear layer to obtain a second intermediate result K.

[0092] S903, processing the context information of each frame of the original audio signal in the time dimension of each original audio signal and the plurality of original audio signals through the third linear layer to obtain a third intermediate result.

[0093] As shown in Figure 8 , the historical audio signal of each frame of the original audio signal in the time dimension of each original audio signal in the plurality of original audio signals , the plurality of original audio signals , and the future audio signal of each frame of the original audio signal in the time dimension of each original audio signal in the plurality of original audio signals

[0094] are processed by the third linear layer to obtain a third intermediate result V.S904. Based on the first intermediate result and the second intermediate result, determine the first correlation between multiple frames of original audio signals with the same time in the multiple original audio signals, and the second correlation between the context information of each of the multiple frames of original audio signals with the same time and each frame of original audio signal in the multiple frames of original audio signals with the same time.

[0095] like Figure 7 As shown, multiplying the first intermediate result Q and the second intermediate result K by a dot product yields A. Figure 7 The A shown represents, for each time moment, the first correlation between the multiple frames of original audio signals corresponding to that time moment, and the second correlation between the context information of each of the multiple frames of original audio signals corresponding to that time moment and each frame of original audio signals within the multiple frames of original audio signals corresponding to that time moment. For example, with Figure 7 Taking 701 as an example, elements 11 to 41 are multiple frames of original audio signals corresponding to the same moment, that is, multiple frames of original audio signals with the same time. The first correlation between these multiple frames of original audio signals with the same time includes the contribution of elements 11 to 41 to element 11, the contribution of elements 11 to 41 to element 21, the contribution of elements 11 to 41 to element 31, and the contribution of elements 11 to 41 to element 41. Because Figure 8 In the diagram 702, elements 10 to 40 are the historical audio signals of elements 11 to 41, respectively (i.e., element 10 is the historical audio signal of element 11, element 20 is the historical audio signal of element 21, and so on). Figure 7 In the diagram 703, elements 12 to 42 are, in turn, the future audio signals of elements 11 to 41 (i.e., element 12 is the future audio signal of element 11, element 22 is the future audio signal of element 21, and so on). Therefore, elements 10 to 40 in 702 and elements 12 to 42 in 703 are the context information of elements 11 to 41. The second correlation described above includes the contributions of elements 10 to 40 and elements 12 to 42 to element 11, the contributions of elements 10 to 40 and elements 12 to 42 to element 21, the contributions of elements 10 to 40 and elements 12 to 42 to element 31, and the contributions of elements 10 to 40 and elements 12 to 42 to element 41. Therefore, Figure 7 The diagram shows A, which includes T sets of data, each set being an array of type C*[(2F+1)C]. When F=1 and C=4, Figure 7 The data sets 701, 702, and 703 shown contain a total of (2F+1)C = 12 elements. Taking the first data set in group T as an example, this first data set can include C rows and (2F+1)C columns. The first row of these C rows is...Figure 8 The 12 elements included in 701, 702, and 703 shown respectively contribute to the degree of contribution of element 11 in 701. The second row in the C row is Figure 8 The 12 elements included in 701, 702, and 703 shown respectively contribute to the degree of contribution of element 21 in 701. Similarly, the calculation method of other groups of data in the T group of data is similar and will not be described here. That is, Figure 8 A shown includes the first correlation and the second correlation as described above.

[0096] S905, according to the third intermediate result, the first correlation, and the second correlation, output the multi-path target audio signal.

[0097] As shown in Figure 10 The third intermediate result V and Figure 10 A shown are multiplied to obtain the multi-path target audio signal H.

[0098] The first attention mechanism of the embodiment calculates the correlation between the multi-path original audio signals in the time dimension and the channel dimension at the same time, so that the first attention mechanism can obtain more fine-grained correlation between channels, and also fully utilizes the context information of each frame of original audio signal in each original audio signal in the time dimension. Therefore, not only the problem that the frame-level cross-channel attention mechanism has poor ability to extract fine-grained channel information is solved, but also the problem that the channel-level cross-channel attention mechanism lacks context information is solved. In addition, the embodiment uses the time delay of the audio acquisition array receiving the audio signal or the voice signal to determine the frame number of the historical audio signal and the frame number of the future audio signal in the context information, so that the length of the context information can be balanced between performance and efficiency.

[0099] Figure 10 The schematic diagram of the audio recognition model provided by another embodiment of the present disclosure is shown in Figure 8As shown, the audio recognition model includes an encoder, a convolutional module, and a decoder. The encoder includes a first attention mechanism, the input of which is multiple original audio signals, each of which includes multiple frames of original audio signals. The first attention mechanism is used to output multiple target audio signals based on the temporal context information of each frame of original audio signal in each original audio signal, the first correlation between the multiple frames of original audio signals with the same time in the multiple original audio signals, and the second correlation between the context information of each frame of original audio signals with the same time and each frame of original audio signals in the multiple frames of original audio signals with the same time. The convolutional module is used to fuse the multiple target audio signals to obtain a single-channel fusion result. The decoder is used to predict the second text information following the first text information based on the multiple target audio signals, the single-channel fusion result, and the first text information currently identified from the multiple original audio signals.

[0100] like Figure 10 The overall structure shown is an audio recognition model, which can also be referred to as a Transformer network, a multi-channel ASR model, or a speech recognition model. This audio recognition model includes an encoder, a convolutional module, and a decoder. The encoder can be repeated N times, meaning the computation process within the encoder can be repeated N times. Similarly, the decoder can be repeated M times, meaning the computation process within the decoder can be repeated M times.

[0101] Specifically, the encoder includes the first attention mechanism as described above, the input of which is multiple raw audio signals, and the first attention mechanism can employ methods such as... Figure 11 The multi-channel target audio signal H is calculated in the manner shown; the specific process will not be elaborated here. This embodiment and subsequent embodiments will focus on the calculation process after the first attention mechanism outputs the multi-channel target audio signal H. For example, in this embodiment, the encoder also includes a Conformer structure (Conformer block), which is a parallel two-body network structure. Processing the multi-channel target audio signal H output by the first attention mechanism through the Conformer structure helps to enhance speech representation learning and has better local and global modeling capabilities. Assuming that the Conformer structure can enhance the multi-channel target audio signal H to obtain an enhanced multi-channel target audio signal, further, the enhanced multi-channel target audio signal can be fused through a convolution module to obtain a single-channel fusion result. Alternatively, the multi-channel target audio signal H output by the first attention mechanism can be directly fused through a convolution module to obtain a single-channel fusion result.

[0102] Optionally, the multi-channel target audio signals are fused to obtain a single-channel fusion result, including: sequentially passing the multi-channel target audio signals through a plurality of convolution layers to obtain the single-channel fusion result. Assuming Figure 11 The input of the first attention mechanism shown in FIG. 8 is an 8-channel original audio signal, the first attention mechanism outputs an 8-channel target audio signal, and the Conformer structure outputs an enhanced 8-channel target audio signal. The convolution module can be a multi-layer convolution module, for example, the convolution module includes five two-dimensional convolution layers as shown in FIG. 8. Sequential processing of the enhanced 8-channel target audio signal through the five two-dimensional convolution layers can gradually reduce the channel dimension, thereby fusing the enhanced 8-channel target audio signal into a single-channel fusion result. Figure 10 Figure 10 The kernal shown in FIG. 8 represents a convolution kernel, and the channel represents a channel. In addition, in some embodiments, the number of input channels of the multi-layer convolution module is fixed. If the actual number of input channels of the multi-layer convolution module is less than the pre-configured maximum channel number, part of the actual input channels needs to be repeated to expand the actual number of input channels of the multi-layer convolution module.

[0103] Further, the enhanced 8-channel target audio signal output by the Conformer structure and the single-channel fusion result output by the convolution module are part of the input of the ordinary attention mechanism (src-attention) in the decoder. For example, the enhanced 8-channel target audio signal output by the Conformer structure and the single-channel fusion result output by the convolution module are spliced in the channel dimension and then transmitted to the ordinary attention mechanism in the decoder.

[0104] ​Assuming that the first text information has been recognized from the 8 original audio signals, the first text information including one or more words, the self-attention mechanism in the decoder can generate a representation vector of the first text information according to the representation vector of each word in the first text information. So that the decoder as a whole can predict the second text information after the first text information according to the enhanced 8 target audio signals, the single-channel fusion result and the representation vector of the first text information. Wherein each word corresponds to a representation vector, and the representation vector of each word can be a 256-dimensional vector. The representation vector of each word can be referred to as Token Embedding, and in addition, the representation vector of each word can also be referred to as a high-dimensional representation at the word level. It can be understood that the present embodiment takes the Chinese conference scenario as an example, and the decoder can recognize the 8 original audio signals word by word, or word by word, or sentence by sentence, which is not specifically limited here, but only illustrative description is taken as an example. In addition, it can be understood that if it is a foreign language conference scenario, for example, an English conference scenario, the decoder can also recognize word by word, character by character, short sentence by short sentence or phrase by phrase.

[0105] Optionally, the decoder comprises a second attention mechanism, a third attention mechanism, a fourth attention mechanism, a fusion layer, and a neural network, the structure of the second attention mechanism is the same as that of the first attention mechanism; the third attention mechanism is used to generate a representation vector of the first text information that has been recognized from the multi-channel original audio signals; the fourth attention mechanism is used to fuse the representation vector of any one of the multi-channel target audio signals and the representation vector of the first text information to obtain a representation vector of a first text sequence corresponding to the any one of the multi-channel target audio signals; fuse the representation vector of the single-channel fusion result and the representation vector of the first text information to obtain a representation vector of a second text sequence corresponding to the single-channel fusion result; the second attention mechanism is used to process the representation vectors of the first text sequences respectively corresponding to the multi-channel target audio signals and the representation vector of the second text sequence, and output representation vectors respectively corresponding to a plurality of third text sequences; the fusion layer is used to fuse the representation vectors respectively corresponding to the plurality of third text sequences to obtain a representation vector of a target text sequence; and the neural network is used to predict the second text information subsequent to the first text information according to the representation vector of the target text sequence.

[0106] Specifically, as Figure 12As shown, the decoder includes a self-attention mechanism, a common attention mechanism, a second attention mechanism, a fusion layer, and a neural network, wherein the structure of the second attention mechanism is the same as that of the first attention mechanism as described above. Here, the self-attention mechanism can be referred to as a third attention mechanism, and the common attention mechanism can be referred to as a fourth attention mechanism. For example, the self-attention mechanism can generate a representation vector of the first text information according to the representation vector of each word that has been recognized at present. As described above, the 8-way target audio signal, the single-channel fusion result, and the representation vector of the first text information are input to the common attention mechanism, so that the common attention mechanism can fuse the representation vector of each of the 8-way target audio signals and the representation vector of the first text information to obtain a representation vector of a first text sequence, thereby obtaining a total of 8 representation vectors of the first text sequences respectively. It can be understood that the representation vector of each target audio signal can be referred to as a high-dimensional representation of an audio sequence of one channel. The representation vector of the first text information can be referred to as a high-dimensional representation of a text sequence that has been recognized at present. The text sequence that has been recognized at present is a sequence of each word that has been recognized at present. In addition, since the main work of the decoder is to convert audio into text, the specific process of fusing the representation vector of the target audio signal and the representation vector of the first text information is to fuse the representation vector of the target audio signal into the representation vector of the first text information, so that the fused result is still a representation vector of a text sequence. In addition, the common attention mechanism can also fuse the representation vector of the single-channel fusion result and the representation vector of the first text information to obtain a representation vector of a second text sequence. Since the single-channel fusion result is the fusion result of the 8-way target audio signal as described above, the representation vector of the single-channel fusion result can be referred to as a high-dimensional representation of an audio sequence of one channel. When fusing the representation vector of the single-channel fusion result and the representation vector of the first text information, the representation vector of the single-channel fusion result can be fused into the representation vector of the first text information, so that the fused result is still a representation vector of a text sequence. That is, the common attention mechanism can output 9 representation vectors of text sequences respectively. Of the 9 text sequences, 8 are first text sequences, and 1 is a second text sequence.Further, the 9 text sequence respectively representation vectors output by the common attention mechanism are taken as the input of the second attention mechanism, so that the second attention mechanism can process the 9 text sequence respectively representation vectors, and since the structure of the second attention mechanism is the same as that of the first attention mechanism, the number of input information and the number of output information of the second attention mechanism are the same, that is, the second attention mechanism can output 9 text sequence respectively representation vectors, but the 9 text sequence output by the second attention mechanism is different from the 9 text sequence input, so the 9 text sequence output by the second attention mechanism is respectively denoted as the third text sequence. Further, Figure 13 The fusion layer in the decoder shown can fuse the 9 third text sequence respectively representation vectors output by the second attention mechanism to obtain the representation vector of the target text sequence. So that the neural network can predict the second text information subsequent to the first text information according to the representation vector of the target text sequence. Wherein, the neural network can be a feed forward neural network (FFN).

[0107] In this embodiment, the first attention mechanism is deployed in the encoder of the audio recognition model, and the second attention mechanism with the same structure as the first attention mechanism is deployed in the decoder, so that the encoder can apply the first attention mechanism to calculate the correlation between the audio signals of multiple channels in the time dimension and the channel dimension. In addition, the decoder applies the second attention mechanism to calculate the correlation between the high-dimensional representations of the text level of multiple channels in the time dimension and the channel dimension. So that the audio recognition model can not only process multi-channel audio signals, but also process multi-channel text sequences, so the accuracy of audio recognition, i.e. converting audio to text, can be significantly improved. In addition, the prior art usually directly reduces the channel dimension by averaging or splicing the multi-channel target audio signals along the time axis to obtain a single-channel fusion result. This direct reduction of the channel dimension will damage the channel-specific information. In this embodiment, the multi-channel target audio signals are sequentially fused through multiple convolution layers to obtain a single-channel fusion result. The channel dimension is gradually reduced in the process of sequentially fusing through multiple convolution layers, rather than being directly reduced. Therefore, this embodiment can effectively avoid damage to the channel-specific information, thereby more effectively fusing the specific information of each channel.

[0108] Figure 10 The audio recognition method flowchart provided for another embodiment of the present disclosure. In this embodiment, according to the multi-channel target audio signal, the single-channel fusion result and the first text information that has been identified from the multi-channel original audio signal, the second text information subsequent to the first text information is predicted, including the following steps:

[0109] S1201, for any one of the multi-path target audio signals, fusing a representation vector of the any one of the multi-path target audio signals and a representation vector of the first text information to obtain a representation vector of a first text sequence corresponding to the any one of the multi-path target audio signals.

[0110] S1202, fusing a representation vector of the single-channel fusion result and a representation vector of the first text information to obtain a representation vector of a second text sequence corresponding to the single-channel fusion result.

[0111] S1203, taking the representation vectors of the first text sequences corresponding to the multi-path target audio signals respectively and the representation vector of the second text sequence as inputs of a second attention mechanism, the structure of the second attention mechanism being the same as that of the first attention mechanism.

[0112] S1204, fusing representation vectors of a plurality of third text sequences output by the second attention mechanism to obtain a representation vector of a target text sequence.

[0113] S1205, predicting second text information subsequent to the first text information according to the representation vector of the target text sequence.

[0114] Specifically, the implementation manners and specific principles of S1201-S1205 can refer to the implementation manners and specific principles of the modules in the decoder as described above.

[0115] It can be understood that the training process and the use process of the audio recognition model as described above can be executed on different devices respectively, or can be executed on the same device. The use process can also be recorded as an inference stage. Since the training process and the use process are two different stages, the audio acquisition array corresponding to the audio recognition model is different in different stages, so that the number of channels of the audio acquisition array corresponding to the audio recognition model can be different in different stages. In order to avoid the dependence of the audio recognition model on the number of channels in the training process, the audio recognition model is obtained by training the following method as shown in the following method: Figure 14

[0116] S1301, determining the number of mask channels according to the total number of channels in the audio acquisition array.

[0117] ​For example, the total number of channels of the audio acquisition array is 8 during the training process. When the audio recognition model is trained, the number of channels of the audio acquisition array corresponding to the audio recognition model may be 4, 5 or 6, etc. during the use process. Therefore, in order to avoid the dependence of the audio recognition model on the number of channels during the training process, the number of mask channels can be randomly determined according to the total number of channels in the audio acquisition array during the training process. For example, a number such as 3 is randomly selected from 1-8 as the number of mask channels.

[0118] S1302, according to the number of mask channels, the number of channels in the audio acquisition array is masked.

[0119] For example, according to the number of mask channels 3, 3 channels are randomly selected from channel 1-channel 8 as mask channels, and the 3 channels are masked, that is, the audio signals on the 3 channels will no longer be input to the first attention mechanism, or the audio signals on the 3 channels will be discarded or shielded. Wherein, each of the 8 channels can be determined as a mask channel with equal probability.

[0120] S1303, the multi-channel original audio signal is obtained from the plurality of channels that are not masked.

[0121] For example, after 3 channels are randomly masked in 8 channels, the remaining 5 channels are not masked, and further, 5 original audio signals are obtained from the remaining 5 channels.

[0122] S1304, the multi-channel original audio signal is input to the audio recognition model as the input of the audio recognition model, so that the audio recognition model outputs the predicted text information.

[0123] For example, the 5 original audio signals are input to the audio recognition model as shown in Figure 14 , for processing, so as to obtain the text information predicted by the audio recognition model.

[0124] S1305, according to the predicted text information and the standard sample text information corresponding to the multi-channel original audio signal, the audio recognition model is trained.

[0125] Because in the training process, the 5 original audio signals can correspond to the standard sample text information, that is, the accurate text comparison, therefore, according to the text information predicted by the audio recognition model and the standard sample text information corresponding to the 5 original audio signals, the audio recognition model can be trained. Specifically, whenever the text information predicted by the audio recognition model constitutes a sentence, according to the sentence recognized by the audio recognition model and the corresponding sentence in the accurate text comparison, the audio recognition model is iteratively trained once.

[0126] It can be understood that, with the difference in the number of mask channels, and / or the difference in the channels to be masked, the geometry of the remaining channels, i.e., the remaining microphones, will be different. For example, the geometry of 8 microphones is a circle, if the number of mask channels is 4, and every other microphone of the 8 microphones is masked, then the geometry of the remaining microphones is a square. If the number of mask channels is 5, then the geometry of the remaining microphones can be a triangle.

[0127] The embodiments of the present disclosure can mask the number of channels in the audio acquisition array according to the number of mask channels in the training stage, and mask the number of channels in the audio acquisition array according to the number of mask channels, so that the corresponding number of channels of the audio recognition model is flexible and variable in the training process, so that the audio recognition model can learn the microphone array data of different channel numbers and geometries in the training process, so that the performance of the audio recognition model after training is not easily affected by the number of channels and / or the geometry of the microphone, and the robustness of the audio recognition model after training to different microphone arrays is improved. Thus, the audio recognition model after training can process microphone arrays of any structure and any number in the use stage.

[0128] In addition, the prior art uses beamforming to model single-channel ASR using spatial information for multi-channel audio signals recorded by a microphone array, but beamforming has various limitations, such as processing multi-channel audio signals in two modules, for example, first using a front-end model to perform beamforming using the spatial information contained in the multi-channel to generate an enhanced single-channel audio, and then identifying the ASR model for us in the back-end. The embodiments can integrate the front-end beamforming module and the back-end ASR module, and use an end-to-end model, i.e., an audio recognition model, to realize the functions of the two modules, and simply and efficiently use the multi-channel audio of the microphone array to improve the speech recognition performance of the model.

[0129] Figure 14 The structure diagram of the audio recognition device provided by the embodiments of the present disclosure is provided. The audio recognition device provided by the embodiments of the present disclosure can execute the processing flow provided by the audio recognition method embodiments, as shown in Figure 15 The audio recognition device 140 includes:

[0130] The acquisition module 141 is configured to acquire a plurality of original audio signals, each original audio signal including a plurality of frames of original audio signals.

[0131] The input module 142 is configured to input the multiple original audio signals as inputs of a first attention mechanism, and the first attention mechanism is configured to output multiple target audio signals according to context information of each frame of original audio signal in each original audio signal in a time dimension, first correlations between multiple frames of original audio signals that are the same in time in the multiple original audio signals, and second correlations of each of the multiple frames of original audio signals that are the same in time with each frame of original audio signal in the multiple frames of original audio signals that are the same in time.

[0132] The fusion module 143 is configured to fuse the multiple target audio signals to obtain a single-channel fusion result.

[0133] The prediction module 144 is configured to predict second text information that is subsequent to first text information according to the multiple target audio signals, the single-channel fusion result, and the first text information that has been identified from the multiple original audio signals.

[0134] Optionally, the context information of each frame of original audio signal in the time dimension includes at least one frame of historical audio signal and at least one frame of future audio signal of the each frame of original audio signal.

[0135] Optionally, the number of frames of the at least one frame of historical audio signal and the number of frames of the at least one frame of future audio signal are the same, and the number of frames is determined by a difference between a time at which a first audio acquisition device in an audio acquisition array receives an audio signal and a time at which a second audio acquisition device receives the audio signal, the first audio acquisition device being an audio acquisition device that receives the audio signal earliest in the audio acquisition array, and the second audio acquisition device being an audio acquisition device that receives the audio signal latest in the audio acquisition array.

[0136] Optionally, the first attention mechanism includes a first linear layer, a second linear layer, and a third linear layer, and when the first attention mechanism outputs the multiple target audio signals according to the context information of each frame of original audio signal in each original audio signal in the time dimension, the first correlations between the multiple frames of original audio signals that are the same in time in the multiple original audio signals, and the second correlations of each of the multiple frames of original audio signals that are the same in time with each frame of original audio signal in the multiple frames of original audio signals that are the same in time, the first attention mechanism is specifically configured to:

[0137] process the multiple original audio signals through the first linear layer to obtain a first intermediate result;

[0138] process the context information of each frame of original audio signal in each original audio signal in the time dimension and the multiple original audio signals through the second linear layer to obtain a second intermediate result;

[0139] context information of each frame of the original audio signals in the time dimension and the multi-channel original audio signals after the third linear layer processing to obtain a third intermediate result;

[0140] According to the first intermediate result and the second intermediate result, determine the first correlation between multiple frames of original audio signals with the same time in the multi-channel original audio signals, and the second correlation between the context information of each frame of original audio signals in the multiple frames of original audio signals with the same time respectively;

[0141] According to the third intermediate result, the first correlation and the second correlation, output the multi-channel target audio signal.

[0142] Optionally, when the fusion module 143 fuses the multi-channel target audio signal to obtain a single-channel fusion result, it is specifically used for:

[0143] After the multi-channel target audio signal is sequentially fused through multiple convolution layers, the single-channel fusion result is obtained.

[0144] Optionally, when the prediction module 144 predicts the second text information subsequent to the first text information according to the multi-channel target audio signal, the single-channel fusion result and the first text information already identified from the multi-channel original audio signal, it is specifically used for:

[0145] For any one of the multi-channel target audio signals, the representation vector of the any one of the multi-channel target audio signals and the representation vector of the first text information are fused to obtain the representation vector of the first text sequence corresponding to the any one of the multi-channel target audio signals;

[0146] The representation vector of the single-channel fusion result and the representation vector of the first text information are fused to obtain the representation vector of the second text sequence corresponding to the single-channel fusion result;

[0147] The representation vectors of the first text sequence corresponding to the multi-channel target audio signal respectively, and the representation vector of the second text sequence are taken as the input of the second attention mechanism, and the structure of the second attention mechanism is the same as that of the first attention mechanism;

[0148] The representation vectors of multiple third text sequences output by the second attention mechanism are fused to obtain the representation vector of the target text sequence;

[0149] According to the representation vector of the target text sequence, the second text information subsequent to the first text information is predicted.

[0150] Figure 15 The audio recognition apparatus of the illustrated embodiment can be used to implement the technical solutions of the above method embodiments, and has similar implementation principles and technical effects, which will not be described here again.

[0151] The internal functions and structures of the audio recognition apparatus are described above, and the apparatus can be implemented as an electronic device. Figure 15 An electronic device embodiment provided by the present disclosure is shown in a structural schematic diagram. As shown in the figure, Figure 15 The electronic device includes a memory 151 and a processor 152.

[0152] The memory 151 is used to store programs. In addition to the above programs, the memory 151 can also be configured to store other various data to support operations on the electronic device. Examples of these data include instructions for any application or method operating on the electronic device, contact data, phonebook data, messages, pictures, videos, etc.

[0153] The memory 151 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0154] The processor 152 is coupled to the memory 151 and executes the programs stored in the memory 151 to:

[0155] Obtain multiple original audio signals, each of which includes multiple frames of original audio signals;

[0156] Take the multiple original audio signals as inputs of a first attention mechanism, and output multiple target audio signals according to the context information of each frame of original audio signal in each original audio signal in the time dimension, the first correlation between multiple frames of original audio signals with the same time in the multiple original audio signals, and the second correlation between the respective context information of the multiple frames of original audio signals with the same time and each frame of original audio signal in the multiple frames of original audio signals with the same time;

[0157] Fuse the multiple target audio signals to obtain a single-channel fusion result;

[0158] According to the multiple target audio signals, the single-channel fusion result, and the first text information that has been recognized from the multiple original audio signals, predict second text information subsequent to the first text information.

[0159] Further, as Figure 15As shown, the electronic device can further include a communication component 153, a power supply component 154, an audio component 155, a display 156, and other components. ​ Only some of the components are shown and described, and it is not meant to imply that the electronic device only includes ​ the components shown.

[0160] The communication component 153 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 153 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 153 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra-WideBand (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0161] The power supply component 154 provides power for various components of the electronic device. The power supply component 154 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device.

[0162] The audio component 155 is configured to output and / or input audio signals. For example, the audio component 155 includes a microphone (MIC) that is configured to receive an external audio signal when the electronic device is in an operational mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 151 or transmitted via the communication component 153. In some embodiments, the audio component 155 also includes a speaker for outputting audio signals.

[0163] The display 156 includes a screen, which can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide, and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect a duration and a pressure related to the touching or sliding action.

[0164] In addition, the embodiments of the present disclosure further provide a computer-readable storage medium having stored thereon a computer program, which is executed by a processor to implement the audio recognition method described in the above embodiments.

[0165] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0166] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An audio recognition method, wherein, The method comprises: obtaining a plurality of original audio signals, each of which comprises a plurality of original audio frames; inputting the plurality of original audio signals into a first attention mechanism, and outputting a plurality of target audio signals according to the context information of each original audio frame in each original audio signal in a time dimension, a first correlation between a plurality of original audio frames with the same time in the plurality of original audio signals, and a second correlation between the context information of each of the plurality of original audio frames with the same time and each original audio frame in the plurality of original audio frames with the same time, wherein the first correlation comprises a contribution degree of the plurality of original audio frames with the same time to each original audio frame, and the second correlation comprises a contribution degree of the context information of each of the plurality of original audio frames with the same time to the original audio frame; fusing the plurality of target audio signals to obtain a single-channel fusion result; predicting second text information subsequent to first text information according to the plurality of target audio signals, the single-channel fusion result, and the first text information.

2. The method of claim 1, wherein, The context information of each original audio frame in a time dimension comprises at least one historical audio frame and at least one future audio frame of the original audio frame.

3. The method of claim 2, wherein, The number of the at least one historical audio frame and the at least one future audio frame is determined by the difference between the time when a first audio acquisition device in an audio acquisition array receives an audio signal and the time when a second audio acquisition device receives the audio signal, the first audio acquisition device being the audio acquisition device that receives the audio signal earliest in the audio acquisition array, and the second audio acquisition device being the audio acquisition device that receives the audio signal latest in the audio acquisition array.

4. The method of claim 1, wherein, The first attention mechanism comprises a first linear layer, a second linear layer, and a third linear layer. Outputting a plurality of target audio signals according to the context information of each original audio frame in each original audio signal in a time dimension, a first correlation between a plurality of original audio frames with the same time in the plurality of original audio signals, and a second correlation between the context information of each of the plurality of original audio frames with the same time and each original audio frame in the plurality of original audio frames with the same time, comprises: processing the plurality of original audio signals through the first linear layer to obtain a first intermediate result; processing the context information of each original audio frame in each original audio signal in a time dimension and the plurality of original audio signals through the second linear layer to obtain a second intermediate result; processing the context information of each original audio frame in each original audio signal in a time dimension and the plurality of original audio signals through the third linear layer to obtain a third intermediate result; determine, according to the first intermediate result and the second intermediate result, a first correlation between multiple frames of original audio signals with a same time in the multiple-channel original audio signals, and a second correlation of context information of each of the multiple frames of original audio signals with each frame of original audio signal in the multiple frames of original audio signals with the same time; output the multiple-channel target audio signals according to the third intermediate result, the first correlation and the second correlation.

5. The method of claim 1, wherein, fuse the multiple-channel target audio signals to obtain a single-channel fusion result, including: fusing the multiple-channel target audio signals through multiple convolution layers in sequence to obtain the single-channel fusion result.

6. The method of claim 1, wherein, predict second text information subsequent to the first text information according to the multiple-channel target audio signals, the single-channel fusion result and the first text information that has been recognized from the multiple-channel original audio signals, including: for any one of the multiple-channel target audio signals, fuse a representation vector of the any one of the multiple-channel target audio signals and a representation vector of the first text information to obtain a representation vector of a first text sequence corresponding to the any one of the multiple-channel target audio signals; fuse a representation vector of the single-channel fusion result and the representation vector of the first text information to obtain a representation vector of a second text sequence corresponding to the single-channel fusion result; input the representation vectors of the first text sequences corresponding to the multiple-channel target audio signals and the representation vector of the second text sequence into a second attention mechanism, a structure of the second attention mechanism being the same as a structure of the first attention mechanism; fuse representation vectors of multiple third text sequences output by the second attention mechanism to obtain a representation vector of a target text sequence; predict the second text information subsequent to the first text information according to the representation vector of the target text sequence.

7. An audio recognition model, wherein, The audio recognition model includes an encoder, a convolution module and a decoder. The encoder includes a first attention mechanism, an input of the first attention mechanism being multiple-channel original audio signals, each of the multiple-channel original audio signals including multiple frames of original audio signals, the first attention mechanism being configured to output multiple-channel target audio signals according to context information of each frame of original audio signal in each of the multiple-channel original audio signals in a time dimension, a first correlation between multiple frames of original audio signals with a same time in the multiple-channel original audio signals, and a second correlation of context information of each of the multiple frames of original audio signals with each frame of original audio signal in the multiple frames of original audio signals with the same time, wherein the first correlation includes a contribution degree of the multiple frames of original audio signals to each frame of original audio signal, and the second correlation includes a contribution degree of context information of each of the multiple frames of original audio signals to the each frame of original audio signal; the convolution module is configured to fuse the multiple-channel target audio signals to obtain a single-channel fusion result; and the decoder is configured to predict second text information subsequent to first text information according to the multiple-channel target audio signals, the single-channel fusion result and the first text information that has been recognized from the multiple-channel original audio signals. The decoder is configured to predict second text information subsequent to the first text information according to the multi-channel target audio signals, the single-channel fusion result, and the first text information that has been recognized from the multi-channel original audio signals.

8. The audio recognition model of claim 7, wherein, The decoder comprises a second attention mechanism, a third attention mechanism, a fourth attention mechanism, a fusion layer, and a neural network, the second attention mechanism has the same structure as the first attention mechanism; The third attention mechanism is configured to generate a representation vector of the first text information that has been recognized from the multi-channel original audio signals; The fourth attention mechanism is configured to, for any one of the multi-channel target audio signals, fuse a representation vector of the any one of the multi-channel target audio signals and a representation vector of the first text information to obtain a representation vector of a first text sequence corresponding to the any one of the multi-channel target audio signals, and fuse a representation vector of the single-channel fusion result and the representation vector of the first text information to obtain a representation vector of a second text sequence corresponding to the single-channel fusion result; The second attention mechanism is configured to process the representation vectors of the first text sequences corresponding to the multi-channel target audio signals respectively and the representation vector of the second text sequence, and output representation vectors of third text sequences respectively; The fusion layer is configured to fuse the representation vectors of the third text sequences respectively to obtain a representation vector of a target text sequence; The neural network is configured to predict the second text information subsequent to the first text information according to the representation vector of the target text sequence.

9. The audio recognition model of claim 7, wherein, The audio recognition model is trained by the following method: According to the total number of channels in the audio acquisition array, the number of mask channels is determined; According to the number of mask channels, the number of channels in the audio acquisition array is masked; The multi-channel original audio signals are obtained from the multiple channels that are not masked; The multi-channel original audio signals are used as inputs of the audio recognition model, so that the audio recognition model outputs predicted text information; According to the predicted text information and standard sample text information corresponding to the multi-channel original audio signals, the audio recognition model is trained.

10. An audio recognition apparatus, wherein, The method comprises: an acquisition module configured to acquire multi-channel original audio signals, each of the original audio signals comprising multiple frames of original audio signals; The input module is configured to input the multiple original audio signals as inputs of a first attention mechanism, and the first attention mechanism is configured to output multiple target audio signals according to context information of each frame of original audio signal in each of the original audio signals in a time dimension, first correlations between multiple frames of original audio signals that are the same in time in the multiple original audio signals, and second correlations of respective context information of the multiple frames of original audio signals that are the same in time with each frame of original audio signal in the multiple frames of original audio signals that are the same in time, wherein the first correlations include contribution degrees of the multiple frames of original audio signals that are the same in time to each frame of original audio signal, and the second correlations include contribution degrees of the respective context information of the multiple frames of original audio signals that are the same in time to the each frame of original audio signal. The fusion module is configured to fuse the multiple target audio signals to obtain a single-channel fusion result. The prediction module is configured to predict second text information that is subsequent to first text information that has been recognized from the multiple original audio signals according to the multiple target audio signals, the single-channel fusion result, and the first text information.

11. An electronic device, comprising: comprise: a memory; a processor; and a computer program; wherein the computer program is stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1-6. The computer program, when executed by the processor, implements the method of any one of claims 1-6.

12. A computer readable storage medium having stored thereon a computer program, wherein, ​

Citation Information

Patent Citations

  • Speech recognition method

    US20120065968A1