An audio information processing method, apparatus and electronic device

CN115831138BActive Publication Date: 2026-09-22LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211230560.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2026-09-22
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

[0003]现有技术中,一般采用神经网络模型进行降噪,如CNN(Convolutional NeuralNetworks,卷积神经网络)和FullConnect(全连接)能够处理局部信息,其只能看到当前时刻前后一小段时间的信息,但不能考虑到长时信息,因此,降噪效果不理想

Benefits of technology

[0047]经由上述的技术方案可知,本申请提供了一种音频信息处理方法,对于音频信息中的各个语音块分别进行处理得到相应的语音块特征,并且对于语音块特征识别得到该语音块所属的场景,确定与目标语音块属于同一场景的历史语音块,将该历史语音块对应的历史块特征与目标语音块对应的目标块特征融合得到融合块特征,基于该融合块特征对于该目标块特征进行降噪处理,得到第一降噪块特征,依次对于该音频信息中的多个语音块分别进行上述过程,基于上述的第一降噪块特征能够第一目标音频信息。由于融合块特征是结合了历史语音块与目标语音块的特征的,该历史语音块与该目标语音块具有相同场景中的音频的特征,该融合块特征是长时信息,基于该融合块特征对于目标块特征进行的降噪处理结合了长时信息,相对于仅采用单独的目标语音块的特征进行降噪处理,降噪效果更优。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115831138B_ABST
    Figure CN115831138B_ABST
Patent Text Reader

Abstract

The application provides an audio information processing method, speech blocks in audio information are processed respectively to obtain corresponding speech block features, a scene to which the speech blocks belong is identified by recognizing the speech block features, historical speech blocks belonging to the same scene as a target speech block are determined, historical block features corresponding to the historical speech blocks are fused with target block features corresponding to the target speech block to obtain fused block features, noise reduction processing is performed on the target block features based on the fused block features to obtain first noise reduction block features, and first target audio information is obtained based on the first noise reduction block features. Since the fused block features are combined features of the historical speech blocks and the target speech block, the historical speech blocks and the target speech block have audio features in the same scene, the fused block features are long-time information, the noise reduction processing performed on the target block features based on the fused block features combines long-time information, and the noise reduction effect is more optimal relative to noise reduction processing performed only by using features of the target speech block alone.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information technology, and more specifically, to an audio information processing method, apparatus, and electronic device. Background Technology

[0002] In scenarios such as phone calls and online meetings, noise can significantly impact the user experience. Therefore, the requirements for noise reduction in these scenarios are extremely high.

[0003] In existing technologies, neural network models are generally used for noise reduction. For example, CNN (Convolutional Neural Networks) and FullConnect can process local information. They can only see information within a short period of time before and after the current moment, but cannot take into account long-term information. Therefore, the noise reduction effect is not ideal. Summary of the Invention

[0004] In view of this, this application provides an audio information processing method, as follows:

[0005] An audio information processing method, comprising:

[0006] Obtain audio information, wherein the audio information contains at least one speech block;

[0007] At least one speech block feature is obtained based on the at least one speech block, and the speech block corresponds to the speech block feature;

[0008] Identify the features of at least one speech block to obtain the scene to which the at least one speech block belongs;

[0009] Obtain historical voice blocks belonging to the same scene as the target voice block, wherein the generation time of the historical voice block is earlier than the generation time of the target voice block;

[0010] The target block feature corresponding to the target speech block is fused with the historical block feature corresponding to the historical speech block to obtain the fused block feature, wherein the target speech block is one of the at least one speech block;

[0011] The first processing model is controlled to perform noise reduction processing on the target block features based on the fused block features to obtain the first noise-reduced block features;

[0012] The first target audio information is obtained based on the features of the first noise reduction block.

[0013] Optionally, in the above method, wherein:

[0014] The first processing model can also process the at least one speech block feature to obtain at least one second noise reduction block feature;

[0015] The second target audio information is obtained based on at least one feature of the second noise reduction block.

[0016] Optionally, in the above method, obtaining at least one speech block feature based on the at least one speech block includes:

[0017] Determine at least two frames of data contained in the target speech block, each frame of data corresponding to the audio information of the target duration, and each frame of data containing features of the target dimension.

[0018] The target speech block features are obtained based on the target dimension features contained in each frame of data in the target speech block.

[0019] Optionally, in the above method, identifying the features of the at least one speech block to obtain the scene to which the at least one speech block belongs includes:

[0020] The target speech block features are used as input features to input the second processing model to obtain the probability that the target speech block belongs to at least two preset scenarios;

[0021] Based on the probability that the target speech block belongs to the first preset scenario, the agreed selection condition is met, and the first preset scenario is selected as the scenario of the target speech block.

[0022] Optionally, in the above method, fusing the target block features corresponding to the target speech block with the historical block features corresponding to the historical speech block to obtain fused block features includes:

[0023] The target block features corresponding to the target speech block and the historical block features corresponding to the historical speech block are fused based on the feature dimension of the frame data to obtain fused block features. Each frame data contains features of the target dimension.

[0024] Optionally, in the above method, the step of performing noise reduction processing on the target block features based on the fused block features to obtain the first noise-reduced block features includes:

[0025] Weighted features are obtained based on the fusion block features;

[0026] The weighted features are fused with the target block features to obtain secondary fused block features;

[0027] The secondary fusion block features are input into the first processing model for noise reduction to obtain the first noise reduction block features.

[0028] Optionally, in the above method, obtaining weighted features based on the fused block features includes:

[0029] The fusion block features are used as the first and second parameters of the third processing model, respectively.

[0030] The target block features are used as the third parameter of the third processing model;

[0031] The third processing model is controlled to obtain weighted features based on the first parameter, the second parameter, and the third parameter.

[0032] Optionally, the above method fuses the weighted features with the target block features to obtain secondary fused block features, including:

[0033] The weighted features and the target block features are concatenated based on the feature dimensions of the frame data to obtain the secondary fused block features;

[0034] or

[0035] The weighted features are stacked with the target block features to obtain secondary fused block features.

[0036] An audio information processing device, comprising:

[0037] An acquisition module is used to acquire audio information, the audio information containing at least one speech block;

[0038] The feature module is used to obtain at least one speech block feature based on the at least one speech block, wherein the speech block corresponds to the speech block feature;

[0039] The recognition module is used to recognize the features of the at least one speech block and obtain the scene to which the at least one speech block belongs;

[0040] The acquisition module is used to acquire historical voice blocks belonging to the same scene as the target voice block, wherein the generation time of the historical voice block is earlier than the generation time of the target voice block;

[0041] The fusion module is used to fuse the target block features corresponding to the target speech block with the historical block features corresponding to the historical speech block to obtain fused block features, wherein the target speech block is one of the at least one speech block;

[0042] The noise reduction module is used to control the first processing model to perform noise reduction processing on the target block features based on the fusion block features to obtain the first noise-reduced block features;

[0043] The audio information module is used to obtain the first target audio information based on the features of the first noise reduction block.

[0044] An electronic device includes: a memory and a processor;

[0045] The memory stores the processing program;

[0046] The processor is used to load and execute the processing program stored in the memory to implement the steps of the audio information processing method as described in any of the preceding claims.

[0047] As can be seen from the above technical solution, this application provides an audio information processing method. Each speech block in the audio information is processed separately to obtain corresponding speech block features. The scene to which the speech block belongs is identified based on the speech block features. Historical speech blocks belonging to the same scene as the target speech block are determined. The historical block features corresponding to the historical speech block are fused with the target block features corresponding to the target speech block to obtain fused block features. Noise reduction processing is performed on the target block features based on the fused block features to obtain first noise-reduced block features. The above process is performed sequentially on multiple speech blocks in the audio information. Based on the first noise-reduced block features, the target audio information can be processed. Since the fused block features combine the features of historical speech blocks and the target speech block, and these historical speech blocks and the target speech block have audio features from the same scene, the fused block features are long-term information. The noise reduction processing of the target block features based on the fused block features incorporates long-term information, resulting in better noise reduction performance compared to using only the features of the target speech block alone. Attached Figure Description

[0048] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0049] Figure 1 This is a flowchart of an embodiment 1 of an audio information processing method provided in this application;

[0050] Figure 2 This is a flowchart of an embodiment 2 of an audio information processing method provided in this application;

[0051] Figure 3 This is a flowchart of an embodiment 3 of an audio information processing method provided in this application;

[0052] Figure 4 This is a flowchart of embodiment 4 of the audio information processing method provided in this application;

[0053] Figure 5 This is a flowchart of embodiment 5 of the audio information processing method provided in this application;

[0054] Figure 6 This is a schematic diagram of a scenario for an audio information processing method provided in this application;

[0055] Figure 7 This is a schematic diagram of an embodiment of an audio information processing device provided in this application. Detailed Implementation

[0056] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0057] like Figure 1 The diagram shown is a flowchart of an embodiment 1 of an audio information processing method provided in this application. The method is applied to an electronic device and includes the following steps:

[0058] Step S101: Obtain audio information;

[0059] The audio information includes at least one speech block.

[0060] The audio information can be real-time audio information or audio information obtained after receiving all audio information.

[0061] The audio information can be switched according to a preset time length to obtain multiple audio blocks.

[0062] If the obtained audio information is not greater than the preset time length, the audio information will be directly processed as a speech block.

[0063] If the obtained audio information is longer than the preset time length, the audio information can be divided into multiple speech blocks.

[0064] For example, the preset time length can be a small value, such as within 500ms (milliseconds), such as 10ms, 30ms, 100ms, 300ms, etc. This application does not impose any restrictions on the specific value of the preset time length.

[0065] Step S102: Obtain at least one speech block feature based on at least one speech block;

[0066] The speech block corresponds to the speech block feature.

[0067] Specifically, the speech blocks contained in the audio information are processed separately to obtain the corresponding speech block features.

[0068] It should be noted that when the audio information is obtained in real time, the audio information is divided into segments based on the real-time received audio information, and the segments are processed in real time to obtain the corresponding audio segment features.

[0069] In practice, the speech block can be processed separately using the short-time Fourier transform (STFT) to obtain speech block features.

[0070] Step S103: Identify the features of the at least one speech block to obtain the scene to which the at least one speech block belongs;

[0071] Specifically, the features of the speech block are identified to determine the scene to which the corresponding speech block belongs.

[0072] This scene can be divided based on the different sounds it contains.

[0073] Specifically, the scenarios include the following sounds: air conditioner noise, clapping, keyboard sounds, baby crying, cat meowing, dog barking, etc.

[0074] Of course, in specific implementations, this scenario is not limited to the scenario example provided in this embodiment, and can also be other scenarios.

[0075] In practice, after identifying the scene to which each speech block belongs, the scene to which the speech block belongs is recorded.

[0076] Step S104: Obtain historical speech blocks belonging to the same scene as the target speech block;

[0077] The historical speech block was generated earlier than the target block.

[0078] Among these steps, historical voice blocks that belong to the same scene as the target voice block are selected from a number of historical voice blocks.

[0079] Specifically, this scenario refers to a sound scenario, that is, a scenario with the same sound, where the noise appearing in the same scenario is similar.

[0080] The number of historical voice blocks belonging to the same scene as the target voice block can be one or more. This application does not impose a limit on the number of historical voice blocks belonging to the same scene as the target voice block.

[0081] The target speech block is the speech block currently being processed.

[0082] In practice, for one or more speech blocks in the audio information, one speech block is selected as the target speech block in the order of its generation / acquisition time.

[0083] In practice, if the audio information is obtained in real time, the latest received voice block is used as the target voice block, and historical voice blocks belonging to the same scene are identified.

[0084] The historical voice block can be a voice block that does not belong to the audio information obtained in step S101. For example, it can be audio information obtained before the audio information obtained in step S101, or it can be a voice block that belongs to the audio information obtained in step S101.

[0085] For example, if the target speech block belongs to a cat meowing scene, then search for historical speech blocks that also belong to the cat meowing scene in the historical speech blocks. There can be one or more historical speech blocks that belong to the same scene as the target speech block.

[0086] Step S105: Fuse the target block features corresponding to the target speech block with the historical block features corresponding to the historical speech block to obtain fused block features;

[0087] The target speech block is one of the at least one speech blocks.

[0088] Among them, the historical speech block is processed to obtain historical block features;

[0089] Then, the speech block features corresponding to the target speech block are fused with the historical block features to obtain the fused block features.

[0090] The fusion block features contain specific information belonging to the same scene. By fusing feature blocks from the same scene, specific information from the same scene can be centrally represented.

[0091] Step S106: Control the first processing model to perform noise reduction processing on the target block features based on the fused block features to obtain the first noise-reduced block features;

[0092] The fusion block feature can centrally represent specific information in its corresponding scene. Therefore, the first processing model performs noise reduction processing on the target block feature corresponding to a speech block in the audio information obtained this time based on the fusion block feature. It can perform noise reduction processing on the target block feature based on the specific information in the scene to which the speech block belongs, and the noise reduction processing effect is better.

[0093] The first processing model can be a CNN (Convolutional Neural Networks) or a FullConnect network, etc.

[0094] The information contained in the fusion of the historical speech block and the current speech block is such that it can be processed by models that process short-term information, such as CNN or FullConnect.

[0095] Specifically, the audio blocks contained in the audio information in step S101 are processed sequentially using steps S103-106 to obtain the corresponding number of first noise reduction block features.

[0096] Step S107: Obtain the first target audio information based on the features of the first noise reduction block.

[0097] The first target audio information is the audio information after noise reduction processing of the obtained audio information.

[0098] In this process, after processing each target feature block to obtain the first noise reduction block feature, they are combined according to the order of the corresponding speech blocks to obtain the first target audio information.

[0099] In practice, each first noise reduction block feature is first subjected to inverse Fourier transform to obtain a segment of noise-reduced audio information corresponding to the first noise reduction block feature. Then, the segments of noise-reduced audio information are combined in sequence to obtain the first target audio information.

[0100] Specifically, the first noise reduction block feature is processed to obtain a noise-reduced audio information segment. Then, each audio block contained in the audio information in step S101 is processed sequentially to obtain a corresponding number of first noise reduction block features. Each of the multiple first noise reduction block features is processed to obtain a corresponding noise-reduced audio information segment. Finally, the first target audio information is obtained by splicing together the noise-reduced audio information segments.

[0101] In specific implementation, the first processing model can also process the at least one speech block feature to obtain at least one second noise reduction block feature; and obtain the second target audio information based on the at least one second noise reduction block feature.

[0102] Specifically, after processing the speech blocks in the audio information to obtain speech block features, the speech block features are directly used as input information into the first processing model, so that the first processing model performs noise reduction processing based solely on the speech block features to obtain the second target information.

[0103] Among them, the noise reduction effect of the first target information is better than that of the second target information.

[0104] In summary, the audio information processing method provided in this embodiment processes each speech block in the audio information to obtain corresponding speech block features. For each speech block feature, the scene to which the speech block belongs is identified. Historical speech blocks belonging to the same scene as the target speech block are determined, and their historical block features are fused with the target block features to obtain fused block features. Based on these fused block features, the target block features are denoised to obtain first denoised block features. This process is repeated for multiple speech blocks in the audio information. Based on the first denoised block features, the target audio information can be denoised. Since the fused block features combine the features of historical and target speech blocks, which share audio features from the same scene, and since the fused block features are long-term information, the denoising process on the target block features incorporates long-term information. Therefore, compared to denoising using only the features of the target speech block alone, the denoising effect is superior.

[0105] like Figure 2 The diagram shown is a flowchart of an embodiment 2 of an audio information processing method provided in this application. The method includes the following steps:

[0106] Step S201: Obtain audio information;

[0107] Step S201 is the same as step S101 in Embodiment 1, and will not be described again in this embodiment.

[0108] Step S202: Determine that the target speech block contains at least two frames of data;

[0109] The target speech block is one of multiple speech blocks in the audio information.

[0110] Each frame of data corresponds to the audio information of the target duration, and each frame of data contains features of the target dimension.

[0111] The audio information contains at least one speech block, and each speech block contains multiple frames of data.

[0112] The audio information is divided into frames according to a target duration, which can be set according to the actual situation. For example, it can be set to approximately 10 to 30 ms per frame.

[0113] A single speech block can contain 1 to 10 frames of data.

[0114] Specifically, each frame of data in the target speech block is subjected to a short-time Fourier transform to obtain the target dimension features corresponding to each frame of data.

[0115] In practice, the audio information can be divided into multiple frames first, and then, based on the number of frames contained in the agreed-upon speech block, the divided multi-frame data can be further divided into multiple speech blocks.

[0116] In practice, since the audio information is continuous, in order to perform noise reduction processing, the information is sampled to obtain a sampling matrix, which represents the audio information.

[0117] As an example, 160 sampling points are set in a frame of data to sample the audio information, resulting in a 1*160 dimensional matrix. Fourier transform is applied to this 1*160 dimensional matrix to obtain a 256-dimensional matrix. That is, the feature corresponding to this frame of data is a 1*256 dimensional matrix.

[0118] Step S203: Based on the features of the target dimension contained in each frame of data in the target speech block, obtain the features of the target speech block;

[0119] Once the number of frames in the target speech block is determined, the target speech block features can be obtained based on the target dimension features contained in each frame.

[0120] The target speech block feature is a matrix obtained by multiplying the number of frames by the target dimension feature in each frame.

[0121] As an example, if a speech block contains C frames of data, and each frame contains F-dimensional features, then the speech block contains a C*F matrix of features, where C and F are positive integers.

[0122] As an example, the feature corresponding to one frame of data is a 1*256 dimensional matrix. If a speech block contains 10 frames of data, then the resulting speech block feature is a 10*256 dimensional matrix.

[0123] Step S204: Identify the features of the at least one speech block to obtain the scene to which the at least one speech block belongs;

[0124] Step S205: Obtain historical speech blocks belonging to the same scene as the target speech block;

[0125] Step S206: Fuse the target block features corresponding to the target speech block with the historical block features corresponding to the historical speech block to obtain fused block features;

[0126] Step S207: Control the first processing model to perform noise reduction processing on the target block features based on the fused block features to obtain the first noise-reduced block features;

[0127] Step S208: Obtain the first target audio information based on the features of the first noise reduction block.

[0128] Steps S204-208 are the same as steps S103-107 in Example 1, and will not be described again in this example.

[0129] In summary, the audio information processing method provided in this embodiment determines that a target speech block contains multiple frames of data, each frame of data contains features of the target dimension, and obtains the target speech block features corresponding to the target speech block based on the number of frames of data contained in the target speech block and the features of the target dimension contained in each frame of data. This embodiment clarifies the process of obtaining the corresponding target speech block features for a target speech block, providing a processing foundation for subsequent processes such as determining the scene to which the target speech block belongs and obtaining fusion block features.

[0130] like Figure 3 The diagram shown is a flowchart of an embodiment 3 of an audio information processing method provided in this application. The method includes the following steps:

[0131] Step S301: Obtain audio information;

[0132] Step S302: Obtain at least one speech block feature based on at least one speech block;

[0133] Steps S301-302 are the same as steps S101-102 in Example 1, and will not be described again in this example.

[0134] Step S303: Input the target speech block features as input features into the second processing model to obtain the probability that the target speech block belongs to at least two preset scenarios;

[0135] In this process, after obtaining the speech block features corresponding to the speech block, a target speech block is determined, and the target speech block is used as the input feature to input the second processing model so that the second processing model can determine the probability that the target speech block belongs to multiple preset scenarios.

[0136] Specifically, the second processing model is a module for detecting sound scenes. It can use CNN or FullConnect, etc. The second processing model can use the same type of model as the first processing model in Example 1, but the parameters of the two are different.

[0137] The target speech block feature is a matrix. If the target speech block contains features of a C*F matrix, then the target speech block feature is a C*F dimensional matrix, where C and F are positive integers.

[0138] The second processing model processes the input features to obtain the probabilities of the input features belonging to multiple preset scenarios.

[0139] For example, the preset scenarios include five noises: air conditioner noise, clapping, keyboard clicks, cat meows, and dog barks. The target speech block feature C*F dimensional matrix is ​​input into the second processing model, and the second processing model outputs the processing results. The probabilities that the target speech block belongs to the above preset scenarios are 75%, 15%, 7%, 2%, and 1%, respectively.

[0140] Step S304: Based on the probability that the target speech block belongs to the first preset scenario meets the agreed selection conditions, select the first preset scenario as the scenario of the target speech block;

[0141] Among the probabilities corresponding to multiple preset scenarios output by the second processing model, one preset scenario that meets the agreed conditions is selected as the scenario to which the target speech block belongs.

[0142] The agreed-upon condition is specifically the one with the highest probability value among these multiple probabilities.

[0143] For example, the probability that the target speech block belongs to one of the five preset scenarios—air conditioner noise, clapping, keyboard sounds, cat meows, and dog barks—is 75%, 15%, 7%, 2%, and 1%, respectively. The air conditioner noise with the highest probability is selected as the scenario to which the target speech block belongs.

[0144] Step S305: Obtain historical speech blocks belonging to the same scene as the target speech block;

[0145] Step S306: Fuse the target block features corresponding to the target speech block with the historical block features corresponding to the historical speech block to obtain fused block features;

[0146] Step S307: Control the first processing model to perform noise reduction processing on the target block features based on the fused block features to obtain the first noise-reduced block features;

[0147] Step S308: Obtain the first target audio information based on the features of the first noise reduction block.

[0148] Steps S305-308 are the same as steps S104-107 in Example 1, and will not be described again in this example.

[0149] In summary, the audio information processing method provided in this embodiment employs a second processing model to process the target speech block features as input features, thereby obtaining the probability that the target speech block belongs to multiple preset scenarios. The first preset scenario whose probability meets the agreed conditions is selected as the scenario to which the target speech block belongs, providing a basis for subsequent selection of historical speech blocks.

[0150] like Figure 4 The diagram shown is a flowchart of an embodiment 4 of an audio information processing method provided in this application. The method includes the following steps:

[0151] Step S401: Obtain audio information;

[0152] Step S402: Obtain at least one speech block feature based on at least one speech block;

[0153] Step S403: Identify the features of the at least one speech block to obtain the scene to which the at least one speech block belongs;

[0154] Step S404: Obtain historical speech blocks belonging to the same scene as the target speech block;

[0155] Steps S401-404 are the same as steps S101-104 in Example 1, and will not be described again in this example.

[0156] Step S405: The target block features corresponding to the target speech block and the historical block features corresponding to the historical speech block are fused based on the feature dimensions of the frame data to obtain fused block features;

[0157] Each frame of data contains features of the target dimension.

[0158] The target speech block contains at least two frames of data, and each historical speech block also contains the same number of frames of data.

[0159] Specifically, the target block features corresponding to the target speech block are fused with the historical block features corresponding to the historical speech block to obtain fused block features.

[0160] Each target block feature is obtained based on the number of frames in the frame data contained in each target block and the features of each frame data. Correspondingly, when the target block features are fused with the historical block features, the feature dimensions of the frame data are used for fusion.

[0161] As an example, the target block feature is a C*F dimensional matrix. If there are (B-1) historical speech blocks, the historical speech block feature is also a C*F dimensional matrix. By concatenating these B matrices along the feature dimensions of the frame data, we obtain a (BC)*F dimensional matrix. This (BC)*F dimensional matrix is ​​the fused block feature, where the values ​​of B, C, and F are positive integers.

[0162] The fused block feature contains all features from the target block feature and the historical block feature. Moreover, since the target speech block and the historical speech block corresponding to the target block feature and the historical block feature belong to the same scene, the fused block feature is a concentrated representation of specific information in the scene. The fused block feature contains long-term information of the scene to which the target speech block belongs in the audio information. Therefore, the first processing model performs noise reduction processing on the target block feature based on the fused block feature. It can combine the specific information in the scene and consider the long-term information to perform noise reduction processing on the target block feature for the scene. Compared with the first processing model that only performs noise reduction processing based on the target block feature, the noise reduction effect is better.

[0163] Step S406: Control the first processing model to perform noise reduction processing on the target block features based on the fused block features to obtain the first noise-reduced block features;

[0164] Step S407: Obtain the first target audio information based on the features of the first noise reduction block.

[0165] Steps S406-407 are the same as steps S106-107 in Example 1, and will not be described again in this example.

[0166] In summary, in the audio information processing method provided in this embodiment, after determining that the target speech block belongs to a historical speech block in the same scene, the historical block features corresponding to the historical speech block are determined. The historical block features and the target block features corresponding to the target speech block are fused based on the feature dimensions of the frame data to obtain fused block features. The fused block features are a concentrated representation of specific information in the scene. The fused block features contain long-term information of the scene to which the target speech block belongs in the audio information. Subsequently, the first processing model performs noise reduction processing on the target block features based on the fused block features. Considering the long-term information, the noise reduction effect is better.

[0167] like Figure 5 The flowchart shown is a fifth embodiment of an audio information processing method provided in this application. The method includes the following steps:

[0168] Step S501: Obtain audio information;

[0169] Step S502: Obtain at least one speech block feature based on at least one speech block;

[0170] Step S503: Identify the features of the at least one speech block to obtain the scene to which the at least one speech block belongs;

[0171] Step S504: Obtain historical speech blocks belonging to the same scene as the target speech block;

[0172] Step S505: The target block feature corresponding to the target speech block and the historical block feature corresponding to the historical speech block are fused based on the feature dimension of the frame data to obtain the fused block feature;

[0173] Steps S501-505 are the same as steps S401-405 in Example 4, and will not be described again in this example.

[0174] Step S506: Obtain weighted features based on the fused block features;

[0175] The fusion block feature contains specific information about the scene to which the target speech block belongs, specifically the long-term information in the audio information corresponding to that scene.

[0176] In this step, a weighted feature is obtained based on the fusion block feature. Specifically, the weighted feature is the long-term information of the scene to which the target speech block belongs.

[0177] The weighted features are obtained by using the third processing model and the features of the fusion block.

[0178] Specifically, the third processing model uses an attention model.

[0179] Specifically, step S506 includes:

[0180] Step S5061: Use the fused block features as the first and second parameters of the third processing model, respectively;

[0181] Step S5062: Use the target block features as the third parameter of the third processing model;

[0182] The third processing model includes three parameters: the first parameter, the second parameter, and the third parameter.

[0183] The third processing model uses an attention model, where the first parameter is the key (denoted as K), the second parameter is the value (denoted as V), and the third parameter is the query (denoted as Q).

[0184] Specifically, the fused block features are used as the first and second parameters of the third processing model, that is, the fused block features are used as K and V respectively, and the target block features are used as the third parameter Q.

[0185] Step S5063: Control the third processing model to obtain weighted features based on the first parameter, the second parameter, and the third parameter.

[0186] The third processing model processes the first, second, and third parameters to obtain weighted features.

[0187] The weighted feature is applied to the target block feature during subsequent processing, with weights added to the target block feature based on the scene to which the target speech block belongs. This allows the first processing model to perform more targeted processing based on the scene when denoising the target block feature.

[0188] The third processing model can be configured with a specific attention formula.

[0189] The attention formula is specifically: Q*K T *V (1)

[0190] Substituting the above-mentioned fusion block features (K and V) and target block features (Q) into the above formula yields a weighted feature.

[0191] As an example, the target block feature is a C*F dimensional matrix. If there are (B-1) historical speech blocks, the fusion block feature is a (BC)*F dimensional matrix. Substituting the above as the fusion block features of K and V, and the target block features of Q, into the above formula (1), the result is a C*F matrix. The weighted feature is this C*F matrix, where the values ​​of B, C, and F are positive integers.

[0192] Step S507: Fuse the weighted features with the target block features to obtain secondary fused block features;

[0193] Specifically, the weighted feature is fused with the target block feature to obtain a secondary fused block feature that incorporates all features of the weighted long-term information and the target block feature.

[0194] Among them, the weighted features and the target block features are matrices with the same structure, and there are multiple ways to fuse them.

[0195] Specifically, step S507 includes:

[0196] Step S5071: Concatenate the weighted features and the target block features based on the feature dimensions of the frame data to obtain secondary fused block features;

[0197] or

[0198] Step S5072: Stack the weighted features and the target block features to obtain secondary fused block features.

[0199] In this process, the weighted features and the target block features are concatenated based on the feature dimensions of the frame data, resulting in a second fusion block feature with doubled frame data dimensions.

[0200] For example, the weighted features are C*F matrices, and the target block features are also C*F matrices. By concatenating the two based on the feature dimensions of the frame data, the resulting secondary fused block features are C*2F matrices.

[0201] In this process, the weighted features are stacked with the target block features, resulting in a doubled number of channels in the secondary fusion block features.

[0202] For example, the weighted features are C*F matrices, and the target block features are also C*F matrices. Stacking the two together results in a secondary fused block feature matrix of 2C*F.

[0203] Step S508: Input the secondary fusion block features into the first processing model for noise reduction processing to obtain the first noise reduction block features;

[0204] Specifically, the secondary fusion block feature is used as input information to the first processing model so that the first processing model can perform noise reduction processing to obtain the first noise reduction block feature.

[0205] The secondary fusion block feature is obtained based on the weighted feature and the target block feature. The weighted feature is long-term information with weighting function. The weighted long-term information represents the specific information of the scene to which the target speech block belongs. The secondary fusion block feature obtained by concatenating the weighted feature and the target block feature contains the weighted long-term information and the target block feature.

[0206] The first processing model performs noise reduction on the input secondary fusion block features to obtain the first noise reduction block features, thereby realizing the noise reduction processing on the target speech block.

[0207] Step S509: Obtain the first target audio information based on the features of the first noise reduction block.

[0208] Step S509 is the same as step S407 in Example 4, and will not be described again in this example.

[0209] In summary, the audio information processing method provided in this embodiment obtains weighted features based on fusion block features. These weighted features contain specific information about the scene to which the target speech block belongs. The weighted features are fused with the target block features to obtain secondary fusion block features. The secondary fusion block features are then input into a first processing model for noise reduction to obtain first noise reduction block features, thereby achieving noise reduction for the target speech block. In this noise reduction process, the secondary fusion block features contain weighted features in addition to the target block features. Correspondingly, during the noise reduction process of the first processing model, the weighted features corresponding to the scene to which the target speech block belongs are combined, and long-term information related to the scene with weighted weights is considered, resulting in a better noise reduction effect.

[0210] like Figure 6The diagram shown is a scenario illustration of an audio information processing method provided in this application. The scenario includes three processing models: the first processing model 601 uses a CNN model for noise reduction, the second processing model 602 uses a CNN model for scene recognition, and the third processing model 603 uses an attention model. The first processing model and the second processing model have different specific parameters and perform different functions.

[0211] The audio information processing procedure in this scenario is as follows:

[0212] Step S601: Input audio information, which contains multiple speech blocks;

[0213] Step S602: Perform Fourier transform on the target speech block to obtain the target block features;

[0214] Step S603: The second processing model identifies the target block features corresponding to the input speech block and outputs the first scene to which the speech block belongs;

[0215] Step S604: Obtain the historical voice block belonging to the first scene from the audio information;

[0216] Step S605: After converting the historical speech block into historical block features, fuse it with the target block features to obtain fused block features, which have long-term information;

[0217] Step S606: The third processing model processes the target block features and fused block features to obtain weighted features;

[0218] Step S607: The weighted feature is concatenated with the target block feature to obtain the secondary fused block feature;

[0219] Step S608: The first processing model processes the features of the secondary fusion block to obtain the noise reduction features;

[0220] Step S609: Performing an inverse Fourier transform on the denoising feature yields the denoised speech block;

[0221] Specifically, steps S602-609 are executed for each speech block in the audio information.

[0222] Step S610: Concatenate the denoised speech blocks according to their corresponding temporal order to obtain the denoised audio information.

[0223] Corresponding to the above-described embodiment of an audio information processing method provided in this application, this application also provides an embodiment of an apparatus for applying the audio information processing method.

[0224] like Figure 7The diagram shown is a structural schematic of an embodiment of an audio information processing device provided in this application. The device includes the following structure: an acquisition module 701, a feature module 702, an identification module 703, an acquisition module 704, a fusion module 705, a noise reduction module 706, and an audio information module 707.

[0225] The obtaining module 701 is used to obtain audio information, which includes at least one speech block;

[0226] The feature module 702 is used to obtain at least one speech block feature based on the at least one speech block, wherein the speech block corresponds to the speech block feature;

[0227] The recognition module 703 is used to recognize the features of the at least one speech block and obtain the scene to which the at least one speech block belongs;

[0228] The acquisition module 704 is used to acquire historical voice blocks belonging to the same scene as the target voice block, wherein the generation time of the historical voice block is earlier than the generation time of the target voice block.

[0229] The fusion module 705 is used to fuse the target block features corresponding to the target speech block with the historical block features corresponding to the historical speech block to obtain fused block features, wherein the target speech block is one of the at least one speech block;

[0230] The noise reduction module 706 is used to control the first processing model to perform noise reduction processing on the target block features based on the fusion block features to obtain the first noise reduction block features;

[0231] The audio information module 707 is used to obtain the first target audio information based on the features of the first noise reduction block.

[0232] Optional, wherein:

[0233] The first processing model can also process the at least one speech block feature to obtain at least one second noise reduction block feature;

[0234] The second target audio information is obtained based on at least one feature of the second noise reduction block.

[0235] Optionally, the feature module is specifically used for:

[0236] Determine at least two frames of data contained in the target speech block, each frame of data corresponding to the audio information of the target duration, and each frame of data containing features of the target dimension.

[0237] The target speech block features are obtained based on the target dimension features contained in each frame of data in the target speech block.

[0238] Optionally, the identification module is specifically used for

[0239] The target speech block features are used as input features to input the second processing model to obtain the probability that the target speech block belongs to at least two preset scenarios;

[0240] Based on the probability that the target speech block belongs to the first preset scenario, the agreed selection condition is met, and the first preset scenario is selected as the scenario of the target speech block.

[0241] Optionally, the fusion module is specifically used for:

[0242] The target block features corresponding to the target speech block and the historical block features corresponding to the historical speech block are fused based on the feature dimension of the frame data to obtain fused block features. Each frame data contains features of the target dimension.

[0243] Optionally, the noise reduction module includes:

[0244] Weighting units are used to obtain weighted features based on the features of the fused block;

[0245] A fusion unit is used to fuse the weighted features with the target block features to obtain secondary fused block features;

[0246] The noise reduction unit is used to input the secondary fusion block features into the first processing model for noise reduction processing to obtain the first noise reduction block features.

[0247] Optionally, the weighting unit is specifically used for:

[0248] The fusion block features are used as the first and second parameters of the third processing model, respectively.

[0249] The target block features are used as the third parameter of the third processing model;

[0250] The third processing model is controlled to obtain weighted features based on the first parameter, the second parameter, and the third parameter.

[0251] Optionally, the fusion unit is specifically used for:

[0252] The weighted features and the target block features are concatenated based on the feature dimensions of the frame data to obtain the secondary fused block features;

[0253] or

[0254] The weighted features are stacked with the target block features to obtain secondary fused block features.

[0255] It should be noted that the functions of each component of the audio information processing device provided in this embodiment are explained in the method embodiment, and will not be repeated in this embodiment.

[0256] In summary, the audio information processing apparatus provided in this embodiment processes each speech block in the audio information to obtain corresponding speech block features. It then identifies the scene to which the speech block belongs based on the speech block features, determines historical speech blocks belonging to the same scene as the target speech block, fuses the historical block features corresponding to the historical speech block with the target block features corresponding to the target speech block to obtain fused block features, and performs noise reduction processing on the target block features based on the fused block features to obtain first noise-reduced block features. This process is repeated for multiple speech blocks in the audio information, and the target audio information can be obtained based on the first noise-reduced block features. Since the fused block features combine the features of historical speech blocks and the target speech block, which share audio features from the same scene, and since the fused block features are long-term information, the noise reduction processing on the target block features based on the fused block features incorporates long-term information. Therefore, compared to using only the features of the target speech block for noise reduction, the noise reduction effect is superior.

[0257] Corresponding to the above-described embodiment of an audio information processing method provided in this application, this application also provides an electronic device and a readable storage medium corresponding to the audio information processing method.

[0258] The electronic device includes: a memory and a processor;

[0259] The memory stores the processing program;

[0260] The processor is used to load and execute the processing program stored in the memory to implement the steps of the audio information processing method as described in any of the preceding claims.

[0261] For details on the specific audio information processing method implemented in this electronic device, please refer to the aforementioned audio information processing method embodiments.

[0262] The readable storage medium stores a computer program that is invoked and executed by a processor to implement the steps of the audio information processing method as described in any one of the preceding claims.

[0263] Specifically, the computer program stored on the readable storage medium executes the audio information processing method, as described in the aforementioned audio information processing method embodiments.

[0264] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. The apparatus provided in the embodiments is described simply because it corresponds to the method provided in the embodiments; relevant parts can be found in the method section.

[0265] The above description of the provided embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features provided herein.

Claims

1. An audio information processing method, comprising: Obtain audio information, wherein the audio information contains at least one speech block; At least one speech block feature is obtained based on the at least one speech block, and the speech block corresponds to the speech block feature; Identify the features of at least one speech block to obtain the scene to which the at least one speech block belongs; Obtain historical voice blocks belonging to the same scene as the target voice block, wherein the generation time of the historical voice block is earlier than the generation time of the target voice block; The target block feature corresponding to the target speech block is fused with the historical block feature corresponding to the historical speech block to obtain the fused block feature, wherein the target speech block is one of the at least one speech block; The first processing model is controlled to perform noise reduction processing on the target block features based on the fused block features to obtain the first noise-reduced block features; The first target audio information is obtained based on the features of the first noise reduction block.

2. The method according to claim 1, wherein: The first processing model can also process the at least one speech block feature to obtain at least one second noise reduction block feature; The second target audio information is obtained based on at least one feature of the second noise reduction block.

3. The method according to claim 1, wherein obtaining at least one speech block feature based on the at least one speech block includes: Determine at least two frames of data contained in the target speech block, each frame of data corresponding to the audio information of the target duration, and each frame of data containing features of the target dimension. The target speech block features are obtained based on the target dimension features contained in each frame of data in the target speech block.

4. The method according to claim 1, wherein identifying the features of the at least one speech block to obtain the scene to which the at least one speech block belongs includes: The target speech block features are used as input features to input the second processing model to obtain the probability that the target speech block belongs to at least two preset scenarios; Based on the probability that the target speech block belongs to the first preset scenario, the agreed selection condition is met, and the first preset scenario is selected as the scenario of the target speech block.

5. The method according to claim 1, wherein fusing the target block feature corresponding to the target speech block with the historical block feature corresponding to the historical speech block to obtain the fused block feature includes: The target block features corresponding to the target speech block and the historical block features corresponding to the historical speech block are fused based on the feature dimension of the frame data to obtain fused block features. Each frame data contains features of the target dimension.

6. The method according to claim 5, wherein the step of performing noise reduction processing on the target block features based on the fused block features to obtain the first noise-reduced block features includes: Weighted features are obtained based on the fusion block features; The weighted features are fused with the target block features to obtain secondary fused block features; The secondary fusion block features are input into the first processing model for noise reduction to obtain the first noise reduction block features.

7. The method according to claim 6, wherein obtaining the weighted features based on the fused block features comprises: The fusion block features are used as the first and second parameters of the third processing model, respectively. The target block features are used as the third parameter of the third processing model; The third processing model is controlled to obtain weighted features based on the first parameter, the second parameter, and the third parameter.

8. The method according to claim 6, wherein the weighted features are fused with the target block features to obtain secondary fused block features, comprising: The weighted features and the target block features are concatenated based on the feature dimensions of the frame data to obtain the secondary fused block features; or The weighted features are stacked with the target block features to obtain secondary fused block features.

9. An audio information processing apparatus, comprising: An acquisition module is used to acquire audio information, the audio information containing at least one speech block; The feature module is used to obtain at least one speech block feature based on the at least one speech block, wherein the speech block corresponds to the speech block feature; The recognition module is used to recognize the features of the at least one speech block and obtain the scene to which the at least one speech block belongs; The acquisition module is used to acquire historical voice blocks belonging to the same scene as the target voice block, wherein the generation time of the historical voice block is earlier than the generation time of the target voice block; The fusion module is used to fuse the target block features corresponding to the target speech block with the historical block features corresponding to the historical speech block to obtain fused block features, wherein the target speech block is one of the at least one speech block; The noise reduction module is used to control the first processing model to perform noise reduction processing on the target block features based on the fusion block features to obtain the first noise-reduced block features; The audio information module is used to obtain the first target audio information based on the features of the first noise reduction block.

10. An electronic device, comprising: Memory, processor; The memory stores the processing program; The processor is used to load and execute the processing program stored in the memory to implement the steps of the audio information processing method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Voice noise reduction method and device, equipment and storage medium

    CN113327626A

  • Speech recognition fast adaptive method assisted by audio scene classification

    CN114464182A