Data processing method and device, storage medium and electronic equipment

By acquiring audio feature data from the microphone array, determining the scene status, and performing scaling modulation and noise reduction processing, the problems of noise and echo in voice interaction are solved, achieving efficient audio data processing and improving voice quality.

CN120708638APending Publication Date: 2025-09-26MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510978353.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

The existing technology has interference noise and echo during the voice interaction process, which affects the user experience and results in poor voice quality improvement effect.

Method used

By obtaining the audio feature data of the audio data set collected by the microphone array, the scene state of each frame data is determined, and the initial audio data is scaled, modulated and denoised based on the scene state to generate a second audio data set, which is suitable for audio processing of microphone arrays of any formation.

Benefits of technology

It improves the versatility and accuracy of audio processing, enhances voice quality, adapts to audio data processing in various scenarios, and reduces interference from echo and noise.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708638A_ABST
    Figure CN120708638A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, a storage medium and electronic equipment, and the method comprises the steps: obtaining a first audio data set collected by a microphone array, obtaining the audio feature data of the first audio data set, enabling the audio data of the first audio data set to be in one-to-one correspondence with the microphone channels of the microphone array, and obtaining the audio feature data of the first audio data set; based on the audio feature data, scene states belonging to the same frame of data in each piece of audio data are determined, initial audio data are obtained, scaling modulation is carried out on data frames of the initial audio data based on each scene state to obtain initial modulation data, and the initial audio data are any audio data in a first audio data set; and modulating the first audio data set based on the initial modulation data to generate a first modulation data set, and performing noise reduction processing on the first modulation data set to obtain a second audio data set. By adopting the specification, the universality and accuracy of audio processing and the data quality of the second audio data set are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of electronic information technology, and in particular to a data processing method, device, storage medium, and electronic device. Background Art

[0002] With the continuous advancement of technology, voice interaction using electronic devices has become an essential means of communication in our daily lives. However, interference noise and echoes can occur during this interaction, impacting the user experience. Therefore, improving voice quality during voice interaction has become an urgent issue. Summary of the Invention

[0003] This specification provides a data processing method, device, storage medium and electronic device. The method can modulate a first audio data set based on the determined scene state of each frame of data and according to the scene state and any initial audio data, and then perform noise reduction processing on the modulated data that highlights the scene state, thereby achieving audio processing that satisfies the audio data of a microphone array of any formation, improving the versatility and accuracy of audio processing, and the data quality of the second audio data set.

[0004] In a first aspect, a data processing method is provided, the method comprising: Obtaining a first audio data set collected by a microphone array, and obtaining audio feature data of the first audio data set, wherein each audio data in the first audio data set corresponds one-to-one to each microphone channel in the microphone array; Based on the audio feature data, determining a scene state belonging to the same frame data in each of the audio data; Acquire initial audio data, and scale and modulate data frames of the initial audio data based on each of the scene states to obtain initial modulated data, where the initial audio data is any audio data in the first audio data set; The first audio data set is modulated based on the initial modulation data to generate a first modulation data set, and the first modulation data set is subjected to noise reduction processing to obtain a second audio data set.

[0005] Through the above technical solution, the audio feature data of the first audio data set collected by the microphone array is extracted, and according to the scene state determined by the audio feature data, the initial audio data in the first audio data set is modulated to generate initial modulation data, and then the first audio data set is modulated according to the initial modulation data to generate the first modulation data set, and then the first modulation data set is subjected to noise reduction processing to obtain the second audio data set, so that according to the determined scene state of each frame data, and after the first audio data set is modulated according to the scene state and any initial audio data, the data highlighting the scene state after modulation is subjected to noise reduction processing, thereby achieving audio processing for the audio data of the microphone array of any formation, improving the versatility and accuracy of audio processing, and the data quality of the second audio data set.

[0006] With reference to the first aspect, in some possible implementations, obtaining audio feature data of the first audio data set includes: Acquire first feature data of the first audio data set, where the first feature data is a feature that retains spatial information of the first audio data; Noise reduction processing is performed on the first feature data to generate audio feature data corresponding to the first audio data set.

[0007] Through the above technical solution, first feature data that retains spatial information is obtained to obtain a scenario that can adapt to multiple microphone arrays, and preliminary echo cancellation and noise reduction processing is performed on the features extracted from the first audio data set to reduce the interference of echoes and noise in subsequent data processing, thereby improving the versatility of audio data processing in multiple scenarios, as well as the echo cancellation and noise reduction processing effects of audio data.

[0008] With reference to the first aspect, in some possible implementations, obtaining first feature data of the first audio data set includes: performing dimension reorganization on the first audio data set based on audio parameters of the first audio data set to obtain first reorganized data, the audio parameters including frequency, time frame length, number of microphone channels, and channel coordinates of each of the microphone channels; Performing feature enhancement on the first reorganized data based on a feedforward network to obtain enhanced data; Performing dimension reorganization on data obtained by concatenating the first audio data set and the enhanced data to generate second reorganized data; performing mean pooling processing on the second reorganized data based on the microphone dimension of the first audio data set to obtain mean data; Obtain reference data corresponding to the first audio data set, perform concatenation and transpose convolution on the reference data and the mean data, and generate first feature data, wherein the reference data and the mean data have the same data dimension.

[0009] In conjunction with the first aspect, in some possible implementations, the dimensionally reorganizing the first audio data set based on the audio parameters of the first audio data set to obtain first reorganized data includes: Obtaining frequency, time frame length, and microphone channel from audio parameters of the first audio data set; Determine the time frame length as the batch dimension, the frequency bin as the embedding dimension, and the number of channels as the sequence dimension; Obtain the channel coordinates of each of the microphone channels, reorganize the first audio data set based on the frequency, the time frame length, the number of channels and the channel coordinates to obtain a spatial attention matrix of a preset dimension, and determine the spatial attention matrix as the first reorganized data.

[0010] Through the above technical solution, based on the first recombined data that retains spatial features, it is ensured that the subsequently obtained audio feature data retains the spatial information of the first audio data set, thereby meeting the subsequent data processing requirements of microphone arrays that can be adapted to various formations, thereby improving the versatility of data processing such as noise reduction on audio data.

[0011] With reference to the first aspect, in some possible implementations, performing noise reduction processing on the first feature data to generate audio feature data corresponding to the first audio data set includes: performing filtering processing on the first feature data based on a narrowband module according to each frequency value in a frequency bin of the first feature data to obtain first noise reduction data, wherein the frequency bin represents the frequency value included in the first feature data; performing filtering processing on the first noise reduction data based on the cross-band module to obtain second noise reduction data; Repeating the noise reduction processing based on the narrowband module and the cross-band module on the second noise reduction data, and determining data obtained after repeating the processing a preset number of times as audio feature data.

[0012] In conjunction with the first aspect, in some possible implementations, determining, based on the audio feature data, the scene state of each audio data belonging to the same frame data includes: Determining second feature data in the audio feature data, and identifying a scene state corresponding to the second feature data, where the second feature data is any frame of data in the audio feature data; Based on the scene states, the scene states of the audio data belonging to the same frame data are determined.

[0013] With reference to the first aspect, in some possible implementations, identifying the scene state corresponding to the second feature data includes: reducing the number of frequencies included in the second characteristic data to a preset number; determining a time domain feature of the second feature data based on a local path of the dual-path cycle module; determining a cross-band feature of the second feature data based on a global path of the dual-path loop module; Determine the scene state corresponding to the second feature data based on the time domain features and the cross-band features according to the fully connected layer module.

[0014] Through the above technical solution, the data characteristics of the second feature data are determined according to the dual-path loop module, and then the scene state of the second feature data is determined according to the fully connected layer, so as to realize the state of the scene state belonging to the same frame data, so as to improve the accuracy and versatility of subsequent noise reduction processing of the audio data.

[0015] In conjunction with the first aspect, in some possible implementations, obtaining initial audio data and scaling and modulating data frames of the initial audio data based on each of the scene states to obtain initial modulated data includes: Obtaining initial audio data in the first audio data set, and determining a data frame sequence in the initial audio data; Based on the data frame sequence, the initial audio data is scaled and modulated in turn according to each of the scene states to obtain initial modulation data corresponding to the initial audio data, where the initial modulation data indicates the scene state corresponding to each data frame of the initial audio data.

[0016] Through the above technical solution, the initial audio data is scaled and modulated according to the scene state to obtain initial modulated data that can indicate the scene state of each data frame, thereby realizing data processing of microphone arrays of any formation and improving the versatility of data noise reduction processing.

[0017] With reference to the first aspect, in certain possible implementations, modulating the first audio data set based on the initial modulation data to generate a first modulated data set, and performing noise reduction processing on the first modulated data set to obtain a second audio data set includes: modulating the first audio data set based on the initial modulation data to obtain a first modulated data set; Performing noise reduction processing on the first modulated data set based on a convolution module, a gated convolution module, and a dual-path loop module to obtain a second modulated data set; Based on the first audio data set and the second modulation data set, a second audio data set corresponding to the first audio data set is obtained.

[0018] Through the above technical solution, the first audio data set is modulated according to the initial modulation data, and then the modulated data is subjected to noise reduction processing to achieve noise reduction processing of the data collected by the microphone array of any formation, thereby improving the echo cancellation effect and noise reduction processing effect of the audio data, as well as the versatility of the noise reduction processing.

[0019] In conjunction with the first aspect, in some possible implementations, obtaining, based on the first audio data set and the second modulation data set, a second audio data set corresponding to the first audio data set includes: performing a bitwise multiplication process on the first audio data set and the second modulation data set to obtain a third modulation data set; Performing point-by-point addition processing on the first audio data set and the third modulation data set to obtain a second audio data set.

[0020] Through the above technical solution, point-by-point multiplication and point-by-point addition processing are performed on the first audio data set, thereby ensuring the data characteristics of the data, reducing the distortion caused by the data calculation, improving the accuracy of audio processing, and the data quality of the second audio data set.

[0021] In a second aspect, a data processing device is provided, the device comprising: a feature acquisition unit, configured to acquire a first audio data set collected by a microphone array, and acquire audio feature data of the first audio data set, wherein each audio data in the first audio data set corresponds one-to-one to each microphone channel in the microphone array; a scene determination unit, configured to determine, based on the audio feature data, a scene state belonging to the same frame data in each of the audio data; a modulation unit, configured to obtain initial audio data, and scale and modulate data frames of the initial audio data based on each of the scene states to obtain initial modulated data, wherein the initial audio data is any audio data in the first audio data set; The data noise reduction unit is configured to modulate the first audio data set based on the initial modulation data to generate a first modulated data set, and perform noise reduction processing on the first modulated data set to obtain a second audio data set.

[0022] In a third aspect, an embodiment of this specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the steps of the above method.

[0023] In a fourth aspect, an embodiment of this specification provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the above method.

[0024] In a fifth aspect, a computer program product is provided, which includes: computer program code, which, when running on a computer, enables the computer to execute the method in the above-mentioned first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions in this specification or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.

[0026] Figure 1 A system architecture diagram of a data processing method provided in an embodiment of this specification; Figure 2 A flowchart of a data processing method provided in an embodiment of this specification; Figure 3 A flowchart of a data processing method provided in an embodiment of this specification; Figure 4 This is a schematic diagram of an example of determining a scene state provided in an embodiment of this specification; Figure 5 This is a schematic diagram illustrating an example of a data noise reduction method provided in an embodiment of this specification; Figure 6 This is a schematic diagram of an example of generating a second audio data set provided in an embodiment of this specification; Figure 7 A schematic diagram of the structure of a data processing device provided in an embodiment of this specification; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0027] To make the features and advantages of this specification more obvious and easy to understand, the technical solutions in this specification are clearly and completely described below in conjunction with the drawings in this specification. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of this specification.

[0028] See Figure 1 , which is a system architecture diagram of a data processing method provided in the embodiment of this specification. Figure 1 As shown, the data processing method provided in the embodiments of this specification can be applied to electronic devices to implement the process of echo cancellation and noise reduction of audio data. The system architecture provided in the embodiments of this specification mainly includes an electronic device 1, a microphone array 20 and a second audio data set 30. The electronic device 1 can be a device with data processing functions, such as a personal computer, a server, etc., or a data processing module with a pre-trained model built in. The microphone array 20 can be an audio acquisition device composed of multiple microphones that implements spatial sampling and signal processing through a specific arrangement. The second audio data set 30 can be an audio data set generated by the electronic device 10 after performing echo cancellation and noise reduction processing on the first audio data set collected by the microphone array 20.

[0029] In the related art, in the process of improving voice quality, if the method adopted is a cascade method, and the sound cancellation and noise reduction are performed according to the echo cancellation module and the array algorithm, there will be a mismatch problem between different modules, resulting in poor improvement of voice quality.

[0030] In an embodiment of the present specification, the electronic device 1 obtains a first audio data set collected by the microphone array 20, obtains audio feature data of the first audio data set, each audio data in the first audio data set corresponds one-to-one to each microphone channel in the microphone array 20, and determines the scene state of each audio data belonging to the same frame data based on the audio feature data, obtains initial audio data, and scales and modulates the data frame of the initial audio data based on each scene state to obtain initial modulated data, the initial audio data is any audio data in the first audio data set, and the first audio data set is modulated based on the initial modulated data to generate a first modulated data set, and the first modulated data set is subjected to noise reduction processing to obtain a second audio data set 30, thereby modulating the first audio data set according to the determined scene state of each frame data and according to the scene state and any initial audio data, and then performing noise reduction processing on the data highlighting the scene state after modulation, thereby achieving audio processing that satisfies the audio data of a microphone array of any formation, improving the versatility and accuracy of audio processing, and the data quality of the second audio data set.

[0031] based on Figure 1 The system architecture shown below will be combined with Figure 2 , the data processing method provided in the embodiments of this specification is introduced in detail.

[0032] See Figure 2 , which is a flow chart of a data processing method provided in the embodiment of this specification. Figure 2 As shown, the method is applied to an electronic device, and the method may include the following steps S101 to S104.

[0033] S101, obtaining a first audio data set collected by a microphone array, and obtaining audio feature data of the first audio data set; In one embodiment, a first audio data set collected by a microphone array is obtained. The microphone array may be an audio collection device composed of multiple microphones that implements spatial sampling and signal processing through a specific arrangement. The first audio data set may be a data set obtained from audio data collected by each microphone in the microphone array. Each audio data in the first audio data set corresponds one-to-one to each microphone channel in the microphone array. The microphone channels correspond to paths for data transmission corresponding to each microphone in the microphone array. Feature extraction is performed on the first audio data set to obtain audio feature data of the first audio data set. The audio feature data may be data that retains features such as spatial information of the first audio data set.

[0034] S102, determining a scene state belonging to the same frame data in each audio data based on the audio feature data; In one embodiment, the scene state of each frame in the audio feature data is identified to determine the scene state belonging to the same frame of audio data. The scene state can be used to indicate the speech scene of the data in the frame, such as double talk (including echo and human voice), close talk (someone is speaking), far talk (only echo), and no talk (no voice or echo).

[0035] It should be noted that, since the microphone array collects data at the same time, the data collected by each microphone in the array collects data at the same time. Therefore, in the same frame, the scene state corresponding to the data collected by each microphone is the same.

[0036] S103, obtaining initial audio data, and scaling and modulating data frames of the initial audio data based on each scene state to obtain initial modulated data; In one embodiment, initial audio data is obtained from a first audio data set. The initial audio data is any audio data in the first audio data set. Based on the determined scene state of each frame of data, scaling modulation is performed on each data frame in the initial audio data to obtain initial modulated data. The data frame may be data corresponding to each frame in the initial audio data, and the length of the data frame may be set according to actual conditions. Scaling modulation may be performed by amplifying or reducing the data frame according to the scene state of the data frame, so as to reflect the scene state of the data frame by amplification or reduction. The initial modulated data may be audio data reflecting the scene state of each data frame.

[0037] S104, modulating the first audio data set based on the initial modulation data to generate a first modulated data set, and performing noise reduction processing on the first modulated data set to obtain a second audio data set; In one embodiment, a first audio data set is modulated based on initial modulation data to generate a first modulated data set. The first modulated data set may be a set consisting of data generated by scaling and modulating the initial modulation data. After noise reduction processing is performed on the first modulated data set, echo and noise in the first modulated data set are eliminated to obtain a second audio data set. The second audio data set may be a set consisting of data that has undergone echo cancellation and noise reduction.

[0038] Specifically, the method of performing noise reduction processing on the first modulated data set can be: combining data processing modules such as a convolution module (ConvBlock), a gated convolution module (GT-Conv2D) and a dual-path recurrent module (Dual-Path Recurrent, DPR) to perform echo cancellation and noise reduction on the first modulated data set.

[0039] It should be noted that since the data collected by each microphone channel has the same scene state in the same frame, the first audio data set can be modulated according to the initial modulation data obtained after modulating any audio data, thereby realizing data processing of data collected by a microphone array with any number of microphones and any formation without affecting the data processing quality, thereby improving the versatility and accuracy of audio data processing.

[0040] In an embodiment of the present specification, audio feature data of a first audio data set collected by a microphone array is extracted, and according to the scene state determined by the audio feature data, initial audio data in the first audio data set is modulated to generate initial modulation data, and then the first audio data set is modulated according to the initial modulation data to generate a first modulation data set, and then the first modulation data set is subjected to noise reduction processing to obtain a second audio data set. Thus, according to the determined scene state of each frame data, and after the first audio data set is modulated according to the scene state and any initial audio data, noise reduction processing is performed on the modulated data that highlights the scene state, thereby achieving audio processing for audio data of a microphone array of any formation, improving the versatility and accuracy of audio processing, and the data quality of the second audio data set.

[0041] See Figure 3 , which is a flow chart of a data processing method provided in the embodiment of this specification. Figure 3 As shown, the method is applied to an electronic device, and the method may include the following steps S201 to S208.

[0042] S201, obtaining a first audio data set collected by a microphone array, and obtaining audio feature data of the first audio data set; In one embodiment, a first audio data set collected by a microphone array is obtained. The microphone array may be an audio collection device composed of multiple microphones that implements spatial sampling and signal processing through a specific arrangement. The first audio data set may be a data set obtained from audio data collected by each microphone in the microphone array. Each audio data in the first audio data set corresponds one-to-one to each microphone channel in the microphone array. The microphone channels correspond to paths for data transmission corresponding to each microphone in the microphone array. Feature extraction is performed on the first audio data set to obtain audio feature data of the first audio data set. The audio feature data may be data that retains features such as spatial information of the first audio data set.

[0043] Furthermore, in order to improve the accuracy of subsequent data processing, echo cancellation and noise reduction processing can be performed in the process of obtaining audio data features. The method of obtaining audio feature data can be: obtaining the first feature data of the first audio data set, performing noise reduction processing on the first feature data, and generating audio feature data corresponding to the first audio data set.

[0044] The method for obtaining the first feature data may be: based on the audio parameters of the first audio data set, dimensionality reorganization of the first audio data set is performed to obtain the first reorganized data. The dimensionality reorganization may be to perform data conversion on the first audio data set according to audio parameters such as the time frame length, frequency bin and number of channels of the first audio data set. The time frame length may be the duration used to describe the first audio data set. The time frame lengths of the audio data of the first audio data set are the same, so the time frame length of any audio data can be determined as the time frame length of the first audio data set. The frequency bin is the frequency value included in each audio data in the first audio data set. The number of channels may be the number of microphone channels included in the first audio data set.

[0045] Specifically, based on the audio parameters of the first audio data set, the first reconstructed data can be obtained by obtaining the frequency, time frame length, and microphone channel from the audio parameters of the first audio data set, determining the time frame length as the batch dimension, determining the frequency bin as the embedding dimension, determining the number of channels (2Q, where Q is the number of channels) as the sequence dimension, obtaining the channel coordinates of each microphone channel, reconstructing the first audio data set based on the frequency, time frame length, number of channels, and channel coordinates of each microphone channel, generating a spatial attention matrix of dimension "2Q*2Q", and determining the spatial attention matrix as the first reconstructed data. It should be noted that because the first reconstructed data is generated based on the channel data and the reconstructed microphone channels, the first reconstructed data can adaptively capture the pairwise interactions between microphones and determine the spatial features of the first audio data set independently of the array geometry of the microphone array. This allows the first reconstructed data, which retains the spatial features, to be used for subsequent data processing on microphone arrays of various configurations, thereby improving the versatility of data processing such as noise reduction on audio data.

[0046] The first reconstructed data is subjected to feature enhancement based on a feedforward network to obtain enhanced data. Feature enhancement can be performed on the first reconstructed data using a feedforward network with layer normalization, thereby introducing nonlinear transformation capabilities to the first reconstructed data. The concatenated data of the first audio data set and the enhanced data are dimensionally reconstructed to generate second reconstructed data. The dimensions of the concatenated first audio data set and the enhanced data are adjusted to the same dimension. The second reconstructed data can be obtained by reconstructing the concatenated data using a learnable upsampling convolution module (US-Cinv2D). The upsampling convolution module can include convolution layers, layer normalization, and parameterized rectified linear unit (PReLU) activation parameters. Mean pooling is performed on the second reconstructed data based on the microphone dimensions of the first audio data set to obtain mean data. Reference data corresponding to the first audio data set is obtained, and the reference data and the mean data are concatenated and transposed convolved to generate first feature data. The reference data and the mean data have the same dimension. Reference data can be used to assist in identifying and separating noise. By providing the algorithm with a "characteristic template of noise" or a "reference basis for pure signals," it can remove noise more accurately.

[0047] The method for generating audio feature data based on the first feature data may include filtering the first feature data using a narrowband module based on each frequency value in a frequency bin of the first feature data to obtain first noise reduction data. The frequency bin represents the frequency value included in the first feature data. The first noise reduction data is filtered using a cross-band module to obtain second noise reduction data. The method for filtering the first feature data using the narrowband module may include filtering the data at each frequency value in the first feature data separately. The method for filtering the first noise reduction data using the cross-band module may include filtering the data at each frequency value of each first feature data simultaneously. The noise reduction processing of the second noise reduction data using the narrowband module and the cross-band module is repeated a preset number of times, and the data obtained after the repetition is repeated a preset number of times is determined as the audio feature data. The specific value of the preset number of times can be set according to actual circumstances. It should be noted that during the repeated execution, the filtering and noise reduction processing is performed in a manner that the narrowband module and the cross-band module are executed in turn.

[0048] It should be noted that the noise reduction processing performed on the first feature data is a preliminary noise reduction performed according to the first-stage model, and then combined with the subsequent post-processing module to perform secondary noise reduction processing, thereby performing noise reduction through dual-module noise reduction, adapting to a variety of scenarios, and improving the versatility of the model for noise reduction of audio data. For example, when there is a new special scenario that needs to be adapted to achieve noise reduction of audio data, preliminary noise reduction can be performed directly according to the first-stage model without the need to retrain the first-stage model. Only fine-tuning of the post-processing module based on a small amount of data is required to complete the adaptation for the new special scenario, thereby improving the efficiency of model training and the versatility of the data processing model.

[0049] S202, determining second feature data in the audio feature data, and identifying a scene state corresponding to the second feature data; In one embodiment, second feature data in the audio feature data is determined, where the second feature data is any frame of data in the audio feature data. A scene state corresponding to the second feature data is identified. The scene state may be used to indicate the speech scene in which the data in the frame exists, such as double talk (including echo and human voice), close talk (someone is speaking), far talk (only echo), and no talk (no voice and no echo).

[0050] Specifically, the method for identifying the scene state corresponding to the second feature data can be: reducing the number of frequencies included in the second feature data to a preset number. The specific value of the preset number can be set according to the actual situation, for example, it can be one-fourth of the original number, etc. Based on the local path of the dual-path loop module, the time domain characteristics of the second feature data are determined. Based on the global path of the dual-path loop module, the cross-band characteristics of the second feature data are determined. The time domain characteristics can be time dependencies within the frequency band characterizing the second feature data, such as the time-varying characteristics of ambient noise, etc. The cross-band characteristics can be cross-band correlations characterizing the second feature data, such as the spatial distribution of the acoustic structure of the scene, etc.

[0051] The scene state corresponding to the second feature data is determined based on the time domain features and cross-band features according to the fully connected layer module. Before determining the scene state according to the fully connected layer module, a convolution module is used to scan the time domain features and cross-band features in the frequency domain and time dimensions. The scanned features are then subjected to maximum pooling dimensionality reduction to reduce the temporal resolution and compress the frequency domain. After that, the features are mapped to a scene probability distribution through the fully connected layer, and the scene state of the second feature data is determined based on the scene probability distribution.

[0052] For example, if the output dimension of the fully connected layer is consistent with the scene category, including double talk, near talk, far talk, and no talk, and the fully connected layer output vector is [0.7, 0.1, 0.1, 0.1], then the probability of determining that the scene state of the second feature data is double talk is 70%. If the probability is greater than a preset threshold (e.g., 50%), then the scene state of the second feature data is determined to be double talk. The specific value of the preset threshold can be set according to actual circumstances.

[0053] S203, based on each scene state, determining the scene state of each audio data belonging to the same frame data; In one embodiment, after determining the scene state of each second feature data, since the microphone array collects data at the same time, the scene state corresponding to the data collected by each microphone in the same frame is the same. Based on the scene state of each second feature data, the scene state of each audio data in the first audio data set belonging to the same frame data can be determined.

[0054] For example, Figure 4 As shown, Figure 4 The audio data includes five frames of data, and the scene state corresponding to each frame of data is determined, and the scene state of each frame of data is determined as the scene state belonging to the same frame of data in each audio data.

[0055] S204, obtaining initial audio data in the first audio data set, and determining a data frame sequence in the initial audio data; S205, scaling and modulating the initial audio data according to each scene state based on the data frame sequence to obtain initial modulated data corresponding to the initial audio data; In one embodiment, initial audio data from a first audio data set is obtained. The initial audio data may be any audio data from the first audio data set. A data frame order of each data frame in the initial audio data is determined. The data frame order may represent a sequential order of data frames in the initial audio data. Based on the data frame order, the initial audio data is scaled and modulated according to the scene state corresponding to each data frame to obtain initial modulated data corresponding to the initial audio data. The initial modulated data may indicate the scene state corresponding to each data frame of the initial audio data.

[0056] It can be understood that, since the scene status of each frame of audio data has been determined, each frame of initial audio data can be scaled and modulated according to the scene status of each frame.

[0057] Specifically, the scaling modulation method for the initial audio data may be to use the FiLM (Feature-wise Linear Modulation) formula shown in formula (1), Formula (1); in, Can be the initial modulation data, It can be the initial audio data, L can be the scene state, and ⊙ can be the element-by-element multiplication.

[0058] The modulation ratio of the scaling modulation of the initial audio data can be set according to actual conditions. For example, the data in a double-talking scenario can be reduced to one-half, and the data in a close-talking scenario can be amplified by one-fold.

[0059] S206, modulating the first audio data set based on the initial modulation data to obtain a first modulated data set; In one embodiment, based on each frame data in the initial modulation data, the data belonging to the same frame in the first audio data set are modulated respectively to obtain a first modulation data set that can highlight the scene state of each frame data.

[0060] Specifically, the first audio data set may be modulated according to the initial modulation data by concatenating the initial modulation data with the first audio data set, and adjusting data characteristics of each audio data in the first audio data set based on the concatenation, thereby obtaining the first modulated data set. The specific concatenation method may be set based on actual needs.

[0061] S207, performing noise reduction processing on the first modulated data set based on the convolution module, the gated convolution module, and the dual-path loop module to obtain a second modulated data set; In one embodiment, noise reduction processing is performed on the first modulated data set based on a convolution module, a gated convolution module, and a dual-path recurrent module according to a preset module order to obtain a second modulated data set. The module order may indicate the order in which the convolution module, the gated convolution module, and the dual-path recurrent module are executed. The specific order can be set according to actual circumstances.

[0062] For example, Figure 5 As shown, Figure 5The first modulated data set is input into the first convolution module (ConvBlock-1) and operated twice. The operation result is input into the first gated convolution module (GT-Conv2D-1) and the second convolution module (Conv Block-2). The first gated convolution module is operated three times, and the operation result is input into the dual-path loop module (DPR) and the second gated convolution module (GT-Conv2D-2). After the dual-path loop module is transported twice, the operation result is input into the second module convolution module (GT-Conv2D-2). The second gated convolution module performs three operations based on the operation results input by the first gated convolution module and the dual-path loop module, and inputs the operation result into the second convolution module (Conv Block-2). The second convolution module (Conv Block-2) performs two operations based on the input of the first convolution module and the second gated convolution module to obtain the second modulated data set. It can be understood that Figure 5 The "X2" in the figure indicates that the module performs two consecutive operations, and the "X3" indicates that the module performs three consecutive operations. The number of operations of each module can be set according to actual conditions.

[0063] Specifically, such as Figure 6 As shown, Figure 6 The following is a flowchart for generating a second audio data set based on a first audio data set. A dynamic spatial encoder performs feature extraction on the first audio data set to obtain mean data that retains spatial features. Reference data corresponding to the first audio data set is then obtained. The reference data is re-dimensionalized to the same dimension as the mean data. The mean data and reference data are then concatenated and denoised to obtain audio feature data. Scene recognition is performed based on the audio feature data to obtain the scene status of each frame of data in the audio feature data. Initial audio data from the first audio data set is obtained, and the initial audio data, scene status, and decoded audio data set are input into a post-processing module. The post-processing module performs scaling modulation on the initial audio data, scene status, and decoded audio data set to obtain a first modulated data set. The post-processing module then performs denoising on the first modulated data set to obtain a second audio data set.

[0064] S208, obtaining a second audio data set corresponding to the first audio data set based on the first audio data set and the second modulation data set; In one embodiment, in order to ensure the data characteristics of the data and reduce the distortion caused by the data after calculation, after obtaining the second adjusted data set, the first audio data set and the second modulation data set are subjected to point-by-point multiplication processing to obtain a third modulation data set. The third modulation data set can be data obtained according to the point-by-point multiplication calculation method, so as to correct the gain distortion of the second modulation data set through the first audio data set and adjust the amplitude ratio of the data.

[0065] The first audio data set and the third modulated data set are then subjected to a point-by-point addition process to obtain a second audio data set. The second audio data set may be a set of data that has undergone echo cancellation and noise reduction. Using the point-by-point addition method, the offset distortion of the second modulated data set is corrected using the first audio data set to eliminate the constant offset of the data.

[0066] In an embodiment of the present specification, by extracting audio feature data of a first audio data set collected by a microphone array, modulating the initial audio data in the first audio data set according to the scene state determined by the audio feature data to generate initial modulation data, and then modulating the first audio data set according to the initial modulation data to generate the first modulation data set, the first modulation data set is subjected to noise reduction processing to obtain a second audio data set, thereby according to the determined scene state of each frame of data, and after modulating the first audio data set according to the scene state and any initial audio data, the data highlighting the scene state after modulation is subjected to noise reduction processing, thereby achieving audio processing that satisfies the requirements of audio data of microphone arrays of any formation, improving the versatility and accuracy of audio processing, and the data quality of the second audio data set. In addition, based on the first recombined data that retains spatial features, the subsequent data processing that can be adapted to microphone arrays of various formations is met, thereby improving the versatility of data processing such as noise reduction on audio data. Furthermore, preliminary echo cancellation and noise reduction processing is performed on the features extracted from the first audio data set, and then combined with subsequent noise reduction processing on the modulated data, the voice data that was over-eliminated in the preliminary noise reduction processing can be repaired, or the echo that was not eliminated cleanly can be eliminated, thereby improving the echo cancellation and noise reduction processing effects of the audio data.

[0067] based on Figure 1 The system architecture shown below will be combined with Figure 7 , the data processing device provided in the embodiment of this specification is introduced in detail. It should be noted that, Figure 7 The data processing device is used to execute the embodiment of this specification Figures 2 to 6 For the convenience of explanation, only the part related to the embodiment of this specification is shown. For the specific technical details not disclosed, please refer to the embodiment of this specification. Figures 2 to 6 The embodiment shown.

[0068] See Figure 7 , is a structural diagram of a data processing device provided in the embodiment of this specification. Figure 7 As shown, the data processing device 1 of the embodiment of this specification may include: a feature acquisition unit 11, a scene determination unit 12, a modulation unit 13 and a data noise reduction unit 14.

[0069] a feature acquisition unit 11, configured to acquire a first audio data set collected by a microphone array, and acquire audio feature data of the first audio data set, wherein each audio data in the first audio data set corresponds one-to-one to each microphone channel in the microphone array; A scene determination unit 12 is configured to determine a scene state belonging to the same frame data in each of the audio data based on the audio feature data; a modulation unit 13, configured to obtain initial audio data, and scale and modulate data frames of the initial audio data based on each of the scene states to obtain initial modulated data, wherein the initial audio data is any audio data in the first audio data set; The data noise reduction unit 14 is configured to modulate the first audio data set based on the initial modulation data to generate a first modulated data set, and perform noise reduction processing on the first modulated data set to obtain a second audio data set.

[0070] Optionally, the feature acquisition unit 11 is further configured to: Acquire first feature data of the first audio data set, where the first feature data is a feature that retains spatial information of the first audio data; Noise reduction processing is performed on the first feature data to generate audio feature data corresponding to the first audio data set.

[0071] Optionally, the feature acquisition unit 11 is further configured to: performing dimension reorganization on the first audio data set based on audio parameters of the first audio data set to obtain first reorganized data, the audio parameters including frequency, time frame length, number of microphone channels, and channel coordinates of each of the microphone channels; Performing feature enhancement on the first reorganized data based on a feedforward network to obtain enhanced data; Performing dimension reorganization on data obtained by concatenating the first audio data set and the enhanced data to generate second reorganized data; performing mean pooling processing on the second reorganized data based on the microphone dimension of the first audio data set to obtain mean data; Obtain reference data corresponding to the first audio data set, perform concatenation and transpose convolution on the reference data and the mean data, and generate first feature data, wherein the reference data and the mean data have the same data dimension.

[0072] Optionally, the feature acquisition unit 11 is further configured to: Obtaining frequency, time frame length, and microphone channel from audio parameters of the first audio data set; Determine the time frame length as the batch dimension, the frequency bin as the embedding dimension, and the number of channels as the sequence dimension; Obtain the channel coordinates of each of the microphone channels, reorganize the first audio data set based on the frequency, the time frame length, the number of channels and the channel coordinates to obtain a spatial attention matrix of a preset dimension, and determine the spatial attention matrix as the first reorganized data.

[0073] Optionally, the feature acquisition unit 11 is further configured to: performing filtering processing on the first feature data based on a narrowband module according to each frequency value in a frequency bin of the first feature data to obtain first noise reduction data, wherein the frequency bin represents the frequency value included in the first feature data; performing filtering processing on the first noise reduction data based on the cross-band module to obtain second noise reduction data; Repeating the noise reduction processing based on the narrowband module and the cross-band module on the second noise reduction data, and determining data obtained after repeating the processing a preset number of times as audio feature data.

[0074] Optionally, the scene determination unit 12 is further configured to: Determining second feature data in the audio feature data, and identifying a scene state corresponding to the second feature data, where the second feature data is any frame of data in the audio feature data; Based on the scene states, the scene states of the audio data belonging to the same frame data are determined.

[0075] Optionally, the scene determination unit 12 is further configured to: reducing the number of frequencies included in the second characteristic data to a preset number; determining a time domain feature of the second feature data based on a local path of the dual-path cycle module; determining a cross-band feature of the second feature data based on a global path of the dual-path loop module; Determine a scene state library corresponding to the second feature data based on the time domain features and the cross-band features according to the fully connected layer module.

[0076] Optionally, the modulation unit 13 is further configured to: Obtaining initial audio data in the first audio data set, and determining a data frame sequence in the initial audio data; Based on the data frame sequence, the initial audio data is scaled and modulated in turn according to each of the scene states to obtain initial modulation data corresponding to the initial audio data, where the initial modulation data indicates the scene state corresponding to each data frame of the initial audio data.

[0077] Optionally, the data noise reduction unit 14 is further configured to: modulating the first audio data set based on the initial modulation data to obtain a first modulated data set; Performing noise reduction processing on the first modulated data set based on a convolution module, a gated convolution module, and a dual-path loop module to obtain a second modulated data set; Based on the first audio data set and the second modulation data set, a second audio data set corresponding to the first audio data set is obtained.

[0078] Optionally, the data noise reduction unit 14 is further configured to: performing a bitwise multiplication process on the first audio data set and the second modulation data set to obtain a third modulation data set; Performing point-by-point addition processing on the first audio data set and the third modulation data set to obtain a second audio data set.

[0079] In an embodiment of the present specification, by extracting audio feature data of a first audio data set collected by a microphone array, modulating the initial audio data in the first audio data set according to the scene state determined by the audio feature data to generate initial modulation data, and then modulating the first audio data set according to the initial modulation data to generate the first modulation data set, the first modulation data set is subjected to noise reduction processing to obtain a second audio data set, thereby according to the determined scene state of each frame of data, and after modulating the first audio data set according to the scene state and any initial audio data, the data highlighting the scene state after modulation is subjected to noise reduction processing, thereby achieving audio processing that satisfies the requirements of audio data of microphone arrays of any formation, improving the versatility and accuracy of audio processing, and the data quality of the second audio data set. In addition, based on the first recombined data that retains spatial features, the subsequent data processing that can be adapted to microphone arrays of various formations is met, thereby improving the versatility of data processing such as noise reduction on audio data. Furthermore, preliminary echo cancellation and noise reduction processing is performed on the features extracted from the first audio data set, and then combined with subsequent noise reduction processing on the modulated data, the voice data that was over-eliminated in the preliminary noise reduction processing can be repaired, or the echo that was not eliminated cleanly can be eliminated, thereby improving the echo cancellation and noise reduction processing effects of the audio data.

[0080] The embodiment of this specification also provides a computer storage medium that can store multiple program instructions, which are suitable for being loaded and executed by a processor as described above. Figures 1-6 The method steps of the embodiment shown, the specific execution process can be found in Figures 1-6 The detailed description of the illustrated embodiment will not be repeated here.

[0081] The embodiment of this specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded and executed by the processor as described above. Figures 1-6 The method steps of the embodiment shown, the specific execution process can be found in Figures 1-6 The detailed description of the illustrated embodiment will not be repeated here.

[0082] See Figure 8 , is a schematic diagram of the structure of an electronic device provided in the embodiment of this specification. Figure 8 As shown, the electronic device 1000 may include: at least one processor 1001, such as a CPU, at least one network interface 1004, an input / output interface 1003, a memory 1005, and at least one communication bus 1002. The communication bus 1002 is used to implement connection and communication between these components. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. Figure 8 As shown, the memory 1004 as a computer storage medium may include an operating system, an input / output interface module, and a data processing application.

[0083] exist Figure 8 In the electronic device 1000 shown, the input / output interface 1003 is mainly used to provide an input interface for the user and obtain data input by the user.

[0084] In one embodiment, the processor 1001 may be configured to call a data processing application stored in the memory 1004 and specifically perform the following operations: Obtaining a first audio data set collected by a microphone array, and obtaining audio feature data of the first audio data set, wherein each audio data in the first audio data set corresponds one-to-one to each microphone channel in the microphone array; Based on the audio feature data, determining a scene state belonging to the same frame data in each of the audio data; Acquire initial audio data, and scale and modulate data frames of the initial audio data based on each of the scene states to obtain initial modulated data, where the initial audio data is any audio data in the first audio data set; The first audio data set is modulated based on the initial modulation data to generate a first modulation data set, and the first modulation data set is subjected to noise reduction processing to obtain a second audio data set.

[0085] Optionally, when executing the acquisition of the audio feature data of the first audio data set, the processor 1001 specifically performs the following operations: Acquire first feature data of the first audio data set, where the first feature data is a feature that retains spatial information of the first audio data; Noise reduction processing is performed on the first feature data to generate audio feature data corresponding to the first audio data set.

[0086] Optionally, when executing the process of obtaining the first feature data of the first audio data set, the processor 1001 specifically performs the following operations: performing dimension reorganization on the first audio data set based on audio parameters of the first audio data set to obtain first reorganized data, the audio parameters including frequency, time frame length, number of microphone channels, and channel coordinates of each of the microphone channels; Performing feature enhancement on the first reorganized data based on a feedforward network to obtain enhanced data; Performing dimension reorganization on data obtained by concatenating the first audio data set and the enhanced data to generate second reorganized data; performing mean pooling processing on the second reorganized data based on the microphone dimension of the first audio data set to obtain mean data; Obtain reference data corresponding to the first audio data set, perform concatenation and transpose convolution on the reference data and the mean data, and generate first feature data, wherein the reference data and the mean data have the same data dimension.

[0087] Optionally, when the processor 1001 performs dimension reorganization on the first audio data set based on the audio parameters of the first audio data set to obtain first reorganized data, the processor 1001 specifically performs the following operations: Obtaining frequency, time frame length, and microphone channel from audio parameters of the first audio data set; Determine the time frame length as the batch dimension, the frequency bin as the embedding dimension, and the number of channels as the sequence dimension; Obtain the channel coordinates of each of the microphone channels, reorganize the first audio data set based on the frequency, the time frame length, the number of channels and the channel coordinates to obtain a spatial attention matrix of a preset dimension, and determine the spatial attention matrix as the first reorganized data.

[0088] Optionally, when performing noise reduction processing on the first feature data to generate audio feature data corresponding to the first audio data set, the processor 1001 specifically performs the following operations: performing filtering processing on the first feature data based on a narrowband module according to each frequency value in a frequency bin of the first feature data to obtain first noise reduction data, wherein the frequency bin represents the frequency value included in the first feature data; performing filtering processing on the first noise reduction data based on the cross-band module to obtain second noise reduction data; Repeating the noise reduction processing based on the narrowband module and the cross-band module on the second noise reduction data, and determining data obtained after repeating the processing a preset number of times as audio feature data.

[0089] Optionally, when the processor 1001 determines the scene state of each piece of audio data belonging to the same frame data based on the audio feature data, it specifically performs the following operations: Determining second feature data in the audio feature data, and identifying a scene state corresponding to the second feature data, where the second feature data is any frame of data in the audio feature data; Based on the scene states, the scene states of the audio data belonging to the same frame data are determined.

[0090] Optionally, when identifying the scene state corresponding to the second feature data, the processor 1001 specifically performs the following operations: reducing the number of frequencies included in the second characteristic data to a preset number; determining a time domain feature of the second feature data based on a local path of the dual-path cycle module; determining a cross-band feature of the second feature data based on a global path of the dual-path loop module; Determine the scene state corresponding to the second feature data based on the time domain features and the cross-band features according to the fully connected layer module.

[0091] Optionally, when the processor 1001 executes the steps of acquiring the initial audio data and scaling and modulating the data frames of the initial audio data based on each of the scene states to obtain the initial modulated data, the processor 1001 specifically performs the following operations: Obtaining initial audio data in the first audio data set, and determining a data frame sequence in the initial audio data; Based on the data frame sequence, the initial audio data is scaled and modulated in turn according to each of the scene states to obtain initial modulation data corresponding to the initial audio data, where the initial modulation data indicates the scene state corresponding to each data frame of the initial audio data.

[0092] Optionally, when the processor 1001 modulates the first audio data set based on the initial modulation data to generate a first modulated data set, and performs noise reduction processing on the first modulated data set to obtain a second audio data set, the processor 1001 specifically performs the following operations: modulating the first audio data set based on the initial modulation data to obtain a first modulated data set; Performing noise reduction processing on the first modulated data set based on a convolution module, a gated convolution module, and a dual-path loop module to obtain a second modulated data set; Based on the first audio data set and the second modulation data set, a second audio data set corresponding to the first audio data set is obtained.

[0093] Optionally, when executing the first audio data set and the second modulation data set to obtain a second audio data set corresponding to the first audio data set, the processor 1001 specifically performs the following operations: performing a bitwise multiplication process on the first audio data set and the second modulation data set to obtain a third modulation data set; Performing point-by-point addition processing on the first audio data set and the third modulation data set to obtain a second audio data set.

[0094] In an embodiment of the present specification, by extracting audio feature data of a first audio data set collected by a microphone array, modulating the initial audio data in the first audio data set according to the scene state determined by the audio feature data to generate initial modulation data, and then modulating the first audio data set according to the initial modulation data to generate the first modulation data set, the first modulation data set is subjected to noise reduction processing to obtain a second audio data set, thereby according to the determined scene state of each frame of data, and after modulating the first audio data set according to the scene state and any initial audio data, the data highlighting the scene state after modulation is subjected to noise reduction processing, thereby achieving audio processing that satisfies the requirements of audio data of microphone arrays of any formation, improving the versatility and accuracy of audio processing, and the data quality of the second audio data set. In addition, based on the first recombined data that retains spatial features, the subsequent data processing that can be adapted to microphone arrays of various formations is met, thereby improving the versatility of data processing such as noise reduction on audio data. Furthermore, preliminary echo cancellation and noise reduction processing is performed on the features extracted from the first audio data set, and then combined with subsequent noise reduction processing on the modulated data, the voice data that was over-eliminated in the preliminary noise reduction processing can be repaired, or the echo that was not eliminated cleanly can be eliminated, thereby improving the echo cancellation and noise reduction processing effects of the audio data.

[0095] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0096] The above disclosure is only a preferred embodiment of this specification, and certainly cannot be used to limit the scope of rights of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.

Claims

1. A data processing method, characterized in that: The method comprises: Obtaining a first audio data set collected by a microphone array, and obtaining audio feature data of the first audio data set, wherein each audio data in the first audio data set corresponds one-to-one to each microphone channel in the microphone array; Based on the audio feature data, determining a scene state belonging to the same frame data in each of the audio data; Acquire initial audio data, and scale and modulate data frames of the initial audio data based on each of the scene states to obtain initial modulated data, where the initial audio data is any audio data in the first audio data set; The first audio data set is modulated based on the initial modulation data to generate a first modulation data set, and the first modulation data set is subjected to noise reduction processing to obtain a second audio data set.

2. The method according to claim 1, characterized in that The obtaining of audio feature data of the first audio data set includes: Acquire first feature data of the first audio data set, where the first feature data is a feature that retains spatial information of the first audio data; Noise reduction processing is performed on the first feature data to generate audio feature data corresponding to the first audio data set.

3. The method according to claim 2, characterized in that The obtaining of first feature data of the first audio data set includes: performing dimension reorganization on the first audio data set based on audio parameters of the first audio data set to obtain first reorganized data, the audio parameters including frequency, time frame length, number of microphone channels, and channel coordinates of each of the microphone channels; Performing feature enhancement on the first reorganized data based on a feedforward network to obtain enhanced data; Performing dimension reorganization on data obtained by concatenating the first audio data set and the enhanced data to generate second reorganized data; performing mean pooling processing on the second reorganized data based on the microphone dimension of the first audio data set to obtain mean data; Obtain reference data corresponding to the first audio data set, perform concatenation and transpose convolution on the reference data and the mean data, and generate first feature data, wherein the reference data and the mean data have the same data dimension.

4. The method according to claim 3, characterized in that The step of performing dimension reorganization on the first audio data set based on the audio parameters of the first audio data set to obtain first reorganized data includes: Obtaining frequency, time frame length, and microphone channel from audio parameters of the first audio data set; Determine the time frame length as the batch dimension, the frequency bin as the embedding dimension, and the number of channels as the sequence dimension; Obtain the channel coordinates of each of the microphone channels, reorganize the first audio data set based on the frequency, the time frame length, the number of channels and the channel coordinates to obtain a spatial attention matrix of a preset dimension, and determine the spatial attention matrix as the first reorganized data.

5. The method according to claim 2, characterized in that The performing noise reduction processing on the first feature data to generate audio feature data corresponding to the first audio data set includes: performing filtering processing on the first feature data based on a narrowband module according to each frequency value in a frequency bin of the first feature data to obtain first noise reduction data, wherein the frequency bin represents the frequency value included in the first feature data; performing filtering processing on the first noise reduction data based on the cross-band module to obtain second noise reduction data; Repeating the noise reduction processing based on the narrowband module and the cross-band module on the second noise reduction data, and determining data obtained after repeating the processing a preset number of times as audio feature data.

6. The method according to claim 1, characterized in that The determining, based on the audio feature data, a scene state belonging to the same frame data in each of the audio data includes: Determining second feature data in the audio feature data, and identifying a scene state corresponding to the second feature data, where the second feature data is any frame of data in the audio feature data; Based on the scene states, the scene states of the audio data belonging to the same frame data are determined.

7. The method according to claim 6, characterized in that The identifying the scene state corresponding to the second feature data includes: reducing the number of frequencies included in the second characteristic data to a preset number; determining a time domain feature of the second feature data based on a local path of the dual-path cycle module; determining a cross-band feature of the second feature data based on a global path of the dual-path loop module; Determine the scene state corresponding to the second feature data based on the time domain features and the cross-band features according to the fully connected layer module.

8. The method according to claim 1, characterized in that The acquiring of the initial audio data, and scaling and modulating the data frames of the initial audio data based on each of the scene states to obtain initial modulated data, includes: Obtaining initial audio data in the first audio data set, and determining a data frame sequence in the initial audio data; Based on the data frame sequence, the initial audio data is scaled and modulated in turn according to each of the scene states to obtain initial modulation data corresponding to the initial audio data, where the initial modulation data indicates the scene state corresponding to each data frame of the initial audio data.

9. The method according to claim 1, characterized in that The step of modulating the first audio data set based on the initial modulation data to generate a first modulation data set, and performing noise reduction processing on the first modulation data set to obtain a second audio data set includes: modulating the first audio data set based on the initial modulation data to obtain a first modulated data set; Performing noise reduction processing on the first modulated data set based on a convolution module, a gated convolution module, and a dual-path loop module to obtain a second modulated data set; Based on the first audio data set and the second modulation data set, a second audio data set corresponding to the first audio data set is obtained.

10. The method according to claim 9, characterized in that The obtaining, based on the first audio data set and the second modulation data set, a second audio data set corresponding to the first audio data set includes: performing a bitwise multiplication process on the first audio data set and the second modulation data set to obtain a third modulation data set; Performing point-by-point addition processing on the first audio data set and the third modulation data set to obtain a second audio data set.

11. A data processing device, characterized in that: The device comprises: a feature acquisition unit, configured to acquire a first audio data set collected by a microphone array, and acquire audio feature data of the first audio data set, wherein each audio data in the first audio data set corresponds one-to-one to each microphone channel in the microphone array; a scene determination unit, configured to determine, based on the audio feature data, a scene state belonging to the same frame data in each of the audio data; a modulation unit, configured to obtain initial audio data, and scale and modulate data frames of the initial audio data based on each of the scene states to obtain initial modulated data, wherein the initial audio data is any audio data in the first audio data set; The data noise reduction unit is configured to modulate the first audio data set based on the initial modulation data to generate a first modulated data set, and perform noise reduction processing on the first modulated data set to obtain a second audio data set.

12. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the steps of the method according to any one of claims 1 to 10.

13. An electronic device comprising: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method according to any one of claims 1 to 10.

14. A computer program product having at least one instruction stored thereon, wherein when the at least one instruction is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.