Data acquisition method and device, electronic equipment and storage medium

By calculating the phase difference features between specified channel audios with location association in a multi-channel microphone array, the problems of sparsity and data redundancy in DOA encoding results are solved, achieving higher-dimensional phase information coverage and more accurate DOA encoding.

CN119724213BActive Publication Date: 2026-05-08VOICEAI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
VOICEAI TECH CO LTD
Filing Date
2024-12-02
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

When multi-channel microphone arrays encode DOA features, the phase difference results between channels are sparse, leading to inaccurate DOA encoding results and data redundancy issues.

Method used

By acquiring multi-channel audio from a microphone array, the phase difference features between specified audio channels with location association are calculated. Fourier transform and phase difference calculation are then used to generate higher-dimensional phase information coverage, avoiding data redundancy.

Benefits of technology

It improves the accuracy and efficiency of DOA encoding results, reduces data redundancy, and enhances the precision and efficiency of DOA encoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119724213B_ABST
    Figure CN119724213B_ABST
Patent Text Reader

Abstract

The application discloses a data acquisition method and device, electronic equipment and storage medium. The method comprises: acquiring audio to be processed collected through a microphone array, the audio to be processed comprising multiple sentences of multi-channel audio; for each sentence of multi-channel audio, acquiring a phase difference feature between at least one group of specified channel audio as target data, the specified channel audio comprising at least two position-associated channel audios. For each sentence of multi-channel audio, the method performs phase difference calculation by incorporating the audio of at least two position-associated microphone channels, which can generate higher-dimensional phase information coverage on one hand, so that the DOA feature obtained after phase difference is more intensive, thereby helping to improve the accuracy of the DOA encoding result; on the other hand, since the position-associated microphone channel audio is used for phase difference calculation, data redundancy can be avoided, thereby helping to improve the efficiency of DOA encoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data acquisition technology, and more specifically, to a data acquisition method, apparatus, electronic device, and storage medium. Background Technology

[0002] Multi-channel microphone arrays can estimate the direction of arrival (DOA) of sound based on target speech extraction (TSE) algorithms using neural networks, thereby estimating the direction and distance of the sound source. During the operation of a multi-channel microphone array, DOA features need to be encoded. However, the DOA features obtained from current inter-channel phase difference methods are relatively sparse, making it difficult to obtain accurate DOA encoding results. Summary of the Invention

[0003] This application proposes a data acquisition method, apparatus, electronic device, and storage medium to improve the above-mentioned problems.

[0004] In a first aspect, embodiments of this application provide a data acquisition method, the method comprising: acquiring audio to be processed collected by a microphone array, the audio to be processed comprising multiple sentences of multi-channel audio; for each sentence of multi-channel audio, acquiring at least one set of phase difference features between specified channel audio as target data, the specified channel audio comprising at least two position-related channel audio.

[0005] Secondly, embodiments of this application provide a data acquisition device, the device comprising: an audio acquisition module for acquiring audio to be processed collected by a microphone array, the audio to be processed including multiple sentences of multi-channel audio; and a data acquisition module for acquiring, for each sentence of multi-channel audio, at least one set of phase difference features between specified channel audio as target data, the specified channel audio including at least two position-related channel audio.

[0006] Thirdly, embodiments of this application provide an electronic device, including: one or more processors; a memory; and one or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more application programs are configured to perform the data acquisition method provided in the first aspect above.

[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing program code, which can be invoked by a processor to execute the data acquisition method provided in the first aspect above.

[0008] This application provides a data acquisition method, apparatus, electronic device, and storage medium. The method acquires audio to be processed, collected by a microphone array. The audio to be processed includes multiple sentences of multi-channel audio. For each sentence of multi-channel audio, at least one set of phase difference features between specified channel audio is acquired as target data. The specified channel audio includes at least two positionally related channel audio. Therefore, for each sentence of multi-channel audio, phase difference calculation is performed by incorporating the audio from at least two positionally related microphone channels. This results in higher-dimensional phase information coverage, leading to denser DOA features after phase difference calculation, thus improving the accuracy of DOA encoding results. Furthermore, using positionally related microphone channel audio for phase difference calculation avoids data redundancy, thereby improving the efficiency of DOA encoding. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A flowchart of a data acquisition method provided in an embodiment of this application is shown.

[0011] Figure 2 An example diagram of the shape of a microphone array provided in an embodiment of this application is shown.

[0012] Figure 3 It shows Figure 1 The flowchart of step S120.

[0013] Figure 4 A structural example diagram of the DOA encoder provided in an embodiment of this application is shown.

[0014] Figure 5 A flowchart of a data acquisition method provided in another embodiment of this application is shown.

[0015] Figure 6 This illustration shows an example of the positional arrangement of different audio channels when the specified audio channel is four channels, as provided in the embodiments of this application.

[0016] Figure 7 A structural block diagram of a data acquisition device provided in an embodiment of this application is shown.

[0017] Figure 8 A structural block diagram of an electronic device provided in an embodiment of this application is shown.

[0018] Figure 9 An embodiment of this application shows a storage unit for storing or carrying program code that implements the data acquisition method according to an embodiment of this application. Detailed Implementation

[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0020] Directional pickup is one of many speech enhancement methods. It refers to picking up the target signal in a mixed signal according to the direction of the sound source. That is, only the sound signal propagating from a specific direction is picked up, while noise and interference signals from other directions are not picked up but attenuated or blocked, thereby achieving the effect of target speech enhancement.

[0021] Multi-channel microphone arrays can estimate the direction of arrival (DOA) of sound based on neural network-based target speech extraction (TSE) algorithms, thereby estimating the direction and distance of the sound source. During the operation of a multi-channel microphone array, DOA features need to be encoded. However, the DOA features obtained from current inter-channel phase difference methods are relatively sparse, making it difficult to obtain accurate DOA encoding results.

[0022] Through long-term research, the inventors discovered that it is possible to acquire audio to be processed, collected by a microphone array, which includes multiple sentences of multi-channel audio. For each sentence of multi-channel audio, at least one set of phase difference features between specified channel audio is obtained as target data. The specified channel audio includes at least two positionally related channel audio. Therefore, for each sentence of multi-channel audio, phase difference calculation is performed by incorporating the audio from at least two positionally related microphone channels. This results in higher-dimensional phase information coverage, leading to denser DOA features after phase difference calculation, thus improving the accuracy of DOA encoding results. Furthermore, using positionally related microphone channel audio for phase difference calculation avoids data redundancy, thereby improving the efficiency of DOA encoding.

[0023] Therefore, in order to improve the above problems, the inventors have proposed the data acquisition method, apparatus, electronic device and storage medium provided in this application, which can generate higher-dimensional phase information coverage, making the DOA features obtained after phase differentiation more dense, thereby helping to improve the accuracy of DOA encoding results and avoid data redundancy, thereby helping to improve the efficiency of DOA encoding.

[0024] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0025] Please see Figure 1 This document illustrates a flowchart of a data acquisition method according to an embodiment of this application. This embodiment provides a data acquisition method applicable to electronic devices. The electronic device in this embodiment can be a smartphone or other mobile communication device with network connectivity; the specific device type is not limited. The method includes:

[0026] Step S110: Acquire the audio to be processed by the microphone array, wherein the audio to be processed includes multiple sentences of multi-channel audio.

[0027] In this embodiment, the microphone array may include multiple microphones, and the audio to be processed acquired by the microphone array may include multiple sentences (which can be understood as several sentences, and the specific number is not limited) of multi-channel audio, that is, each sentence of the audio to be processed is multi-channel reverberant audio (the reverberant audio may include various environmental noises). The specific type of the audio to be processed is not limited, for example, it may be user voice, singing, video audio, etc.

[0028] In this embodiment, the specific shape of the microphone array is not limited; for example, it can be a ring (circular) microphone array, a rectangular microphone array, or a microphone array of other shapes. This embodiment is described using a circular microphone array as an example.

[0029] For example, please refer to Figure 2 This diagram illustrates an example shape of a microphone array provided in an embodiment of this application. Figure 2 As shown, the microphone array is a circular array, including 9 microphone channels. Eight of these channels correspond to the ring of the microphone array, and the 9th microphone channel is located at the center of the ring; for ease of description, it can be simply referred to as the center microphone. Optionally, the audio captured by the center microphone can be called the center channel audio. The audio captured by each microphone channel can be multi-channel reverberant audio.

[0030] Step S120: For each sentence of multi-channel audio, obtain at least one set of phase difference features between specified channel audio as target data, wherein the specified channel audio includes at least two channel audios that are positionally associated.

[0031] Each multi-channel audio segment comprises multiple audio channels, and a designated audio channel can include at least two positionally related audio channels from these multiple audio channels. For example, assuming a multi-channel audio segment includes 10 audio channels, the designated audio channel can be two positionally related audio channels from these 10 audio channels, three positionally related audio channels from these 10 audio channels, or four positionally related audio channels from these 10 audio channels.

[0032] In this embodiment, location association can be understood as adjacent or close in location. For example, as described above... Figure 2 Taking the microphone array shown as an example, since the microphone array has a center microphone, in this way, position association can be understood as the positional correspondence between the audio collected by any microphone on the ring and the audio of the center channel, or position association can be understood as the positional correspondence between the audio collected by any two adjacent microphones on the ring (optionally, position association can also be determined in this way when the microphone array does not have a center microphone).

[0033] To minimize data redundancy, phase difference processing can be performed on the audio to be processed. Specifically, for each multi-channel audio sentence, at least one set of phase difference features between specified audio channels can be obtained as target data.

[0034] If only audio from a specific direction needs to be considered, the phase difference characteristics between a set of specified audio channels corresponding to that direction can be obtained. For example, using... Figure 2 Taking the microphone array shown as an example, when the specified channel audio is two channels, the phase difference characteristics between the audio collected by microphone 1 (which can be any one of 2, 3, 4, 5, 6, 7, or 8) and microphone 9 can be obtained.

[0035] Optionally, if it is necessary to focus on audio from multiple directions, phase difference features can be obtained between multiple sets of specified audio channels corresponding to multiple directions (which can be several specific directions or omnidirectional). For example, continuing with... Figure 2 Taking the microphone array shown as an example, when the specified channel audio is still two channels of audio, the phase difference characteristics between the audio collected by microphone 1 and microphone 9, the phase difference characteristics between the audio collected by microphone 2 and microphone 9, the phase difference characteristics between the audio collected by microphone 3 and microphone 9, the phase difference characteristics between the audio collected by microphone 4 and microphone 9, the phase difference characteristics between the audio collected by microphone 5 and microphone 9, the phase difference characteristics between the audio collected by microphone 6 and microphone 9, the phase difference characteristics between the audio collected by microphone 7 and microphone 9, and the phase difference characteristics between the audio collected by microphone 8 and microphone 9 can be obtained.

[0036] It should be noted that the examples above are for illustrative purposes only and do not constitute a limitation on the number of audio channels specified. For example, continuing with... Figure 2Taking the microphone array shown as an example, if the specified channel audio is still three channels, and it is necessary to focus on the audio in a specific direction, the phase difference characteristics between the audio collected by microphone 1, microphone 9, and microphone 2 can be obtained; if it is necessary to focus on the omnidirectional audio, the phase difference characteristics between the audio collected by microphone 1, microphone 9, and microphone 2, the phase difference characteristics between the audio collected by microphone 2, microphone 9, and microphone 3, the phase difference characteristics between the audio collected by microphone 3, microphone 9, and microphone 4, the phase difference characteristics between the audio collected by microphone 4, microphone 9, and microphone 5, the phase difference characteristics between the audio collected by microphone 5, microphone 9, and microphone 6, the phase difference characteristics between the audio collected by microphone 6, microphone 9, and microphone 7, the phase difference characteristics between the audio collected by microphone 7, microphone 9, and microphone 8, and the phase difference characteristics between the audio collected by microphone 8, microphone 9, and microphone 1 can be obtained.

[0037] Optionally, for each multi-channel audio sentence, the process of obtaining at least one set of phase difference features between specified audio channels is described below.

[0038] Please see Figure 3 In one implementation method, step S120 may specifically include:

[0039] Step S121: Perform a Fourier transform on each multi-channel audio segment to obtain the real part parameters and imaginary part parameters corresponding to each audio channel.

[0040] As one implementation method, a Fourier transform can be performed on each multi-channel audio segment to obtain the real and imaginary parameters corresponding to each channel. The dimensions of each multi-channel audio segment after the Fourier transform are [N, C, F, T], where N represents the number of training samples, C represents the number of channels, F represents the frequency dimension, and T represents the time dimension.

[0041] The specific principles and implementation process of the Fourier transform will not be elaborated here.

[0042] Step S122: Based on the corresponding real part parameters and imaginary part parameters, obtain at least one set of first phase difference and second phase difference between specified channel audio, wherein the second phase difference is 90 degrees sine difference from the first phase difference.

[0043] As one implementation method, at least one set of first phase difference and second phase difference between specified channel audio can be obtained based on the corresponding real part parameters and imaginary part parameters. In this implementation method, the second phase difference is 90 degrees sine difference from the first phase difference.

[0044] As one implementation method, when the specified audio channel consists of two audio channels, the first phase difference and the second phase difference between at least one set of specified audio channels can be obtained using the following formula:

[0045] IPD(W k1 W k2 )={cos(arctan2(R k1 ,I k1 )-arctan2(R k2 ,I k2 )),sin(arctan2(R k1 ,I k1 )-arctan2(R k2 ,I k2 ))}

[0046] Among them, cos(arctan2(R) k1 ,I k1 )-arctan2(R k2 ,I k2 The first phase difference is represented by sin(arctan2(R). k1 ,I k1 )-arctan2(R k2 ,I k2 R represents the second phase difference. k1 R k2 Characterizing the real part parameter, I k1 I k2 Characterizing the imaginary part parameter, IPD(W) at this time k1 W k2 This characterizes the phase difference parameters between specified channels.

[0047] Optional, IPD(W) k1 W k2 The dimensions of ) are (N, C) IPD ,2*F,T), where N represents the number of training samples, C IPD The number of pairs between channels obtained from the calculation is represented by F, which represents the frequency dimension, and T represents the time dimension.

[0048] Optionally, when the specified channel audio is two channels, the specified channel audio can be the two channels associated with the aforementioned position, or it can be two channels selected according to specified rules.

[0049] Optionally, the specified rule can be to select the audio captured by the two microphone channels that are physically closest in the microphone array as the specified channel audio, or it can be to select the audio captured by the two microphone channels that are physically farthest in the microphone array as the specified channel audio, or it can be to use the audio captured by any two microphone channels in the microphone array as the specified channel audio (in this case, all microphones in the microphone array will form pairs, and each pair of microphones will be a specified channel). For example, in a microphone array... Figure 2 In the case of the circular array shown, each audio segment includes 9 microphone channels, which will be denoted as W for ease of description. k ={W k1 W k2 ,…,W k9 At this point, each of the eight microphones located on the ring can be paired with the center microphone to form a designated channel in pairs. In this way, eight designated channels can be obtained, i.e., C=8.

[0050] Step S123: Concatenate the first phase difference and the second phase difference to obtain the phase difference feature, which is used as the target data.

[0051] In this embodiment, to avoid data redundancy and reduce subsequent data processing pressure, the first phase difference and the second phase difference obtained above can be concatenated to obtain a phase difference feature, which is used as the target data. The size of the phase difference feature obtained after concatenating the first phase difference and the second phase difference is still (N, C). IPD ,2*F,T).

[0052] In this embodiment, the obtained target data can be used for DOA feature encoding. Specifically, the target data can be input into the DOA (Direction of Arrival) encoder, which can first encode the target data and then compress (reduce dimensionality).

[0053] The DOA encoder may include a first coding layer, a second coding layer, and a third coding layer, and the first coding layer, the second coding layer, and the third coding layer form a bottleneck structure (e.g., Figure 4 As shown), the output of the first coding layer enters the second coding layer, and the output of the second coding layer enters the third coding layer.

[0054] Optionally, the target data in this embodiment may include the number of channels and the frequency dimension, and each audio segment of the audio to be processed includes a label, which can be used to guide the model to learn the audio processing operations required for different pickup areas. In this approach, during the dimensionality reduction of the number of channels and the frequency dimension, one implementation method is to reduce the number of channels to a first dimension and the frequency dimension to a second dimension in the first coding layer of the DOA encoder; then, in the second coding layer of the DOA encoder, reduce the frequency dimension output by the first coding layer to a second dimension while maintaining the number of channels output by the first coding layer. That is, the second coding layer may not perform dimensionality reduction on the number of channels, but simply transmit the number of channels output by the first coding layer to the third coding layer; then, in the third coding layer of the DOA encoder, increase the number of channels transmitted by the second coding layer to a first dimension and reduce the frequency dimension output by the second coding layer to a second dimension.

[0055] The first dimension is determined by the number of pairs of channel audio data formed by the specified channel audio data, and the second dimension is determined by the number of sampling points based on the Fourier transform. Optionally, the number of pairs of channel audio data formed by the specified channel audio data is even.

[0056] In some other implementations, the number of pairs of channel audios formed by the specified channel audios can also be odd, in which case the initial frequency dimension is F; while if the number of pairs of channel audios formed by the specified channel audios is even, the initial frequency dimension can be 2F.

[0057] Optionally, in one specific implementation, the number of channels can be reduced by a factor of 4 (first dimension) and the frequency dimension can be reduced by a factor of 2 (second dimension) in the first coding layer; the frequency dimension of the output of the first coding layer can be reduced by a factor of 2 (second dimension) in the second coding layer while maintaining the number of channels output by the first coding layer; and in the third coding layer, the number of channels transmitted by the second coding layer can be increased by a factor of 4 (first dimension) and the frequency dimension of the output of the second coding layer can be reduced by a factor of 2 (second dimension) again.

[0058] In the DOA encoder, the frequency dimension is a redundant part, so it can be reduced in multiple dimensions. After the number of channels is reduced in dimension, it will be raised back to the original dimension to facilitate the complete output audio channel characteristics.

[0059] In this embodiment, the second coding layer can be expanded to n layers (n≥1). Optionally, each of the n layers only performs dimensionality reduction processing on the frequency dimension. Optionally, each layer can reduce the frequency dimension to half that of the previous layer.

[0060] In this embodiment, the number of channels output by the third coding layer can be 4 or an even number divisible by 4. Optionally, the minimum number of channels output by the third coding layer can be 4. These 4 channels can enhance modeling of four different right-angled directions in the plane. For example, connecting two microphones with a line, where the two lines intersect, constructs a coordinate system that can be used to enhance modeling of four different directions in the plane.

[0061] In this embodiment, the number of channels and frequency dimension output by the third coding layer of the DOA encoder can be used as the output parameters of the DOA encoder.

[0062] Furthermore, the output parameters of the DOA encoder can be input into a specified network layer of the neural network to be trained, so as to start training the neural network from the specified network layer and obtain the target neural network.

[0063] Here, the specified network layer is the hidden layer at a specified position. The specified position can be understood as the network layer in the neural network to be trained that corresponds to the position of the third encoding layer of the DOA encoder. For example, the specified position can be the third or fourth network layer of the neural network to be trained. Both the third and fourth network layers are hidden layers.

[0064] Specifically, when inputting the output parameters of the DOA encoder into a specified network layer of the neural network to be trained, the output parameters of the DOA encoder can be concatenated with the input features of the specified network layer to obtain reference input features. These reference input features are then used as new input features for the specified network layer. All hidden layers of the neural network to be trained, starting from the current hidden layer, are updated. Finally, the updated neural network is used as the target neural network.

[0065] Furthermore, noise reduction can be performed on the audio to be processed using a target neural network. The target neural network trained in this embodiment has the ability to perform different audio processing on audio from different pickup areas. The audio can originate from a pickup area or a non-pickup area. If the audio originates from a pickup area, the target neural network model can perform speech signal enhancement processing on the audio based on the label carried by the sample audio (in this case, the label is either sample audio or reverberant audio). Conversely, if the audio originates from a non-pickup area, the target neural network model can perform suppression processing on the audio based on the label carried by the sample audio (in this case, the label is silent audio).

[0066] As one implementation method, noise reduction can be performed on the audio to be processed using a target neural network. Specifically, the target neural network can remove reverberation (or reflected sound) from the audio, thereby achieving noise reduction, de-reverberation, and directional sound pickup. The audio to be processed can be of the same type as the sample audio but with different content.

[0067] The data acquisition method provided in this embodiment acquires audio to be processed collected by a microphone array. The audio to be processed includes multiple sentences of multi-channel audio. For each sentence of multi-channel audio, at least one set of phase difference features between specified channel audio is acquired as target data. The specified channel audio includes at least two positionally related channel audio. Thus, for each sentence of multi-channel audio, phase difference calculation is performed by incorporating the audio from at least two positionally related microphone channels. On the one hand, this can generate higher-dimensional phase information coverage, making the DOA features obtained after phase difference more dense, thereby helping to improve the accuracy of DOA encoding results. On the other hand, since phase difference calculation is performed using positionally related microphone channel audio, data redundancy can be avoided, thereby helping to improve the efficiency of DOA encoding.

[0068] Please see Figure 5 The diagram illustrates a flowchart of a data acquisition method according to another embodiment of this application. This embodiment provides a data acquisition method that can be applied to electronic devices, and the method includes:

[0069] Step S210: Acquire the audio to be processed by the microphone array, wherein the audio to be processed includes multiple sentences of multi-channel audio.

[0070] The specific implementation of step S210 can be referred to the relevant description of step S110 in the foregoing embodiments, and will not be repeated here.

[0071] Step S220: Perform a Fourier transform on each multi-channel audio segment to obtain the real part parameters and imaginary part parameters corresponding to each audio channel.

[0072] The specific implementation of step S220 can be referred to the relevant description of step S121 in the foregoing embodiments, and will not be repeated here.

[0073] Step S230: Based on the corresponding real part parameters and imaginary part parameters, obtain at least one set of first phase difference and second phase difference between specified channel audio, wherein the second phase difference is 90 degrees sine difference from the first phase difference.

[0074] The specific implementation of step S230 can be referred to the relevant description of step S122 in the foregoing embodiments, and will not be repeated here.

[0075] In this embodiment, the method of obtaining the first phase difference and the second phase difference between at least one set of specified channel audio will differ depending on the number of specified channel audios. As one implementation, in the process of obtaining the first phase difference and the second phase difference between at least one set of specified channel audio based on the corresponding real and imaginary parameters, the number of channel audios included in the specified channel audio can be determined first, and then the first phase difference and the second phase difference between at least one set of said number of channel audios can be obtained based on the corresponding real and imaginary parameters.

[0076] Optionally, the number of audio channels included in the specified audio channel in this embodiment is not limited. This application describes the specified audio channel as including two audio channels, three audio channels, and four audio channels respectively.

[0077] In this scenario, where a specified audio channel includes two audio channels (which can be understood as first-order audio channels), if the audio to be processed acquired by the microphone array includes the center channel audio, then the specified audio channel can include the center channel audio and any other channel audio. In this approach, if the number of specified audio channels is a group, a phase difference can be performed between the other channel audio and the center channel audio based on the real and imaginary parameters corresponding to the other channel audio and the center channel audio, yielding a phase difference result. The cosine value of the phase difference result is then obtained as the first phase difference, and the sine value of the phase difference result is obtained as the second phase difference.

[0078] Optionally, if the audio to be processed acquired by the microphone array does not include the central channel audio, the specified channel audio can include any two adjacent channel audios. In this way, the phase difference is performed on the two adjacent channel audios based on the real part parameters and imaginary part parameters corresponding to each of the two adjacent channel audios to obtain the phase difference result.

[0079] Optionally, if the audio to be processed acquired by the microphone array does not include the central channel audio, the specified channel audio can also include the audio of two channels at any position. In this way, the phase difference can be performed on the two channels at any position based on the real part parameters and imaginary part parameters corresponding to the two channels at that position, and the phase difference result can be obtained.

[0080] When a specified audio channel includes two audio channels, as one implementation method, the phase difference feature can be obtained according to the following formula:

[0081] IPD(W k1 W k2)=concat{cos(arctan2(R k1 ,I k1 )-arctan2(R k2 ,I k2 )),sin(arctan2(R k1 ,I k1 )-arctan2(R k2 ,I k2 ))}.

[0082] Among them, IPD(W) k1 W k2 The first phase difference is characterized by its phase difference feature, and the second phase difference is characterized by its concatenation along the F (frequency) axis.

[0083] Optionally, the above formula only illustrates one set of cases (i.e., 1-channel audio and 2-channel audio). If we use... Figure 2 Taking the 9-channel circular microphone array shown as an example, the audio acquired by the center channel of the microphone array can be used as the center channel audio of the specified channel. In this way, the following 8 sets of phase difference features (which can be simply referred to as IPD features) can be constructed:

[0084] {W k1 W k9},{W k2 W k9},{W k3 W k9},{W k4 W k9},{W k5 W k9},{W k6 W k9},{W k7 W k9},{W k8 W k9}

[0085] In this embodiment, when the specified channel audio includes three audio channels (which can be understood as second-order audio channels), if the audio to be processed acquired by the microphone array includes the center channel audio, the specified channel audio can include the center channel audio and any two other adjacent channel audios. In this way, if the number of specified channel audios is a group, each of the two adjacent channel audios can be phase-differed with the center channel audio based on the real and imaginary parameters corresponding to the other two adjacent channel audios and the real and imaginary parameters corresponding to the center channel audio, to obtain two phase difference results to be processed; and the difference between the two phase difference results to be processed is obtained to obtain the target phase difference result; then the cosine value of the target phase difference result is obtained as the first phase difference; and the sine value of the target phase difference result is obtained as the second phase difference.

[0086] In other words, this method performs phase difference on the result of phase difference between each of the two adjacent audio channels and the center channel, in order to avoid data redundancy caused by data overlap or intersection.

[0087] Optionally, if the audio to be processed acquired by the microphone array does not include the center channel audio, the specified channel audio can include any three adjacent channel audios. In this way, the middle channel audio of these three adjacent channel audios can be used as the center channel audio of the specified channel audio. In this method, the calculation of the first phase difference and the second phase difference is the same as the calculation method when the audio to be processed acquired by the microphone array includes the center channel audio, and will not be repeated here.

[0088] Optionally, if the audio to be processed acquired by the microphone array does not include the center channel audio, the specified channel audio can also include the three channels audio at any position. In this case, the calculation method for the first phase difference and the second phase difference is the same as when the audio to be processed acquired by the microphone array includes the center channel audio. For example, assuming the three channels audio are channel 1, channel 2, and channel 3, the phase difference between channel 1 and channel 2 will be calculated and denoted as the first phase difference to be processed, and the phase difference between channel 2 and channel 3 will be calculated and denoted as the second phase difference to be processed. Then, the phase difference between the first phase difference to be processed and the second phase difference to be processed will be calculated to obtain the target phase difference result.

[0089] When a specified audio channel includes three audio channels, as one implementation method, the phase difference characteristics can be obtained according to the following formula:

[0090] IPD(W k1 W k2 W k3)=concat{cos(arctan2(R k1 ,I k1 )-2*arctan2(R k2 ,I k2 )+arctan2(R k3 ,I k3 )),

[0091] sin(arctan2(R k1 ,I k1 )-2*arctan2(R k2 ,I k2 )+arctan2(R k3 ,I k3 ))}

[0092] Among them, cos(arctan2(R) k1 ,I k1 )-2*arctan2(R k2 ,I k2 )+arctan2(R k3 ,I k3 The first phase difference is represented by sin(arctan2(R). k1 ,I k1 )-2*arctan2(R k2 ,I k2 )+arctan2(R k3 ,I k3 The second phase difference (IPD) is represented by the second phase difference (W). k1 W k2 W k3 The first phase difference is characterized by its phase difference feature, and the second phase difference is characterized by its concatenation along the F (frequency) axis.

[0093] Optionally, the above formula only illustrates one set of cases (i.e., 1-channel audio, 2-channel audio, and 3-channel audio), and uses 2-channel audio as the center channel audio of the specified channel audio. However, if... Figure 2 Taking the 9-channel circular microphone array shown as an example, the audio acquired by the center channel of the microphone array can be used as the center channel audio of the specified channel. In this way, the following 8 sets of phase difference features (which can be simply referred to as hoIPD (high-order inter-microphone phase difference) features) can be constructed:

[0094] {W k1 W k9 W k2},{Wk2 W k9 W k3},{W k3 W k9 W k4},{W k4 W k9 W k5},{W k5 W k9 W k6},{W k6 W k9 W k7},{W k7 W k9 W k8},{W k8 W k9 W k1}

[0095] In this embodiment, when the specified channel audio includes four channel audios (which can be understood as third-order channel audios), if the microphone array includes two center channel audios (for example, when the microphone array is a concentric circle array layout, there will be two center channel audios), then the specified channel audios may include the two center channel audios and any two other channel audios that are adjacent in position. In this approach, if the number of specified audio channels is a group, then based on the real and imaginary parameters corresponding to the other audio channels that are adjacent to each other, and the real and imaginary parameters corresponding to the two center audio channels, each audio channel in the two adjacent audio channels is phase-differentiated with the center audio channel that is closer in distance, to obtain a first phase difference result and a second phase difference result to be processed; and based on the real and imaginary parameters corresponding to the two center audio channels, the two center audio channels are phase-differiated to obtain a third phase difference result to be processed; then, the first phase difference result and the third phase difference result to be processed are subtracted to obtain a first reference phase difference result; and the second phase difference result and the third phase difference result to be processed are subtracted to obtain a second reference phase difference result; then, the first reference phase difference result and the second reference phase difference result are subtracted to obtain a target phase difference result; then, the cosine value of the target phase difference result is obtained as the first phase difference; and the sine value of the target phase difference result is obtained as the second phase difference.

[0096] In other words, assuming the four audio channels are channel 1, channel 2, channel 3, and channel 4, and the arrangement of channel 1, channel 2, channel 3, and channel 4 is as follows: Figure 6 As shown, Figure 6 The two rings shown represent a microphone array. First, the phase difference between audio channels 1 and 2 is calculated, denoted as the first phase difference result to be processed. Then, the phase difference between audio channels 3 and 4 is calculated, denoted as the second phase difference result to be processed. Next, the phase difference between the first and third phase difference results is calculated, denoted as the first reference phase difference result. Finally, the phase difference between the second and third phase difference results is calculated, denoted as the second reference phase difference result. Finally, the phase difference between the first and second reference phase difference results is calculated to obtain the target phase difference result.

[0097] By calculating multiple phase difference results, the correlation between different audio channels in close proximity can be established, which enables the effective integration of phase information from multiple different microphones at the feature level. Furthermore, by performing phase difference on the phase difference results multiple times, data redundancy can be reduced, which helps to reduce the learning pressure on the subsequent DOA encoder and neural network.

[0098] Optionally, if the microphone array does not include (two) center channel audio, the specified channel audio can include any four adjacent channel audios. In this case, the middle two of these four adjacent channel audios can be used as the center channel audio of the specified channel audio. In this method, the calculation of the first phase difference and the second phase difference is the same as when the microphone array includes two center channel audios, and will not be repeated here.

[0099] Optionally, if the microphone array does not include (two) center channel audio, the specified channel audio can also include four channel audios at any position. In this case, the first phase difference and the second phase difference are calculated in the same way as when the microphone array includes two center channel audios.

[0100] When a specified audio channel includes four audio channels, as one implementation method, the phase difference characteristics can be obtained according to the following formula:

[0101] IPD(W k1 W k2 W k3 W k4 )=concat{cos((artan2(R k1 I k1 )-arctan2(R k2 I k2 ))-2*(arctan2(R k2 Ik2 ) - arctan2(R k3 , I k3 )) + (arctan2(R k2 , I k3 ) - arctan2(R k4 , I k4 ))), sin((arctan2(R k1 , I k1 )) - arctan2(R k2 , I k2 )) - 22 * (arctan2(R k2 , I k2 )) - arctan2(R k3 , I k3 )) + (arctan2(R k3 , I k3 )) - arctan2(R k4 , I k4 )))} = concat{cos(arctan2(R k1 , I k1 )) - 3 * arctan2(R k2 , I k2 + 3 * arctan2(R k3 , I k3 )) - arctan2(R k4 , k4 )), sin(arctan2(R k1 , I k1 )) - 3 * arctan2(R k2 , I k2 )) + 3 * arctan2(R k3 , I k3 )) - arctan2(R k4 , I k4 ))}

[0102] Where, cos((arctan2(R k1 , I k1 )) - arctan2(R k2 , I k2 )) - 2 * (arctan2(R k2 , I k2 )) - arctan2(R k3 , I k3 )) + (arctan2(R k3 , I k3 )) - arctan2(R k4 , I k4 ))

[0103] Characterizing the first phase difference, sin((arctan2(R) k1 I k1 )-arctan2(R k2 I k2 ))-2*(arctan2(R k2 I k2 )-arctan2(R k3 I k3 ))+(arctan2(R k3 I k3 )-arctan2(R k4 I k4 )))

[0104] Characterizing the second phase difference, IPD(W) k1 W k2 W k3 W k4 The first phase difference is characterized by its phase difference feature, and the second phase difference is characterized by its concatenation along the F (frequency) axis.

[0105] Optionally, the above formula only illustrates one set (i.e., 1-channel audio, 2-channel audio, 3-channel audio, and 4-channel audio), and uses 2-channel audio and 3-channel audio as the center channel audio of the specified channel audio. Here, 2-channel audio and 3-channel audio can be the two center channel audios of a concentric circle microphone array, or they can be the two center channel audios located between 1-channel audio and 4-channel audio. There is no limitation here.

[0106] It is worth noting that, if it is necessary to reduce data redundancy, the implementation method of this application can give priority to using at least two location-related audio channels as designated audio channels. This method allows the DOA encoder to learn the differences between different audio channels more efficiently based on the structural relationship between different audio channels. If the structural relationship between different audio channels is not considered when selecting the designated audio channel, at least two audio channels can be arbitrarily selected as the designated audio channel. This method allows the DOA encoder to learn the differences between different audio channels more comprehensively based on the structural relationship between different audio channels.

[0107] Step S240: The first phase difference and the second phase difference are concatenated to obtain the phase difference feature, which is used as the target data. The specified channel audio includes at least two channel audios with positional association.

[0108] The specific implementation of step S240 can be referred to the relevant description of step S123 in the foregoing embodiments, and will not be repeated here.

[0109] The data acquisition method provided in this embodiment acquires audio to be processed collected by a microphone array, the audio to be processed including multiple sentences of multi-channel audio; performs Fourier transform on each sentence of multi-channel audio to obtain the real part parameters and imaginary part parameters corresponding to each channel audio; based on the corresponding real part parameters and imaginary part parameters, obtains at least one set of first phase difference and second phase difference between specified channel audio, the second phase difference being 90 degrees sine difference from the first phase difference; concatenates the first phase difference and the second phase difference to obtain the phase difference feature, which is used as target data, the specified channel audio including at least two positionally related channel audio. Therefore, for each sentence of multi-channel audio, by incorporating the audio from at least two positionally related microphone channels for phase difference calculation, on the one hand, higher-dimensional phase information coverage can be generated, making the DOA features obtained after phase difference more dense, thereby helping to improve the accuracy of DOA encoding results; on the other hand, since phase difference calculation is performed using positionally related microphone channel audio, data redundancy can be avoided, thereby helping to improve the efficiency of DOA encoding.

[0110] Meanwhile, by introducing a phase difference algorithm between high-order (i.e., second-order and third-order) channel audio, the phase information of multiple different microphones is effectively integrated at the feature level, expanding the time delay information. Furthermore, the IPD features obtained after phase difference between high-order channel audio can enable the DOA encoder to learn more comprehensively, thereby reducing the learning pressure on the DOA encoder and subsequent neural networks.

[0111] Please see Figure 7 This is a structural block diagram of a data acquisition device provided in an embodiment of this application. This embodiment provides a data acquisition device 300, which can operate in an electronic device. The device 300 includes: an audio acquisition module 310 to be processed and a data acquisition module 320.

[0112] The audio acquisition module 310 is used to acquire audio to be processed collected by a microphone array, wherein the audio to be processed includes multiple sentences of multi-channel audio.

[0113] The data acquisition module 320 is used to acquire at least one set of phase difference features between specified channel audio for each sentence of multi-channel audio, as target data, wherein the specified channel audio includes at least two channel audios with positional association.

[0114] As one implementation, the data acquisition module 320 can be used to perform Fourier transform on each multi-channel audio to obtain the real part parameters and imaginary part parameters corresponding to each channel audio; based on the corresponding real part parameters and imaginary part parameters, to obtain at least one set of first phase difference and second phase difference between specified channel audio, wherein the second phase difference differs from the first phase difference by a 90-degree sine wave; and to concatenate the first phase difference and the second phase difference to obtain the phase difference feature, which is used as target data.

[0115] In this embodiment, if the audio to be processed acquired by the microphone array includes the center channel audio, the designated channel audio includes the center channel audio and any other channel audio. In this way, if the number of designated channel audios is a group, the data acquisition module 320 can specifically be used to perform phase difference between the arbitrary other channel audio and the center channel audio based on the real and imaginary parameters corresponding to the arbitrary other channel audio and the real and imaginary parameters corresponding to the center channel audio, to obtain a phase difference result; obtain the cosine value of the phase difference result as the first phase difference; and obtain the sine value of the phase difference result as the second phase difference.

[0116] Wherein, if the audio to be processed acquired by the microphone array does not include the central channel audio, the specified channel audio includes the audio of any two adjacent channels.

[0117] In this embodiment, if the audio to be processed acquired by the microphone array includes the center channel audio, the specified channel audio includes the center channel audio and any two other channel audios that are adjacent in position. In this way, if the number of specified channel audios is a group, the data acquisition module 320 can specifically be used to perform phase difference between each of the two adjacent channel audios and the center channel audio based on the real and imaginary parameters corresponding to the other two adjacent channel audios and the real and imaginary parameters corresponding to the center channel audio, to obtain two phase difference results to be processed; subtract the two phase difference results to obtain a target phase difference result; obtain the cosine value of the target phase difference result as the first phase difference; and obtain the sine value of the target phase difference result as the second phase difference.

[0118] If the audio to be processed acquired by the microphone array does not include the center channel audio, and the specified channel audio includes any three adjacent channel audios, then the middle channel audio of the three adjacent channel audios can be used as the center channel audio of the specified channel audio.

[0119] In this embodiment, if the microphone array includes two center channel audios, the designated channel audio includes the two center channel audios and any two other channel audios that are adjacent in position. In this way, if the number of designated channel audios is a group, the data acquisition module 320 can specifically be used to perform phase difference analysis on each of the two adjacent channel audios and the center channel audio with the closest interval, based on the real and imaginary parameters corresponding to the two other channel audios and the real and imaginary parameters corresponding to the two center channel audios, to obtain a first phase difference result and a second phase difference result to be processed; based on the real and imaginary parameters corresponding to the two center channel audios, the data acquisition module 320 can perform phase difference analysis on each of the two adjacent channel audios and the center channel audio with the closest interval, to obtain a first phase difference result and a second phase difference result to be processed; based on the real and imaginary parameters corresponding to the two center channel audios, the data acquisition module 320 can perform phase difference analysis on the two center channel audios. The first phase difference result is obtained by performing phase difference analysis on the frequency to obtain a third phase difference result to be processed; the first phase difference result to be processed and the third phase difference result to be processed are subtracted to obtain a first reference phase difference result; the second phase difference result to be processed and the third phase difference result to be processed are subtracted to obtain a second reference phase difference result; the first reference phase difference result and the second reference phase difference result are subtracted to obtain a target phase difference result; the cosine value of the target phase difference result is obtained as the first phase difference; the sine value of the target phase difference result is obtained as the second phase difference.

[0120] If the audio to be processed acquired by the microphone array does not include the center channel audio, and the specified channel audio includes any four adjacent channel audios, then in this case, the middle two channel audios of the four adjacent channel audios can be used as the center channel audio of the specified channel audio.

[0121] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the above-described device and module can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0122] In the several embodiments provided in this application, the coupling between modules can be electrical, mechanical, or other forms of coupling.

[0123] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0124] Please see Figure 8Based on the data acquisition method and apparatus described above, this application also provides an electronic device 100 capable of executing the aforementioned data acquisition method. The electronic device 100 includes a memory 102 and one or more (only one shown in the figure) processors 104 coupled to each other, with a communication line connecting the memory 102 and the processors 104. The memory 102 stores a program capable of executing the contents of the aforementioned embodiments, and the processors 104 can execute the program stored in the memory 102.

[0125] The processor 104 may include one or more processing cores. The processor 104 connects to various parts within the electronic device 100 using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 102, and by calling data stored in the memory 102. Optionally, the processor 104 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 104 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 104 and may be implemented separately using a communication chip.

[0126] The memory 102 may include random access memory (RAM) or read-only memory (ROM). The memory 102 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 102 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the aforementioned embodiments. The data storage area may also store data created by the electronic device 100 during use (such as phonebook data, audio and video data, chat log data, etc.).

[0127] Please refer to Figure 9This diagram illustrates a structural block diagram of a computer-readable storage medium provided in an embodiment of this application. The computer-readable storage medium 400 stores program code that can be called by a processor to execute the methods described in the above method embodiments.

[0128] The computer-readable storage medium 400 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. Optionally, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has storage space for program code 410 that performs any of the method steps described above. This program code can be read from or written to one or more computer program products. The program code 410 may be compressed, for example, in a suitable form.

[0129] In summary, the data acquisition method, apparatus, electronic device, and storage medium provided in this application acquire audio to be processed by a microphone array, the audio to be processed including multiple sentences of multi-channel audio; for each sentence of multi-channel audio, at least one set of phase difference features between specified channel audio is acquired as target data, the specified channel audio including at least two positionally related channel audio. Thus, for each sentence of multi-channel audio, by incorporating the audio from at least two positionally related microphone channels for phase difference calculation, on the one hand, higher-dimensional phase information coverage can be generated, making the DOA features obtained after phase difference more dense, thereby helping to improve the accuracy of DOA encoding results; on the other hand, since phase difference calculation is performed using positionally related microphone channel audio, data redundancy can be avoided, thereby helping to improve the efficiency of DOA encoding.

[0130] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A data acquisition method, characterized in that, The method includes: Acquire audio to be processed by a microphone array, the audio to be processed including multiple sentences of multi-channel audio; Perform a Fourier transform on each multi-channel audio sentence to obtain the real part parameters and imaginary part parameters corresponding to each audio channel; Determine the number of audio channels included in a specified audio channel; Based on the corresponding real part parameters and imaginary part parameters, obtain at least one set of the number of channel audios, a first phase difference and a second phase difference, wherein the second phase difference is 90 degrees sine wave difference from the first phase difference; The first phase difference and the second phase difference are concatenated to obtain the phase difference feature, which is used as the target data. The target data is used for direction of arrival coding. Where, in the case that the number of specified channel audios is a set, the audio to be processed acquired by the microphone array includes the center channel audio, and the specified channel audio includes the center channel audio and any two other channel audios in adjacent positions, the step of obtaining the first phase difference and the second phase difference between at least one set of the specified number of channel audios based on the corresponding real part parameters and imaginary part parameters includes: performing phase difference analysis on each of the two adjacent channel audios and the center channel audio based on the real part parameters and imaginary part parameters corresponding to each of the other two adjacent channel audios and the real part parameters and imaginary part parameters corresponding to the center channel audio, to obtain two phase difference results to be processed; subtracting the two phase difference results to be processed to obtain a target phase difference result; obtaining the cosine value of the target phase difference result as the first phase difference; and obtaining the sine value of the target phase difference result as the second phase difference. Wherein, if the audio to be processed acquired by the microphone array does not include the center channel audio acquired by the center microphone of the microphone array, and the specified channel audio includes any three adjacent channel audios, the center channel audio of the specified channel audio is the middle channel audio of the three adjacent channel audios.

2. The method according to claim 1, characterized in that, When the number of specified channel audios is a set, and if the audio to be processed acquired by the microphone array includes the center channel audio, and the specified channel audio includes the center channel audio and any other channel audio, obtaining the first phase difference and the second phase difference between at least one set of the specified number of channel audios based on the corresponding real part parameter and imaginary part parameter includes: Based on the real and imaginary parameters corresponding to any other audio channel and the real and imaginary parameters corresponding to the center audio channel, a phase difference is performed between the audio of any other audio channel and the audio of the center audio channel to obtain a phase difference result. Obtain the cosine value of the phase difference result and use it as the first phase difference; Obtain the sine value of the phase difference result, and use it as the second phase difference.

3. The method according to claim 2, characterized in that, If the audio to be processed acquired by the microphone array does not include the center channel audio, the specified channel audio includes the audio of any two adjacent channels.

4. The method according to claim 1, characterized in that, When the number of specified channel audios is a set, and if the microphone array includes two center channel audios, and the specified channel audios include the two center channel audios and any two other channel audios in adjacent positions, obtaining the first phase difference and the second phase difference between at least one set of the specified number of channel audios based on the corresponding real part parameter and imaginary part parameter includes: Based on the real and imaginary parameters of the other audio channels that are adjacent to each other, and the real and imaginary parameters of the two central audio channels, each audio channel in the two adjacent audio channels is phase-divided with the central audio channel that is closer in distance to obtain the first phase difference result and the second phase difference result to be processed. Based on the real and imaginary parameters corresponding to the two center channel audios, phase difference is performed on the two center channel audios to obtain a third phase difference result to be processed. The first phase difference result to be processed is subtracted from the third phase difference result to be processed to obtain the first reference phase difference result; The second phase difference result to be processed is subtracted from the third phase difference result to be processed to obtain the second reference phase difference result; The target phase difference result is obtained by subtracting the first reference phase difference result from the second reference phase difference result. Obtain the cosine value of the target phase difference result and use it as the first phase difference; Obtain the sine value of the target phase difference result, and use it as the second phase difference.

5. The method according to claim 4, characterized in that, If the audio to be processed acquired by the microphone array does not include the center channel audio acquired by the center microphone of the microphone array, and the specified channel audio includes the audio of any four adjacent channels, the method further includes: The two middle channels of the four adjacent audio channels are taken as the center channel of the designated audio channel.

6. A data acquisition device, characterized in that, The device includes: The audio to be processed module is used to acquire the audio to be processed collected by the microphone array, wherein the audio to be processed includes multiple sentences of multi-channel audio; The data acquisition module is used to perform Fourier transform on each multi-channel audio sentence to obtain the real part parameters and imaginary part parameters corresponding to each channel audio; determine the number of channel audios included in a specified channel audio; based on the corresponding real part parameters and imaginary part parameters, obtain at least one set of first phase difference and second phase difference between the number of channel audios, wherein the second phase difference differs from the first phase difference by a 90-degree sine wave; and concatenate the first phase difference and the second phase difference to obtain phase difference features, which are used as target data for direction-of-arrival encoding. The data acquisition module is configured to, when the number of specified channel audios is a group, the audio to be processed acquired by the microphone array includes the center channel audio, and the specified channel audio includes the center channel audio and any two other channel audios in adjacent positions, perform phase difference analysis on each of the two adjacent channel audios and the center channel audio, based on the real and imaginary parameters corresponding to the other two adjacent channel audios and the real and imaginary parameters corresponding to the center channel audio, to obtain two phase difference results to be processed; subtract the two phase difference results to obtain a target phase difference result; obtain the cosine value of the target phase difference result as the first phase difference; obtain the sine value of the target phase difference result as the second phase difference; wherein, if the audio to be processed acquired by the microphone array does not include the center channel audio acquired by the center microphone of the microphone array, and the specified channel audio includes any three adjacent channel audios, the center channel audio of the specified channel audio is the middle channel audio of the three adjacent channel audios.

7. An electronic device, characterized in that, Includes one or more processors and memory; One or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs being configured to perform the method of any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores program code, wherein the program code, when executed by a processor, performs the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN118366468A