Multi-channel Noise Data Simulation Method, Device, Equipment and Storage Medium
By collecting sound signals in real three-dimensional space, determining multi-path parameters and establishing target three-dimensional space, the problem of time-consuming and labor-consuming multi-channel noise data simulation is solved, and fast and efficient noise data simulation is achieved, and R&D efficiency is improved.
Patent Information
- Application Number
- CN202311525441.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-15
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2043-11-15
AI Technical Summary
In the prior art, the simulation process of multi-channel noise data is time-consuming and labor-intensive, and it is difficult to implement quickly and efficiently.
By acquiring sound signals in the real three-dimensional space based on the voice acquisition array, determining multipath parameters, establishing target three-dimensional space, and simulating multi-channel noise data using target simulation parameters.
It realizes fast and effective noise data simulation, saves R&D time and manpower, and improves R&D efficiency.
Smart Images

Figure CN117634157B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of signal processing, and in particular, to a multi-channel noise data simulation method, apparatus, device, and storage medium. Background Art
[0002] In current vehicle cockpits, a microphone array composed of multiple microphones is generally used for voice calls, speech recognition, etc. In the actual environment, it generally faces the influence of noise sources such as road noise, wind noise, and rain noise. When performing array signal processing and training neural networks, a large amount of multi-channel environmental noise is required as a training database. However, since collecting noise data using the microphone array in the vehicle cockpit requires a large amount of time and human resources.
[0003] Therefore, how to quickly and efficiently implement multi-channel noise data simulation is an urgent problem to be solved currently. Summary of the Invention
[0004] This application provides a multi-channel noise data simulation method, apparatus, device, and storage medium, aiming to at least solve one of the technical problems in the related art to some extent.
[0005] In a first aspect, this application provides a multi-channel noise data simulation method, including:
[0006] Collecting a first sound signal in a real three-dimensional space based on a voice collection array, where the real three-dimensional space includes at least one real sound source device;
[0007] Determining multipath parameters between the voice collection array and each real sound source device according to the first sound signal;
[0008] Establishing a target three-dimensional space according to the real three-dimensional space, where the target three-dimensional space is used to simulate multi-channel noise data;
[0009] Determining target simulation parameters related to the target three-dimensional space according to multiple multipath parameters; and
[0010] Determining multi-channel noise data of the real three-dimensional space based on the target simulation parameters.
[0011] In a second aspect, this application provides a multi-channel noise data simulation apparatus, including:
[0012] An acquisition module, configured to collect a first sound signal in a real three-dimensional space based on a voice collection array, where the real three-dimensional space includes at least one real sound source device;
[0013] A first determination module, configured to determine multipath parameters between the voice acquisition array and each of the real sound source devices according to the first sound signal;
[0014] A building module, configured to build a target three-dimensional space according to the real three-dimensional space, where the target three-dimensional space is used to simulate multi-channel noise data;
[0015] A second determination module, configured to determine target simulation parameters related to the target three-dimensional space according to the plurality of multipath parameters; and
[0016] A third determination module, configured to determine multi-channel noise data of the real three-dimensional space based on the target simulation parameters.
[0017] In a third aspect, the present application provides an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement a multi-channel noise data simulation method.
[0018] In a fourth aspect, the present application provides a computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute a multi-channel noise data simulation method.
[0019] In a fifth aspect, the present application provides a computer program product, including a computer program, and the computer program is executed by a processor to perform a multi-channel noise data simulation method.
[0020] The multi-channel noise data simulation method, device, equipment and storage medium provided by the present application first collect a first sound signal in a real three-dimensional space based on a voice acquisition array, where the real three-dimensional space includes at least one real sound source device, and then determine multipath parameters between the voice acquisition array and each real sound source device according to the first sound signal, and then build a target three-dimensional space according to the real three-dimensional space, where the target three-dimensional space is used to simulate multi-channel noise data, and then determine target simulation parameters related to the target three-dimensional space according to the plurality of multipath parameters; and determine multi-channel noise data of the real three-dimensional space based on the target simulation parameters. Thus, based on the real three-dimensional space, the target three-dimensional space can be simulated, and the target three-dimensional space and the multipath parameters can be used to simulate noise data, realizing fast and effective noise data simulation, saving R & D time, improving R & D efficiency, and saving the time and manpower for actually recording noise.
[0021] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. Description of the Drawings
[0022] The accompanying drawings here are incorporated into and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0023] Figure 1 It is a schematic flowchart of a multi-channel noise data simulation method shown according to the first embodiment of this application;
[0024] Figure 2 It is a schematic diagram of the setting of a real sound source device in a cockpit shown according to the first embodiment of this application;
[0025] Figure 3 It is a schematic diagram of estimated multipath parameters shown according to the first embodiment of this application;
[0026] Figure 4 It is a schematic flowchart of a multi-channel noise data simulation method shown according to the second embodiment of this application;
[0027] Figure 5 It is a schematic diagram of a frequency-domain smoothing window shown according to the second embodiment of this application;
[0028] Figure 6 It is a schematic diagram of a two-dimensional matrix corresponding to the frequency-domain smoothing window shown according to the second embodiment of this application;
[0029] Figure 7 It is a schematic diagram of a sound source signal shown according to the second embodiment of this application;
[0030] Figure 8 It is a schematic diagram of an optimized dictionary shown according to the second embodiment of this application;
[0031] Figure 9 It is a schematic flowchart of a multi-channel noise data simulation method shown according to the third embodiment of this application;
[0032] Figure 10 It is a schematic diagram of a target three-dimensional space shown according to the third embodiment of this application;
[0033] Figure 11 It is a schematic diagram of the setting of a real sound source device in a cockpit shown according to the third embodiment of this application;
[0034] Figure 12 It is an impulse response diagram shown according to the third embodiment of this application;
[0035] Figure 13 It is a block diagram of a multi-channel noise data simulation device shown according to this application;
[0036] Figure 14 It shows a block diagram of an exemplary electronic device suitable for implementing the embodiments of this application.
[0037] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and there will be a more detailed description hereinafter. These drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments
[0038] The embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application, and should not be construed as a limitation of the present application. On the contrary, the embodiments of the present application include all changes, modifications and equivalents that fall within the spirit and scope of the appended claims.
[0039] It should be noted that the execution subject of the multi-channel noise data simulation method in this embodiment can be a multi-channel noise data simulation device, which can be implemented in software and / or hardware, and this device can be configured in an electronic device, and the electronic device can include but is not limited to a terminal, a server, a vehicle-mounted computer, etc., which are not limited herein.
[0040] Figure 1 is a schematic flowchart of a multi-channel noise data simulation method shown according to the first embodiment of the present application, as Figure 1 shown, the method includes:
[0041] S101: Collect a first sound signal in a real three-dimensional space based on a voice collection array, where the real three-dimensional space includes at least one real sound source device.
[0042] Among them, the real three-dimensional space refers to the three-dimensional space in the real environment where the object is located, which is a way to describe the physical world. In this space, a spatial rectangular coordinate system can be established, and each point has three independent coordinate components: x, y, and z. These three coordinates can describe the position of a point in three directions.
[0043] In the embodiments of the present disclosure, the space inside the vehicle cockpit can be used as the real three-dimensional space, or it can also be any three-dimensional space such as a room, an aircraft cockpit, a ship's cabin, etc., which is not limited herein.
[0044] Among them, the real sound source device can be a sound source generating device in the real environment, which can generate a sound signal for experiments and tests. Among them, the real sound source device can be any sound-emitting device such as a speaker, a sound system, a diaphragm, a piezoelectric ceramic, a controlled sound source, etc. in the real environment, and it is not limited herein.
[0045] Among them, the real three-dimensional space includes at least one real sound source device, that is, one or more real sound source devices. For example, if the real three-dimensional space is a car cockpit, the real sound source device is a speaker. As Figure 2 shown, a sound source, that is, a speaker, can be respectively arranged at the 5 positions of 1, 2, 3, 4, and 5.
[0046] Optionally, sound signals such as white noise or swept-frequency signals can be played through the real sound source device, which is not limited herein.
[0047] Among them, the first sound signal can be the sound signal played by the real sound source device, such as white noise or swept-frequency signal, which is not limited herein.
[0048] Among them, the voice collection array can be a sensor array for collecting sound signals.
[0049] Optionally, the voice collection array can be a microphone array, which is a device composed of multiple microphones. Among them, each microphone can be arranged together in a specific geometric shape to form a whole for capturing and collecting sound signals, thereby enhancing the ability to locate and separate sound sources and improving the quality and clarity of voice collection. Compared with a single microphone, the voice collection array can achieve sound source localization through the time difference and sound pressure difference between microphones and suppress interference such as noise and echo.
[0050] Among them, the voice collection array can include a linear array, a circular array or a distributed array, which is not limited herein. In the embodiments of the present disclosure, the voice collection devices can be arranged in a non-uniform layout, which is not limited herein.
[0051] S102: Determine the multipath parameters between the voice collection array and each real sound source device according to the first sound signal.
[0052] Optionally, the voice collection array can include multiple voice collection devices, such as 2, 3, 4... N.
[0053] Among them, the voice collection device can be used to collect sound signals, such as a microphone.
[0054] It should be noted that the voice collection device can be used to capture and record sound signals. The voice collection device includes microphones, microphones, voice recognition devices, etc. Among them, a microphone is a device that records sound by converting sound waves into electrical signals and is composed of a vibration sensor, an amplifier and a processor. The microphone can capture sound signals at a specific position according to factors such as the sound direction and distance and output corresponding electrical signals.
[0055] Among them, the multipath parameters can include delay, amplitude, reverberation time, etc., which is not limited herein.
[0056] Specifically, multiple first signal models can be established based on the first sound signal. Then, based on the multiple first signal models, the inverse transformation of the cross-power spectral density function of the first sound signal received by the speech acquisition array can be performed to obtain the initial cross-correlation function related to the first sound signal. Subsequently, based on the initial cross-correlation function, the multipath parameters between the speech acquisition array and each real sound source device can be determined.
[0057] Among them, the first signal model is used to describe the sound signal received by the speech acquisition device from the real sound source device, and the transmission function is used to model the loss and distortion of the sound signal.
[0058] Among them, each first signal model is used to model the reception of the sound signal generated by the real sound source device by the corresponding speech acquisition device, and parameters such as the distance, direction, and sound intensity between the sound source and the microphone can be estimated, so as to better understand and process the speech signal.
[0059] It should be noted that the reception of the sound signal generated by the real sound source device by different speech acquisition devices is usually different, so the corresponding first signal models can also be different.
[0060] Next, the embodiments of the present disclosure will be described by taking 2 speech acquisition devices (speech acquisition device 1 and speech acquisition device 2) included in the speech acquisition array as an example.
[0061] For example, x 1 (t) = α 1 s(t - τ 1 ) + n 1 (t) (Equation 1) can be used as the first signal model corresponding to speech acquisition device 1, and x 2 (t) = α 2 s(t - τ 2 ) + n 2 (t) (Equation 2) can be used as the first signal model corresponding to speech acquisition device 2.
[0062] Among them, α 1 and α 2 are the attenuation factors corresponding to speech acquisition devices 1 and 2 respectively (the attenuation factor is caused by path and material absorption), x 1 (t) and x 2 (t) are the data received by speech acquisition devices 1 and 2 respectively, n 1 (t) and n 2 (t) are the ambient noises received by speech acquisition devices 1 and 2 respectively, τ 1 and τ 2 are the time delays of the same sound source signal reaching speech acquisition devices 1 and 2 respectively.
[0063] Optionally, a frequency-domain transformation may first be performed on each first signal model to obtain a second signal model, and then a cross-power spectral density function of the first sound signal may be determined based on the multiple second signal models.
[0064] For example, performing a frequency-domain transformation on the first signal models corresponding to Equation 1 and Equation 2, such as performing a discrete-time Fourier transform, may respectively obtain a second signal model corresponding to Equation 1: and a second signal model corresponding to Equation 2
[0065] Among them, the cross-power spectral density function is used to describe the frequency-domain mutual relationship between two signals, and it measures the amplitude and phase relationships of the two signals at different frequencies.
[0066] Specifically, the cross-power spectral density function may be obtained according to Equation 3 and Equation 4
[0067] Among them,
[0068] Furthermore, an inverse transformation may be performed on the cross-power spectral density function so as to obtain that is, the initial cross-correlation function.
[0069] Among them, the initial cross-correlation function may be a function directly obtained by performing an inverse transformation on the cross-power spectral density function.
[0070] Such as Figure 3 as shown, the amplitude value A corresponding to point B may be an estimated value of the amplitude in the multipath parameters, the time E corresponding to point C may be an estimated value of the time delay, the amplitude value after point F is approximately 0, and the time length corresponding to EF may be an estimated value of the reverberation time, which is not limited herein.
[0071] As a possible implementation manner, a cross-correlation function matrix may first be constructed based on the initial cross-correlation function, and then eigenvalue analysis may be performed on the cross-correlation function matrix based on the beamforming method to evaluate the multipath channel. Specifically, a beamformer may be constructed first, and then the beamformer may be used to weight the initial cross-correlation function corresponding to each pair of microphones, and the result of the weighted sum may be used as the beam output. Finally, the beam outputs are summed for all pairs of microphones to obtain the beam output of the entire signal. Furthermore, the sidelobes appearing in the output waveform after beamforming may be analyzed, and parameters such as multipath time delay and multipath amplitude may be estimated by analyzing the positions and sizes of the sidelobes.
[0072] Alternatively, the least squares method can also be used. First, based on the initial cross-correlation function, a cross-correlation function matrix is constructed. Then, by modeling the cross-correlation function matrix as a system of linear equations, the multipath parameters are obtained by minimizing the sum of squared errors. For example, assuming there are M voice acquisition devices, the size of the cross-correlation function matrix is M×M, and thus it can be mapped to a vector. Then, representing this vector as the sum of a multipath parameter vector and an interference noise vector, the multipath parameters are solved by minimizing the sum of the squares of the interference noise vector.
[0073] It should be noted that there are many methods for estimating multipath parameters through the initial cross-correlation function. For example, the system identification method can also be used, which will not be elaborated here.
[0074] S103: Establish a target three-dimensional space according to the real three-dimensional space.
[0075] Among them, the target three-dimensional space is used to simulate multi-channel noise data.
[0076] It should be noted that the target three-dimensional space can be a simulated virtual space constructed based on the spatial dimension parameters of the real three-dimensional space, and virtual sound source devices can be arranged at multiple set positions in the target three-dimensional space. The target three-dimensional space can be a simulated three-dimensional model of the physical shape of the real three-dimensional space, and a virtual sound source array is arranged.
[0077] For example, if the real three-dimensional space is the space of an automobile cockpit, and the space of an automobile cockpit is usually approximated as a frustum shape, then when constructing the target three-dimensional space, a frustum-shaped simulated three-dimensional model can be first constructed according to the relevant dimension parameters of the automobile cockpit space, and then a virtual sound source array is arranged in this simulated three-dimensional model to form the target three-dimensional space.
[0078] S104: Determine the target simulation parameters related to the target three-dimensional space according to multiple multipath parameters.
[0079] Among them, the target simulation parameters are used to simulate the reflection coefficient of the sound signal to the boundary of the target three-dimensional space.
[0080] In the embodiments of the present disclosure, the reflection coefficient can characterize the reflection ability of sound on the space boundary, and its value range can be between 0 and 1, specifically depending on the characteristics of the space boundary material and the transmission situation of the sound signal. Usually, the space boundary of the target three-dimensional space will reflect the sound signals generated at different positions to different degrees.
[0081] Specifically, if there are multiple virtual sound source devices, the first multipath parameter corresponding to each virtual sound source device can be determined, and then the target simulation parameter corresponding to each virtual sound source device can be determined according to the first multipath parameter.
[0082] Among them, the virtual sound source device corresponds to the real sound source device, and the sound signals generated by each real sound source device can be received by one or more voice collection devices. For example, the virtual sound source device A corresponds to the real sound source device A1, and the sound signal generated by the real sound source device A1 can be simultaneously received by the microphone array X (including microphone 1, microphone 2, and microphone 3).
[0083] Further, the device can select the multipath parameter corresponding to the real sound source device A1 and the microphone array X as the first multipath parameter K1 among multiple path parameters, and then can adjust the multipath parameter K2 of the virtual sound source device A and the virtual microphone array X1 in the target three-dimensional space so that K2 is consistent with K1.
[0084] Further, when adjusting K2 so that K2 is approximately the same as K1, it can be to control the difference between K2 and K1 to be less than a preset threshold. If the difference between K2 and K1 is less than the preset threshold, then it can be considered that K2 and K1 are consistent.
[0085] Among them, there can be multiple first multipath parameters, such as time delay, amplitude, and reverberation time. In the embodiments of the present disclosure, it can be to adjust each first multipath parameter to be consistent with the multipath parameters of the virtual sound source device and the virtual microphone array. For example, control the difference between the time delay corresponding to K2 and the time delay corresponding to K1, the amplitude corresponding to K2 and the amplitude corresponding to K1, and the reverberation time corresponding to K2 and the reverberation time corresponding to K1 to be less than the corresponding preset thresholds, then it can be determined that K2 and K1 are consistent.
[0086] Alternatively, it can also be to control the difference between the reverberation time corresponding to K2 and the reverberation time corresponding to K1 to be less than the preset threshold, then it can be determined that K2 and K1 are consistent. The same applies to time delay and amplitude, which will not be elaborated here.
[0087] Further, when the first multipath parameter is consistent with the multipath parameter of the virtual microphone array, the device can use the reflection coefficient of the sound signal obtained by simulation calculation to the boundary of the target three-dimensional space as the target simulation parameter.
[0088] Optionally, by using acoustic simulation software, such as using the mirror sound source model for simulation, the first multipath parameter can be adjusted to be consistent with the multipath parameter of the virtual microphone array, so as to obtain the target simulation parameter by simulation calculation.
[0089] S105: Determine the multi-channel noise data of the real three-dimensional space based on the target simulation parameter.
[0090] Specifically, first, the voice collection array can receive only the sound signal generated by any real sound source device, and this sound signal can be used as single-channel noise data. Then, based on the target simulation parameters, the impulse response of each voice collection device when receiving any real sound source device can be determined. As a possible implementation, based on a pre-constructed noise simulation model, and inputting the target simulation parameters, as well as the relevant parameters of the real three-dimensional space and the target three-dimensional space into this noise simulation model, the impulse response between the voice collection device and any real sound source device can thus be obtained.
[0091] Then, the impulse response corresponding to each voice collection device and the single-channel noise data can be convolved, so that the multi-channel noise data of each voice collection device and this any real sound source device in the real three-dimensional space can be obtained.
[0092] The multi-channel noise data simulation method, device, equipment and storage medium provided by this application first collect the first sound signal in the real three-dimensional space based on the voice collection array, where the real three-dimensional space includes at least one real sound source device. Then, according to the first sound signal, determine the multipath parameters between the voice collection array and each real sound source device. Then, according to the real three-dimensional space, establish a target three-dimensional space, where the target three-dimensional space is used to simulate multi-channel noise data. Then, according to multiple multipath parameters, determine the target simulation parameters related to the target three-dimensional space; and based on the target simulation parameters, determine the multi-channel noise data of the real three-dimensional space. Thus, based on the real three-dimensional space, the target three-dimensional space can be simulated, and the target three-dimensional space and the multipath parameters can be used to simulate the noise data, realizing fast and effective noise data simulation, saving the R & D time, improving the R & D efficiency, and saving the time and manpower for actually recording noise.
[0093] Figure 4 is a schematic flowchart of the multi-channel noise data simulation method shown in the second embodiment of this application, as Figure 4 shown, this method includes:
[0094] S201: Collect the first sound signal in the real three-dimensional space based on the voice collection array, where the real three-dimensional space includes at least one real sound source device.
[0095] S202: According to the first sound signal, establish multiple first signal models, where each first signal model is used to model the reception situation of the corresponding voice collection device for the sound signal generated by the real sound source device.
[0096] S203: According to multiple first signal models, perform an inverse transformation on the cross-power spectral density function of the first sound signal received by the voice collection array to obtain the initial cross-correlation function related to the first sound signal.
[0097] It should be noted that the specific implementation manners of steps S201 - S203 may refer to the above - mentioned embodiments, and will not be elaborated here.
[0098] S204: Perform frequency - domain segmented smoothing on the initial cross - correlation function to obtain the target cross - correlation function.
[0099] It should be noted that frequency - domain segmented smoothing can be used to reduce noise in the frequency domain or suppress unwanted frequency components. By dividing the signal into multiple frequency - domain segments and smoothing the signal within each segment, the purpose of reducing noise or removing unwanted frequency components can be achieved.
[0100] Optionally, the initial cross - correlation function can be first divided into multiple segments, each segment containing a certain number of spectral data. Then, the data within each frequency - domain segment is smoothed (methods such as moving average, weighted average, median filtering, etc. can be used), so that the target cross - correlation function can be obtained.
[0101] It should be noted that the effect of frequency - domain segmented smoothing is affected by the number of segments divided, the amount of data within each segment, and the selected smoothing method. Different parameter selections may produce different smoothing effects.
[0102] As a possible implementation manner, the smoothing window information can be first determined, where the smoothing window information includes: window width and window shift. Then, the initial cross - correlation function can be subjected to frequency - domain segmented smoothing according to the smoothing window information to obtain the target cross - correlation function.
[0103] In frequency - domain segmented smoothing, the window width can be the number of data within each frequency - domain segment, which affects the size of the frequency - domain segment, that is, how many spectral data points are included in each segment. A larger window width can provide better frequency resolution, and a smaller window width can provide higher frequency - domain resolution. The window shift can be the step size of each sliding window, and the window shift affects the overlapping degree between adjacent frequency - domain segments. A larger window shift can increase the calculation efficiency, and a smaller window shift can provide better smoothing effects.
[0104] In the embodiments of the present disclosure, the window width and window shift can be parameter values determined by reasonable selection in advance, and will not be limited here.
[0105] Among them, the initial cross - correlation function is
[0106] Specifically, the frequency domain can be divided into L segments, Each segment corresponds to a frame in the time domain. Among them, the window width and window shift can be preset values. For example, the window width B Φ can be 128, and the window shift M Φ can be 64. The examples here are only for illustrative purposes and are not limitations.
[0107] Among them, the target cross-correlation function can be the following formula:
[0108]
[0109] Among them, Φ(ω) represents a symmetric smoothing window, L represents the number of segments obtained by smoothing and segmenting the frequency domain, l represents any one of the L segments, and τ represents the time delay.
[0110] As Figure 5 shown Figure 5 is a schematic diagram of a frequency-domain smoothing window.
[0111] S205: Determine the multipath parameters according to the target cross-correlation function.
[0112] Specifically, the eigenvector corresponding to the maximum eigenvalue of the target cross-correlation function can be determined first, and then the multipath parameters can be determined based on the eigenvector.
[0113] Furthermore, when determining the time-delay estimate value, the first parameter that makes the norm of the eigenvector the largest can be obtained first, and then the time-delay estimate value can be determined based on the first parameter, the eigenvector, and a pre-constructed time-delay estimation model.
[0114] Specifically, a two-dimensional matrix can be determined first according to the target cross-correlation function, hereinafter denoted as This two-dimensional matrix is the stacking of cross-correlation functions corresponding to different segments in the frequency domain, Figure 6 is the image of its absolute value, which shows the two-dimensional matrix corresponding to the frequency-domain smoothing window.
[0115] After that, can be subjected to singular value decomposition, so that a diagonal matrix (denoted as S) can be obtained. Among them, the eigenvalues of the diagonal matrix are arranged in descending order on the diagonal of S. After that, the first eigenvector of the diagonal matrix can be extracted, denoted as S1. Among them, S1 is the eigenvector with the largest eigenvalue and is the principal eigenvector of the matrix
[0116] Furthermore, based on the formula: β = argmax|S1|, the value β (the first parameter) that makes the norm of the eigenvector S1 the largest can be calculated, and then the real part of the eigenvector S1 can be extracted, denoted as real(S1). After that, the sign function sign() is applied to real(S1) to obtain the sign sequence of real(S1). Multiply real(S1) by its corresponding sign sequence to obtain a piecewise-smooth matched filter estimate value, denoted as R fs_mf , finally, the first parameter β can be used as an index to determine the time delay corresponding to β as the estimated time delay, which is used to estimate the true time delay of the direct wave.
[0117] Among them, the pre-constructed time delay estimation model can be R fs_mf = real(S1)·sign(S1(β)).
[0118] Optionally, the inner product operation can be performed based on the eigenvector S1 and the waveform of the signal source, so as to obtain the amplitude estimation value of the direct wave.
[0119] Thus, the accuracy and precision of the time delay estimation can be improved through the low-rank approximation of the matrix and the extraction of the main eigenvector. Among them, the operations of using the sign function and element-wise multiplication can smoothly estimate the response of the matched filter, avoid jumps in time, and improve the accuracy of the estimation result.
[0120] It should be noted that the above calculation method can be an estimation calculation of the multipath parameters of the direct wave. The direct wave is the sound wave directly received by the voice acquisition array from the real sound source device without acoustic wave reflection. The multipath parameter estimation in the case of reverberant waves will be described below. The reverberant wave is the reflected sound wave generated by phenomena such as reflection, refraction, and scattering of sound in an enclosed space.
[0121] Optionally, if it is necessary to estimate each multipath parameter corresponding to the reverberant wave, the following formula can be used for calculation:
[0122] Equation 1: minξ 1 + ξ 2 + ξ 3 ;
[0123] Equation 2: ||α|| 1 ≤ ξ 2 , ||α|| 2 ≤ ξ 3 .
[0124] Among them, X is the first sound signal, that is, the received data of the voice acquisition device, ξ 1 , ξ 2 , ξ 3 are intermediate variables, ||·|| 1 represents the 1-norm, ||·|| 2 represents the 2-norm, α is the estimated multipath parameter, and S is the optimized dictionary, which is composed of the sound source signal through different time delay combinations. As Figure 7 shown, Figure 7 is a schematic diagram of a sound source signal.
[0125] As Figure 8 shown, Figure 8It is a schematic diagram of an optimized dictionary. Among them, the first column is obtained by delaying the sound source signal by a time delay τ est , τ est can be determined according to step S205, Figure 8 The "padding with zeros" in is because the entire possible range of multipath time delays needs to cover the true multipath time delay. For example, if the reverberation time is 200 ms, the time delay range covered by the optimized dictionary S should be greater than 200 ms, that is, the length of "padding with zeros" should be greater than 200 ms.
[0126] S206: Establish a target three-dimensional space according to the real three-dimensional space, where the target three-dimensional space is used to simulate multi-channel noise data;
[0127] S207: Determine target simulation parameters related to the target three-dimensional space according to multiple multipath parameters; and
[0128] S208: Determine the multi-channel noise data of the real three-dimensional space based on the target simulation parameters.
[0129] It should be noted that the specific implementation manners of steps S206 - S208 can refer to the above embodiments and will not be elaborated here.
[0130] In the embodiments of the present disclosure, first, a first sound signal in the real three-dimensional space is collected based on a voice acquisition array, then multiple first signal models are established according to the first sound signal, and then, according to the multiple first signal models, an inverse transform is performed on the cross-power spectral density function of the first sound signal received by the voice acquisition array to obtain an initial cross-correlation function related to the first sound signal. After that, frequency-domain segmented smoothing processing is performed on the initial cross-correlation function to obtain a target cross-correlation function. According to the target cross-correlation function, multipath parameters are determined, then a target three-dimensional space is established according to the real three-dimensional space, and then, according to the multiple multipath parameters, target simulation parameters related to the target three-dimensional space are determined. Finally, multi-channel noise data of the real three-dimensional space are determined based on the target simulation parameters. Thus, frequency-domain segmented smoothing processing can be used to reduce noise in the frequency domain or suppress unwanted frequency components. By simulating noise data using a single-channel noise source and acoustic path parameters, various multi-channel noise data can be simulated, achieving fast noise data simulation, and saving the time and manpower for actual noise recording.
[0131] Figure 9 is a schematic flowchart of a multi-channel noise data simulation method shown in the third embodiment of the present application. As Figure 9 shown, the method includes:
[0132] S301: Collect a first sound signal in the real three-dimensional space based on a voice acquisition array, where the real three-dimensional space includes at least one real sound source device.
[0133] S302: Determine the multipath parameters between the voice collection array and each real sound source device according to the first voice signal.
[0134] It should be noted that the specific implementation methods of steps S301 and S302 can refer to the above embodiments and will not be elaborated here.
[0135] S303: Determine the spatial dimension parameters of the real three-dimensional space.
[0136] Among them, the spatial dimension parameters can be boundary parameters such as the length, width, and height of the real three-dimensional space, which are not limited here.
[0137] For example, if the real three-dimensional space is the vehicle cockpit space, when determining the three-dimensional space dimension parameters of the vehicle cockpit space, some key parameters related to the vehicle cockpit space need to be considered. For example, they can be the following types:
[0138] Length: The distance in the front-back direction of the cockpit space, usually based on the vehicle driving direction, from the front to the back.
[0139] Width: The distance in the left-right direction of the cockpit space, usually the width perpendicular to the vehicle driving direction.
[0140] Height: The distance in the up-down direction of the cockpit space, usually referring to the height from the floor to the roof.
[0141] Specifically, these dimension parameters can be determined by measuring the actual dimensions of the vehicle cockpit space. Or, they can also be determined based on the vehicle's specifications and technical parameters, which contain the dimension information of the cockpit space.
[0142] It should be noted that the spatial dimension parameters of the cockpit space may vary due to factors such as vehicle models, vehicle uses, and design styles.
[0143] S304: Establish an initial three-dimensional space according to the spatial dimension parameters.
[0144] Among them, the initial three-dimensional space can be a simulated virtual space constructed based on the spatial dimension parameters of the real three-dimensional space. The initial three-dimensional space can be a simulation three-dimensional model of the physical shape of the real three-dimensional space.
[0145] For example, if the real three-dimensional space is the car cockpit space, usually the car cockpit space is approximately frustum-shaped, then an inscribed cuboid with a frustum shape can be constructed based on the spatial dimension parameters corresponding to the car cockpit space, as Figure 10 shown.
[0146] It should be noted that the above example is only an illustrative explanation and does not limit the present disclosure.
[0147] S305: Determine the simulation type of the noise data, and determine the target sound source position according to the simulation type, where the target sound source position is used to indicate the setting position of the virtual sound source device.
[0148] Among them, the virtual sound source device may be a virtual sound source device, such as a virtual speaker, a virtual audio, etc., which is not limited here.
[0149] It should be noted that when determining the simulation type of the noise data and the target sound source position, the specific application scenario and requirements need to be considered. For different types of noise data, there are corresponding different simulation types. For example, in the field of vehicles or transportation, common noise types include vibration noise, wind noise, mechanical noise, road noise, etc., so the simulation types can also be vibration noise, wind noise, mechanical noise, road noise, etc., which are not limited here.
[0150] In addition, for some specific types of noise data, they are usually generated by sound sources at specific positions. Therefore, in the embodiments of the present disclosure, considering this situation, before arranging the sound source device in the initial three-dimensional space, first, according to the simulation type of the noise data, determine the respective setting positions corresponding to the virtual sound source devices of the noise data of each simulation type, that is, determine each target sound source position.
[0151] For example, there are 2 simulation types of noise data, namely type A and type B. For the virtual sound source device a that can generate type A noise data, it can be set at positions a1 and a2 in the initial three-dimensional space. For the virtual sound source device b that can generate type B noise data, it can be set at positions b1 and b2 in the initial three-dimensional space, which is not limited here.
[0152] Among them, positions a1 and a2 are the target sound source positions associated with simulation type A, and positions b1 and b2 are the target sound source positions associated with simulation type B.
[0153] For example, for the simulation of vibration noise, it is necessary to determine the specific position of setting the vibration noise source according to the vehicle structure and vibration characteristics. Specifically, the vibration information of each vehicle component can be obtained through actual measurement or simulation analysis, and the position of the noise source can be determined according to the vibration transmission path. For example, in the vehicle chassis system, the vibration situation at the rear wheel vehicle joint can be estimated based on the rotational speed and steering angle information of the drive shaft, and the position of the vibration noise source can be determined accordingly.
[0154] As Figure 2 shown, Figure 2 is a schematic diagram of the sound source position of the chassis proposed in the embodiments of the present disclosure, Figure 2 where 1, 2, 3, 4, and 5 in it are respectively the vibration noise sources set on the vehicle chassis.
[0155] For example, for the simulation of wind noise, factors such as the external air flow of the vehicle and the shape and size of components such as window glass need to be considered. Specifically, computational fluid dynamics analysis tools can be used to simulate the external air flow field of the vehicle, and based on this, the location of the wind noise source can be determined. Additionally, simulation software can be used to simulate the vibration characteristics of components such as window glass, and based on the vibration transmission path, the location of the noise source can be determined. When determining the location of the target sound source, the location and direction of the sound source need to be considered. Specifically, it can be determined according to the actual noise source location inside the vehicle or through simulation analysis. For example, when simulating multi-channel wind noise, virtual noise sources can be arranged according to the position of the window frame and the position of the window center, and the sound level and frequency characteristics of the wind noise can be determined based on parameters such as vehicle speed and wind speed.
[0156] As Figure 11 shown, Figure 11 FIG. is a schematic diagram of the position of wind noise proposed by an embodiment of the present disclosure, Figure 11 where 1, 2, 3, 4, 5, and 6 in are respectively virtual noise sources arranged at the window positions.
[0157] It should be noted that when determining the simulation type of the noise data and the location of the target sound source, it needs to match the actual application scenario and requirements to ensure the reliability and accuracy of the simulation results.
[0158] S306: Set a virtual sound source device at the position indicated by the target sound source location in the initial three-dimensional space to establish a target three-dimensional space.
[0159] Among them, the target three-dimensional space can be the virtual three-dimensional space after arranging the virtual sound source device in the initial three-dimensional space.
[0160] Optionally, corresponding virtual voice collection devices can also be arranged in the target three-dimensional space, that is, a virtual microphone array can be arranged.
[0161] It should be noted that after determining the target sound source location of the virtual sound source device in the initial three-dimensional space, the target sound source location can be represented using a coordinate system or the point coordinates in the 3D space, and the virtual sound source device can be set at the corresponding target sound source location.
[0162] S307: Determine the real sound source device corresponding to the virtual sound source device from at least one real sound source device.
[0163] It should be noted that each virtual sound source device in the target three-dimensional space corresponds to a real sound source device in the real three-dimensional space.
[0164] Specifically, for example, if it is necessary to determine the target simulation parameters corresponding to the virtual sound source device A1 currently, it is necessary to first determine the real sound source device A2 corresponding to the virtual sound source device A1 from at least one real sound source device.
[0165] S308: Determine the first multipath parameter between the corresponding real sound source device and the voice acquisition array from multiple multipath parameters.
[0166] Among them, the first multipath parameter can be the multipath parameter corresponding to any virtual sound source device, and this multipath parameter can be the multipath parameter between the real sound source device corresponding to any virtual sound source device and the voice acquisition array.
[0167] For example, if there are 3 real sound source devices, namely S1, S2, and S3. Among them, the multipath parameters between S1, S2, S3 and the voice acquisition array are y1, y2, and y3 respectively. If the virtual sound source devices corresponding to S1, S2, and S3 are x1, x2, and x3 respectively, then the multipath parameters corresponding to x1, x2, and x3 are y1, y2, and y3 respectively. That is, y1 is the first multipath parameter corresponding to x1, y2 is the first multipath parameter corresponding to x2, and y3 is the first multipath parameter corresponding to x3, which is not limited here.
[0168] S309: Determine the target simulation parameters related to the target three-dimensional space according to the first multipath parameter.
[0169] Among them, the target simulation parameter can be the reflection coefficient related to the spatial boundary of the target three-dimensional space.
[0170] In the embodiments of the present disclosure, the reflection coefficient can characterize the reflection ability of sound on the spatial boundary, and its value range can be between 0 and 1, specifically depending on the characteristics of the spatial boundary material and the transmission situation of the sound signal. Usually, the spatial boundary of the target three-dimensional space will reflect the sound signals generated at different positions to different degrees.
[0171] As a possible implementation manner, an acoustic simulation software can be used. For example, the mirror sound source model is used for simulation to adjust the multipath time delay and attenuation amplitude of the target three-dimensional space to make it as close as possible to the actual value, that is, to meet the first multipath parameter. It can be understood that this process can be iterated and adjusted multiple times until the error value between the multipath parameter corresponding to the target three-dimensional space and the first multipath parameter is less than the preset threshold. At this time, the reflection coefficient related to the spatial boundary of the target three-dimensional space can be used as the target simulation parameter.
[0172] As another possible implementation, the second multipath parameter can be determined first according to the first multipath parameter and a preset perturbation parameter, and then the third multipath parameter in the target three-dimensional space is adjusted until the difference between the third multipath parameter and the second multipath parameter meets the preset condition, and the target simulation parameter related to the target three-dimensional space is determined.
[0173] It should be noted that, in order to expand the dataset and cover the actual impulse response, in the embodiments of the present disclosure, a certain perturbation can be set for the multipath parameter. Among them, the perturbation parameter can be a preset value determined according to actual experience, such as 20% or 30%, which is not limited herein.
[0174] For example, a 20% perturbation can be set for the time delay in the first multipath parameter, and a 30% perturbation can be set for the attenuation amplitude in the first multipath parameter, which is not limited herein.
[0175] Among them, the second multipath parameter can be an extended multipath parameter determined after giving a certain perturbation to the first multipath parameter.
[0176] Among them, the third multipath parameter can be a simulated multipath parameter.
[0177] It should be noted that by adjusting the simulated third multipath parameter to be close to the second multipath parameter until the difference between the third multipath parameter and the second multipath parameter is less than the preset threshold, it can be considered that the error value between the third multipath parameter and the second multipath parameter is small and basically the same. At this time, the reflection coefficient related to the spatial boundary of the target three-dimensional space can be used as the target simulation parameter.
[0178] Optionally, if the difference between the third multipath parameter and the second multipath parameter is less than the preset threshold, it can be determined that the difference between the third multipath parameter and the second multipath parameter meets the preset condition. Among them, the preset threshold can be determined according to experience and is not limited herein.
[0179] For example, if the multipath parameter is the reverberation time, the preset threshold is 1.5 ms, the reverberation time of the third multipath parameter is 39 ms, and the reverberation time of the second multipath parameter is 40 ms, it can be considered that the third multipath parameter and the second multipath parameter are basically the same, and the difference meets the preset condition, that is, less than 1.5 ms.
[0180] S310: Determine the single-channel noise data of the real three-dimensional space.
[0181] Among them, the single-channel noise data is the noise signal data of one channel, representing the noise response at a specific position or direction. For example, if only one accelerometer is installed in a mechanical device, then the device measures the vibration signal at a specific position during machine operation, so single-channel vibration noise data is obtained.
[0182] Optionally, each voice collection device can be used to collect the sound signal of any real sound source device, so as to obtain single-channel noise data in the real three-dimensional space. For example, if a swept-frequency signal W is emitted by any real sound source device, and the noise data received by microphone 1 and microphone 2 are W1 and W2 respectively, then W1 is the single-channel noise data received by microphone 1, and W2 is the single-channel noise data received by microphone 2.
[0183] S311: Determine the impulse response in the real three-dimensional space based on the target simulation parameters.
[0184] Among them, the impulse response can usually be expressed as a function with time as the independent variable and spatial position as the parameter.
[0185] Optionally, based on a preset noise simulation model, the impulse response in the real three-dimensional space can be determined according to the target simulation parameters, the target three-dimensional space, and the real three-dimensional space.
[0186] Among them, the preset noise simulation model can be a pre-constructed mathematical formula for determining the amplitude corresponding to each sampling point. In the embodiments of the present disclosure, the pre-constructed noise simulation model can be the first formula or the second formula.
[0187] As a possible implementation, the amplitude corresponding to each sampling point can be calculated through the following first formula, so as to determine the impulse response corresponding to each voice collection device when receiving the sound signal of any sound source.
[0188] Among them, the first formula is:
[0189]
[0190] Among them, p=(q,j,k) is a three-element combination, and each element can take a value of 0 or 1, so as to form a set P={(q,j,k):q,j,k∈{0,1}}.
[0191] It should be noted that when each element (q, j, k) of p is 1, it means that the mirror image in that direction is included in the calculation. Since there are multiple reflected mirror images of the sound source, in order to include the multiple reflected mirror images in the calculation, the parameter R m =[2m x L x ,2m y L y ,2m z L z .
[0192] Among them, Lx, L y , Lz are the lengths of the real three-dimensional space in the x, y, and z directions, and m=(m x,m y ,m z ) is a three - element combination, m x ,m y ,m z ∈{-N,N}, representing all possible reflections, N is the number of sampling points, r = [x, y, z] is the position of the voice acquisition array in the real three - dimensional space, x, y, and z respectively represent the horizontal, vertical, and vertical coordinates, r s = [x s ,y s ,z s is the position of the target sound source, n represents the nth sampling point, f s is the sampling frequency.
[0193] Among them, the position from the virtual sound source device to the virtual voice acquisition device is expressed as d = ||R p + R m ||, R p = [x s - x + 2qx, y s - y + 2jy, z s - z + 2kz] represents the position of the target sound source (q, j, k are elements in the above p), ||·|| represents the modulus operation, is the time delay, c is the speed of sound, β represents the target simulation parameter, LFP{·} is a low - pass filter.
[0194] As another possible implementation method, the amplitude corresponding to each sampling point can be calculated through the following second formula, so as to determine the impulse response corresponding to each voice acquisition device when receiving the sound signal of any sound source.
[0195] Among them, the second formula is
[0196] It should be noted that the impulse function δ(t - τ) in the second formula can be replaced by δ LPF (t), and the meanings of the various parameters in the second formula can refer to the above first formula.
[0197] Among them, T w represents the signal width, f c represents the cut - off frequency of the low - pass filter.
[0198] As Figure 12 shown, the impulse response diagrams of four microphones (voice acquisition devices) and a sound source (real sound source device) are shown.
[0199] S312: Determine the multi - channel noise data according to the single - channel noise data and the impulse response in the real three - dimensional space.
[0200] Among them, the multi-channel noise data can be noise signal data collected from different positions or directions in a certain system or environment, usually including multiple channels. The data of each channel represents the noise response of the system at different positions or directions.
[0201] As a possible implementation method, convolution operation can be used to convolve the single-channel noise data with the impulse response of the real three-dimensional space to obtain multi-channel noise data.
[0202] Specifically, convolution operation can be used to convolve the single-channel noise data with the impulse response. In the time domain, convolution can be transformed into the frequency domain to improve the calculation efficiency, then perform the multiplication operation in the frequency domain, and inverse-transform the result back to the time domain.
[0203] In the embodiments of the present disclosure, first, a first sound signal in the real three-dimensional space is collected based on a voice collection array. Then, according to the first sound signal, the multipath parameters between the voice collection array and each real sound source device can be determined. Then, the spatial size parameters of the real three-dimensional space are determined. According to the spatial size parameters, an initial three-dimensional space is established. Then, the simulation type of the noise data can be determined, and the target sound source position can be determined according to the simulation type. Then, a virtual sound source device is set at the position indicated by the target sound source position in the initial three-dimensional space to establish a target three-dimensional space. The real sound source device corresponding to the virtual sound source device is determined from at least one real sound source device. From multiple multipath parameters, the first multipath parameter between the corresponding real sound source device and the voice collection array is determined. According to the first multipath parameter, the target simulation parameters related to the target three-dimensional space are determined. Then, the single-channel noise data of the real three-dimensional space is determined. Then, according to the target three-dimensional space and the target simulation parameters, the impulse response of the real three-dimensional space is determined. Finally, according to the single-channel noise data and the impulse response of the real three-dimensional space, the multi-channel noise data is determined. Thus, by collecting the first sound signal in the real three-dimensional space based on the voice collection array and combining the simulation method, the generation of multi-channel noise data can be realized. The multi-channel noise in the real environment can be simulated, and more accurate three-dimensional space information can be obtained. This is very helpful for various application scenarios that require the use of multi-channel data, such as sound source localization, speech enhancement, intelligent microphone arrays, etc.
[0204] In some scenarios, such as in application platforms like mobile phones, headphones, speakers, and automotive cockpits, it is often necessary to estimate the multipath parameters of a sound source or the impulse response of a channel. The higher the signal-to-noise ratio, the higher the parameter estimation accuracy. However, in daily life, different types of noise often appear in different frequency bands, resulting in a low signal-to-noise ratio in local frequency bands and affecting the accuracy of impulse response parameter estimation. The commonly used cross-correlation algorithm or matched filtering algorithm in the field of signal processing estimates parameters through full-band information, and this estimation method has low accuracy when the signal-to-noise ratio is low. This application combines frequency-domain segmented matched filtering and matrix singular value decomposition to achieve a higher-precision channel impulse response estimation algorithm.
[0205] Figure 13 is a block diagram of a multi-channel noise data simulation device shown according to this application, as Figure 13 shown. The multi-channel noise data simulation device 400 includes:
[0206] An acquisition module 410, configured to acquire a first sound signal in a real three-dimensional space based on a voice acquisition array, where the real three-dimensional space includes at least one real sound source device;
[0207] A first determination module 420, configured to determine the multipath parameters between the voice acquisition array and each real sound source device according to the first sound signal;
[0208] A building module 430, configured to build a target three-dimensional space according to the real three-dimensional space, where the target three-dimensional space is used to simulate multi-channel noise data;
[0209] A second determination module 440, configured to determine target simulation parameters related to the target three-dimensional space according to multiple multipath parameters; and
[0210] A third determination module 450, configured to determine the multi-channel noise data of the real three-dimensional space based on the target simulation parameters.
[0211] Optionally, the building module is specifically configured to:
[0212] Determine the spatial dimension parameters of the real three-dimensional space;
[0213] Build an initial three-dimensional space according to the spatial dimension parameters;
[0214] Determine the simulation type of the noise data, and determine the target sound source position according to the simulation type, where the target sound source position is used to indicate the setting position of a virtual sound source device;
[0215] Set a virtual sound source device at the position indicated by the target sound source position in the initial three-dimensional space to build the target three-dimensional space.
[0216] Optionally, the second determination module includes:
[0217] A first determination unit, configured to determine a real sound source device corresponding to the virtual sound source device from at least one real sound source device;
[0218] A second determination unit, configured to determine a first multipath parameter between the corresponding real sound source device and the voice acquisition array from multiple multipath parameters;
[0219] A third determination unit, configured to determine a target simulation parameter related to the target three-dimensional space based on the first multipath parameter.
[0220] Optionally, the third determination unit is specifically configured to:
[0221] Determine a second multipath parameter according to the first multipath parameter and a preset perturbation parameter;
[0222] Adjust the third multipath parameter of the target three-dimensional space until the difference between the third multipath parameter and the second multipath parameter meets a preset condition, and determine a target simulation parameter related to the target three-dimensional space,
[0223] wherein the target simulation parameter is a reflection coefficient related to the spatial boundary of the target three-dimensional space.
[0224] Optionally, the third determination module includes:
[0225] A fourth determination unit, configured to determine single-channel noise data of the real three-dimensional space;
[0226] A fifth determination unit, configured to determine an impulse response of the real three-dimensional space based on the target simulation parameter;
[0227] A sixth determination unit, configured to determine the multi-channel noise data according to the single-channel noise data and the impulse response of the real three-dimensional space.
[0228] Optionally, the fifth determination unit is specifically configured to:
[0229] Based on a preset noise simulation model, determine an impulse response of the real three-dimensional space according to the target simulation parameter, the target three-dimensional space, and the real three-dimensional space.
[0230] Optionally, the voice acquisition array includes: a plurality of voice acquisition devices; wherein, the first determination module includes:
[0231] A seventh determination unit, configured to establish a plurality of first signal models according to the first sound signal, where each of the first signal models is used to model the reception of the sound signal generated by the real sound source device by the corresponding voice collection device;
[0232] An eighth determination unit, configured to perform an inverse transformation on the cross-power spectral density function of the first sound signal received by the voice collection array according to the plurality of first signal models, so as to obtain an initial cross-correlation function related to the first sound signal;
[0233] A ninth determination unit, configured to determine multipath parameters between the voice collection array and each of the real sound source devices according to the initial cross-correlation function.
[0234] Optionally, the eighth determination unit is further configured to:
[0235] Perform a frequency-domain transformation on each of the first signal models to obtain a second signal model;
[0236] Determine the cross-power spectral density function of the first sound signal according to the plurality of second signal models.
[0237] Optionally, the ninth determination unit includes:
[0238] An acquisition unit, configured to perform frequency-domain segmented smoothing processing on the initial cross-correlation function to obtain a target cross-correlation function;
[0239] A tenth determination unit, configured to determine the multipath parameters according to the target cross-correlation function.
[0240] Optionally, the acquisition unit is specifically configured to:
[0241] Determine smoothing window information, where the smoothing window information includes: window width and window shift;
[0242] Perform frequency-domain segmented smoothing processing on the initial cross-correlation function according to the smoothing window information to obtain the target cross-correlation function.
[0243] Optionally, the tenth determination unit includes:
[0244] An eleventh determination unit, configured to determine an eigenvector corresponding to the maximum eigenvalue of the target cross-correlation function;
[0245] A twelfth determination unit, configured to determine the multipath parameters based on the eigenvector.
[0246] Optionally, the twelfth determination unit is specifically configured to:
[0247] Obtain a first parameter that maximizes the norm of the eigenvector;
[0248] Determine a delay estimation value based on the first parameter, the feature vector, and a pre-constructed delay estimation model.
[0249] The multi-channel noise data simulation method, device, equipment, and storage medium provided in this application first collect a first sound signal in a real three-dimensional space based on a voice acquisition array, where the real three-dimensional space includes at least one real sound source device. Then, according to the first sound signal, determine the multipath parameters between the voice acquisition array and each real sound source device. Then, based on the real three-dimensional space, establish a target three-dimensional space, where the target three-dimensional space is used to simulate multi-channel noise data. Then, according to multiple multipath parameters, determine target simulation parameters related to the target three-dimensional space; and based on the target simulation parameters, determine the multi-channel noise data of the real three-dimensional space. Thus, based on the real three-dimensional space, a target three-dimensional space can be simulated, and the target three-dimensional space and multipath parameters can be used to simulate noise data, achieving fast and effective noise data simulation, saving R & D time, improving R & D efficiency, and saving the time and manpower for actual recording of noise.
[0250] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium, and a computer program product.
[0251] Figure 14 It is a block diagram of an electronic device shown according to the present application. For example, the electronic device 600 can be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0252] Referring to Figure 14 , the electronic device 600 may include one or more of the following components: a processing component 602, a memory 604, a power component 606, a multimedia component 608, an audio component 610, an input / output (I / O) interface 612, a sensor component 614, and a communication component 616.
[0253] The processing component 602 generally controls the overall operation of the electronic device 600, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate the interaction between the multimedia component 608 and the processing component 602.
[0254] The memory 604 is configured to store various types of data to support the operation of the electronic device 600. Examples of such data include instructions for any application or method operating on the electronic device 600, contact data, phone book data, messages, pictures, videos, and the like. The memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0255] The power supply component 606 provides power to various components of the electronic device 600. The power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 600.
[0256] The multimedia component 608 includes a touch display screen that provides an output interface between the electronic device 600 and the user. In some embodiments, the touch display screen may include a liquid crystal display (LCD) and a touch panel (TP). The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the electronic device 600 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0257] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 604 or transmitted via the communication component 616.
[0258] In some embodiments, the audio component 610 further includes a speaker for outputting audio signals.
[0259] The I / O interface 612 provides an interface between the processing component 602 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a start button, and a lock button.
[0260] The sensor assembly 614 includes one or more sensors for providing a status assessment of various aspects of the electronic device 600. For example, the sensor assembly 614 can detect the on / off state of the electronic device 600, the relative positioning of components, such as the display and keypad of the electronic device 600. The sensor assembly 614 can also detect a change in the position of the electronic device 600 or a component of the electronic device 600, the presence or absence of user contact with the electronic device 600, the orientation or acceleration / deceleration of the electronic device 600, and a change in the temperature of the electronic device 600. The sensor assembly 614 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 614 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 614 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0261] The communication component 616 is configured to facilitate communication between the electronic device 600 and other devices in a wired or wireless manner. The electronic device 600 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 616 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0262] In an exemplary embodiment, the electronic device 600 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-mentioned multi-channel noise data simulation method.
[0263] In an exemplary embodiment, a computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, and the above instructions can be executed by a processor 920 of the electronic device 600 to complete the above method. Optionally, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0264] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include the common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0265] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A multi-channel noise data simulation method, characterized in that, it includes the following steps: Collect a first sound signal in a real three-dimensional space based on a voice collection array, where the real three-dimensional space includes at least one real sound source device; Determine the multipath parameters between the voice collection array and each real sound source device according to the first sound signal; Establish a target three-dimensional space according to the real three-dimensional space, where the target three-dimensional space is used to simulate multi-channel noise data; Determine target simulation parameters related to the target three-dimensional space according to multiple multipath parameters; and Determine the multi-channel noise data of the real three-dimensional space based on the target simulation parameters; The determining the target simulation parameters related to the target three-dimensional space according to multiple multipath parameters includes: Determine the real sound source device corresponding to the virtual sound source device from at least one real sound source device; Determine the first multipath parameter between the corresponding real sound source device and the voice collection array from multiple multipath parameters; Determine target simulation parameters related to the target three-dimensional space based on the first multipath parameter; The determining the target simulation parameters related to the target three-dimensional space based on the first multipath parameter includes: Determine a second multipath parameter according to the first multipath parameter and a preset perturbation parameter; Adjust the third multipath parameter of the target three-dimensional space until the difference between the third multipath parameter and the second multipath parameter meets a preset condition, and determine the target simulation parameters related to the target three-dimensional space, where the target simulation parameter is the reflection coefficient related to the spatial boundary of the target three-dimensional space; The determining the multi-channel noise data of the real three-dimensional space based on the target simulation parameters includes: Determine the single-channel noise data of the real three-dimensional space; Determine the impulse response of the real three-dimensional space based on the target simulation parameters; Determine the multi-channel noise data according to the single-channel noise data and the impulse response of the real three-dimensional space; The voice collection array includes: a plurality of voice collection devices; where the determining the multipath parameters between the voice collection array and each real sound source device according to the first sound signal includes: Establish a plurality of first signal models according to the first sound signal, where each first signal model is used to model the reception of the sound signal generated by the real sound source device by the corresponding voice collection device; Inverse-transform the cross-power spectral density function of the first sound signal received by the voice collection array according to multiple first signal models to obtain an initial cross-correlation function related to the first sound signal; Determine the multipath parameters between the voice collection array and each real sound source device according to the initial cross-correlation function.
2. The method according to claim 1, characterized in that, the establishing a target three-dimensional space according to the real three-dimensional space includes: Determine the spatial dimension parameters of the real three-dimensional space; Establish an initial three-dimensional space according to the spatial dimension parameters; Determine the simulation type of the noise data, and determine the target sound source position according to the simulation type, where the target sound source position is used to indicate the installation position of the virtual sound source device; Install a virtual sound source device at the position indicated by the target sound source position in the initial three-dimensional space to establish the target three-dimensional space.
3. The method according to claim 2, characterized in that, The determining the impulse response of the real three-dimensional space based on the target simulation parameters includes: Based on a preset noise simulation model, determine the impulse response of the real three-dimensional space according to the target simulation parameters, the target three-dimensional space and the real three-dimensional space.
4. The method according to claim 1, characterized in that, Before performing an inverse transform on the cross-power spectral density function of the first sound signal received by the voice acquisition array according to the plurality of first signal models to obtain an initial cross-correlation function related to the first sound signal, further comprising: Perform a frequency-domain transform on each of the first signal models to obtain a second signal model; Determine the cross-power spectral density function of the first sound signal according to the plurality of second signal models.
5. The method according to claim 1, characterized in that, The determining the multipath parameters between the voice acquisition array and each of the real sound source devices according to the initial cross-correlation function includes: Perform frequency-domain segmented smoothing processing on the initial cross-correlation function to obtain a target cross-correlation function; Determine the multipath parameters according to the target cross-correlation function.
6. The method according to claim 5, characterized in that, The performing frequency-domain segmented smoothing processing on the initial cross-correlation function to obtain a target cross-correlation function includes: Determine smoothing window information, where the smoothing window information includes: window width and window shift; Perform frequency-domain segmented smoothing processing on the initial cross-correlation function according to the smoothing window information to obtain the target cross-correlation function.
7. The method according to claim 5, characterized in that, The determining the multipath parameters according to the target cross-correlation function includes: Determine the eigenvector corresponding to the maximum eigenvalue of the target cross-correlation function; Determine the multipath parameters based on the eigenvector.
8. The method according to claim 7, characterized in that, The determining the multipath parameters based on the eigenvector includes: Obtain a first parameter that maximizes the norm of the eigenvector; Determine a time delay estimation value based on the first parameter, the eigenvector and a pre-constructed time delay estimation model.
9. A multi-channel noise data simulation device, characterized in that, comprises the following steps: An acquisition module, configured to acquire a first sound signal in a real three-dimensional space based on a voice acquisition array, where the real three-dimensional space includes at least one real sound source device; A first determination module, configured to determine the multipath parameters between the voice acquisition array and each of the real sound source devices according to the first sound signal; A construction module, configured to construct a target three-dimensional space according to the real three-dimensional space, where the target three-dimensional space is used to simulate multi-channel noise data; A second determination module, configured to determine target simulation parameters related to the target three-dimensional space according to the multiple multipath parameters; and A third determination module, configured to determine multi-channel noise data of the real three-dimensional space based on the target simulation parameters; The second determination module includes: A first determination unit, configured to determine a real sound source device corresponding to the virtual sound source device from at least one real sound source device; A second determination unit, configured to determine a first multipath parameter between the corresponding real sound source device and the voice acquisition array from the multiple multipath parameters; A third determination unit, configured to determine target simulation parameters related to the target three-dimensional space based on the first multipath parameter; The third determination unit is further configured to determine a second multipath parameter according to the first multipath parameter and a preset perturbation parameter; Adjust the third multipath parameter of the target three-dimensional space until the difference between the third multipath parameter and the second multipath parameter meets a preset condition, and determine the target simulation parameters related to the target three-dimensional space, wherein the target simulation parameter is a reflection coefficient related to the spatial boundary of the target three-dimensional space; The third determination module includes: A fourth determination unit, configured to determine single-channel noise data of the real three-dimensional space; A fifth determination unit, configured to determine an impulse response of the real three-dimensional space based on the target simulation parameters; A sixth determination unit, configured to determine the multi-channel noise data according to the single-channel noise data and the impulse response of the real three-dimensional space; The first determination module includes: A seventh determination unit, configured to establish a plurality of first signal models according to the first sound signal, wherein each first signal model is used to model the reception of the sound signal generated by the real sound source device by the corresponding voice acquisition device; An eighth determination unit, configured to perform an inverse transform on the cross-power spectral density function of the first sound signal received by the voice acquisition array according to the plurality of first signal models, so as to obtain an initial cross-correlation function related to the first sound signal; A ninth determination unit, configured to determine a multipath parameter between the voice acquisition array and each real sound source device according to the initial cross-correlation function.
10. An electronic device, characterized in that it includes: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1-8.
12. A computer program product, characterized in that it includes a computer program, and when the computer program is executed by a processor, it implements the method according to any one of claims 1-8.
Citation Information
Patent Citations
Multi-channel speech enhancement method and device
CN113030862A
Sound source orientation positioning method, device and equipment , and computer readable storage medium
CN113687305A
Acoustic channel digital twinning method, device and system facing underground tunnel environment
CN115019825A