Sound field partitioning method and device, equipment, storage medium and program product

By determining the impulse response data from the loudspeaker to the control point, using a psychoacoustic model and a deep neural network to train the filter coefficient model, and combining it with the loudspeaker array, the sound field zoning is optimized, solving the problems of poor user experience and zoning effect deviation in the existing technology, and achieving a better sound field zoning effect.

CN121865162APending Publication Date: 2026-04-14IFLYTEK (SUZHOU) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-08
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing sound field zoning technology cannot meet personalized audio needs, resulting in a poor user experience and discrepancies between objective and subjective zoning effects.

Method used

By determining the impulse response data from the loudspeaker to the control point, a psychoacoustic model and a deep neural network are used to train the filter coefficient model. This is combined with a loudspeaker array to achieve sound field zoning, taking into account changes in audio signals and human hearing characteristics to optimize the zoning effect.

Benefits of technology

It achieves better sound field zoning, improves user experience, resolves the discrepancy between objective and subjective zoning effects, and meets the personalized audio needs of different passengers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121865162A_ABST
    Figure CN121865162A_ABST
Patent Text Reader

Abstract

The invention provides a sound field partitioning method and device, equipment, a storage medium and a program product. The method comprises the following steps: determining impulse response data from a loudspeaker to a control point; the loudspeaker is one loudspeaker in the loudspeaker array; the control point is one of a plurality of control points corresponding to the target bright area or the target dark area; obtaining a filter coefficient corresponding to the audio signal according to the audio signal output by the loudspeaker and the impulse response data; filtering the audio signal by using the filter coefficient to obtain a target audio to be played; the sound field formed by playing the target audio through the loudspeaker comprises a target bright area and a target dark area. The method can improve the sound field partitioning effect and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound field control technology, and in particular to a sound field zoning method, apparatus, device, storage medium, and program product. Background Technology

[0002] In real-world applications, taking driving scenarios as an example, passengers or drivers in different seats often have different audio needs during a ride. For instance, the driver may need to listen to navigation audio, while passengers may need to listen to music or watch movies. Furthermore, front-seat passengers and rear-seat passengers also have different audio needs.

[0003] Sound field control technology precisely manipulates sound wave interference to control the energy distribution of sound in three-dimensional space, thereby achieving functions such as directional sound propagation, area muting, and environmental simulation. Sound field control technology allows sound waves to superimpose in a designated space (such as the space above the driver's seat), making the sound louder to the human ear in that space, and cancel each other out in other spaces (such as the space above the passenger seats), making the sound quieter. While current sound field control technology can achieve sound field zoning, the zoning effect often fails to meet expectations, impacting the user experience. Summary of the Invention

[0004] This invention provides a sound field zoning method, apparatus, device, storage medium, and program product to address the shortcomings of related technologies, such as poor sound field zoning effects and the need to improve user experience.

[0005] This invention provides a sound field partitioning method, comprising: determining impulse response data from a loudspeaker to a control point; the loudspeaker being one loudspeaker in a loudspeaker array; the control point being one of multiple control points corresponding to a target bright area or a target dark area; obtaining filter coefficients corresponding to the audio signal based on the audio signal output by the loudspeaker and the impulse response data; filtering the audio signal using the filter coefficients to obtain the target audio to be played; the sound field formed by the target audio being played through the loudspeaker includes a target bright area and a target dark area.

[0006] According to a sound field zoning method provided by the present invention, determining the impulse response data from a loudspeaker to a control point includes: playing a sweep frequency signal through a loudspeaker, acquiring the received signal corresponding to the sweep frequency signal at the control point through a pickup device; and performing cross-correlation calculation on the sweep frequency signal and the received signal to obtain the impulse response data from the loudspeaker to the control point.

[0007] According to a sound field partitioning method provided by the present invention, filter coefficients corresponding to the audio signal are obtained based on the audio signal output by the loudspeaker and the impulse response data. The method includes: obtaining acoustic perception features corresponding to the audio signal through a psychoacoustic model based on the audio signal output by the loudspeaker and the impulse response data; inputting the acoustic perception features into a pre-trained filter coefficient model; and outputting filter coefficients corresponding to the audio signal through the filter coefficient model.

[0008] According to a sound field partitioning method provided by the present invention, acoustic perception features corresponding to the audio signal are obtained through a psychoacoustic model based on the audio signal output by the loudspeaker and the impulse response data. The method includes: performing convolution operation on the audio signal output by the loudspeaker and the impulse response data to obtain the signal at the control point; normalizing and calibrating the signal at the control point, and extracting the frequency domain features through a preset sub-band filter bank; and performing frequency domain masking and time domain masking on the frequency domain features to obtain the acoustic perception features.

[0009] According to a sound field partitioning method provided by the present invention, the sub-band division method of the sub-band filter bank is a Bark band; and / or, frequency domain masking of frequency domain features includes: applying a diffusion function to the frequency domain features to simulate the frequency domain masking effect; and time domain masking of frequency domain features includes: simulating time domain forward masking through a low-pass infinite impulse response filter, simulating time domain backward masking through a low-pass finite impulse response filter, outputting a psychological masking curve, and using the psychological masking curve as an acoustic perception feature.

[0010] According to a sound field partitioning method provided by the present invention, before inputting acoustic perception features into a pre-trained filter coefficient model and outputting filter coefficients corresponding to the audio signal through the filter coefficient model, the method further includes: determining a first loss term based on at least one of the filter coefficients output in the current training step, the transfer function between the speaker and the control point, and target impulse response data; the first loss term is used to characterize the degree of reconstruction error of the target bright area; the target impulse response data is obtained based on at least one impulse response data under at least one sound effect scenario; or, the target impulse response data is a standard Dirac function; the transfer function is in the frequency domain form of the impulse response data; a preset sound quality evaluation model is invoked to evaluate the audio signal and the target audio to be played to determine a second loss term; the filter coefficient model is trained based on the preset loss function; the loss function includes the first loss term and / or the second loss term.

[0011] According to the sound field partitioning method provided by the present invention, the loss function further includes a third loss term; before inputting the acoustic perception features into a pre-trained filter coefficient model and outputting the filter coefficients corresponding to the audio signal through the filter coefficient model, the method further includes: obtaining the frequency domain form of the audio signal; obtaining the transfer function from the loudspeaker to the control point; the transfer function is the frequency domain form of the impulse response data; obtaining a single target signal of a single loudspeaker at a control point based on the filter coefficients output in the current training step, the frequency domain form of the audio signal, and the transfer function; and obtaining a first superposition of multiple loudspeakers included in the loudspeaker array at a control point based on the single target signal. The signal; based on the first superimposed signal, the acoustic perception feature corresponding to a control point is obtained through a psychoacoustic model; based on the acoustic perception feature corresponding to a control point and multiple first control points included in the target bright area, the first acoustic perception feature corresponding to the target bright area is obtained; based on the acoustic perception feature corresponding to a control point and multiple second control points included in the target dark area, the second acoustic perception feature corresponding to the target dark area is obtained; based on the first acoustic perception feature, the second acoustic perception feature, and the ratio of the number of control points, a third loss term is determined; wherein, the ratio of the number of control points is the ratio of the number of control points included in the target dark area to the number of control points included in the target bright area.

[0012] According to the sound field partitioning method provided by the present invention, the loss function further includes adjustment factors corresponding to the first loss term, the second loss term, and the third loss term, respectively; based on the preset loss function, training the filter coefficient model includes: using the adjustment factors as weights to perform a weighted summation of the first loss term, the second loss term, and the third loss term to obtain the loss value of the loss function; and updating the trainable parameters in the filter coefficient model according to the loss value.

[0013] The present invention also provides a sound field zoning device, comprising: an impulse response module for determining impulse response data from a loudspeaker to a control point; the loudspeaker being one of a loudspeaker in a loudspeaker array; the control point being one of a plurality of control points corresponding to a target bright area and a target dark area; a filter coefficient prediction module for obtaining filter coefficients corresponding to the audio signal based on the audio signal output by the loudspeaker and the impulse response data; and a filtering module for filtering the audio signal using the filter coefficients to obtain the target audio to be played; the sound field formed by the target audio being played through the loudspeaker includes a target bright area and a target dark area.

[0014] The present invention also provides a sound field zoning device, including a speaker array and a processor; the speaker array includes multiple speakers; each of the multiple speakers is used to output an audio signal; the processor is used to determine impulse response data from the speaker to a control point; based on the audio signal output by the speaker and the impulse response data, filter coefficients corresponding to the audio signal are obtained; the audio signal is filtered using the filter coefficients to obtain the target audio to be played; wherein, the speaker is one of the speakers in the speaker array; the control point is one of the multiple control points corresponding to the target bright area and the target dark area; the sound field formed by the target audio being played through the speaker includes the target bright area and the target dark area.

[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the sound field partitioning methods described above.

[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the sound field partitioning method as described above.

[0017] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements any of the sound field partitioning methods described above.

[0018] The sound field zoning method, apparatus, device, storage medium, and program product provided by this invention first determine the impulse response data from the speaker to the control point. Then, based on the audio signal output by the speaker and the impulse response data, filter coefficients corresponding to the audio signal are obtained. The audio signal is then convolved with the corresponding filter coefficients (equivalent to filtering) before being played through the speaker, forming a sound field including bright and dark areas. The filter coefficients are determined based on the audio signal output by the speaker and the impulse response data. In other words, when predicting the filter coefficients, the combined influence of the audio signal played by the speaker and the impulse response data on the filter coefficients is considered. Different filter coefficients are obtained for different audio signals played by the speaker. In practical applications, the input audio in each frame is constantly changing. Most current control algorithms rarely consider the differences in zoning effects caused by the audio itself. However, the solution adopted in this invention takes into account the impact of audio signal differences on the zoning effect, achieving better zoning results and significantly improving the user experience. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is a schematic flowchart of the sound field zoning method provided by the present invention.

[0021] Figure 2 This is a schematic diagram of the process for obtaining IR data in one embodiment of the sound field partitioning method provided by the present invention.

[0022] Figure 3 This is a schematic diagram illustrating the acquisition of filter coefficients in some embodiments of the sound field partitioning method provided by the present invention.

[0023] Figure 4 This is a schematic diagram of some embodiments of the sound field partitioning method provided by the present invention during the training phase.

[0024] Figure 5 This is a schematic diagram illustrating the sound field zoning method provided by the present invention, which extracts acoustic perception features through a psychoacoustic model.

[0025] Figure 6 This is a schematic diagram of data processing in the application stage of the sound field zoning method provided by the present invention.

[0026] Figure 7 This is a schematic diagram of data processing in the application stage of the sound field zoning method provided by the present invention.

[0027] Figure 8 This is a schematic diagram of the sound field zoning device provided by the present invention.

[0028] Figure 9 This is a schematic diagram of the sound field zoning device provided by the present invention.

[0029] Figure 10 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0031] The implementation of sound field control technology mainly relies on loudspeaker arrays and complex digital signal processing algorithms to achieve sound wave interference. Sound wave interference includes constructive interference and destructive interference. Constructive interference controls the meeting of the peaks of two sound waves to enhance the sound. Destructive interference controls the meeting of the peak of one sound wave with the trough of another sound wave to cancel each other out and weaken the sound. In this way, by precisely controlling the timing (phase) and intensity (amplitude) of the sound emitted by each loudspeaker in the array, sound waves can be made to superimpose at some points in space (becoming louder) and cancel each other out at other points (becoming quieter), creating independent auditory experiences for different users within a shared space.

[0032] Sound field control technology can be applied to a variety of practical scenarios. Taking in-vehicle audio systems as an example, with the rapid development of in-vehicle audio systems, users' demands for intelligence, personalization, and contextualization are gradually increasing, expecting to provide a more comfortable and private listening experience. Users in different spaces within the vehicle often have significantly different audio preferences during their journey. For example, the driver may need to listen to navigation prompts, the front passenger may want to listen to pop music, and the rear passengers may want to watch movies. Traditional car audio systems cannot meet these personalized needs, and some passengers may experience discomfort during the process.

[0033] Sound field zoning technology aims to overcome the limitations of traditional audio systems, providing users with a more personalized listening experience. This technology is suitable for various usage scenarios, such as family cars and commercial vehicles. In different scenarios, each seat occupant can enjoy personalized audio content according to their own needs, without being disturbed by others. Whether the driver is focused on listening to navigation voice prompts or children in the back seat are watching entertainment videos, sound field zoning ensures that everyone gets the best listening experience. This technology not only improves the quality of the riding experience but also drives the evolution of car audio systems to a higher level. With further refinement and popularization of the technology, sound field zoning is expected to become an important indicator of overall vehicle performance.

[0034] In related technologies, sound field zoning technology can be implemented through two types of schemes. The first type involves directional loudspeakers, which are designed and manufactured to achieve naturally high isolation. The second type is a zoning control method based on loudspeaker arrays, which uses vehicle loudspeakers to play inverse audio to cancel out sound waves leaking from bright areas to dark areas, thus achieving the zoning effect. In the second type of scheme, the control algorithms used can be, for example, Acoustic Contrast Control (ACC), Pressure Matching with the Least Squares Criterion (PM-LS), and Acoustic Contrast Control-Pressure Matching (ACC-PM) algorithms. These algorithms are mainly solved under constraints such as sound contrast level, sound field matching accuracy, and loudspeaker array power to achieve the desired zoning effect.

[0035] Among them, the first type of solution based on directional loudspeakers is currently in the early research stage. The related products have poor performance in the in-vehicle reverberation field and have disadvantages such as poor durability and high product cost.

[0036] Regarding zonal control based on speaker arrays, some solutions use audio phase cancellation, but in practical applications, the isolation between different sound zones is often insufficient, resulting in only a basic effect. Furthermore, in terms of control algorithms, there are too many control variables in multi-objective scenarios. How to solve the optimization problem of multiple objectives and the selection of control parameters are key factors restricting the performance of zonal control algorithms. Currently, there is no mature solution with an ideal zonal effect.

[0037] During their research, the inventors discovered that in real-world applications, the audio input for each frame is constantly changing. Most control algorithms rarely consider the differences in partitioning effects caused by the audio itself, meaning they haven't discovered the impact of audio signal differences on partitioning effects.

[0038] Furthermore, in actual testing, there is a certain discrepancy between the partitioning effect (isolation) that humans can subjectively perceive and the partitioning effect (isolation) that is objectively calculated based on signals. Under the same objective isolation, the human perception results are different. That is, even if the same algorithm is used to set the same partitions, users will have different subjective feelings. This technical problem has not yet been solved.

[0039] In view of this, embodiments of the present invention provide a sound field partitioning method to solve the above problems.

[0040] This invention provides a sound field zoning method that can be applied to various scenarios such as smart cars, museums / exhibition halls, and open offices. For example, in a smart car scenario, the driver can listen to navigation, the front passenger can listen to music, and the rear passengers can watch cartoons without interfering with each other. In a museum / exhibition hall scenario, narration automatically plays when someone walks to an exhibit, and the sound disappears when they leave. In an open office scenario, "sound bubbles" are formed above workstations for private calls without disturbing other colleagues' normal work.

[0041] Specifically, Figure 1 This is one of the flowcharts illustrating the sound field zoning method provided by the present invention, such as... Figure 1 As shown, the method includes: Step 101: Determine the impulse response data from the speaker to the control point.

[0042] In this context, a speaker is one speaker in a speaker array, which is a group of speaker units arranged in a specific geometric shape. In automotive audio control applications, speaker arrays can be installed in a ceiling array configuration: a small set of full-range or mid-to-high-frequency speaker units is independently installed on the headliner for each seat (or row of seats). For example, 2-4 speakers are installed directly above the driver's and passenger's heads. Alternatively, speaker arrays can be positioned on either side of the headrests, such as directly embedded in the headrests of the driver and passenger seats, typically one speaker on each side, forming an independent stereo system.

[0043] Control points are virtual locations pre-defined in space for designing and optimizing the sound field. For example, control points can be positioned around the listener's head, with multiple control points evenly or randomly distributed within a predefined area of ​​the listener's head (a cube or spherical space). Alternatively, control points can be placed near the user's ears: for an optimal personal listening experience, control points can be precisely positioned at the listener's left and right ears. Or, in some applications (such as public address systems), where the bright area might be a large region (like the entire driver's seat), control points can be distributed on a plane, for example, in a grid-like even distribution to ensure uniform coverage of the entire area.

[0044] A bright area refers to a physical space region where the energy of a target sound field signal is intentionally amplified and concentrated by controlling the sound waves of a loudspeaker array. A dark area refers to a physical space region where the energy of a target sound field signal is intentionally suppressed and canceled by controlling the sound waves of a loudspeaker array. In this embodiment of the invention, the target bright area is the bright area that meets the expected requirements, and the target dark area is the dark area that meets the expected requirements. The control point in step 101 is one of multiple control points corresponding to the target bright area or the target dark area.

[0045] In some embodiments, determining the impulse response data from the speaker to the control point can be achieved using methods such as... Figure 2 The process shown is implemented as follows: Step 1011: Play the sweep frequency signal through the speaker, and collect the received signal corresponding to the sweep frequency signal at the control point through the pickup device.

[0046] A swept signal is a sine wave whose frequency changes at a constant rate over time. For example, it starts at 20 Hz (the lowest frequency that humans can hear) and linearly increases to 20 kHz (the highest frequency that humans can hear) within a few seconds.

[0047] After determining the spatial locations (coordinates) of multiple control points included in the bright and dark areas respectively, a pickup device can be set up at each control point. For example, the pickup device is a microphone. After the speaker emits a sweep frequency signal, the signal received at this location is collected by the microphone. For ease of description, the signal that the sweep frequency signal reaches the control point and is picked up by the microphone is defined as the received signal.

[0048] Step 1012: Perform cross-correlation calculation on the swept frequency signal and the received signal to obtain the impulse response data from the loudspeaker to the control point.

[0049] Cross-correlation is used to measure the similarity of two signals at different time offsets. For example, assuming the swept signal and the received signal are discrete signals or discrete signals obtained by processing analog signals, the cross-correlation R_xy[m] at a delay of m for two discrete signals, swept signal x[n] and received signal y[n], is calculated as follows: R_xy[m] = Σ (x[n]·y[n + m]); m represents the delay, and R_xy[m] represents the cross-correlation result. The cross-correlation process is to align the two signals, multiply them, and then sum them for each possible delay m, which is equivalent to performing deconvolution.

[0050] Impulse response (IR) data refers to the complete time-dependent output characteristics of a system in response to an ideal instantaneous impulse input.

[0051] It should be noted that step 1012 is only an exemplary way to obtain impulse response data. In other embodiments, considering situations such as seat adjustment and in-vehicle environment adjustment that may occur in actual scenarios, the received signals collected from multiple actual in-vehicle environments can be cross-correlated with the audio signals played by the speakers to obtain IR data under multiple actual scenarios. Based on the multiple IR data, a comprehensive IR data is obtained, and the comprehensive IR data is used as the input for subsequent processing.

[0052] Step 102: Obtain the filter coefficients corresponding to the audio signal based on the audio signal output by the speaker and the impulse response data.

[0053] In some embodiments, such as Figure 3 As shown, the filter coefficients corresponding to the audio signal are obtained based on the audio signal and impulse response data. This can be achieved by first obtaining the signal at the control point (i.e., the sound wave signal received by the control point) based on the audio signal output by the speaker and the impulse response data, and then inputting the signal at the control point into a preset psychoacoustic model. Through the psychoacoustic model, the acoustic perception features corresponding to the audio signal are obtained.

[0054] Psychoacoustic models are mathematical models that simulate the characteristics of human auditory perception. Their processing simulates how the human ear and brain "perceive" sound, rather than simply measuring the physical properties of sound. Psychoacoustic models reveal the limits of human hearing and, by utilizing these limits, achieve highly efficient compression in audio technology that "deceives the ear," compensating for discrepancies between physical acoustics and subjective hearing. Through the processing of psychoacoustic models, the processed signal more closely matches the subjective auditory characteristics of humans.

[0055] For example, the signal at the control point can be obtained by convolving the audio signal and the impulse response data (IR) output by the speaker.

[0056] It should be noted that the convolution performed on the sound signal here is different from the convolution operation in a Convolutional Neural Network (CNN). The convolution in CNN is for images, while the convolution here is for sound wave signals.

[0057] Specifically, the steps for convolving audio signals and impulse response data can include: Folding, Shifting, Multiplication, and Integration. Folding involves flipping and shifting the IR along the time axis, essentially translating the flipped IR along the time axis so that the two signals overlap at different time points. Multiplication involves multiplying corresponding points of the two signals at each overlap location. Integration involves integrating the result of the multiplication (or summing in the discrete case) to obtain the convolution value at the current shift amount.

[0058] After obtaining the signal at the control point, the signal is input into the psychoacoustic model to extract acoustic perception features. Then, the acoustic perception features are input into a pre-trained filter coefficient model, which outputs the filter coefficients corresponding to the audio signal.

[0059] The filter coefficient model can be trained based on a deep neural network (DNN), such as one or more of the following deep neural networks: Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Generative Adversarial Network (GAN), Transformer, or other neural networks based on the Transformer architecture.

[0060] The acoustic signal at the control point is obtained by convolving the original audio signal output by the speaker with the impulse response IR. This acoustic signal contains the original audio features and impulse response features, as well as the acoustic perception features processed by the psychoacoustic model. It retains the original audio and impulse response features. Through training, the DNN can learn which filter coefficients should be used for different audio and impulse response features. Using these filter coefficients to filter the original audio signal yields the best zoning effect. Therefore, in practical applications, the trained DNN can predict the filter coefficients corresponding to the audio signal currently played by the speaker, and then use these filter coefficients to filter the audio signal to be played by the speaker before playback.

[0061] Step 103: Filter the audio signal using filter coefficients to obtain the target audio to be played.

[0062] Filtering audio signals using filter coefficients can be achieved by convolving the audio signal with the filter coefficients. For details on the specific method of convolution, please refer to the convolution of audio signals with IR.

[0063] Thus, the target audio obtained through the above steps, when played through a speaker, forms a sound field that includes both the target bright area and the target dark area.

[0064] In related technologies, the input signal is basically the impulse response signal of the application scenario (e.g., inside a car). This signal can only represent the transmission path from the speaker to the microphone in the in-vehicle environment, without considering the influence of the audio signal itself on the zoning effect during actual use. This invention considers that changes in the input audio signal itself will affect the zoning isolation effect. Each frame of audio signal is used as input, and a pre-trained DNN model is used to predict filter coefficients, thereby achieving the optimal zoning effect. Furthermore, some embodiments using psychoacoustic models to extract acoustic perception features and inputting these features into the DNN for prediction further solve the problem of discrepancies between objective zoning and human subjective perception of zoning. Psychoacoustic models can extract psychological masking curves to simulate human auditory characteristics, making the objective separation of bright and dark areas more closely match human subjective perception. Therefore, sound field zoning can achieve a zoning effect that is perceptible to humans, improving the user experience.

[0065] The following are some specific embodiments. To avoid confusion, these specific embodiments will be described from two phases: the training phase and the application phase.

[0066] Training phase: like Figure 4 As shown, the sound field partitioning method provided in this specific embodiment includes the following process during the training phase: Step 401: Determine the impulse response data from the speaker to the control point.

[0067] First, prepare the input data, which includes the audio signal. and shock response data . Indicates the first i The audio signal output by each speaker. For example, it could be a single frame of audio, which could be audio from various multimedia data such as music, movies, and animations. Assuming the speaker array includes... I One speaker, 1≤ i ≤ I .

[0068] Impulse response data Let represent the impulse response data from the i-th speaker to the i-th control point, where Li represents the i-th speaker and Mi represents the microphone at the i-th control point.

[0069] To determine the impulse response data from the speaker to the control point, a sweep frequency signal can be played sequentially through the speaker, and a microphone at the control point picks up the sound (receives the sweep frequency signal). The received signal is simply referred to as the received signal. Then, the cross-correlation between the sweep frequency signal and the received signal is performed to obtain the impulse response data. .

[0070] Optionally, considering actual scenarios such as seat adjustment and in-vehicle environment adjustment, multiple IR data can be obtained through cross-correlation, and then the multiple IR data can be calculated to obtain comprehensive IR data.

[0071] Step 402: Based on the audio signal and impulse response data output by the speaker, obtain the acoustic perception features corresponding to the audio signal through a psychoacoustic model.

[0072] During the training phase, the audio signal output by the speaker and the impulse response data are convolved to obtain the signal at the control point. In other words, the audio signal is... With impulse response Convolution is performed to obtain the signals at the control points. Then, the signals at the control points are used to extract acoustic perception features through a psychoacoustic model.

[0073] like Figure 5 As shown, for example, the psychoacoustic model (AM) may specifically include a normalization calibration module, a sub-band filter bank filtering module, a frequency domain masking module, and a time domain masking module, which are used to perform normalization calibration, sub-band filter bank filtering, frequency domain masking, and time domain masking on the input signal at the control point, respectively.

[0074] Specifically, in combination Figure 5 By using psychoacoustic models, acoustic perception features corresponding to audio signals can be obtained, which can be the signals at control points. Normalization calibration is performed, and frequency domain features are extracted using a preset sub-band filter bank. Next, frequency domain features are masked in both the frequency domain and the time domain to obtain acoustic sensing features.

[0075] The normalization calibration of the signal at the control point can be performed by normalizing the audio amplitude, unifying the volume, or maximizing the loudness. For example, normalization calibration may include the following process: Scan the signal at the control point to find its maximum absolute value (peak value) or calculate its overall loudness; Define the target value: For example, for peak normalization, the target value is typically the maximum value in the range of -1.0 to 1.0 (such as 1.0 or 0 dBFS). For loudness normalization, the target value is a specific full-scale loudness unit (LUFS) value, such as -14 LUFS.

[0076] Calculate gain: Peak gain = Target peak / Current peak; Loudness gain = Target LUFS - Current LUFS; Multiplying each sampled value of the signal at the entire control point by the calculated gain yields the normalized and calibrated signal.

[0077] Then, frequency domain features are extracted using a subband filter bank. A subband filter bank is a system that decomposes a broadband signal into multiple narrowband sub-signals according to frequency. Its goal is to decompose a complex, complete signal containing multiple frequencies into several simpler "subband" signals with different frequency ranges.

[0078] For example, the subband division method of the subband filter bank can be Bark band. In other words, the subband filter bank is a subband filter bank obtained by using the Bark band division method, that is, the subband division method of the subband filter bank refers to the Bark band division method.

[0079] Bark bands are nonlinear frequency scales based on the critical bandwidth of the human ear. The critical bandwidth is defined as the minimum frequency range at a specific center frequency that the human ear can distinguish between two pure tones. Within this bandwidth, sounds interfere with each other; beyond this bandwidth, they sound independent. The core rule of the Bark band scale is to divide the entire audible frequency range (0–20 kHz) into 24 Bark bands, each with a width approximately equal to a critical bandwidth. In other words, on the Bark scale, each Bark band is perceived as having "equal width" by the human ear. Regardless of whether the band is in low or high frequencies, the "perceptual load" or "perceptual width" it creates for the human auditory system is similar. For example, in the low-frequency region (e.g., Bark 1), the bandwidth is only 100 Hz; in the mid-high frequency region (e.g., Bark 10), the bandwidth is approximately 270 Hz; and in the high-frequency region (e.g., Bark 20), the bandwidth exceeds 2000 Hz. This division method is designed to accommodate the human ear's greater sensitivity to low frequencies and less sensitivity to high frequencies.

[0080] Next, the frequency domain masking effect is simulated by using a diffusion function to achieve frequency domain masking processing of the signal.

[0081] For example, in psychoacoustic models, the diffusion function can be a piecewise, asymmetric ramp function. For a center frequency of... fm The masking tone, which is present at any frequency f Increased masking threshold caused by the location S ( f (Unit: dB) can be described by a function, which has different slopes in different frequency ranges: Low frequency side ( f < fm : A relatively steep negative slope (e.g., -27 dB / Bark). This indicates that the masking effect decays rapidly towards lower frequencies.

[0082] High frequency side ( f > fm : A relatively gentle positive slope (e.g., +10 dB / Bark). This indicates that the masking effect decays very slowly towards higher frequencies and has a wider range of influence.

[0083] This asymmetrical shape corresponds precisely to the psychoacoustic characteristic that "low-frequency sounds are more likely to mask high-frequency sounds."

[0084] Next, the frequency-domain masked signal is subjected to time-domain masking. For example, in this embodiment, time-domain masking of frequency-domain features can be achieved by first simulating forward time-domain masking using a low-pass Infinite Impulse Response (IIR) filter, and then simulating backward time-domain masking using a low-pass Finite Impulse Response (FIR) filter, outputting a psychological masking curve, which is then used as an acoustic perception feature.

[0085] Step 403: Input the acoustic perception features into the deep neural network, and output the filter coefficients corresponding to the audio signal through the deep neural network.

[0086] Then, the features are input into the network model to calculate the filter coefficients. , making the filter coefficients The threshold for satisfying the optimization function is used; the acoustic perception features here are calculated using a psychoacoustic model. Step 404: Based on the preset loss function, calculate the loss value corresponding to the filter coefficients output in the current training step, and optimize the trainable parameters in the DNN according to the loss value.

[0087] A training step is one of the many iterations in training a DNN.

[0088] Step 405: Determine whether the DNN meets the convergence condition. If yes, end the training; otherwise, return to step 401 and continue to the next training step.

[0089] The convergence condition can be that the loss value is less than a preset threshold, or the number of training steps reaches a preset upper limit (e.g., 100,000 steps).

[0090] For example, the loss can be minimized by iteratively optimizing the loss, and training can be completed when the loss is less than a preset threshold or after a specified number of training iterations.

[0091] For example, the predefined loss function (or objective optimization function) Loss is specifically formulated as follows: (1) in, , , The adjustment factor can be a preset real number. + + By setting , , The value of can adjust the weight of each loss item, that is, control the proportion of each loss item in the loss value, and thus affect the degree of influence of each loss item on the prediction result.

[0092] in, Indicates the first loss item. Indicates the second loss item. This indicates the third loss item.

[0093] The level of sound contrast perceived by the human ear represents how many dB the energy in the bright area is higher than the energy in the dark area. The specific calculation formula is as follows: (2) in, This indicates the number of control points in the visible area (i.e., the target visible area). This indicates the number of control points (microphones) in the dark area. and This indicates that the microphone (control point) is in a dark area.

[0094] Indicates the current audio signal The frequency domain representation (frequency domain characteristics), for example, for Performing a Fourier transform or a fast Fourier transform yields the frequency domain form. This represents the transfer function from a loudspeaker to a control point. The transfer function can be the frequency domain representation of IR data, that is, by performing a Fourier transform or a fast Fourier transform on the IR data, the frequency domain form of the IR data can be obtained. Represents the filter coefficients. This represents the audio signal output by a single speaker at a control point, determined by the filter coefficients. The filtered signal at this control point is then superimposed with the signals from all the speakers in the speaker array at this control point. This is the signal at this control point when all speakers are working.

[0095] Then through psychoacoustic models The signal at this control point is extracted when all loudspeakers are working, and the corresponding acoustic perception features are extracted, which yields the perception masking curve (i.e., psychological masking curve) at a control point.

[0096] The acoustic perception features (first acoustic perception features) of the entire bright area are obtained by summing the psychological masking curves corresponding to all control points (or a portion of all control points) included in the bright area. The acoustic perception features (second acoustic perception features) of the entire dark area are obtained by summing the psychological masking curves corresponding to all control points (or a portion of all control points) included in the dark area.

[0097] Divide the first acoustic perception feature corresponding to the bright area by the second acoustic perception feature corresponding to the dark area, then perform energy calculations, and multiply by the number of control points in the dark area. Number of control points in the Ming area The ratio of these values ​​yields the final result. For example, the logarithmic operation 20log in formula (2) 10 This is an example of energy calculation.

[0098] First loss item The specific calculation formula for the degree of reconstruction error of the target sound field in the bright area is as follows: (3) in, This represents the desired target sound field at a certain control point (i.e., the comprehensive IR data mentioned above). The target sound field can be IR data obtained under a specific in-vehicle sound effect scenario, a standard Dirac function, or modulated IR data obtained under various sound effect scenarios. For example, the specific modulation method can be to filter the IR data using a filter to make the IR data meet the desired sound field effect. The filter can be one of those filters that can adjust the amplitude or time delay.

[0099] This is a constraint on the sound quality variation of bright-area signals, used to control the sound quality loss between the input audio signal and the filtered audio signal. The specific calculation formula is as follows: (4) Here, 'net' represents some neural network-based sound quality evaluation models. It is necessary to select a differentiable sound quality evaluation model in order to obtain the gradient of the loss function. For example, it can be one or more of the following sound quality evaluation models: Mean Opinion Score Network (MOSNet) and Deep Noise Suppression Mean Opinion Score (DNSMOS).

[0100] Traditional audio quality metrics, such as Perceptual Evaluation of Speech Quality (PESQ) or Short-Time Objective Intelligibility (STOI), are not differentiable everywhere. When using gradient descent to update the model's parameters, the gradient may not be obtained. This invention employs a differentiable audio quality evaluation model for evaluation.

[0101] It should be noted that the loss functions shown in formulas (1) to (4) above are only specific examples. In practical applications, those skilled in the art can make adaptive modifications based on the above examples and related descriptions in this specification to obtain other specific implementations of loss functions.

[0102] For example, in some embodiments, the adjustment factor may not be set, and the expression for the loss function is as follows: (6) Alternatively, in some embodiments, the loss function includes only at least one of the first to third loss terms described above. For example, the expression for the loss function is as follows: (7) Correspondingly, if the loss function does not include a third loss term Therefore, in the aforementioned steps, the step of extracting acoustic perception features using a psychoacoustic model can be assumed by default. That is, after convolving the audio signal with the impulse response data, it can be directly input into the DNN. In this way, at least by predicting the filter coefficients corresponding to different audio signals, the impact of the differences in the audio signals themselves on the partitioning effect can be reduced.

[0103] or, (8) or, (9) Based on the above exemplary description of the loss function, it can be seen that in this embodiment of the invention, the loss function can be determined in the following manner: Based on the filter coefficients output at the current training step Transfer function between loudspeaker and control point Target impulse response data At least one of them, determine the first loss term.

[0104] The first loss term characterizes the degree of reconstruction error in the target's bright area, and the target impulse response data... It is obtained from at least one impulse response data in at least one sound effect scenario, for example, by modulating at least one impulse response data. Alternatively, target impulse response data This is the standard Dirac function.

[0105] This is just one example; in practice, it can be modified to obtain various other specific implementations of the first loss term. For example, instead of summing all control points in the bright area, random sampling can be performed from multiple control points, and the calculation can be performed only on a subset of the control points.

[0106] Then, a preset sound quality evaluation model is invoked to evaluate the audio signal and the target audio to be played to determine the second loss term. Based on the preset loss function, a deep neural network is trained. The loss function includes the first loss term and / or the second loss term. For example, as shown in the example in formula (7) above, the loss function may include only the first loss term and the second loss term.

[0107] Optionally, the loss function may also include a third loss term. Based on the above exemplary description, the third loss term can be obtained in the following manner: Acquire audio signals (e.g.) The frequency domain form of () The process involves obtaining the transfer function from the loudspeaker to the control point; as mentioned above, the transfer function is in the frequency domain form of the impulse response data. Then, based on the filter coefficients output by the DNN at the current training step (e.g., the nth step). The frequency domain form of audio signals ( ), transfer function ( ), to obtain a single target signal from a single loudspeaker at a single control point, for example The signal at the control point is the audio signal output from a single target signal, which is filtered by the corresponding filter coefficients.

[0108] Next, based on a single target signal, the first superimposed signal of the multiple speakers in the speaker array at a control point is obtained, for example, by calculation. The resulting signal is the first superimposed signal.

[0109] Based on the first superimposed signal, the acoustic perception features corresponding to a control point are obtained through a psychoacoustic model, for example... In this context, AM represents the extraction of acoustic perception features from the first superimposed signal. The specific process for extracting acoustic perception features from the first superimposed signal can be found in the above embodiments regarding the process of extracting acoustic perception features using a psychoacoustic model, for example, referring to... Figure 5 And its textual description.

[0110] Next, based on the acoustic perception features corresponding to a control point and multiple first control points included in the target bright area, the first acoustic perception features corresponding to the target bright area are obtained. For example, the multiple first control points can be all the control points included in the bright area, or they can be a subset of control points sampled from all the control points. For example, 6 or 8 out of 10 control points can be selected as the first control points each time to reduce computational load. This means summing up multiple first control points in the bright area. In other embodiments, weights can be further assigned to the multiple control points, and a weighted summation of the multiple first control points can be performed to obtain the overall first acoustic perception features of the bright area.

[0111] Similarly, based on the acoustic sensing features corresponding to a control point and the multiple second control points included in the target dark area, the second acoustic sensing features corresponding to the target dark area are obtained. The method for determining the acoustic sensing features of the dark area can refer to the method for the bright area, and will not be elaborated further.

[0112] Next, the third loss term is determined based on the ratio of the first acoustic perception feature, the second acoustic perception feature, and the number of control points.

[0113] The control point ratio is the ratio of the number of control points in the target dark area to the number of control points in the target bright area. For example, This represents the ratio of the number of control points in the target's dark area to the number of control points in the target's light area.

[0114] It should be noted that energy calculations can be further performed on the first and second acoustic perception features. Based on the results of the energy calculations, and combined with the ratio of the number of control points, a third loss term can be obtained. .

[0115] Based on the above example, we can also deduce that the loss function also includes adjustment factors corresponding to the first, second, and third loss terms, for example... , , As a adjustment factor, the loss function may or may not have an adjustment factor. Setting an adjustment factor can enhance the control over the impact of each loss term on the overall loss.

[0116] For example, by using the adjustment factor as a weight, the first loss term, the second loss term, and the third loss term are summed in a weighted manner to obtain the loss value of the loss function; based on the loss value, the trainable parameters in the deep neural network are updated.

[0117] Based on the above explanation, as Figure 6 As shown, the optimization process during the training phase can be summarized as follows: Figure 6 The data processing and training process is shown below.

[0118] The human ear perception curve at the control point in the bright area is the first acoustic perception feature mentioned above; the human ear perception curve at the control point in the dark area is the second acoustic perception feature mentioned above.

[0119] The signals at the control points are defined as the signals at each control point in the visible area.

[0120] The signal at the control point in the dark area is the signal at the control point corresponding to each control point in the dark area.

[0121] After the above training, the trained DNN can be obtained.

[0122] Application phase (testing phase): like Figure 7 As shown, for different partition configurations and corresponding scenarios in the application, the impulse response data from the speaker to each control point is first determined. According to the audio signal The DNN model obtained through the above training phase predicts Z control filter parameters. , recorded as , With the input audio signal Convolution is performed, which is essentially the process of filtering the audio signal. This produces the target audio that will be played through the speaker. When the target audio is played through the speaker, the corresponding control points can be partitioned.

[0123] In summary, the embodiments of the present invention calculate the signal of each frame of audio data at the control point, and then extract features based on the psychoacoustic model, which can effectively improve the subjective perception of the zoning effect. Secondly, by using perception-based optimization objectives and constraints on sound quality, the subjective perception of the zoning effect can be more obvious under the same sound quality impairment. Furthermore, by using the loss function to constrain and solve aspects such as bright area sound field, sound quality and zoning isolation, the weights of each part can be adjusted more intuitively, and more constraints can be flexibly added.

[0124] In related technologies, the input is basically the impulse response signal inside the vehicle. This signal can only characterize the transmission path from the speaker to the microphone in the in-vehicle environment, without considering the audio signal itself during actual use. The solution proposed in this embodiment of the invention takes into account that the changes in the input audio itself will affect the effect of the partition isolation. It takes each frame of audio signal as input and trains the model through data to achieve the best effect. In addition, compared with various traditional control algorithms, the solution proposed in this embodiment of the invention simplifies the solution process and reduces control parameters by using model training.

[0125] Furthermore, compared to other optimization functions (loss functions) in related technologies, which only optimize acoustic contrast at the control points in the bright and dark regions, the scheme proposed in this invention goes a step further. It considers the factors of human ear perception, obtains the human ear perception curves in the bright and dark regions through a psychoacoustic model, and optimizes the human ear perception acoustic contrast based on this, i.e., PACC_loss. Secondly, considering the changes in sound quality, it optimizes the sound quality constraints on the input audio and the audio after the input audio convolution filter coefficients, i.e., SQ_loss. Finally, it also considers the desired target sound field and includes sound field reconstruction error constraints, i.e., PM_loss.

[0126] And, as Figure 3 As shown, a psychoacoustic model is added to the input feature extraction to represent the perceptual attributes of the human ear. Combined with the subsequent perception optimization function, a better perceptual isolation can be achieved.

[0127] The sound field partitioning device provided by the present invention is described below. The sound field partitioning device described below can be referred to in correspondence with the sound field partitioning method described above.

[0128] Figure 8 This is a schematic diagram of the sound field zoning device provided by the present invention, as shown below. Figure 8 As shown, the sound field zoning device includes: The impulse response module 810 is used to determine the impulse response data from the loudspeaker to the control point; the loudspeaker is one of the loudspeakers in the loudspeaker array; the control point is one of multiple control points corresponding to the target bright area and the target dark area.

[0129] The filter coefficient prediction module 820 is used to obtain the filter coefficients corresponding to the audio signal based on the audio signal output by the speaker and the impulse response data.

[0130] The filtering module 830 is used to filter the audio signal using filter coefficients to obtain the target audio to be played; the sound field formed by the target audio being played through the speaker includes the target bright area and the target dark area.

[0131] The device provided in this embodiment of the invention takes into account that changes in the input audio signal itself will affect the effect of partition isolation. It takes each frame of audio signal as input and predicts filter coefficients through a pre-trained DNN model, thereby achieving the best partitioning effect and improving the user experience.

[0132] Based on any of the above embodiments, the impulse response module is specifically used for: playing a sweep frequency signal through a speaker, acquiring the received signal corresponding to the sweep frequency signal at the control point through a pickup device; performing cross-correlation calculation on the sweep frequency signal and the received signal to obtain the impulse response data from the speaker to the control point.

[0133] Based on any of the above embodiments, the filter coefficient prediction module includes an acoustic perception feature extraction unit and a prediction unit. The acoustic perception feature extraction unit is used to: obtain acoustic perception features corresponding to the audio signal based on the audio signal output by the speaker and impulse response data, using a psychoacoustic model. The prediction unit is specifically used to: input the acoustic perception features into a pre-trained deep neural network, and output filter coefficients corresponding to the audio signal via the deep neural network.

[0134] Based on any of the above embodiments, the acoustic perception feature extraction unit is specifically used to: perform convolution operation on the audio signal output by the loudspeaker and the impulse response data to obtain the signal at the control point; normalize and calibrate the signal at the control point, and extract the frequency domain features through a preset sub-band filter bank; and perform frequency domain masking and time domain masking on the frequency domain features to obtain acoustic perception features.

[0135] The subband filter bank is implemented based on the Bark band division method.

[0136] And / or, the acoustic perception feature extraction unit is specifically used to: apply a diffusion function to the frequency domain features to simulate the frequency domain masking effect; simulate forward masking in the time domain through a low-pass infinite impulse response filter, simulate backward masking in the time domain through a low-pass finite impulse response filter, output a psychological masking curve, and use the psychological masking curve as an acoustic perception feature.

[0137] Based on any of the above embodiments, the device further includes a training module, configured to: determine a first loss term based on at least one of the filter coefficients output in the current training step, the transfer function between the speaker and the control point, and the target impulse response data; the first loss term is used to characterize the degree of reconstruction error in the target bright area; the target impulse response data is obtained based on at least one impulse response data under at least one sound effect scenario; or, the target impulse response data is a standard Dirac function; the transfer function is in the frequency domain form of the impulse response data; a preset sound quality evaluation model is invoked to evaluate the audio signal and the target audio to be played to determine a second loss term; a deep neural network is trained based on the preset loss function; the loss function includes the first loss term and / or the second loss term.

[0138] Based on any of the above embodiments, the loss function further includes a third loss term; the training module is further configured to: obtain the frequency domain form of the audio signal; obtain the transfer function from the loudspeaker to the control point; the transfer function is the frequency domain form of the impulse response data; obtain a single target signal of a single loudspeaker at a control point based on the filter coefficients output in the current training step, the frequency domain form of the audio signal, and the transfer function; obtain a first superimposed signal of multiple loudspeakers included in the loudspeaker array at a control point based on the single target signal; obtain the acoustic perception feature corresponding to a control point through a psychoacoustic model based on the first superimposed signal; obtain the first acoustic perception feature corresponding to the target bright area based on the acoustic perception feature corresponding to a control point and the multiple first control points included in the target bright area; obtain the second acoustic perception feature corresponding to the target dark area based on the acoustic perception feature corresponding to a control point and the multiple second control points included in the target dark area; determine the third loss term based on the ratio of the first acoustic perception feature and the second acoustic perception feature and the number of control points; wherein, the ratio of the number of control points is the ratio of the number of control points included in the target dark area to the number of control points included in the target bright area.

[0139] Based on any of the above embodiments, the loss function further includes adjustment factors corresponding to the first loss term, the second loss term, and the third loss term, respectively; the training module is further configured to: use the adjustment factors as weights to perform a weighted summation of the first loss term, the second loss term, and the third loss term to obtain the loss value of the loss function; and update the trainable parameters in the deep neural network according to the loss value.

[0140] Figure 9 This is a schematic diagram of the sound field zoning device provided by the present invention, as shown below. Figure 9 As shown, the sound field zoning device includes a loudspeaker array 910 and a processor 920; The speaker array 910 includes multiple speakers; each of the multiple speakers is used to output an audio signal.

[0141] The processor 920 is used to determine the impulse response data from the speaker to the control point; based on the audio signal output by the speaker and the impulse response data, it obtains the filter coefficients corresponding to the audio signal; the filter coefficients are used to filter the audio signal to obtain the target audio to be played; wherein, the speaker is one of the speakers in the speaker array; the control point is one of the multiple control points corresponding to the target bright area and the target dark area; the sound field formed by the target audio being played through the speaker includes the target bright area and the target dark area.

[0142] The device provided in this embodiment of the invention, based on the fact that changes in the input audio signal itself affect the partition isolation effect, takes each frame of audio signal as input, and predicts filter coefficients through a pre-trained DNN model, thereby achieving a better partitioning effect and improving the user experience.

[0143] Figure 10 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 10 As shown, the electronic device may include: a processor 1010, a communications interface 1020, a memory 1030, and a communication bus 1040, wherein the processor 1010, the communications interface 1020, and the memory 1030 communicate with each other via the communication bus 1040. The processor 1010 can call logical instructions in the memory 1030 to execute a sound field zoning method, which includes: Determine the impulse response data from the loudspeaker to the control point; the loudspeaker is one loudspeaker in a loudspeaker array; the control point is one of multiple control points corresponding to the target bright area or the target dark area; obtain the filter coefficients corresponding to the audio signal based on the audio signal output by the loudspeaker and the impulse response data; use the filter coefficients to filter the audio signal to obtain the target audio to be played; the sound field formed by the target audio played through the loudspeaker includes the target bright area and the target dark area.

[0144] Furthermore, the logical instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0145] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the sound field partitioning method provided by the above methods, the method comprising: Determine the impulse response data from the loudspeaker to the control point; the loudspeaker is one loudspeaker in a loudspeaker array; the control point is one of multiple control points corresponding to the target bright area or the target dark area; obtain the filter coefficients corresponding to the audio signal based on the audio signal output by the loudspeaker and the impulse response data; use the filter coefficients to filter the audio signal to obtain the target audio to be played; the sound field formed by the target audio played through the loudspeaker includes the target bright area and the target dark area.

[0146] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the sound field partitioning method provided by the methods described above, the method comprising: Determine the impulse response data from the loudspeaker to the control point; the loudspeaker is one loudspeaker in a loudspeaker array; the control point is one of multiple control points corresponding to the target bright area or the target dark area; obtain the filter coefficients corresponding to the audio signal based on the audio signal output by the loudspeaker and the impulse response data; use the filter coefficients to filter the audio signal to obtain the target audio to be played; the sound field formed by the target audio played through the loudspeaker includes the target bright area and the target dark area.

[0147] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0148] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. Such computer software products can be stored in computer-readable storage media, such as ROM / RAM, magnetic disks, optical disks, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A sound field partitioning method, characterized in that, include: Determine the impulse response data from the loudspeaker to the control point; The loudspeaker is one of the loudspeakers in a loudspeaker array; The control point is one of multiple control points corresponding to the target's bright area or dark area; Based on the audio signal output by the speaker and the impulse response data, the filter coefficients corresponding to the audio signal are obtained; The audio signal is filtered using the filter coefficients to obtain the target audio to be played; the sound field formed by the target audio played through the speaker includes the target bright area and the target dark area.

2. The method according to claim 1, characterized in that, Determine the impulse response data from the speaker to the control point, including: A sweep frequency signal is played through the speaker, and the received signal corresponding to the sweep frequency signal is acquired at the control point by a pickup device. The cross-correlation operation is performed on the swept frequency signal and the received signal to obtain the impulse response data from the loudspeaker to the control point.

3. The sound field zoning method according to claim 1, characterized in that, Based on the audio signal output by the speaker and the impulse response data, filter coefficients corresponding to the audio signal are obtained, including: Based on the audio signal output by the speaker and the impulse response data, acoustic perception features corresponding to the audio signal are obtained through a preset psychoacoustic model. The acoustic sensing features are input into the filter coefficient model, and the filter coefficient model outputs the filter coefficients corresponding to the audio signal.

4. The sound field zoning method according to claim 3, characterized in that, Based on the audio signal output by the speaker and the impulse response data, acoustic perception features corresponding to the audio signal are obtained through a psychoacoustic model, including: The audio signal output by the speaker and the impulse response data are convolved to obtain the signal at the control point; The signal at the control point is normalized and calibrated, and the frequency domain features are extracted through a preset sub-band filter bank. The frequency domain features are masked in both the frequency domain and the time domain to obtain acoustic sensing features.

5. The sound field zoning method according to claim 4, characterized in that, The subband division method of the subband filter bank is Bark band; And / or, Frequency domain masking of the frequency domain features includes: applying a diffusion function to the frequency domain features to simulate a frequency domain masking effect; The frequency domain features are masked in the time domain by: simulating forward masking in the time domain using a low-pass infinite impulse response filter, simulating backward masking in the time domain using a low-pass finite impulse response filter, outputting a psychological masking curve, and using the psychological masking curve as the acoustic perception feature.

6. The sound field zoning method according to any one of claims 3-5, characterized in that, Before inputting the acoustic perception features into the filter coefficient model and outputting the filter coefficients corresponding to the audio signal through the filter coefficient model, the method further includes: A first loss term is determined based on at least one of the filter coefficients output in the current training step, the transfer function between the speaker and the control point, and the target impulse response data; the first loss term is used to characterize the degree of reconstruction error of the target bright area; the target impulse response data is obtained based on at least one impulse response data under at least one sound effect scenario; or, the target impulse response data is a standard Dirac function; the transfer function is the frequency domain form of the impulse response data; A preset sound quality evaluation model is invoked to evaluate the audio signal and the target audio to be played to determine a second loss term; The filter coefficient model is trained based on a preset loss function; the loss function includes the first loss term and / or the second loss term.

7. The sound field zoning method according to claim 6, characterized in that, The loss function also includes a third loss term; Before inputting the acoustic perception features into the filter coefficient model and outputting the filter coefficients corresponding to the audio signal through the filter coefficient model, the method further includes: Obtain the frequency domain form of the audio signal; Obtain the transfer function from the loudspeaker to the control point; the transfer function is in the frequency domain form of the impulse response data. Based on the filter coefficients output in the current training step, the frequency domain form of the audio signal, and the transfer function, a single target signal from a single loudspeaker at a control point is obtained. Based on the single target signal, a first superimposed signal of multiple loudspeakers included in the loudspeaker array at a single control point is obtained; Based on the first superimposed signal, the acoustic perception features corresponding to a control point are obtained through the psychoacoustic model. Based on the acoustic sensing features corresponding to a control point and the multiple first control points included in the target bright area, the first acoustic sensing features corresponding to the target bright area are obtained. Based on the acoustic sensing features corresponding to a control point and the multiple second control points included in the target dark area, the second acoustic sensing features corresponding to the target dark area are obtained. A third loss term is determined based on the first acoustic sensing feature, the second acoustic sensing feature, and the ratio of the number of control points; wherein the ratio of the number of control points is the ratio of the number of control points included in the target dark area to the number of control points included in the target bright area.

8. The sound field zoning method according to claim 7, characterized in that, The loss function also includes adjustment factors corresponding to the first loss term, the second loss term, and the third loss term, respectively; The filter coefficient model is trained based on a preset loss function, including: Using the adjustment factor as a weight, the first loss term, the second loss term, and the third loss term are weighted and summed to obtain the loss value of the loss function; Based on the loss value, update the trainable parameters in the filter coefficient model.

9. A sound field zoning device, characterized in that, include: The impulse response module is used to determine the impulse response data from the speaker to the control point; The loudspeaker is one of the loudspeakers in a loudspeaker array; The control point is one of multiple control points corresponding to the target's bright area and dark area; The filter coefficient prediction module is used to obtain the filter coefficients corresponding to the audio signal based on the audio signal output by the speaker and the impulse response data. The filtering module is used to filter the audio signal using the filter coefficients to obtain the target audio to be played; the sound field formed by the target audio being played through a speaker includes the target bright area and the target dark area.

10. A sound field zoning device, characterized in that, Includes speaker arrays and processors; The speaker array includes multiple speakers; each of the multiple speakers is used to output an audio signal. The processor is used to determine impulse response data from the speaker to the control point; Based on the audio signal output by the speaker and the impulse response data, filter coefficients corresponding to the audio signal are obtained; the audio signal is filtered using the filter coefficients to obtain the target audio to be played; wherein, the speaker is one of the speakers in the speaker array; the control point is one of multiple control points corresponding to the target bright area and the target dark area; the sound field formed by the target audio played through the speaker includes the target bright area and the target dark area.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the sound field partitioning method as described in any one of claims 1 to 8.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the sound field partitioning method as described in any one of claims 1 to 8.

13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the sound field partitioning method as described in any one of claims 1 to 8.