Control filter coefficient determination method and device, electronic equipment and storage medium

By optimizing the filter coefficients through neural networks, the problems of poor audio quality and sound energy leakage in the design of personal vocal zones in existing technologies have been solved, achieving efficient audio privacy protection.

CN121664150APending Publication Date: 2026-03-13BEIJING CO WHEELS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-11
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies, when implementing personal voice zones, result in a deterioration in audio quality in the target listening area and significant sound energy leakage in non-target listening areas, failing to effectively protect audio privacy.

Method used

By iteratively optimizing the control filter coefficients through a neural network, suitable control filter coefficients are generated based on the frequency response information and acoustic contrast differences between the target and non-target listening areas, ensuring high acoustic contrast and spectral flatness.

Benefits of technology

It achieves the goal of improving audio quality in the target listening area while ensuring high sound contrast, reducing sound energy leakage in non-target listening areas, protecting audio privacy, and realizing the design of a personal voice zone in the car.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121664150A_ABST
    Figure CN121664150A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a control filter coefficient determination method and device, equipment and a medium. The method comprises the following steps: acquiring first frequency response information corresponding to a target listening area and listening sound contrast according to a control filter coefficient; total difference information is determined according to the frequency response difference between the first frequency response information and the target frequency response information and the sound contrast difference between the listening sound contrast and the target sound contrast, the neural network is updated according to the total difference information to generate a next round of control filter coefficient, iteration execution is carried out, and the next round of control filter coefficient is obtained. The frequency response and the sound contrast are used as targets, the appropriate control filter coefficient is iteratively generated through the neural network, the flatness of the frequency spectrum of the target sound listening area is considered while the high sound contrast is guaranteed, the audio quality of the target sound listening area is guaranteed, and the sound quality of the target sound listening area is improved. The sound energy leaked in the non-target listening area is reduced, and then the privacy of the audio is protected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic device cooling technology, and in particular to a method for determining control filter coefficients, a device for determining control filter coefficients, an electronic device, and a readable storage medium. Background Technology

[0002] With the rapid development of smart cockpits, in-cabin sound control technology has received widespread attention. By assigning a unique set of filter coefficients to the speakers, the output sound energy of the entire speaker array can be concentrated in a specific area while remaining relatively quiet in another area. This sound field zoning control technology is called Personal Sound Zone (PSZ).

[0003] like Figure 1 As shown, within the vehicle space, a personal voice zone allows for different sound energy intensities in adjacent spatial areas. The target listening area is the bright area, and the non-target listening area is the dark area. Ultimately, while ensuring the sound energy of the target listening area, the sound energy of the non-target listening area is reduced, thereby ensuring listening privacy. To achieve this goal, a digital filter can be designed based on the impulse response of the environment, so that the original audio signal is filtered before driving the speaker, thereby improving the sound contrast between the target and non-target listening areas. Sound contrast refers to the contrast of sound energy between the target and non-target listening areas. The higher the sound contrast, the greater the difference in sound energy, the easier it is to form a personal voice zone, and the less audio interference to the non-target listening area.

[0004] Currently, common methods for achieving personalized audio zones include acoustic contrast control (ACC) and pressure matching (PM). However, these methods often sacrifice audio quality to achieve high acoustic contrast. Furthermore, these methods require matrix inversion operations when designing control filters, typically necessitating a reduction in acoustic contrast for stability and reconstruction accuracy. This is because both acoustic contrast control and pressure matching are convex optimization methods, making it difficult to find a suitable balance between spectral flatness and acoustic contrast. These shortcomings result in: degraded audio quality in the target listening zone, losing original audio detail and quality, potentially causing passenger dissatisfaction; and significant leakage of sound energy in non-target listening zones, easily interfering with the driving experience of passengers in these zones, and failing to protect audio privacy. Summary of the Invention

[0005] The technical problem to be solved by the embodiments of the present invention is to provide a method, device, electronic device, readable storage medium and vehicle for determining control filter coefficients, so as to solve the problems that the audio quality deteriorates in the target listening area, loses the original details and quality of the audio, and is easy to cause passenger dissatisfaction; the sound energy leakage in the non-target listening area is still large, which can easily interfere with the driving experience of passengers in the non-target listening area, and cannot protect the privacy of the audio.

[0006] To address the above problems, the present invention provides a method for determining control filter coefficients, the method comprising:

[0007] Based on the output control filter coefficients, the first frequency response information corresponding to the target listening area and the listening sound contrast are obtained; wherein, the first frequency response information is determined based on the first sound signal received by the target listening area, and the listening sound contrast is the sound contrast between the first sound signal and the second sound signal received by the non-target listening area.

[0008] Based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the heard sound contrast and the target sound contrast, the total difference information is determined.

[0009] Based on the total difference information, the neural network is updated to generate the control filter coefficients for the next round of output. This process is repeated iteratively until the termination condition is met, resulting in the final generated control filter coefficients.

[0010] Optionally, the first frequency response information includes a first amplitude response and a first phase response, the target frequency response information includes a target amplitude response and a target phase response, the frequency response difference includes an amplitude response difference and a phase response difference, and the method further includes:

[0011] Calculate the amplitude response difference between the first amplitude response and the target amplitude response, and the phase response difference between the first phase response and the target phase response.

[0012] Optionally, determining the total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the acoustic contrast difference between the heard sound contrast and the target sound contrast, includes:

[0013] The total difference information is determined based on the frequency response difference and the corresponding preset frequency response weight, as well as the acoustic contrast difference and the corresponding preset acoustic contrast weight.

[0014] Optionally, the preset sound contrast weight includes a preset first sound contrast weight when the sound is in a preset frequency band and a preset second sound contrast weight when the sound is not in a preset frequency band, wherein the preset first sound contrast weight is greater than the preset second sound contrast weight.

[0015] Optionally, determining the total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the acoustic contrast difference between the heard sound contrast and the target sound contrast, further includes:

[0016] If the first sound signal is above a preset frequency, the total difference information is determined to be zero.

[0017] Optionally, the neural network is a fully connected network with a depth of four layers, including an input layer, two hidden layers, and an output layer; before obtaining the first frequency response information corresponding to the target listening area and the sound contrast based on the output control filter coefficients, the method further includes:

[0018] The initial frequency response information is input into the input layer; wherein the initial frequency response information is determined based on the initial control filter coefficients;

[0019] After processing by the two hidden layers, the control filter coefficients for the first round are output by the output layer;

[0020] The step of updating the neural network based on the total difference information to generate the control filter coefficients for the next round of output includes:

[0021] The total difference information is backpropagated to the neural network to update the parameters in the neural network;

[0022] After processing by the two hidden layers, the output layer outputs the control filter coefficients for the next round.

[0023] Optionally, the initial frequency response information includes frequency response information determined based on the initial control filter coefficients, and frequency response information obtained after different perturbation processes.

[0024] Optionally, the initial frequency response information includes frequency response information determined based on the initial control filter coefficients, and frequency response information obtained after different perturbation processes.

[0025] Optionally, obtaining the first frequency response information corresponding to the target listening area and the sound contrast based on the output control filter coefficients includes:

[0026] The first sound signal is determined based on the reference audio signal, the transfer function from the speaker to the microphone in the target listening area, and the control filter coefficients of the output; the second sound signal is determined based on the reference audio signal, the transfer function from the speaker to the microphone in the non-target listening area, and the control filter coefficients of the output; the reference audio signal is an audio signal that includes the frequency band to be controlled.

[0027] The first frequency response information is determined based on the first sound signal;

[0028] The sound contrast is determined based on the first sound signal and the second sound signal.

[0029] Optionally, the termination condition includes the total difference information being less than a preset threshold.

[0030] The present invention also provides a control filter coefficient determination device, the device comprising:

[0031] The information acquisition module is used to acquire first frequency response information and sound contrast corresponding to the target listening area based on the output control filter coefficients; wherein, the first frequency response information is determined based on the first sound signal received by the target listening area, and the sound contrast is the sound contrast between the first sound signal and the second sound signal received by the non-target listening area.

[0032] The difference determination module is used to determine total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the listening sound contrast and the target sound contrast.

[0033] The coefficient generation module is used to update the neural network based on the total difference information to generate the control filter coefficients for the next round of output. The process is iterated until the termination condition is met, and the final generated control filter coefficients are obtained.

[0034] Optionally, the first frequency response information includes a first amplitude response and a first phase response, the target frequency response information includes a target amplitude response and a target phase response, the frequency response difference includes an amplitude response difference and a phase response difference, and the device further includes:

[0035] The difference calculation module is used to calculate the difference in amplitude response between the first amplitude response and the target amplitude response, as well as the difference in phase response between the first phase response and the target phase response.

[0036] Optionally, the difference determination module includes:

[0037] The first difference determination submodule is used to determine the total difference information based on the frequency response difference and the corresponding preset frequency response weight, as well as the acoustic contrast difference and the corresponding preset acoustic contrast weight.

[0038] Optionally, the preset sound contrast weight includes a preset first sound contrast weight when the sound is in a preset frequency band and a preset second sound contrast weight when the sound is not in a preset frequency band, wherein the preset first sound contrast weight is greater than the preset second sound contrast weight.

[0039] Optionally, the difference determination module further includes:

[0040] The second difference determination submodule is used to determine the total difference information as zero when the first sound signal is above a preset frequency.

[0041] Optionally, the neural network is a fully connected network with a depth of four layers, including an input layer, two hidden layers, and an output layer; before obtaining the first frequency response information corresponding to the target listening area and the sound contrast based on the output control filter coefficients, the method further includes:

[0042] The initial frequency response information is input into the input layer; wherein the initial frequency response information is determined based on the initial control filter coefficients;

[0043] After processing by the two hidden layers, the control filter coefficients for the first round are output by the output layer;

[0044] The step of updating the neural network based on the total difference information to generate the control filter coefficients for the next round of output includes:

[0045] The total difference information is backpropagated to the neural network to update the parameters in the neural network;

[0046] After processing by the two hidden layers, the output layer outputs the control filter coefficients for the next round.

[0047] Optionally, the initial frequency response information includes frequency response information determined based on the initial control filter coefficients, and frequency response information obtained after different perturbation processes.

[0048] Optionally, the information acquisition module includes:

[0049] The signal determination submodule is used to determine the first sound signal based on a reference audio signal, the transfer function from the speaker to the microphone in the target listening area, and the control filter coefficients of the output; and to determine the second sound signal based on the reference audio signal, the transfer function from the speaker to the microphone in the non-target listening area, and the control filter coefficients of the output; the reference audio signal is an audio signal that includes the frequency band to be controlled;

[0050] The information determination submodule is used to determine the first frequency response information based on the first sound signal;

[0051] The contrast determination submodule is used to determine the sound contrast based on the first sound signal and the second sound signal.

[0052] Optionally, the termination condition includes the total difference information being less than a preset threshold.

[0053] This invention also discloses an electronic device, characterized in that it includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0054] Memory, used to store computer programs;

[0055] When a processor executes a program stored in memory, it implements the method steps described above.

[0056] This invention also discloses a readable storage medium, which, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to perform the method steps described above.

[0057] According to an embodiment of the present invention, by obtaining the first frequency response information and the sound contrast of the target listening area based on the output control filter coefficients, the first frequency response information is determined based on the first sound signal received by the target listening area, and the sound contrast is the sound contrast between the first sound signal and the second sound signal received by the non-target listening area. Based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the sound contrast and the target sound contrast, total difference information is determined. Based on the total difference information, the neural network is updated to generate the control filter coefficients for the next round of output. This process is iteratively executed until a termination condition is met, resulting in the final generated control filter coefficients. This allows for the iterative generation of suitable control filter coefficients through a neural network, targeting both frequency response and sound contrast. This ensures high sound contrast while maintaining the flatness of the spectrum in the target listening area, guaranteeing audio quality in the target listening area, reducing sound energy leakage in the non-target listening area, protecting audio privacy, and achieving audio non-interference between the target and non-target listening areas. This efficiently and quickly realizes the design of a personal audio zone within the vehicle. Attached Figure Description

[0058] Figure 1 A schematic diagram illustrating the generation principle of individual vocal registers is shown;

[0059] Figure 2 A flowchart illustrating the steps of a control filter coefficient determination method according to an embodiment of the present invention is shown.

[0060] Figure 3 A schematic diagram of the process for generating control filter coefficients based on a neural network is shown.

[0061] Figure 4 A flowchart illustrating the steps of a control filter coefficient determination method according to another embodiment of the present invention is shown.

[0062] Figure 5 A schematic diagram illustrating the principle of determining the control filter coefficients is shown;

[0063] Figure 6 This diagram illustrates a structural block diagram of an embodiment of a control filter coefficient determination device according to one embodiment of the present invention.

[0064] Figure 7 A block diagram of an electronic device for controlling the determination of filter coefficients is shown according to an exemplary embodiment. Detailed Implementation

[0065] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0066] Reference Figure 2 The diagram illustrates a flowchart of a method for determining control filter coefficients according to an embodiment of the present invention, which may specifically include the following steps:

[0067] Step 101: Based on the output control filter coefficients, obtain the first frequency response information corresponding to the target listening area and the sound contrast ratio; wherein, the first frequency response information is determined based on the first sound signal received by the target listening area, and the sound contrast ratio is the sound contrast ratio between the first sound signal and the second sound signal received by the non-target listening area.

[0068] In some embodiments of the present invention, the original audio signal is filtered and then used to drive a loudspeaker, according to the control filter coefficients. The loudspeaker can achieve different sound energy intensities in adjacent regions. The region with the desired strong sound energy is denoted as the target listening region, and the region with the desired weak sound energy is denoted as the non-target listening region.

[0069] Alternatively, the output control filter coefficients can be generated by a neural network, which is then used to optimize the control filter coefficients.

[0070] In some embodiments of the present invention, the sound signal received in the target listening area is denoted as the first sound signal. The sound signal received in a non-target listening area is denoted as the second sound signal.

[0071] In some embodiments of the present invention, frequency response information refers to information characterizing the relationship between steady-state output and input at a given frequency, specifically including the functional relationship between the ratio of output to input amplitude and input frequency, and the functional relationship between output to input phase difference and input frequency. The frequency response information determined based on the first sound signal is denoted as the first frequency response information.

[0072] Alternatively, the first frequency response information can be obtained from the first sound signal through a Fast Fourier Transform. Any applicable method can be used to determine this, and the embodiments of the present invention do not impose any limitations on it.

[0073] In some embodiments of the present invention, the acoustic contrast ratio is the acoustic contrast ratio between the first sound signal and the second sound signal, specifically the quotient between the acoustic energy of the first sound signal and the acoustic energy of the second sound signal.

[0074] For example, according to the frequency domain acoustic contrast formula, when the number of target and non-target listening areas is the same, the acoustic contrast between the target and non-target listening areas in the frequency domain can be expressed as:

[0075]

[0076] Where w{f} is the frequency response of the control filter at frequency block f, w H {f} is the conjugate transpose of w{f}, G B {f} is the frequency response of the transfer function from the loudspeaker to the target point in the target listening area at frequency block f. It is G B The conjugate transpose of {f}, G D {f} is the frequency response of the transfer function from the loudspeaker to the target point in the non-target listening area at frequency block f. It is G D The conjugate transpose of {f}.

[0077] Step 102: Determine the total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the heard sound contrast and the target sound contrast.

[0078] In some embodiments of the present invention, the target frequency response information refers to a preset frequency response information that serves as the target. The target acoustic contrast ratio refers to a preset acoustic contrast ratio that serves as the target.

[0079] In some embodiments of the present invention, the difference between the first frequency response information and the target frequency response information is determined and denoted as the frequency response difference. The difference between the listening sound contrast and the target sound contrast is determined and denoted as the sound contrast difference. The frequency response difference is used to keep the frequency response within the target listening area as flat as possible, ensuring the audio quality of the target listening area. The target frequency response is the result of a reference audio signal after pure delay filtering and linear attenuation. The sound contrast difference is used to ensure good sound contrast between the target listening area and the non-target listening area.

[0080] For example, a multi-objective cost function can be designed to train a neural network and obtain the control filter coefficients for PSZ. This cost function consists of several parts. The first part is the Euclidean distance between the first frequency response information corresponding to the target listening area and the target frequency response information, which serves as the frequency response difference. The second part is the Euclidean distance between the listening sound contrast and the target sound contrast, which serves as the sound contrast difference.

[0081] In some embodiments of the present invention, the specific implementation methods for determining the total difference information based on the frequency response difference and the acoustic contrast difference can include various approaches. For example, the total difference information can be obtained by weighted summation of the frequency response difference and the acoustic contrast difference. The total difference information can be determined using any applicable cost function, and the embodiments of the present invention do not impose any limitations on this.

[0082] Step 103: Update the neural network based on the total difference information to generate the control filter coefficients for the next round of output. Iterate until the termination condition is met to obtain the control filter coefficients generated in the last round.

[0083] In some embodiments of the present invention, the parameters of the neural network are updated based on the total difference information to obtain an updated neural network. The updated neural network is then used to generate new control filter coefficients, which serve as the control filter coefficients for the next round of output. Based on these control filter coefficients for the next round of output, steps 101 to 103 are executed iteratively until the termination condition is met, and the final generated control filter coefficients are obtained.

[0084] In one optional embodiment of the present invention, the termination condition may be that the total difference information is less than a preset threshold, or any other applicable termination condition, and the embodiments of the present invention do not limit this.

[0085] For example, such as Figure 3 The diagram illustrates the process of generating control filter coefficients based on a neural network. Based on the control filter coefficients W, a first sound signal and a second sound signal are obtained through the loudspeaker transmission path G. The first frequency response information is obtained through a Fast Fourier Transform. A cost function is constructed based on the objective function (i.e., target frequency response information) for the bright area (i.e., the target listening area) and the target sound contrast. The total difference information obtained from the cost function is fed into the backpropagation process to update the parameters of the neural network. The updated neural network is then used to generate the control filter coefficients for the next round of output.

[0086] According to an embodiment of the present invention, by obtaining the first frequency response information and the sound contrast of the target listening area based on the output control filter coefficients, the first frequency response information is determined based on the first sound signal received by the target listening area, and the sound contrast is the sound contrast between the first sound signal and the second sound signal received by the non-target listening area. Based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the sound contrast and the target sound contrast, total difference information is determined. Based on the total difference information, the neural network is updated to generate the control filter coefficients for the next round of output. This process is iteratively executed until a termination condition is met, resulting in the final generated control filter coefficients. This allows for the iterative generation of suitable control filter coefficients through a neural network, targeting both frequency response and sound contrast. This ensures high sound contrast while maintaining the flatness of the spectrum in the target listening area, guaranteeing audio quality in the target listening area, reducing sound energy leakage in the non-target listening area, protecting audio privacy, and achieving audio non-interference between the target and non-target listening areas. This efficiently and quickly realizes the design of a personal audio zone within the vehicle.

[0087] This invention can be used for the design of personal sound zones in enclosed spaces such as automobiles, and has a wide range of applications. The frequency response of the PSZ control filter obtained through the traditional ACC method is augmented and used as input to train a neural network model. Using a designed cost function as the objective, the neural network model extracts the features of the acoustic inverse system, thereby generating a suitable control filter. This filter then drives the loudspeaker to produce sound, constructing both the target and non-target listening zones, thus achieving privacy in audio playback.

[0088] In an optional embodiment of the present invention, the first frequency response information includes a first amplitude response and a first phase response, the target frequency response information includes a target amplitude response and a target phase response, the frequency response difference includes an amplitude response difference and a phase response difference, and may further include: calculating the amplitude response difference between the first amplitude response and the target amplitude response, and the phase response difference between the first phase response and the target phase response.

[0089] Frequency response includes amplitude response and phase response. The amplitude response is the real part, and the phase response is the imaginary part. By comparing the amplitude and phase responses separately, the first frequency response information includes the first amplitude response and the first phase response, while the target frequency response information includes the target amplitude response and the target phase response.

[0090] The difference between the first amplitude response and the target amplitude response is calculated and denoted as the amplitude response difference. The difference between the first phase response and the target phase response is calculated and denoted as the phase response difference. For example, the Euclidean distance between the first amplitude response and the target amplitude response is calculated, and the Euclidean distance between the first phase response and the target phase response is also calculated. The frequency response difference consists of two parts: the amplitude response difference and the phase response difference.

[0091] By comparing the amplitude response and phase response separately, the original details and quality of the audio are preserved because the difference in phase response is taken into account, thereby improving the audio quality of the target listening area.

[0092] In an optional embodiment of the present invention, a specific implementation of determining the total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the acoustic contrast difference between the listening sound contrast and the target sound contrast may include: determining the total difference information based on the frequency response difference and the corresponding preset frequency response weight, and the acoustic contrast difference and the corresponding preset acoustic contrast weight.

[0093] The weights pre-set for frequency response differences are denoted as preset frequency response weights. The weights pre-set for acoustic contrast differences are denoted as preset acoustic contrast weights. The product of the frequency response difference and its corresponding preset frequency response weight, and the product of the acoustic contrast difference and its corresponding preset acoustic contrast weight, are calculated. These two products are then added together to obtain the total difference information. Specifically, any applicable preset frequency response weight and preset acoustic contrast weight can be used according to actual needs; this embodiment of the invention does not impose any limitations on this. By adjusting the preset frequency response weights and preset acoustic contrast weights, different application scenarios and requirements can be flexibly adapted.

[0094] In an optional embodiment of the present invention, the preset sound contrast weight includes a preset first sound contrast weight when the sound is in a preset frequency band and a preset second sound contrast weight when the sound is not in a preset frequency band, wherein the preset first sound contrast weight is greater than the preset second sound contrast weight.

[0095] The preset sound contrast weights can be set in segments according to the frequency of the sound. Specifically, it can be divided into a preset first sound contrast weight when the sound is in a preset frequency band, and a preset second sound contrast weight when the sound is not in a preset frequency band. The preset first sound contrast weight is greater than the preset second sound contrast weight. By setting a greater weight for the preset frequency band, a stronger distinction in sound contrast can be achieved within the preset frequency band.

[0096] The preset frequency band refers to frequencies within a preset range, while the non-preset frequency band refers to frequencies outside the preset range. These can be set according to actual needs, and this embodiment of the invention does not impose any limitations on them. For example, the effective frequency band of PSZ is typically below 2kHz, therefore additional weight needs to be added to this band to ensure the effect of low-frequency sound contrast. Generally, a sound contrast of 20dB in the 0-2kHz frequency range means that the two sound registers are basically isolated, thus ensuring a high sound contrast.

[0097] In an optional embodiment of the present invention, a specific implementation of determining the total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the listening sound contrast and the target sound contrast, may further include: determining the total difference information as zero when the first sound signal is above a preset frequency.

[0098] To ensure good sound contrast in the low-frequency range, it's difficult to make the loudspeaker's operating frequency range very high. Furthermore, the acoustic transfer function of the system in the high-frequency range is easily affected by scattering from objects, leading to instability. To protect the loudspeaker and avoid frequency response instability in the high-frequency range, a penalty term is added to frequencies above a preset frequency, and PSZ (Power-Side Array) is not applied. Therefore, a piecewise function is used to determine the total difference information. When the first sound signal is above the preset frequency, the total difference information is directly set to zero. When the first sound signal is not above the preset frequency, the total difference information is determined in the same way. The preset frequency can be set to 3kHz, or any other applicable frequency value can be used as needed; this embodiment of the invention does not impose any restrictions on this.

[0099] For example, such as Figure 3 As shown, a penalty term for the high-frequency fitting effect is added to the cost function, that is, when the first sound signal is above the preset frequency, the result of the cost function is directly determined to be zero.

[0100] In an optional embodiment of the present invention, the neural network is a fully connected network with a depth of four layers. The neural network includes an input layer, two hidden layers, and an output layer. The input layer takes frequency response information as input, and the output layer takes the generated control filter coefficients as output.

[0101] To reduce the complexity of the neural network, a fully connected network with a depth of four layers was used. The input layer of the network contains the frequency response information of the control filter obtained by the ACC method, which includes the real and imaginary parts.

[0102] After processing through two hidden layers, the output layer yields the generated control filter coefficients. The training architecture used is as follows: Figure 3 As shown. A ReLU (Rectified Linear Unit) activation function can be added after a fully connected network to improve the network's non-linear expressive power. Stochastic gradient descent is used during training to minimize the loss function. Layer normalization and Dropout (a regularization method) are used to prevent overfitting.

[0103] Before obtaining the first frequency response information corresponding to the target listening area and the sound contrast based on the output control filter coefficients, the method may further include: inputting initial frequency response information into the input layer; wherein the initial frequency response information is determined based on the initial control filter coefficients; and after processing by the two hidden layers, the output layer outputs the control filter coefficients for the first round.

[0104] Initially, the frequency response information determined based on the initial control filter coefficients is denoted as the initial frequency response information. The initial frequency response information can be determined using the ACC method based on the initial control filter coefficients. This initial frequency response information is then input to the input layer. The input layer passes the input data to the hidden layers, and after processing by two hidden layers, the output layer outputs the generated control filter coefficients.

[0105] One specific method for updating the neural network based on the total difference information to generate the control filter coefficients for the next round of output may include: backpropagating the total difference information to the neural network to update the parameters in the neural network; and outputting the control filter coefficients for the next round by the output layer after processing by the two hidden layers.

[0106] In each iteration, the backpropagation algorithm is used to propagate the total difference information obtained above back into the neural network, and the parameters in the neural network are updated according to the total difference information.

[0107] After processing through two more hidden layers, the parameters in the neural network have been updated, so the control filter coefficients output by the output layer for the next round will change.

[0108] In an optional embodiment of the present invention, the initial frequency response information includes frequency response information determined based on the initial control filter coefficients, and frequency response information obtained after different perturbation processes.

[0109] Furthermore, based on the frequency response information determined by the initial control filter coefficients, different perturbations are added to obtain more frequency response information. Both the initial frequency response information determined by the control filter coefficients and the frequency response information obtained after different perturbation processes are used as the initial frequency response information input to the neural network. For example, based on the initial control filter coefficients, the frequency response information can be determined; in this process, different signal-to-noise ratios are added to obtain other frequency response information different from the above. This pre-trained input data can reduce the number of parameters and improve the network's convergence speed and generalization ability.

[0110] Reference Figure 4 The diagram illustrates a flowchart of a method for determining control filter coefficients according to another embodiment of the present invention, which may specifically include the following steps:

[0111] Step 201: Determine the first sound signal based on the reference audio signal, the transfer function from the speaker to the microphone in the target listening area, and the control filter coefficients of the output; determine the second sound signal based on the reference audio signal, the transfer function from the speaker to the microphone in the non-target listening area, and the control filter coefficients of the output; the reference audio signal is an audio signal that includes the frequency band to be controlled.

[0112] In some embodiments of the present invention, the reference audio signal refers to the original audio signal used as a reference, such as an audio source reference signal. The reference audio signal is an audio signal containing the frequency range to be controlled, such as Gaussian white noise. The microphone includes a microphone for the target listening area and a microphone for the non-target listening area.

[0113] In some embodiments of the present invention, such as Figure 5 The diagram illustrates the principle of determining the control filter coefficients. The design of the control filter in a personal voice zone depends on the secondary path, namely the transfer function from the loudspeaker to the microphone in the target listening area (bright zone) / non-target listening area (dark zone). There are multiple transfer functions, each corresponding to different loudspeakers and microphones. In the m-th listening area (target listening area / non-target listening area), the result of the loudspeaker's action can be expressed as:

[0114]

[0115] Among them, g s,m (n) represents the transfer function from the s-th loudspeaker to the m-th listening zone, x(n) represents the audio source reference signal, and w s (n) represents the control filter coefficients required for the PSZ system.

[0116] In some embodiments of the present invention, such as Figure 5 As shown, based on the reference audio signal, the transfer function from the speaker to the microphone in the target listening area, and the output control filter coefficients, a first sound signal emitted by the speaker and picked up by the microphone can be simulated. Similarly, based on the reference audio signal, the transfer function from the speaker to the microphone in the non-target listening area, and the output control filter coefficients, a second sound signal emitted by the speaker and picked up by the microphone can be simulated. For example, initially, the initial values ​​(frequency response information) of the ACC offline design are input into the neural network to obtain the output control filter coefficients. Then, based on the reference audio signal, the transfer function from the speaker to the microphone, and the output control filter coefficients, the first and second sound signals are simulated. This process is iteratively executed to continuously optimize the control filter coefficients.

[0117] Step 202: Determine the first frequency response information based on the first sound signal.

[0118] In this embodiment of the invention, the specific implementation of this step can be found in the description of the foregoing embodiments, and will not be repeated here.

[0119] Step 203: Determine the sound contrast based on the first sound signal and the second sound signal.

[0120] In this embodiment of the invention, the specific implementation of this step can be found in the description of the foregoing embodiments, and will not be repeated here.

[0121] Step 204: Determine the total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the heard sound contrast and the target sound contrast.

[0122] In this embodiment of the invention, the specific implementation of this step can be found in the description of the foregoing embodiments, and will not be repeated here.

[0123] Step 205: Update the neural network based on the total difference information to generate the control filter coefficients for the next round of output. Iterate until the termination condition is met to obtain the control filter coefficients generated in the last round.

[0124] In this embodiment of the invention, the specific implementation of this step can be found in the description of the foregoing embodiments, and will not be repeated here.

[0125] According to an embodiment of the present invention, a first sound signal is determined based on a reference audio signal, the transfer function of the microphone from the speaker to the target listening area, and the output control filter coefficients; a second sound signal is determined based on the reference audio signal, the transfer function of the microphone from the speaker to the non-target listening area, and the output control filter coefficients; the reference audio signal is an audio signal containing the frequency band to be controlled; the first frequency response information is determined based on the first sound signal; the listening contrast is determined based on the first sound signal and the second sound signal; and the listening contrast is determined based on the frequency response difference between the first frequency response information and the target frequency response information, and the listening contrast... The acoustic contrast difference between the target acoustic contrast and the target acoustic contrast is used to determine the total difference information. Based on the total difference information, the neural network is updated to generate the control filter coefficients for the next round of output. This process is iteratively executed until the termination condition is met, resulting in the final generated control filter coefficients. This allows for the iterative generation of suitable control filter coefficients through the neural network, targeting both frequency response and acoustic contrast. This ensures high acoustic contrast while also maintaining the flatness of the spectrum in the target listening area, guaranteeing audio quality in the target listening area, reducing sound energy leakage in non-target listening areas, and thus protecting audio privacy. This achieves audio non-interference between the target and non-target listening areas, thereby efficiently and quickly realizing the design of personal audio zones within the vehicle.

[0126] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0127] Reference Figure 6 The diagram illustrates a structural block diagram of a control filter coefficient determination device according to another embodiment of the present invention, which may specifically include the following modules:

[0128] The information acquisition module 301 is used to acquire first frequency response information and sound contrast corresponding to the target listening area based on the output control filter coefficients; wherein, the first frequency response information is determined based on the first sound signal received by the target listening area, and the sound contrast is the sound contrast between the first sound signal and the second sound signal received by the non-target listening area.

[0129] The difference determination module 302 is used to determine total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the listening sound contrast and the target sound contrast.

[0130] The coefficient generation module 303 is used to update the neural network according to the total difference information to generate the control filter coefficients for the next round of output, and iterate until the termination condition is reached to obtain the control filter coefficients generated in the last round.

[0131] Optionally, the first frequency response information includes a first amplitude response and a first phase response, the target frequency response information includes a target amplitude response and a target phase response, the frequency response difference includes an amplitude response difference and a phase response difference, and the device further includes:

[0132] The difference calculation module is used to calculate the difference in amplitude response between the first amplitude response and the target amplitude response, as well as the difference in phase response between the first phase response and the target phase response.

[0133] Optionally, the difference determination module includes:

[0134] The first difference determination submodule is used to determine the total difference information based on the frequency response difference and the corresponding preset frequency response weight, as well as the acoustic contrast difference and the corresponding preset acoustic contrast weight.

[0135] Optionally, the preset sound contrast weight includes a preset first sound contrast weight when the sound is in a preset frequency band and a preset second sound contrast weight when the sound is not in a preset frequency band, wherein the preset first sound contrast weight is greater than the preset second sound contrast weight.

[0136] Optionally, the difference determination module further includes:

[0137] The second difference determination submodule is used to determine the total difference information as zero when the first sound signal is above a preset frequency.

[0138] Optionally, the neural network is a fully connected network with a depth of four layers, including an input layer, two hidden layers, and an output layer; the device further includes:

[0139] An input module is used to input initial frequency response information into the input layer before obtaining the first frequency response information corresponding to the target listening area and the listening contrast based on the output control filter coefficients; wherein the initial frequency response information is determined based on the initial control filter coefficients;

[0140] The processing module is used to process the two hidden layers and output the control filter coefficients of the first round from the output layer;

[0141] The coefficient generation module includes:

[0142] The parameter update submodule is used to backpropagate the total difference information to the neural network and update the parameters in the neural network.

[0143] The processing module is also used to output the control filter coefficients for the next round from the output layer after processing by the two hidden layers.

[0144] Optionally, the initial frequency response information includes frequency response information determined based on the initial control filter coefficients, and frequency response information obtained after different perturbation processes.

[0145] Optionally, the information acquisition module includes:

[0146] The signal determination submodule is used to determine the first sound signal based on a reference audio signal, the transfer function from the speaker to the microphone in the target listening area, and the control filter coefficients of the output; and to determine the second sound signal based on the reference audio signal, the transfer function from the speaker to the microphone in the non-target listening area, and the control filter coefficients of the output; the reference audio signal is an audio signal that includes the frequency band to be controlled;

[0147] The information determination submodule is used to determine the first frequency response information based on the first sound signal;

[0148] The contrast determination submodule is used to determine the sound contrast based on the first sound signal and the second sound signal.

[0149] Optionally, the termination condition includes the total difference information being less than a preset threshold.

[0150] According to an embodiment of the present invention, by obtaining the first frequency response information and the sound contrast of the target listening area based on the output control filter coefficients, the first frequency response information is determined based on the first sound signal received by the target listening area, and the sound contrast is the sound contrast between the first sound signal and the second sound signal received by the non-target listening area. Based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the sound contrast and the target sound contrast, total difference information is determined. Based on the total difference information, the neural network is updated to generate the control filter coefficients for the next round of output. This process is iteratively executed until a termination condition is met, resulting in the final generated control filter coefficients. This allows for the iterative generation of suitable control filter coefficients through a neural network, targeting both frequency response and sound contrast. This ensures high sound contrast while maintaining the flatness of the spectrum in the target listening area, guaranteeing audio quality in the target listening area, reducing sound energy leakage in the non-target listening area, protecting audio privacy, and achieving audio non-interference between the target and non-target listening areas. This efficiently and quickly realizes the design of a personal audio zone within the vehicle.

[0151] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0152] Figure 7 This is a structural block diagram illustrating an electronic device 700 for controlling the determination of filter coefficients according to an exemplary embodiment. For example, the electronic device 700 may be a vehicle-mounted computer, a computer, a digital broadcasting terminal, a messaging device, a game console, a medical device, a fitness device, etc.

[0153] Reference Figure 7 The electronic device 700 may include one or more of the following components: a processing component 702, a memory 704, a power supply component 706, a multimedia component 708, an audio component 710, an input / output (I / O) interface 712, a sensor component 714, and a communication component 716.

[0154] Processing component 702 typically controls the overall operation of electronic device 700, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the methods described above. Furthermore, processing component 702 may include one or more modules to facilitate interaction between processing component 702 and other components. For example, processing component 702 may include a multimedia module to facilitate interaction between multimedia component 708 and processing component 702.

[0155] Memory 704 is configured to store various types of data to support the operation of device 700. Examples of this data include instructions for any application or method operating on electronic device 700, contact data, phonebook data, messages, pictures, videos, etc. Memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0156] Power supply component 706 provides power to various components of electronic device 700. Power supply component 706 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 700.

[0157] Multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 708 includes a front-facing camera and / or a rear-facing camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0158] Audio component 710 is configured to output and / or input audio signals. For example, audio component 710 includes a microphone (MIC) configured to receive external audio signals when electronic device 700 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 704 or transmitted via communication component 716. In some embodiments, audio component 710 also includes a speaker for outputting audio signals.

[0159] I / O interface 712 provides an interface between processing component 702 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0160] Sensor assembly 714 includes one or more sensors for providing state assessments of various aspects of electronic device 700. For example, sensor assembly 714 may detect the on / off state of device 700, the relative positioning of components such as the display and keypad of electronic device 700, changes in position of electronic device 700 or a component of electronic device 700, the presence or absence of user contact with electronic device 700, orientation or acceleration / deceleration of electronic device 700, and temperature changes of electronic device 700. Sensor assembly 714 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 714 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 714 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.

[0161] Communication component 716 is configured to facilitate wired or wireless communication between electronic device 700 and other devices. Electronic device 700 can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, communication component 716 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 716 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0162] In an exemplary embodiment, the electronic device 700 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0163] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 704 including instructions, which can be executed by a processor 720 of an electronic device 700 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0164] A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a terminal's processor, enables the terminal to perform a control filter coefficient determination method, the method comprising:

[0165] Based on the output control filter coefficients, the first frequency response information corresponding to the target listening area and the listening sound contrast are obtained; wherein, the first frequency response information is determined based on the first sound signal received by the target listening area, and the listening sound contrast is the sound contrast between the first sound signal and the second sound signal received by the non-target listening area.

[0166] Based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the heard sound contrast and the target sound contrast, the total difference information is determined.

[0167] Based on the total difference information, the neural network is updated to generate the control filter coefficients for the next round of output. This process is repeated iteratively until the termination condition is met, resulting in the final generated control filter coefficients.

[0168] Optionally, the first frequency response information includes a first amplitude response and a first phase response, the target frequency response information includes a target amplitude response and a target phase response, the frequency response difference includes an amplitude response difference and a phase response difference, and the method further includes:

[0169] Calculate the amplitude response difference between the first amplitude response and the target amplitude response, and the phase response difference between the first phase response and the target phase response.

[0170] Optionally, determining the total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the acoustic contrast difference between the heard sound contrast and the target sound contrast, includes:

[0171] The total difference information is determined based on the frequency response difference and the corresponding preset frequency response weight, as well as the acoustic contrast difference and the corresponding preset acoustic contrast weight.

[0172] Optionally, the preset sound contrast weight includes a preset first sound contrast weight when the sound is in a preset frequency band and a preset second sound contrast weight when the sound is not in a preset frequency band, wherein the preset first sound contrast weight is greater than the preset second sound contrast weight.

[0173] Optionally, determining the total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the acoustic contrast difference between the heard sound contrast and the target sound contrast, further includes:

[0174] If the first sound signal is above a preset frequency, the total difference information is determined to be zero.

[0175] Optionally, the neural network is a fully connected network with a depth of four layers, including an input layer, two hidden layers, and an output layer; before obtaining the first frequency response information corresponding to the target listening area and the sound contrast based on the output control filter coefficients, the method further includes:

[0176] The initial frequency response information is input into the input layer; wherein the initial frequency response information is determined based on the initial control filter coefficients;

[0177] After processing by the two hidden layers, the control filter coefficients for the first round are output by the output layer;

[0178] The step of updating the neural network based on the total difference information to generate the control filter coefficients for the next round of output includes:

[0179] The total difference information is backpropagated to the neural network to update the parameters in the neural network;

[0180] After processing by the two hidden layers, the output layer outputs the control filter coefficients for the next round.

[0181] Optionally, the initial frequency response information includes frequency response information determined based on the initial control filter coefficients, and frequency response information obtained after different perturbation processes.

[0182] Optionally, the initial frequency response information includes frequency response information determined based on the initial control filter coefficients, and frequency response information obtained after different perturbation processes.

[0183] Optionally, obtaining the first frequency response information corresponding to the target listening area and the sound contrast based on the output control filter coefficients includes:

[0184] The first sound signal is determined based on the reference audio signal, the transfer function from the speaker to the microphone in the target listening area, and the control filter coefficients of the output; the second sound signal is determined based on the reference audio signal, the transfer function from the speaker to the microphone in the non-target listening area, and the control filter coefficients of the output; the reference audio signal is an audio signal that includes the frequency band to be controlled.

[0185] The first frequency response information is determined based on the first sound signal;

[0186] The sound contrast is determined based on the first sound signal and the second sound signal.

[0187] Optionally, the termination condition includes the total difference information being less than a preset threshold.

[0188] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0189] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0190] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0191] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0192] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0193] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0194] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0195] The foregoing has provided a detailed description of a control filter coefficient determination method, a control filter coefficient determination device, an electronic device, and a readable storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A method for determining the coefficients of a control filter, characterized in that, The method includes: Based on the output control filter coefficients, the first frequency response information corresponding to the target listening area and the listening sound contrast are obtained; wherein, the first frequency response information is determined based on the first sound signal received by the target listening area, and the listening sound contrast is the sound contrast between the first sound signal and the second sound signal received by the non-target listening area. Based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the heard sound contrast and the target sound contrast, the total difference information is determined. Based on the total difference information, the neural network is updated to generate the control filter coefficients for the next round of output. This process is repeated iteratively until the termination condition is met, resulting in the final generated control filter coefficients.

2. The method according to claim 1, characterized in that, The first frequency response information includes a first amplitude response and a first phase response; the target frequency response information includes a target amplitude response and a target phase response; the frequency response difference includes an amplitude response difference and a phase response difference; the method further includes: Calculate the amplitude response difference between the first amplitude response and the target amplitude response, and the phase response difference between the first phase response and the target phase response.

3. The method according to claim 1, characterized in that, The step of determining the total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the heard sound contrast and the target sound contrast, includes: The total difference information is determined based on the frequency response difference and the corresponding preset frequency response weight, as well as the acoustic contrast difference and the corresponding preset acoustic contrast weight.

4. The method according to claim 3, characterized in that, The preset sound contrast weights include a preset first sound contrast weight when the sound is in a preset frequency band and a preset second sound contrast weight when the sound is not in a preset frequency band, wherein the preset first sound contrast weight is greater than the preset second sound contrast weight.

5. The method according to claim 1, characterized in that, The step of determining the total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the heard sound contrast and the target sound contrast, further includes: If the first sound signal is above a preset frequency, the total difference information is determined to be zero.

6. The method according to claim 1, characterized in that, The neural network is a fully connected network with a depth of four layers, including an input layer, two hidden layers, and an output layer; before obtaining the first frequency response information corresponding to the target listening area and the sound contrast based on the output control filter coefficients, the method further includes: The initial frequency response information is input into the input layer; wherein the initial frequency response information is determined based on the initial control filter coefficients; After processing by the two hidden layers, the control filter coefficients for the first round are output by the output layer; The step of updating the neural network based on the total difference information to generate the control filter coefficients for the next round of output includes: The total difference information is backpropagated to the neural network to update the parameters in the neural network; After processing by the two hidden layers, the output layer outputs the control filter coefficients for the next round.

7. The method according to claim 6, characterized in that, The initial frequency response information includes frequency response information determined based on the initial control filter coefficients, and frequency response information obtained after different perturbation processes.

8. The method according to claim 1, characterized in that, The step of obtaining the first frequency response information corresponding to the target listening area and the sound contrast based on the output control filter coefficients includes: The first sound signal is determined based on the reference audio signal, the transfer function from the speaker to the microphone in the target listening area, and the control filter coefficients of the output; the second sound signal is determined based on the reference audio signal, the transfer function from the speaker to the microphone in the non-target listening area, and the control filter coefficients of the output; the reference audio signal is an audio signal that includes the frequency band to be controlled. The first frequency response information is determined based on the first sound signal; The sound contrast is determined based on the first sound signal and the second sound signal.

9. The method according to claim 1, characterized in that, The termination condition includes the total difference information being less than a preset threshold.

10. A control filter coefficient determination device, characterized in that, The device includes: The information acquisition module is used to acquire first frequency response information and sound contrast corresponding to the target listening area based on the output control filter coefficients; wherein, the first frequency response information is determined based on the first sound signal received by the target listening area, and the sound contrast is the sound contrast between the first sound signal and the second sound signal received by the non-target listening area. The difference determination module is used to determine total difference information based on the frequency response difference between the first frequency response information and the target frequency response information, and the sound contrast difference between the listening sound contrast and the target sound contrast. The coefficient generation module is used to update the neural network based on the total difference information to generate the control filter coefficients for the next round of output. The process is iterated until the termination condition is met, and the final generated control filter coefficients are obtained.

11. The apparatus according to claim 10, characterized in that, The first frequency response information includes a first amplitude response and a first phase response; the target frequency response information includes a target amplitude response and a target phase response; the frequency response difference includes an amplitude response difference and a phase response difference; the device further includes: The difference calculation module is used to calculate the difference in amplitude response between the first amplitude response and the target amplitude response, as well as the difference in phase response between the first phase response and the target phase response.

12. The apparatus according to claim 10, characterized in that, The difference determination module includes: The first difference determination submodule is used to determine the total difference information based on the frequency response difference and the corresponding preset frequency response weight, as well as the acoustic contrast difference and the corresponding preset acoustic contrast weight.

13. The apparatus according to claim 10, characterized in that, The difference determination module further includes: The second difference determination submodule is used to determine the total difference information as zero when the first sound signal is above a preset frequency.

14. The apparatus according to claim 10, characterized in that, The neural network is a fully connected network with a depth of four layers, including an input layer, two hidden layers, and an output layer; the device further includes: An input module is used to input initial frequency response information into the input layer before obtaining the first frequency response information corresponding to the target listening area and the listening contrast based on the output control filter coefficients; wherein the initial frequency response information is determined based on the initial control filter coefficients; The processing module is used to process the two hidden layers and output the control filter coefficients of the first round from the output layer; The coefficient generation module includes: The parameter update submodule is used to backpropagate the total difference information to the neural network and update the parameters in the neural network. The processing module is also used to output the control filter coefficients for the next round from the output layer after processing by the two hidden layers.

15. The apparatus according to claim 10, characterized in that, The information acquisition module includes: The signal determination submodule is used to determine the first sound signal based on a reference audio signal, the transfer function from the speaker to the microphone in the target listening area, and the control filter coefficients of the output; and to determine the second sound signal based on the reference audio signal, the transfer function from the speaker to the microphone in the non-target listening area, and the control filter coefficients of the output; the reference audio signal is an audio signal that includes the frequency band to be controlled; The information determination submodule is used to determine the first frequency response information based on the first sound signal; The contrast determination submodule is used to determine the sound contrast based on the first sound signal and the second sound signal.

16. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the method described in any one of claims 1-9.

17. A readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device is able to perform the method as described in any one of claims 1-9.