Responsive audio equalizer adjustment method, system and device and medium

Through the responsive audio equalizer adjustment method, the audio equalizer parameters are automatically adjusted using spectrum characteristics and interfering sound characteristics, solving the problem of time-consuming and professional knowledge in setting traditional audio equalizers, achieving more efficient and personalized audio output.

CN119943089APending Publication Date: 2025-05-06HANSONG NANJING TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510021013.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The settings of traditional audio equalizers require manual adjustments by users, which are time-consuming and require users to have certain audio knowledge, making it difficult to adapt to changes in different audio content or listening environments.

Method used

A responsive audio equalizer adjustment method is provided. By obtaining the spectrum characteristics of the audio to be tuned and the interference sound characteristics in the environment, adjusting parameters are automatically determined and the audio equalizer is controlled for parameter adjustment.

Benefits of technology

It realizes automatic optimization of the parameter settings of the audio equalizer, improves the sound quality and user experience, and makes the audio output more accurate and personalized, adapting to different environments and user preferences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119943089A_ABST
    Figure CN119943089A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a response type audio equalizer adjusting method, system and device and a medium, and the method comprises the steps: obtaining a to-be-adjusted audio, and determining the spectrum characteristics of the to-be-adjusted audio; determining interference sound features based on the sound data; determining adjustment parameters based on the spectrum features and the interference sound features; and based on the adjustment parameter, controlling the tuning unit to adjust the parameter of the audio equalizer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of equalizer adjustment technology, and in particular to a responsive audio equalizer adjustment method, system, device and medium. Background Art

[0002] The equalizer of audio equipment is an audio processing device that can be used to adjust the volume level of audio signals in different frequency ranges, thereby changing the spectral characteristics of the audio. The equalizer is often adjusted directly through buttons, with a single adjustment method, and requires users to have a high level of music theory knowledge to make appropriate adjustments.

[0003] Therefore, a responsive audio equalizer adjustment method, system, device and medium are provided to help automatically optimize the parameter settings of the audio equalizer and improve the sound quality and user experience. Summary of the invention

[0004] One or more embodiments of the present specification provide a responsive audio equalizer adjustment method, the method comprising: obtaining an audio to be adjusted and determining a frequency spectrum feature of the audio to be adjusted; determining an interfering sound feature based on sound data; determining an adjustment parameter based on the frequency spectrum feature and the interfering sound feature; and controlling a tuning unit to adjust a parameter of an audio equalizer based on the adjustment parameter.

[0005] One or more embodiments of the present specification provide a responsive audio equalizer adjustment system, the system comprising: a first determination module, configured to obtain an audio to be adjusted and determine a frequency spectrum feature of the audio to be adjusted; a second determination module, configured to determine an interfering sound feature based on sound data; a third determination module, configured to determine an adjustment parameter based on the frequency spectrum feature and the interfering sound feature; and a control module, configured to control a tuning unit to adjust a parameter of an audio equalizer based on the adjustment parameter.

[0006] One or more embodiments of the present specification provide a responsive audio equalizer adjustment device, characterized in that the device includes at least one processor and at least one memory; the at least one memory is used to store computer instructions; the at least one processor is used to execute at least part of the computer instructions to implement the responsive audio equalizer adjustment method.

[0007] One or more embodiments of the present specification provide a computer-readable storage medium, wherein the storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer executes the responsive audio equalizer adjustment method. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] This specification will be further described in the form of exemplary embodiments, which will be described in detail by the accompanying drawings. These embodiments are not restrictive, and in these embodiments, the same number represents the same structure, wherein:

[0009] Figure 1 is an application scenario diagram of a responsive audio equalizer adjustment system according to some embodiments of this specification;

[0010] Figure 2 is a structural schematic diagram of a responsive audio equalizer adjustment system according to some embodiments of this specification;

[0011] Figure 3 is an exemplary flow chart of a responsive audio equalizer adjustment method according to some embodiments of this specification;

[0012] Figure 4 is an exemplary schematic diagram of a prediction model according to some embodiments of this specification;

[0013] Figure 5 is an exemplary schematic diagram of an adjustment model according to some embodiments of this specification;

[0014] Figure 6 This is an exemplary schematic diagram of determining user preference data according to some embodiments of this specification. DETAILED DESCRIPTION

[0015] In order to more clearly illustrate the technical solutions of the embodiments of this specification, the following is a brief introduction to the drawings required for the description of the embodiments. Obviously, the drawings described below are only some examples or embodiments of this specification. For ordinary technicians in this field, this specification can also be applied to other similar scenarios based on these drawings without creative work. Unless it is obvious from the language environment or otherwise explained, the same reference numerals in the figures represent the same structure or operation.

[0016] It should be understood that the "system", "device", "unit" and / or "module" used herein are a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the words can be replaced by other expressions.

[0017] As shown in this specification and claims, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural. Generally speaking, the terms "comprises" and "includes" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0018] Flowcharts are used in this specification to illustrate the operations performed by the system according to the embodiments of this specification. It should be understood that the preceding or following operations are not necessarily performed precisely in order. Instead, the steps may be processed in reverse order or simultaneously. At the same time, other operations may also be added to these processes, or one or more operations may be removed from these processes.

[0019] An audio equalizer can adjust the gain of frequency ranges such as bass, midrange, and treble in music, thereby changing the timbre and sound quality of the music. By adjusting the audio equalizer, you can enhance or weaken the sound of a specific frequency, making the music sound clearer, more balanced, or adapt to different auditory environments. However, traditional audio equalizers are based on hardware DSP (digital signal processor) to implement parameter settings. The settings of audio equalizers are usually fixed and require users to manually adjust them to adapt to different audio content or listening environments. However, manual adjustment is not only time-consuming, but also requires users to have certain audio knowledge. It is difficult for ordinary users to operate and cannot respond to environmental changes or personal preference changes instantly.

[0020] In view of this, some embodiments of the present specification provide a responsive audio equalizer adjustment method, system, device and medium, which automatically optimize the parameter settings of the audio equalizer by real-time monitoring and analysis of the audio content being played and the user's environment. This not only adapts to the user's personal preferences, but also can cope with environmental changes, thereby improving the sound quality and user experience, and making the audio output more accurate and personalized.

[0021] Figure 1 It is a schematic diagram of an application scenario of a responsive audio equalizer adjustment system according to some embodiments of this specification.

[0022] The application scenario 100 of the responsive audio equalizer adjustment system (hereinafter referred to as the application scenario 100) can determine the audio equalizer parameters by implementing the method and / or process disclosed in this specification. In some embodiments, the application scenario 100 can correspond to a spatial scene where one or more speakers are installed. For example, the system 200 can be applied to KTV, cinemas, home theaters, etc. In some embodiments, the system 200 can be used to determine the adjustment parameters based on the spectral characteristics of the audio to be adjusted and the interfering sound characteristics.

[0023] In some embodiments, Figure 1 As shown, the application scenario 100 may include a storage device 110 , a network 120 , a processor 130 , an audio device 140 , a user terminal 150 , and a sound pickup device 160 .

[0024] The storage device 110 is used to store data, instructions and / or any other information. The storage device 110 may include one or more storage components, each of which may be an independent device or part of another device. In some embodiments, the storage device 110 may include a random access memory (RAM), a read-only memory (ROM), a removable memory, etc. or any combination thereof. In some embodiments, the storage device 110 may be connected to the network 120 to communicate with one or more components in the application scenario 100. In some embodiments, the storage device 110 may be part of the processor 130.

[0025] The network 120 may include any suitable network capable of facilitating information and / or data exchange of the application scenario 100. In some embodiments, one or more components of the application scenario 100 (e.g., the storage device 110, the processor 130, the audio device 140, the user terminal 150, etc.) may exchange information and / or data with one or more components of the application scenario 100 via the network 120.

[0026] The processor 130 is used to process information and / or data related to the application scenario 100. In some embodiments, the processor 130 can process data, information and / or processing results obtained from other devices or system components, and execute program instructions based on these data, information and / or processing results to perform one or more functions described in this specification. For example, the processor 130 can obtain the audio to be adjusted and determine the spectral characteristics of the audio to be adjusted; determine the interfering sound characteristics based on the sound data; determine the adjustment parameters based on the spectral characteristics and the interfering sound characteristics; and control the tuning unit to adjust the parameters of the audio equalizer based on the adjustment parameters.

[0027] The audio device 140 is a device for playing sound. The audio device 140 may include multiple types. For example, it can be divided into main power amplifier audio, monitoring audio and return listening audio according to the purpose of the audio device; for another example, it can be divided into full-band audio, bass audio and subwoofer audio according to the playback frequency; for another example, it can be divided into closed audio, bass reflex audio, transmission line audio, etc. according to the box structure. In some embodiments, the audio device 140 may include multiple components such as an equalizer, an amplifier, a speaker, a DVD, etc. In some embodiments, the audio device 140 can be connected to the network 120 to communicate with one or more components of the application scenario 100.

[0028] The user terminal 150 refers to one or more terminal devices or software used by the user. The user may refer to an operator or an administrator. In some embodiments, the user terminal 150 may interact with other components (such as the processor 130, etc.) in the application scenario 100 via the network 120. For example, the user terminal 150 may send adjustment parameters to the processor 130 via the network 120. The above examples are only used to illustrate the wide range of the user terminal 150 device range and are not intended to limit its range.

[0029] In some embodiments, the user terminal 150 may include a mobile device 150 - 1 , a computer 150 - 2 , a laptop computer 150 - 3 , etc. or any combination thereof.

[0030] In some embodiments, the user terminal 150 may include a display component (eg, a display screen, etc.), an interactive component (eg, a mouse, a keyboard, a touch screen, etc.), and the like.

[0031] The sound pickup device 160 is a device for capturing and converting sound wave signals into electronic signals. For example, the sound pickup device 160 can be used to collect sound data in a space. In some embodiments, the sound pickup device 160 may include a sound card, a voice recorder, an array microphone, or other environmental monitoring equipment. The sound pickup device 160 can be installed at a preset position in the space. For example, in places such as conference rooms, classrooms, recording studios, and live broadcast rooms, the sound pickup device 160 can be installed on the ceiling, wall, or desktop to capture the sound in the room. In some embodiments, the sound pickup device 160 can be an independent device or a part of other devices. For example, the sound pickup device 160 is a built-in microphone on a user terminal 150, a speaker device 140, or a wearable device.

[0032] It is worth noting that the application scenario 100 is provided for illustrative purposes only and is not intended to limit the scope of this specification. For those of ordinary skill in the art, various changes and modifications can be made according to the description of this specification. For example, the application scenario 100 can also include a database, an information source, etc. For another example, the application scenario 100 can be implemented on other devices to achieve similar or different functions. However, these changes and modifications will not deviate from the scope of this specification.

[0033] Figure 2 It is a structural schematic diagram of a responsive audio equalizer adjustment system shown in some embodiments of this specification.

[0034] like Figure 2 As shown, the responsive audio equalizer adjustment system 200 may include a first determination module 210 , a second determination module 220 , a third determination module 230 , and a control module 240 .

[0035] The first determination module 210 is configured to obtain the audio to be adjusted and determine the frequency spectrum characteristics of the audio to be adjusted;

[0036] The second determination module 220 is configured to determine the interfering sound feature based on the sound data.

[0037] In some embodiments, the second determination module is further configured to: acquire sound data based on an acquisition cycle; separate the sound data to obtain at least one single sound signal; identify at least one single sound signal to determine the spectral characteristics of a single interfering sound; and determine the interfering sound characteristics based on the spectral characteristics of the single interfering sound.

[0038] The third determination module 230 is configured to determine the adjustment parameter based on the frequency spectrum characteristics and the interference sound characteristics.

[0039] In some embodiments, the third determination module is further configured to: determine whether the sound data satisfies the user preference data; and in response to a negative result, determine the adjustment parameter based on the interfering sound characteristics, the spectrum characteristics, and the user preference data.

[0040] The control module 240 is configured to control the tuning unit to adjust the parameters of the audio equalizer based on the adjustment parameters.

[0041] In some embodiments, the responsive audio equalizer adjustment system 200 also includes a fourth determination module 250, which is configured to: determine the user's corresponding reference preference data and build a preference database based on the user's historical adjustment data and its corresponding historical spectral characteristics and historical interference sound characteristics; determine the user's preference data based on the spectral characteristics of the audio to be adjusted, the interference sound characteristics of the sound data, and the preference database.

[0042] For more information about the aforementioned modules and their functions, see Figure 3-Figure 6 and its related description.

[0043] In some embodiments, part or all of the first determination module 210 , the second determination module 220 , the third determination module 230 , the control module 240 , and the fourth determination module 250 may be integrated into a processor.

[0044] It should be noted that the above description of the audio equipment equalizer adjustment system and its modules is only for convenience of description and cannot limit this specification to the scope of the embodiments. It is understandable that for those skilled in the art, after understanding the principle of the system, it is possible to arbitrarily combine the modules or form a subsystem to connect with other modules without deviating from this principle.

[0045] Figure 3is an exemplary flow chart of a responsive audio equalizer adjustment method according to some embodiments of the present specification.

[0046] In some embodiments, process 300 may be implemented based on a responsive audio equalizer adjustment system or processor. Figure 3 As shown, process 300 includes the following steps:

[0047] Step 310: Acquire the audio to be adjusted and determine the frequency spectrum characteristics of the audio to be adjusted.

[0048] Audio to be adjusted refers to audio files that are waiting to be adjusted or directly used for playback in audio processing or sound equipment. For example, audio to be adjusted can include music playlists, radio live broadcasts, online audio streaming services, recorded voice messages, or any other type of audio content.

[0049] In some embodiments, the audio to be adjusted may be an audio stream composed of a series of single audio signals. The single audio signals may come from different sound sources, such as guitar, bass, drums, keyboard, lead vocals, etc. The single audio signal is a component of the audio to be adjusted. The single audio signal is an audio segment from a specific sound source that has not been affected by the environment or physically transformed by the playback system.

[0050] In some embodiments, the audio to be adjusted includes an audio file currently being played. Currently being played means being played through a speaker, headphones or other audio equipment. Currently refers to the moment when the adjustment parameters need to be determined or a user-defined time period reasonably close to the moment (for example, 1 minute, 2 minutes or 5 minutes before the current moment).

[0051] In some embodiments, the processor may obtain the audio to be adjusted in a variety of ways. For example, the processor may obtain the audio to be adjusted through a storage device. For another example, the processor may download the audio to be adjusted through a third-party website (e.g., an audio library or an online resource website, etc.).

[0052] Spectral features are parameters used to describe and quantify the characteristics of audio signals. For example, the spectral features of the audio to be adjusted may include the frequency, phase, amplitude, timbre, etc. of the audio to be adjusted. Timbre is a parameter used to distinguish audio to be adjusted from different sources. For example, timbre may include the harmonics of the audio to be adjusted.

[0053] In some embodiments, the frequency spectrum characteristics of the audio to be adjusted are used to describe the distribution of the audio to be adjusted on different frequency components. For example, the frequency spectrum characteristics of the audio to be adjusted may include an amplitude spectrum and a phase spectrum. The amplitude spectrum represents the intensity of the audio to be adjusted on each frequency component, and the phase spectrum represents the phase information of the audio to be adjusted on each frequency component.

[0054] In some embodiments, the processor can determine the spectral characteristics of the audio to be adjusted in a variety of ways. For example, based on the audio to be adjusted, the processor can convert the audio to be adjusted from the time domain to the frequency domain through spectrum analysis, determine the various frequency components of the audio to be adjusted in the frequency domain as the frequency components of the audio to be adjusted; determine the intensity of the audio to be adjusted on each frequency component by calculating the absolute value or square value of the complex number corresponding to each frequency component in the frequency domain; determine the phase information of the audio to be adjusted on each frequency component by calculating the phase angle of the complex number corresponding to each frequency component in the frequency domain. Spectral analysis includes but is not limited to short-time Fourier transform, wavelet transform, etc.

[0055] In some embodiments, the processor may perform preprocessing based on the audio to be adjusted to obtain the preprocessed audio to be adjusted, and perform spectrum analysis based on the preprocessed audio to be adjusted. The preprocessing includes but is not limited to framing, windowing, discretization, etc.

[0056] Step 320, determining interference sound characteristics based on the sound data.

[0057] The sound data refers to various sound signals captured from a specific environment when the audio file is played. The specific environment may refer to the environment where the user is located.

[0058] In some embodiments, the sound data corresponds to the audio to be adjusted. For example, the sound data is acquired by collecting the sound data at the same time point or the same time period when the audio to be adjusted is played.

[0059] In some embodiments, the sound data may include multiple types of sound signals. For example, the sound data may include the audio to be adjusted, and other sound signals. Other sound signals include but are not limited to outdoor sound, indoor sound, background noise, etc.

[0060] Outdoor sound refers to the sound signals outside that come from nature, human activities or equipment, etc. For example, environmental sounds can include water flow, animal calls, and the sound of machines running in industrial production.

[0061] Indoor sound refers to the sound signals in indoor places such as homes, offices, and shops. For example, indoor environmental sounds may include the operating sounds of appliances such as air conditioners, refrigerators, and printers, and the echo signals of audio to be tuned.

[0062] Background noise refers to low-decibel sound signals that are constantly present. For example, background noise can include distant traffic noise, the hum of people talking, or the hum of an air conditioner or fan.

[0063] In some embodiments, the sound data may be an audio stream composed of a series of single sound signals. The single sound signal may come from different sound sources, such as the audio to be adjusted, outdoor sound, indoor sound, background noise, etc. For more information about the single sound signal, please refer to the relevant description below.

[0064] In some embodiments, the processor may obtain sound data in a variety of ways. For example, the processor may detect the environment where the user is located based on the sound pickup device by periodic continuous acquisition to obtain sound data. The periodic continuous acquisition refers to continuous acquisition at intervals of a certain period of time.

[0065] In some embodiments, the intervals of the acquisition time may be distributed at equal intervals. For example, the processor may control at least one microphone to acquire sound data at the 1st second, the 2nd second, ..., the 10th second. The intervals of the acquisition time may be determined by system default or artificial preset. In some embodiments, the acquisition time intervals may also be distributed at non-equal intervals, for example, when the playback system starts to play the audio to be adjusted, the sound data is collected.

[0066] The interfering sound feature refers to a frequency spectrum feature of the interfering sound in the sound data.

[0067] Interference sound refers to a sound signal in sound data that affects the playback effect (eg, quality and clarity, etc.) of the audio to be adjusted.

[0068] In some embodiments, the interfering sound may be an audio stream composed of a series of single interfering sounds. The single interfering sound may come from different sound sources, such as sound sources of outdoor environment, indoor environment, background noise, etc.

[0069] It should be noted that a single audio signal is the original component of the audio to be adjusted, a single interference sound is a component of the interference sound in the environment, and a single sound signal is any independent sound component in the sound data. A single sound signal can be a part of the audio to be adjusted or a part of the interference sound in the environment.

[0070] In some embodiments, the processor may determine the interfering sound features based on the sound data in a variety of ways. For example, the processor may determine the spectral features of the sound data through spectrum analysis based on the preprocessed sound data; the spectral features of the audio to be adjusted are removed from the spectral features of the sound data, and the remaining spectral features are the determined interfering sound features. Exemplarily, in the frequency domain, the frequency components of the audio to be adjusted are subtracted from the frequency components of the sound data to determine the frequency components of the interfering sound.

[0071] In some embodiments, the interference sound characteristics include spectral characteristics of a single interference sound. The processor can obtain sound data based on an acquisition cycle; separate the sound data to obtain at least one single sound signal; identify at least one single sound signal to determine the spectral characteristics of the single interference sound; and determine the interference sound characteristics based on the spectral characteristics of the single interference sound.

[0072] The acquisition period refers to the time interval between two consecutive acquisitions of sound data. The acquisition period affects the adjustment frequency of the adjustment parameters. The longer the acquisition period, the lower the adjustment frequency of the adjustment parameters. For example, the acquisition period is per second, per minute, etc. In some embodiments, the acquisition period can be a system default value, a manually preset value, or determined based on experience or experiments. The adjustment frequency refers to the number of times the parameters of the audio equalizer are adjusted within a specified time period.

[0073] In some embodiments, the processor may obtain acquisition information of two adjacent acquisitions; calculate the change rate of the acquisition information based on the difference between the two adjacent acquisition information and the time difference between the acquisition times of the two adjacent acquisition information; and determine the acquisition period based on the change rate of the acquisition information and the second preset rule. The acquisition time refers to the moment when a certain sound wave signal is captured.

[0074] The collected information refers to data related to the sound data collected by the sound collection device. For example, the collected information can be the sound data directly, or can be the interference sound features corresponding to the sound data.

[0075] Exemplarily, the second preset rule is: the greater the rate of change of the collected information, the shorter the collection period. For example, the processor can calculate the difference in amplitude of the sound data collected twice on the same frequency component, and determine the amplitude change rate of each frequency component based on the ratio of the difference between the two and the time difference. For another example, the processor can convert the phase information of each frequency component of the sound data collected twice into a complex number form; calculate the difference between the two complex numbers, and pass the real and imaginary parts of the difference between the two complex numbers as parameters to the atan2 function to determine the phase difference on the corresponding frequency component; based on the ratio of the above phase difference to the time difference, determine the phase change rate of each frequency component.

[0076] The two adjacent acquisitions refer to the process of acquiring sound data in two adjacent acquisition cycles on the time axis. For example, the two adjacent acquisition cycles may be any two adjacent acquisition cycles before the current acquisition cycle, or may be the two acquisition cycles closest to the current acquisition cycle.

[0077] It should be noted that a large rate of change in the collected information indicates possible sudden noise or environmental changes, and it is necessary to increase the frequency of collecting sound data and adjust the parameters of the audio equalizer in a timely manner to ensure the user's auditory experience.

[0078] In some embodiments, the processor may evaluate the interference degree of the audio to be adjusted based on the interference sound characteristics; and determine the sampling period based on the interference degree.

[0079] The interference level refers to the degree to which the audio to be adjusted is affected by environmental noise or other interference sources during transmission or playback.

[0080] In some embodiments, the interference degree may be expressed as a numerical value (eg, interference degree, interference value, etc.) or a level (eg, interference level).

[0081] In some embodiments, the processor may calculate the similarity between each single interfering sound in the interfering sound feature and the specific feature of the audio to be adjusted; and determine the interference degree according to the similarity of each specific feature and the amplitude of each single interfering sound by a first preset rule. Exemplarily, the first preset rule is: the higher the similarity of the specific feature, the greater the amplitude of the single interfering sound, and the higher the interference degree.

[0082] Specific features are feature vectors used to distinguish different sounds. For example, specific features include vectors composed of parameters such as timbre, frequency, and phase. It should be noted that specific features do not include amplitude, which affects the volume of the sound signal. The volume of the same sound signal is different at different locations.

[0083] Calculating the similarity of specific features can be achieved by calculating Euclidean distance, cosine distance, etc.

[0084] Exemplarily, the first preset rule is determined based on formula (1):

[0085] B=Σ(k1*si+k2*Ai)(1), where B is the interference degree of the audio to be adjusted; k1 and k2 are the first coefficient and the second coefficient respectively; si is the similarity of the specific characteristics between the i-th single interfering sound and the audio to be adjusted; Ai is the amplitude of the i-th single interfering sound.

[0086] The first coefficient and the second coefficient are coefficients greater than 0, and the first coefficient and the second coefficient can be determined based on experiments or experience.

[0087] In some embodiments, the processor may determine the interference characteristics of the future audio playback in the future time period based on the future spectrum characteristics and the future interference characteristics through acoustic calculation; and determine the degree of interference based on the interference characteristics. For more information about this embodiment, please refer to Figure 4 Related description.

[0088] In some embodiments, the sampling period is negatively correlated with the interference degree. For example, the higher the interference degree, the shorter the sampling period.

[0089] In some embodiments of the present specification, by dynamically adjusting the sampling period, it is possible to better adapt to changing environmental conditions. In a high-noise environment, by shortening the sampling period, it is possible to respond to noise changes more quickly and update the adjustment parameters in a timely manner, thereby effectively reducing the impact of noise on audio quality. In a relatively quiet environment, extending the sampling period can reduce unnecessary computing burdens and save resources.

[0090] In some embodiments, when the audio device starts playing an audio file, the processor controls the sound pickup device at the start of each collection cycle to start collecting sound data of the user's environment.

[0091] A single sound signal refers to a sound signal from a single sound source collected by a sound pickup device at a specific time point or time period. For example, a single sound signal may include the call of a specific type of animal (such as a bird or a wolf howl), the sound produced by a person speaking, etc. In an acoustic environment, sound is the result of the superposition of sound signals from multiple sound sources.

[0092] In some embodiments, the processor can represent the spectral characteristics of the sound data through a spectrogram; determine the fundamental frequencies of various sound sources based on the peaks in the spectrogram; determine the harmonics of a sound source based on the integer multiple frequencies of the fundamental frequency of the sound source; and determine single sound signals of different sound sources in the sound data based on the fundamental frequencies and harmonics of each sound source.

[0093] In some embodiments, the processor may separate single sound signals of different sound sources in the sound data based on the sound data by using source separation techniques such as Independent Component Analysis (ICA) or Non-negative Matrix Factorization (NMF).

[0094] Single interfering sound refers to an interfering sound signal from a single sound source at a specific time point or time period. For example, single interfering sound may include but is not limited to environmental noise (such as the sound of machinery running, the sound of people talking, etc.), white noise (such as rain, the sound of air conditioning running, etc.), echo signal of the audio to be adjusted, etc.

[0095] In some embodiments, the processor may identify at least one single sound signal in a variety of ways to determine the spectral characteristics of a single interfering sound. For example, the processor may generate a corresponding search vector based on a combination of the spectral characteristics of the audio to be adjusted and the spectral characteristics of a single sound signal corresponding to the audio to be adjusted, search in an interfering sound database based on the retrieval vector, determine a combination that meets the matching condition, and use the single sound signal in the combination as a single interfering sound, and use the spectral characteristics of the single sound signal in the combination as the spectral characteristics of the single interfering sound.

[0096] The matching condition may refer to a judgment condition for determining a target vector. The matching condition may include that the vector distance from the search vector is less than a distance threshold, etc. There are many methods for calculating vector distance, such as Euclidean distance, cosine distance, etc.

[0097] The interference sound database is a database used to store, index and query vectors. The database can store multiple reference vectors. The reference vector includes the spectral features of the historical audio and the spectral features of a single historical interference sound corresponding to the historical audio.

[0098] In some embodiments, the interference sound database may be preset. For example, the interference sound database may include a plurality of preset frequency spectrum features of historical audio and frequency spectrum features of one or more single interference sounds corresponding to each of the plurality of preset frequency spectrum features of historical audio.

[0099] In some embodiments, when the user uses the audio equalizer, the processor can build an interference sound database based on historical adjustment data. For example, the processor can obtain the historical audio to be adjusted each time it is played, and determine the historical single interference sound corresponding to the historical audio to be adjusted. Exemplarily, the processor can use the historical single sound signal with the largest energy drop after adjustment as the historical single interference sound corresponding to the historical audio to be adjusted each time the user actively adjusts the parameters of the audio equalizer, and store the spectral features corresponding to the historical audio to be adjusted and the spectral features corresponding to the historical single interference sound in association.

[0100] It should be noted that, since a single sound signal includes the audio to be adjusted and its echo signal, and the similarity of the spectral characteristics between the audio to be adjusted and its echo signal is high (e.g., the similarity of timbre, frequency, etc. is very high), the echo signal of the audio to be adjusted is not stored in the interference sound database. If the echo signal of the audio to be adjusted is also stored in the interference sound database, the system may mistakenly identify the audio to be adjusted itself as interference sound during actual processing.

[0101] In some embodiments, through the above method, the processor can determine the frequency spectrum characteristics of one or more single interfering sounds corresponding to the audio to be adjusted.

[0102] A single interfering sound from the same source will have different effects when playing different audios to be adjusted. If the effect on the audio to be adjusted is very small, it will not be considered as the single interfering sound corresponding to the audio to be adjusted. By building an interfering sound database through historical adjustment records, the accuracy of the determined single interfering sound can be improved.

[0103] Among them, the historical adjustment record refers to the record of historical adjustment data. The historical adjustment data is the data of the parameters of the audio equalizer that the user actively adjusted in the past. For example, the historical adjustment record may include multiple historical adjustment data, each of which includes parameter settings, adjustment timestamps, audio content identifiers, historical sound data before adjustment, historical sound data after adjustment, etc. Parameter settings refer to the specific parameters of the audio equalizer adjusted by the user. Audio content identifiers refer to the specific audio files or audio streams that are adjusted, reflecting the personalized needs of users when listening to different types of music, movies, podcasts, etc.

[0104] In some embodiments, the processor may determine the interfering sound features by comparing the spectral features of one or more single interfering sounds of the audio to be adjusted obtained through the above retrieval with the spectral features of the corresponding echo signal. The processor may measure the similarity between each of the single sound signals and the audio to be adjusted through a cross-correlation function, and determine the one with the greatest similarity as the echo signal of the audio to be adjusted.

[0105] In some embodiments of the present specification, by separating sound data into multiple single sound signals, different interference sources can be accurately identified and located; by identifying the single sound signal corresponding to the audio to be adjusted, it helps to accurately identify the corresponding interfering sound characteristics, and helps to eliminate or weaken interference of specific frequencies. Compared with traditional one-size-fits-all noise reduction, it is more efficient and will not damage the quality of the audio to be adjusted.

[0106] Step 330: determining adjustment parameters based on the frequency spectrum characteristics of the audio to be adjusted and the interfering sound characteristics.

[0107] The adjustment parameters are data for adjusting and setting the parameters of the audio equalizer, wherein the adjustment parameters may include the gain or attenuation of each frequency component, as well as the filter type and bandwidth, etc.

[0108] In some embodiments, the processor may determine the adjustment parameters in a variety of ways based on the spectral features of the audio to be adjusted and the interfering sound features. For example, the processor may generate a corresponding search vector based on the spectral features of the audio to be adjusted and the interfering sound features, search in the user adjustment database, and determine the adjustment parameters. Exemplarily, the processor determines the adjustment parameters based on a method similar to the above method of retrieving the spectral features of a single interfering sound.

[0109] The user adjustment database is a database used to store, index and query information related to the user's adjustment data. The user adjustment database can store multiple reference vectors and historical adjustment parameters corresponding to each reference vector. The reference vector can include historical spectrum features, historical interference sound features, etc.

[0110] In some embodiments, the user adjustment database can be constructed based on historical data. For example, each time a user actively adjusts the parameters of an audio equalizer, the historical frequency spectrum characteristics of the historical audio to be adjusted, the historical interference sound characteristics of the historical sound data, and the historical adjustment parameters corresponding to the adjustment are combined into historical adjustment data, and the historical frequency spectrum characteristics and historical interference sound characteristics in the historical adjustment data are used as reference vectors, and the reference vectors are associated with the corresponding historical adjustment parameters and stored.

[0111] It should be noted that when determining the adjustment parameters, the interfering sound features correspond to the spectral features of the audio to be adjusted. For example, the interfering sound features are spectral features of the interfering sound collected at the same time point or time period when the audio to be adjusted is played.

[0112] In some embodiments, the processor may determine whether the sound data satisfies the user preference data; and in response to a negative result, determine the adjustment parameters based on the interfering sound characteristics, the frequency spectrum characteristics of the audio to be adjusted, and the user preference data.

[0113] User preference data refers to the specific needs and preferences of the user for the playback effect of the audio to be adjusted. For example, the user preference data may include the volume, timbre, noise threshold, etc. that the user prefers. The preferred timbre is used to express the sound that the user wants to be gained (such as human voice, music, treble part, mid-range part, bass part, etc.) and the sound that is attenuated (such as background music) in the audio to be adjusted. The noise threshold refers to the maximum noise decibel that the user can accept.

[0114] In some embodiments, the processor may determine the user preference data in a variety of ways. For example, the processor may directly ask the user about their preferences for the playback effect by setting an interface or conducting a questionnaire survey.

[0115] In some embodiments, the processor may obtain adjustment parameters preset by the user, and determine the user preference data based on the preset adjustment parameters.

[0116] The pre-set adjustment parameters directly reflect the user's preference for audio playback effects, but there are some limitations. The pre-set adjustment parameters do not take into account the impact of environmental factors on the playback effect, so the actual playback effect may not achieve the ideal playback effect of the adjustment parameters. The processor can perform statistical analysis based on historical adjustment records to determine user preference data. For more information on determining user preference data, see Figure 6Related description.

[0117] Whether the sound data meets the user preference data refers to the degree of matching between the frequency spectrum characteristics of the sound data and the user preference data. For example, whether the sound data meets the user preference data may include: whether the frequency distribution of the sound data matches the timbre preferred by the user, whether the ambient noise exceeds the noise threshold acceptable to the user, etc.

[0118] In some embodiments, the processor can determine whether the sound data meets the user preference data based on a variety of methods. For example, the processor can generate an actual sound waveform based on the sound data collected by the microphone array; generate an ideal sound waveform based on the user preference data and the audio to be adjusted; determine whether the sound data meets the user preference data based on the similarity between the ideal sound waveform and the actual sound waveform; if the similarity is greater than a similarity threshold, then determine that the sound data meets the user preference data.

[0119] Among them, the similarity threshold can be a system preset value or a system default value. The similarity threshold can also be determined based on historical adjustment records. For example, each time the user completes an adjustment, the similarity between the actual sound waveform after adjustment and the preset sound waveform is calculated; the average of the similarities corresponding to each adjustment is taken as the similarity threshold.

[0120] The preset sound waveform is the signal waveform corresponding to the audio to be adjusted under the adjustment parameters preset by the user. Under different adjustment parameters, the audio to be adjusted corresponds to different signal waveforms.

[0121] The ideal sound waveform refers to the signal waveform of the audio to be adjusted without external interference, distortion and noise.

[0122] The actual sound waveform refers to the signal waveform diagram presented after the audio to be adjusted is played under specific environment and equipment conditions.

[0123] In some embodiments, the processor can process each single sound signal in the audio to be adjusted through audio processing software or programming libraries (such as Python's Librosa, Soundfile, etc.), and adjust the audio characteristics of each single sound signal in the audio to be adjusted (such as the gain, amplitude, dynamic range, etc. of a certain frequency band) to make the audio to be adjusted consistent with the user preference data; the processor can compose the audio to be adjusted based on each adjusted single sound signal, and establish a corresponding time domain waveform diagram as the ideal sound waveform of the audio to be adjusted.

[0124] In some embodiments, the processor can divide the audio to be adjusted into audio segments of different frequency bands (such as low frequency, medium frequency and high frequency, etc.) based on the spectral characteristics of the audio to be adjusted; determine the interference degree of each audio segment; adjust the audio characteristics of each audio segment according to the interference degree of each audio segment to simulate the interference of the actual environment on the audio segment; generate a corresponding actual sound waveform based on each adjusted audio segment. For example, if the interference degree of an audio segment of the audio to be adjusted is 20%, the amplitude of the audio segment is proportionally reduced by 20%. Exemplarily, the processor can determine the interference degree of each audio segment based on formula (1).

[0125] In some embodiments, when the similarity is less than the similarity threshold, it is determined that the sound data does not meet the user preference data, and the processor can determine the adjustment parameters based on the actual sound waveform and the ideal sound waveform. For example, the processor can determine the difference between the actual sound waveform and the ideal sound waveform in each frequency band; based on the difference in each frequency band, generate corresponding adjustment parameters. For example, in the high frequency band, if the amplitude of the actual sound waveform is attenuated by 20% compared to the ideal sound waveform, the processor will use the compensation amount of the amplitude in this frequency band as an adjustment parameter to make the actual sound waveform closer to the ideal sound waveform.

[0126] The processor may determine the compensation amount based on a ratio of an amplitude of an ideal sound waveform in a certain frequency band to an amplitude of an actual sound waveform in the same frequency band.

[0127] In some embodiments, when the similarity is greater than a similarity threshold, it is determined that the sound data satisfies the user preference data, that is, the current playback effect meets the user's preference, and at this time, there is no need to adjust the parameters of the audio equalizer.

[0128] In some embodiments of the present specification, by dynamically determining adjustment parameters, not only can a personalized, high-quality audio experience be provided, but also efficient, energy-saving and healthy audio output can be achieved in different environments, greatly improving the user experience; by considering user preference data, the audio output can be adjusted, such as enhancing bass, improving clarity or adjusting the tone balance to make it more in line with personal hearing preferences.

[0129] In some embodiments, the processor may determine echo characteristics based on sound data; determine interfered audio in future played audio based on the echo characteristics and future played audio; determine echo attenuation parameters based on the interfered audio and echo characteristics; and determine adjustment parameters based on the echo attenuation parameters.

[0130] The echo feature refers to the feature related to the sound reflection phenomenon generated by the physical environment (such as a room or outdoor space) during the audio signal output process.

[0131] In some embodiments, the echo characteristics include a time difference between a designated sound signal and its echo signal, and an energy loss between the designated sound signal and its echo signal.

[0132] The designated sound signal refers to a single sound signal corresponding to each single audio signal in the audio to be adjusted. Among them, the single audio signal is the sound signal in the creation or recording stage, and the single sound signal is the sound signal captured in the actual environment. The single sound signal is the manifestation of the single audio signal in the real world and has undergone physical processes such as air propagation, reflection, and absorption.

[0133] The time difference refers to the time interval between the acquisition time of a given sound signal and its echo signal. The time difference depends on the propagation path length of the sound wave of the given sound signal in the environment.

[0134] Energy loss refers to the attenuation of the energy of the echo signal relative to the energy of the corresponding designated sound signal. In some embodiments, energy loss is related to the attenuation of the designated sound signal when propagating in the medium, and increases with the increase of propagation distance.

[0135] In some embodiments, the processor can separate a designated sound signal and its corresponding echo signal from the sound data; calculate the energy of the echo signal and the energy of the designated sound signal based on the spectral characteristics of the designated sound signal and its corresponding echo signal; determine the energy loss between the designated sound signal and its echo signal based on the ratio of the energy of the echo signal to the energy of the designated sound signal; determine the acquisition time of the designated sound signal and the acquisition time of its echo signal based on the sound pickup device, and determine the difference between the two acquisition times as the time difference. For more information about separation, please refer to Figure 2 The relevant description above.

[0136] Future-playing audio refers to audio content that has not yet started playing but is scheduled to play in the future.

[0137] In some embodiments, the future-played audio and the audio to be adjusted may belong to the same audio file, or may belong to different audio files.

[0138] In some embodiments, the processor can obtain the play queue and predict the future play audio to be played in the future time period based on the order of the play queue.

[0139] The future time period refers to a period after the current time. The future time period is associated with the future played audio and is the time interval when the future played audio is about to be played or is planned to be played.

[0140] The play queue includes a list of audio files that are scheduled to be played. In some embodiments, the play queue can be automatically generated based on the user's selection, playlist, random play logic or recommendation algorithm.

[0141] Interfered audio refers to an audio clip that is easily affected by external environmental factors (such as echo, noise, etc.) during the audio playback process, resulting in reduced sound quality or poor listening experience.

[0142] In some embodiments, the processor may, based on the time difference between the designated sound signal and its echo signal, use the audio segment corresponding to the moment when the echo is generated in the future played audio as the interfered audio.

[0143] The echo generation time refers to the time point when a specified sound signal is emitted from a sound source and is captured again by a sound pickup device (eg, a microphone) after being reflected.

[0144] The echo reduction parameter is a parameter for reducing or counteracting the influence of the echo signal.

[0145] In audio processing, the echo reduction parameter may include a parameter for adjusting the amplitude of a specified sound signal to compensate for the sound quality loss or interference caused by the echo signal.

[0146] In some embodiments, the processor may calculate the amplitude of the echo signal based on the energy loss in the echo feature, and based on the amplitude of the echo signal, increase the amplitude of the interfered audio to offset the impact of the echo signal. The increase in amplitude may be proportional to the amplitude of the echo signal to ensure that the compensation effect is maximized without causing audio overload.

[0147] In some embodiments, the energy loss is proportional to the square of the amplitude of the echo signal. The greater the energy loss, the greater the square of the amplitude of the echo signal.

[0148] In some embodiments, the processor may merge and update the echo attenuation parameter with the previously determined adjustment parameter to obtain an updated adjustment parameter.

[0149] In some embodiments of the present specification, if there is an echo signal in the environment, and without changing the environment, its interference with the playback effect is continuous and difficult to change, while in theory the interference between sounds is mutual, and for the interfered sound signal, by enhancing its amplitude, the influence of the echo signal on the playback effect can be offset; by determining the echo attenuation parameters, the system can specifically adjust the audio to be adjusted, effectively suppress the echo, and improve the clarity and auditory experience of the audio to be adjusted.

[0150] In some embodiments, the processor may determine the adjustment parameter based on the estimated adjustment effect corresponding to at least one candidate adjustment parameter. Figure 5 Related description.

[0151] Step 340: Based on the adjustment parameters, control the tuning unit to adjust the parameters of the audio equalizer.

[0152] The tuning unit is a component used to adjust and optimize the audio signal. For example, the tuning unit may include a compressor, an audio equalizer, a limiter, etc.

[0153] In some embodiments, the tuning unit may be a hardware device or a software program.

[0154] The parameters of the audio equalizer are used to adjust the performance (eg, amplitude, frequency response, etc.) of the audio signal at different frequency components.

[0155] In some embodiments, the parameters of the audio equalizer include, but are not limited to, frequency, gain, filter type, bandwidth, etc.

[0156] In some embodiments, the processor may generate corresponding control instructions based on the adjustment parameters and send them to the tuning unit; the tuning unit may adjust the gain, center frequency, bandwidth, or change the filter type of the audio equalizer based on the control instructions.

[0157] In some embodiments of the present specification, the parameters of the audio equalizer are dynamically adjusted through the spectral characteristics of the audio to be adjusted and the characteristics of the ambient noise, so as to meet the preferences and needs of the audience in different scenarios and provide a more personalized audio experience; in a noisy environment, by identifying and analyzing the characteristics of the ambient noise, the parameters of the audio equalizer can be adjusted in a targeted manner to reduce the impact of noise on the audio signal, thereby improving the clarity of speech or music and the overall audio quality; the system can maintain the optimal state of audio output in different environments and scenarios, and can automatically adapt to the best sound quality whether in an outdoor concert, a noisy cafe or a quiet library.

[0158] It should be noted that the above description of the relevant process is only for example and explanation, and does not limit the scope of application of this specification. For those skilled in the art, various modifications and changes can be made to the process under the guidance of this specification. However, these modifications and changes are still within the scope of this specification.

[0159] Figure 4 is an exemplary schematic diagram of a prediction model according to some embodiments of the present specification.

[0160] In some embodiments, Figure 4As shown, the processor can predict the future interference feature 430 through the prediction model 420 based on the interference sound feature 410-1; determine the interference feature 440 of the future played audio 410-2 in the future time period through acoustic calculation based on the future spectrum feature 410-5 and the future interference feature 430; and determine the degree of interference 450 of the future played audio based on the interference feature 440.

[0161] The future interference characteristics refer to the frequency spectrum characteristics of the future interference sounds.

[0162] The future interference sound refers to the interference sound in the environment when the future playback audio is played in the future time period.

[0163] A predictive model is a model used to predict future disturbance characteristics.

[0164] In some embodiments, the prediction model is a machine learning model. For example, the prediction model may include any one or combination of a convolutional neural network (CNN) model, a neural network (NN) model, or other custom model structures.

[0165] In some embodiments, the input of the prediction model includes a sequence of interference sound features, and the output of the prediction model includes future interference features.

[0166] The interference sound feature sequence refers to a sequence composed of interference sound features corresponding to the sound data collected in multiple consecutive historical collection periods.

[0167] It should be noted that when the next collection cycle is reached, the previous collection cycle is the historical collection cycle. Continuous means that the collection cycles are arranged continuously on the time axis, such as the first collection cycle, the second collection cycle, the third collection cycle, etc.

[0168] In some embodiments, Figure 4 As shown, the input of the prediction model 420 also includes future playback audio 410 - 2 , echo features 410 - 3 , and adjustment parameters 410 - 4 .

[0169] For more information about playing audio, echo characteristics, and adjusting parameters in the future, see Figure 3 The steps 320 and 330 are described in detail.

[0170] In some embodiments of the present specification, by incorporating future played audio, echo features, and adjustment parameters into the input of the prediction model, the prediction model can consider the audio content and the potential impact of dynamic adjustment on interference, such as specific echo patterns that may be caused by the frequency characteristics of the audio, thereby more accurately predicting future interference features.

[0171] In some embodiments, the prediction model can be trained based on a large number of first training samples with a first label through various feasible methods. For example, parameter updating can be performed based on the gradient descent method. An exemplary training process includes: inputting a plurality of first training samples with a first label into an initial prediction model, constructing a loss function through the results of the first label and the initial prediction model, and iteratively updating the parameters of the initial prediction model based on the loss function through gradient descent or other methods. When the preset conditions are met, the model training is completed and a trained prediction model is obtained. Among them, the preset conditions may be that the loss function converges, the number of iterations reaches a threshold, etc.

[0172] In some embodiments, the first training sample may include multiple groups of training samples, each group of training samples includes at least a sample interference sound feature sequence of a first period of time. The first training sample may be obtained based on historical data.

[0173] In some embodiments, the first label may include the actual interference sound feature of the second time period corresponding to the sample interference sound feature sequence. The first label may be obtained by a processor or manual annotation. For example, the processor may determine the corresponding actual interference sound feature based on the historical sound data of the second time period, and determine it as the first label. The first time period and the second time period are historical time periods, and the first time period is before the second time period.

[0174] In some embodiments, when the input of the prediction model includes future playback audio, echo features, and adjustment parameters, the first training sample may also include sample future playback audio, sample echo features, and sample adjustment parameters. The sample future playback audio refers to the audio that needs to be played in the second time period; the sample echo features refer to the echo features corresponding to the audio that needs to be played in the second time period; and the sample adjustment parameters include the adjustment parameters corresponding to the first time period.

[0175] The future spectrum feature refers to the spectrum feature of the audio played in the future. For example, the future spectrum feature may include the frequency distribution of the audio played in the future.

[0176] Interference features are features related to interference reactions between sounds. For example, interference features may include a single sound signal being interfered with, the degree of interference, and the source of interference. The source of interference refers to the source of the interfering sound that interferes with the single sound signal.

[0177] Acoustic computing refers to the algorithms or mappings used to determine interference characteristics.

[0178] In some embodiments, the acoustic calculation may be implemented based on steps S31 to S33.

[0179] Step S31, determining a target pairing signal based on a frequency difference between a future single interfering sound and a future single audio signal.

[0180] The future single interference sounds are components of the future interference sounds, and each future single interference sound originates from a different sound source associated with the future interference sounds.

[0181] The future single audio signal is a component of the future played audio, and each future single audio signal originates from a sound source associated with the future audio to be adjusted.

[0182] In some embodiments, the processor performs spectrum analysis on each future single audio signal and each future single interfering sound, and converts them into frequency domain representation; based on the results of the spectrum analysis, extracts the fundamental frequency of each future single audio signal and each future single interfering sound; pairs each future single audio signal with each future single interfering sound, and illustratively, if there are N future single audio signals and M future single interfering sounds, N×M pairings will be generated; for each pair of paired signals, calculate the difference in fundamental frequency between the two; based on multiple pairs of paired signals, select paired signals whose frequency difference is less than a preset threshold as target paired signals, and the target paired signals are close enough in frequency and may affect each other during playback. The preset threshold can be a system preset value or a system default value.

[0183] Step S32: for the target pairing signal, determine the interference type based on the phase difference.

[0184] Interference type describes the nature of the interaction between two or more sound signals when they meet. For example, interference types include constructive interference and destructive interference.

[0185] Constructive interference refers to the interference reaction between paired signals that are superimposed and enhanced. Destructive interference refers to the interference reaction between paired signals that cancel each other out.

[0186] Phase difference refers to the phase difference between paired signals at the same moment or time.

[0187] In some embodiments, when the phase difference of the target paired signal is an even multiple of π, it indicates that the interference type is constructive interference; if the phase difference of the target paired signal is an odd multiple of π, it indicates that the interference type is destructive interference.

[0188] Step S33, determining the interference degree based on the interference type.

[0189] In some embodiments, for target pairing signals that interfere destructively, the processor may determine the interference degree (hereinafter referred to as the first interference degree) by calculating the ratio of the absolute difference in amplitude between the target pairing signals and the amplitude of the future single audio signal in the target pairing signals.

[0190] In some embodiments, for constructively interfering target pairing signals, the processor may determine the interference degree (hereinafter referred to as the second interference degree) by calculating the ratio of the sum of the amplitudes of the target pairing signals and the amplitude of the future single audio signal in the target pairing signals.

[0191] In some embodiments, when the phase difference between the target pairing signals does not belong to the above two cases, for example, the phase difference is any value between 0 and π; the processor can calculate the product of the amplitude of the future single interfering sound in the target pairing signal and the cosine value of the corresponding phase difference; and calculate the absolute value of the sum of the above product and the amplitude of the future single audio signal in the target pairing signal; determine the degree of interference based on the ratio of the above sum value to the amplitude of the future single audio signal in the target pairing signal.

[0192] In some embodiments, the processor may determine the interference degree of different parts of the future audio playback (such as the high-pitched part, the middle-pitched part, the low-pitched part, or the voice parts of different people, etc.) respectively. For example, the greater the interference degree of the future interference sound on a certain part of the future audio playback, the higher the interference degree of the future audio playback of the part.

[0193] In some embodiments, when sounds of different frequencies propagate in space, they may be subject to different types and degrees of interference. The processor may divide the future played audio into future audio segments of different frequency bands according to the frequency range, such as the treble part (high frequency band), the mid-range part (mid-frequency band) and the bass part (low frequency band); for each future single interfering sound, determine its interference type and degree with a future audio segment of a certain frequency band; for a future audio segment of each frequency band, accumulate the degrees of interference caused by all future single interfering sounds on the frequency band to determine the degree of interference of the future audio segment.

[0194] In some embodiments of the present specification, by predicting future interference characteristics, it is helpful to make adjustments in advance, rather than just reacting to current or past interference, so that the system can pre-optimize the audio signal, reduce the interference suffered by future audio playback during actual playback, and improve the audio quality; through the prediction model, it is possible to learn complex patterns from historical data and predict the type, intensity and frequency distribution of future environmental noise, so as to more accurately estimate the degree of interference of future audio playback, which helps to adapt to complex and dynamic environments.

[0195] Figure 5 is an exemplary schematic diagram of an adjustment model according to some embodiments of the present specification.

[0196] In some embodiments, the processor can generate at least one candidate adjustment parameter based on the deviation between the actual sound waveform corresponding to the sound data and the ideal sound waveform corresponding to the user preference data; determine the estimated adjustment effect corresponding to the candidate adjustment parameter through the adjustment model; and determine the adjustment parameter based on the estimated adjustment effect corresponding to at least one candidate adjustment parameter.

[0197] Deviation refers to the difference between the actual sound waveform and the ideal sound waveform. For example, the deviation may include a difference in volume, a difference in frequency response, a difference in timbre, etc.

[0198] In some embodiments, the processor may determine the deviation between the actual sound waveform and the ideal sound waveform in a variety of ways. For example, the processor may obtain the loudness values ​​of the actual sound waveform and the ideal sound waveform based on the actual sound waveform and the ideal sound waveform through a loudness measurement algorithm; determine the difference in volume between the actual sound waveform and the ideal sound waveform based on the loudness value of the actual sound waveform and the loudness value of the ideal sound waveform.

[0199] For another example, the processor can obtain a spectrum graph corresponding to an actual sound waveform (hereinafter referred to as the actual spectrum graph) and a spectrum graph corresponding to an ideal sound waveform (hereinafter referred to as the ideal spectrum graph) based on spectrum analysis; based on the amplitude of the spectrum graphs of the two at each frequency point, the difference in frequency response is quantified through a spectrum distance metric (such as Mean Square Error, MSE).

[0200] For another example, the processor determines the timbre feature vectors of the actual sound waveform and the ideal sound waveform; based on the timbre feature vectors of the two, the vector distance is calculated to determine the difference in timbre. The timbre features include but are not limited to Mel-frequency cepstral coefficients (MFCCs), zero crossing rate, spectral kurtosis, etc. The method of calculating the vector distance includes measurement methods such as Euclidean distance and Mahalanobis distance.

[0201] The candidate adjustment parameters refer to various parameters or parameter combinations to be determined as adjustment parameters, and various parameter combinations can set the frequency response, dynamic range, gain, delay, etc. of the audio to be adjusted.

[0202] In some embodiments, the processor can accumulate or subtract the adjustment step size based on the historical adjustment parameters, execute multiple times, and obtain multiple candidate adjustment parameters. The adjustment direction is determined according to the size relationship between the actual sound waveform and the ideal sound waveform. For example, when an audio characteristic (such as frequency response, volume, etc.) of the actual sound waveform is lower than the ideal sound waveform, the adjustment direction is accumulation; conversely, if it is higher than the ideal sound waveform, the adjustment direction is subtraction. The historical adjustment parameter can be the adjustment parameter of the most recent adjustment.

[0203] The adjustment step size refers to the amount of change in the adjustment direction. The adjustment step size can be determined based on the similarity between the actual sound waveform and the ideal sound waveform. For example, the adjustment step size is negatively correlated with the above similarity. The higher the similarity, the smaller the adjustment step size. The adjustment step size is the distance between two adjacent candidate adjustment parameters. For more information about the similarity between the actual sound waveform and the ideal sound waveform, please refer to Figure 3 Related description.

[0204] The adjustment model is a model for determining the estimated adjustment effects corresponding to the candidate adjustment parameters.

[0205] In some embodiments, the adjustment model is a machine learning model. For example, the adjustment model may include any one or combination of a convolutional neural network (CNN) model, a neural network (NN) model, or other custom model structures.

[0206] In some embodiments, Figure 5 As shown, the input of the adjustment model 520 may include interference sound characteristics 410-1, spectral characteristics 510-1 of the audio to be adjusted, actual sound waveform 510-2, candidate adjustment parameters 510-3, user preference data 510-4, and future interference characteristics 510-5, and the output may include an estimated adjustment effect 530 corresponding to the candidate adjustment parameters 510-3.

[0207] In some embodiments, the interfering sound features of the input adjustment model may correspond to the frequency spectrum features of the audio to be adjusted one by one. For example, the interfering sound features of the input adjustment model are acquired at the same time point or the same time period when the audio to be adjusted is played.

[0208] In some embodiments, Figure 5 As shown, the adjustment model 520 includes an environmental feature layer 521 and an effect prediction layer 522 .

[0209] The environmental characteristics layer is a model used to determine the characteristics of environmental impacts.

[0210] In some embodiments, the environmental feature layer may be a machine learning model, for example, the environmental feature layer may be a convolutional neural network (CNN), etc.

[0211] In some embodiments, Figure 5 As shown, the input of the environment feature layer 521 may include the interfering sound feature 410 - 1 , the spectrum feature 510 - 1 of the audio to be adjusted, and the actual sound waveform 510 - 2 ; and the output may include the environment impact feature 521 - 1 .

[0212] Environmental impact characteristics are characteristics that describe the impact of the environment on the propagation of sound. For example, environmental impact characteristics may include the geometry and size of the room, background noise, sound absorption and reflection characteristics, etc.

[0213] The effect prediction layer is a model used to determine the estimated adjustment effects corresponding to the candidate adjustment parameters.

[0214] In some embodiments, the effect prediction layer may be a machine learning model, for example, the effect prediction layer may be, for example, a deep neural network (DNN), etc.

[0215] In some embodiments, Figure 5 As shown, the input of the effect prediction layer 522 includes candidate adjustment parameters 510-3, environmental impact characteristics 521-1, spectral characteristics 510-1 of the audio to be adjusted, and user preference data 510-4, and the output includes the estimated adjustment effect 530 corresponding to the candidate adjustment parameters.

[0216] The estimated adjustment effect refers to the degree of improvement of the playback effect of the audio to be adjusted relative to the user preference data after the audio equalizer is adjusted based on the candidate adjustment parameters.

[0217] In some embodiments, Figure 5 As shown, the input of the effect prediction layer 522 also includes future interference features 510 - 5 .

[0218] For more information on future interference characteristics, see Figure 3 Related description.

[0219] In some embodiments of the present specification, by adding future interference features, the model can make predictions based on a more complete set of information, including the current environmental state, expected interference changes, and the characteristics of the audio signal itself, which helps the model to more accurately estimate the adjustment effect and reduce prediction errors.

[0220] In some embodiments, the adjustment model can be obtained by jointly training the environment feature layer and the effect prediction layer based on a large number of second training samples with second labels. The training samples used for joint training include sample environment impact features, spectral features of sample audio to be adjusted, sample actual sound waveform, sample adjustment parameters, sample environment impact features, and sample user preference data. The second training samples can be obtained based on historical data.

[0221] In some embodiments, the second label is the historical adjustment effect corresponding to the sample adjustment parameter, and the second label can be obtained based on manual or automatic annotation. For example, the processor can generate the actual sound waveform of the adjusted historical sound data and the ideal sound waveform corresponding to the sample user preference data based on the historical adjustment record, and determine the historical adjustment effect based on the similarity between the actual sound waveform and the ideal sound waveform, and use it as the second label.

[0222] An exemplary joint training process: input the sample interference sound features, the spectral features of the sample audio to be adjusted, and the sample actual sound waveform in the second training sample into the initial environmental feature layer to obtain the environmental impact features output by the initial environmental feature layer; input the output of the initial environmental feature layer, the spectral features of the sample audio to be adjusted, the sample candidate adjustment parameters, the sample environmental impact features, and the sample user preference data into the initial effect prediction layer to obtain the estimated adjustment effect corresponding to the candidate adjustment parameters; construct a loss function based on the output and label of the initial effect prediction layer, and update the parameters of the initial environmental feature layer and the initial effect prediction layer at the same time until the preset conditions are met and the training is completed. The preset conditions may be that the loss function is less than a threshold, converges, or the training cycle reaches a threshold.

[0223] The joint training of the environmental feature layer and the effect prediction layer helps to solve the problem of difficulty in obtaining labels when training the adjustment model alone, improve the training efficiency of the adjustment model, and reduce the difficulty of training.

[0224] In some embodiments, the processor may select, from among a plurality of candidate adjustment parameters, a parameter with the best estimated adjustment effect as the adjustment parameter.

[0225] In some embodiments of the present specification, by adjusting the model, the algorithm is allowed to estimate the corresponding adjustment effect before actually applying the adjustment parameters, which helps to avoid invalid or negative adjustments and ensures that each adjustment can effectively improve the audio quality; moreover, the model takes into account real-time interference sound characteristics and user preference data, so that the model can learn the correspondence between the ever-changing environmental noise, each user's unique auditory preferences and the estimated adjustment effect, thereby improving the accuracy of the output estimated adjustment effect.

[0226] Figure 6 This is an exemplary schematic diagram of determining user preference data according to some embodiments of this specification.

[0227] In some embodiments, Figure 6As shown, the processor can determine the user's corresponding reference preference data 622 and build a preference database 623 based on the user's historical adjustment data 610-1 and its corresponding historical spectrum characteristics 610-3 and historical interference sound characteristics 610-4; and determine the user's preference data 630 based on the spectrum characteristics 610-5 of the audio to be adjusted, the interference sound characteristics 610-6 of the sound data and the preference database 623.

[0228] For more information about historical reconciliation data, see Figure 3 Related description.

[0229] For more information about historical adjustment data, reference preference data, preference database, spectrum characteristics, and interference sound characteristics, see Figure 3 Related description.

[0230] The historical spectrum characteristics refer to the spectrum characteristics of the historical audio to be adjusted.

[0231] The historical interference sound characteristics refer to the frequency spectrum characteristics of the historical interference sounds.

[0232] It should be noted that when constructing the preference database, the historical interference sound features correspond to the historical spectrum features. For example, the interference sound features are the spectrum features of the historical interference sounds collected at the same time point or the same time period when the historical audio to be adjusted is played.

[0233] The preference database is a database used to store, index and query vectors. The preference database can store multiple reference vectors and reference preference data corresponding to each reference vector. The reference vector can include historical spectrum features, historical interference sound features, etc.

[0234] The reference preference data may reflect the user's preference settings under different audio to be adjusted and environmental conditions.

[0235] In some embodiments, the preference database is constructed based on a large amount of historical data. For example, the processor can determine the corresponding historical adjustment parameters based on the historical spectrum characteristics and the historical interference sound characteristics; after adjusting the parameters of the audio equalizer based on the historical adjustment parameters, the adjusted historical sound data is obtained; based on the changes in the spectrum characteristics of each single sound signal in the adjusted historical sound data relative to the historical spectrum characteristics, the changes in amplitude, frequency response, etc. are determined, and then the reference preference data is determined. For example, the processor can determine the reference volume preferred by the user based on the changes in the amplitude on each frequency component, and determine the reference timbre preferred by the user based on the changes in the frequency response, etc. Exemplarily, when the user tends to increase the gain in the low frequency band, it indicates that the user prefers a warmer and more powerful timbre; the larger the amplitude, the louder the volume preferred by the user.

[0236] In some embodiments, the processor may associate and store historical spectrum features, historical interference sound features, and corresponding reference preference data to construct a preference database.

[0237] In some embodiments, Figure 6 As shown, the processor can determine the actual playback effect 621 corresponding to each piece of historical adjustment data based on the historical environmental sound data 610-2; based on the actual playback effect 621 corresponding to each piece of historical adjustment data, merge multiple historical adjustment data 610-1 that meet similar conditions to obtain merged reference preference data 622; and based on the merged reference preference data 622, construct a preference database 623.

[0238] For more information about historical reconciliation data, see Figure 3 Related description.

[0239] The historical sound data refers to the sound data corresponding to the historical audio to be adjusted in the past time period.

[0240] The actual playback effect refers to the playback effect of the historical audio to be adjusted after the user adjusts the audio parameters of the historical audio to be adjusted. The actual playback effect may include the actual playback effect of each frequency band.

[0241] In some embodiments, the actual playback effect is determined based on the historical sound data. For example, the processor may determine the corresponding historical interference sound features based on the adjusted historical sound data, and generate the actual sound waveform corresponding to the historical sound data as the actual playback effect. Figure 3 The actual sound waveform corresponding to the historical sound data is generated in a similar manner to that of generating the actual sound waveform.

[0242] The similarity condition is a judgment condition for merging multiple historical adjustment data. For example, the similarity condition may include that the similarity of the historical spectrum characteristics, historical interference sound characteristics, and actual playback effects of two historical adjustment data are greater than their respective corresponding thresholds. The threshold may be a system preset value or a system default value.

[0243] In some embodiments, the processor can filter out multiple historical adjustment data from all historical adjustment records whose similarities of historical spectrum features, historical interference sound features, and actual playback effects are all greater than their corresponding thresholds as a data set; and merge the individual historical adjustment data of the above multiple data sets that meet similar conditions into a new historical adjustment data through a merging rule. The new historical adjustment data will replace the original multiple historical adjustment data in the above data set to reduce the redundancy of the preference database.

[0244] The merging rules include but are not limited to determining the actual playback effect in the new historical adjustment data based on statistical methods such as average, median, mode, weighted average and replacement. For example, the processor uses the average method to calculate the average of the actual playback effects in all data sets as the actual playback effect of the new historical adjustment data. For another example, the processor can use the historical spectrum features and historical interference sound features of any historical adjustment data in the data set as the historical spectrum features and historical interference sound features in the new historical adjustment data.

[0245] In some embodiments, the processor may insert new historical adjustment data into the preference database and delete the original multiple historical adjustment data in the corresponding data set.

[0246] In some embodiments of the present specification, the same or similar spectral characteristics and interfering sound characteristics of the audio to be adjusted may correspond to the adjustment situations of different users in the historical adjustment records, thereby resulting in the spectral characteristics and interfering sound characteristics of the same audio to be adjusted corresponding to multiple reference preference data. By evaluating the interfering sound characteristics and the spectral characteristics of the audio to be adjusted and the actual playback effects, and merging them based on the actual playback effects, the preference database can be streamlined to improve the retrieval efficiency. Only the merged data is retained in the preference database, which improves the data consistency and query efficiency and reduces the redundancy of the database.

[0247] In some embodiments, the processor may construct a corresponding search vector based on the spectral characteristics of the audio to be adjusted and the interfering sound characteristics of the sound data, and search in the preference database to determine the user preference data. Figure 3 The user preference data is determined by retrieving the spectral features of a single interfering sound in a similar manner.

[0248] In some embodiments, the processor may, in response to retrieving a plurality of different reference preference data, determine user preference change data based on a construction time of the reference preference data; and determine user preference data based on the preference change data.

[0249] The construction time refers to the timestamp when the reference vector and its corresponding reference preference data are created, updated or modified in the preference database.

[0250] In some embodiments, the processor may automatically generate the build time. For example, whenever a reference vector and its corresponding reference preference data are created or updated, the processor automatically records the current system time as the build time.

[0251] Preference change data refers to the changes in reference preference data.

[0252] In some embodiments, the processor may perform curve fitting on the plurality of reference preference data and their construction time, and determine a change function of the reference preference data as the preference change data. The curve fitting may be implemented by the least squares method, the Levenberg-Marquardt algorithm, the genetic algorithm, etc.

[0253] In some embodiments, the processor may determine the user preference data based on the change function and the current time. For example, the processor may substitute the current time into the change function to calculate the function value as the user preference data.

[0254] In some embodiments of the present specification, when multiple different reference preference data are retrieved, it means that the user's preference is not fixed, but changes with time, situation or personal needs. By determining the preference change data based on the construction time, the system can gain insight into the historical evolution of user preferences and determine that the user preference data in the current state is more in line with the user's actual situation.

[0255] In some embodiments of the present specification, by analyzing the user's historical adjustment data and its corresponding historical spectral characteristics and historical interference sound characteristics, the user's preference settings under different audio content and environmental conditions can be identified, which is conducive to the system building a personalized preference database for each user; the construction of the preference database is based on the analysis of a large amount of historical data and the decision-making method based on empirical data, which helps to subsequently determine the accuracy and reliability of the adjustment parameters, reduce the possibility of erroneous adjustment, and ensure the consistency and stability of the audio output quality.

[0256] One or more embodiments of the present specification also provide a responsive audio equalizer adjustment device, including a processing device, wherein the processing device is used to execute a responsive audio equalizer adjustment method as described in any of the above embodiments.

[0257] One or more embodiments of the present specification also provide a computer-readable storage medium, wherein the storage medium stores computer instructions. When a computer reads the computer instructions in the storage medium, the computer runs a responsive audio equalizer adjustment method as described in any of the above embodiments.

[0258] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only for example and does not constitute a limitation of this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements and corrections to this specification. Such modifications, improvements and corrections are suggested in this specification, so such modifications, improvements and corrections still belong to the spirit and scope of the exemplary embodiments of this specification.

[0259] At the same time, this specification uses specific words to describe the embodiments of this specification. For example, "one embodiment", "an embodiment", and / or "some embodiments" refer to a certain feature, structure or characteristic related to at least one embodiment of this specification. Therefore, it should be emphasized and noted that "one embodiment" or "an embodiment" or "an alternative embodiment" mentioned twice or more in different positions in this specification does not necessarily refer to the same embodiment. In addition, certain features, structures or characteristics in one or more embodiments of this specification can be appropriately combined.

[0260] In addition, unless explicitly stated in the claims, the order of the processing elements and sequences described in this specification, the use of alphanumeric characters, or the use of other names are not intended to limit the order of the processes and methods of this specification. Although the above disclosure discusses some invention embodiments that are currently considered useful through various examples, it should be understood that such details are only for illustrative purposes, and the attached claims are not limited to the disclosed embodiments. On the contrary, the claims are intended to cover all modifications and equivalent combinations that are consistent with the essence and scope of the embodiments of this specification. For example, although the system components described above can be implemented by hardware devices, they can also be implemented only by software solutions, such as installing the described system on an existing server or mobile device.

[0261] Similarly, it should be noted that in order to simplify the description disclosed in this specification and thus help understand one or more embodiments of the invention, in the above description of the embodiments of this specification, multiple features are sometimes combined into one embodiment, figure or description thereof. However, this disclosure method does not mean that the features required by the subject matter of this specification are more than the features mentioned in the claims. In fact, the features of the embodiments are less than all the features of the single embodiment disclosed above.

[0262] In some embodiments, numbers describing the number of components and attributes are used. It should be understood that such numbers used in the description of the embodiments are modified by the modifiers "about", "approximately" or "substantially" in some examples. Unless otherwise specified, "about", "approximately" or "substantially" indicate that the numbers are allowed to vary by ±20%. Accordingly, in some embodiments, the numerical parameters used in the specification and claims are approximate values, which may change according to the required features of individual embodiments. In some embodiments, the numerical parameters should take into account the specified significant digits and adopt the general method of retaining digits. Although the numerical domains and parameters used to confirm the breadth of their range in some embodiments of this specification are approximate values, in specific embodiments, the setting of such numerical values ​​is as accurate as possible within the feasible range.

[0263] Each patent, patent application, patent application publication, and other materials, such as articles, books, specifications, publications, documents, etc., cited in this specification are hereby incorporated by reference in their entirety. Except for application history documents that are inconsistent with or conflicting with the contents of this specification, documents that limit the broadest scope of the claims of this specification (currently or later attached to this specification) are also excluded. It should be noted that if the descriptions, definitions, and / or use of terms in the materials attached to this specification are inconsistent or conflicting with the contents described in this specification, the descriptions, definitions, and / or use of terms in this specification shall prevail.

[0264] Finally, it should be understood that the embodiments described in this specification are only used to illustrate the principles of the embodiments of this specification. Other variations may also fall within the scope of this specification. Therefore, as an example and not a limitation, alternative configurations of the embodiments of this specification may be considered consistent with the teachings of this specification. Accordingly, the embodiments of this specification are not limited to the embodiments explicitly introduced and described in this specification.

Claims

1. A responsive audio equalizer adjustment method, characterized in that: The method comprises: Acquire the audio to be adjusted, and determine the frequency spectrum characteristics of the audio to be adjusted; Based on the sound data, determining interference sound characteristics; Determining adjustment parameters based on the frequency spectrum characteristics and the interference sound characteristics; Based on the adjustment parameters, the tuning unit is controlled to adjust the parameters of the audio equalizer.

2. The method according to claim 1, characterized in that The interfering sound feature includes a frequency spectrum feature of a single interfering sound, and determining the interfering sound feature based on the sound data includes: Based on the collection cycle, obtain sound data; Separating the sound data to obtain at least one single sound signal; Identifying the at least one single sound signal and determining a frequency spectrum feature of the single interfering sound; The interfering sound feature is determined based on the frequency spectrum feature of the single interfering sound.

3. The method according to claim 1, characterized in that The determining of adjustment parameters based on the interference sound characteristics and spectrum characteristics includes: determining whether the sound data satisfies the user preference data; and In response to no, the adjustment parameter is determined based on the interfering sound feature, the frequency spectrum feature of the audio to be adjusted, and the user preference data.

4. The method according to claim 3, characterized in that The method further comprises: Based on the user's historical adjustment data and its corresponding historical spectrum characteristics and historical interference sound characteristics, determining reference preference data corresponding to the user, and building a preference database; The user preference data is determined based on the frequency spectrum characteristics of the audio to be adjusted, the interfering sound characteristics of the sound data, and the preference database.

5. A responsive audio equalizer adjustment system, characterized in that: The system comprises: A first determining module is configured to obtain the audio to be adjusted and determine the frequency spectrum characteristics of the audio to be adjusted; A second determination module is configured to determine an interfering sound feature based on the sound data; A third determination module is configured to determine an adjustment parameter based on the frequency spectrum feature and the interference sound feature; The control module is configured to control the tuning unit to adjust the parameters of the audio equalizer based on the adjustment parameters.

6. The system according to claim 5, characterized in that The second determining module is further configured to: Based on the collection cycle, obtain sound data; Separating the sound data to obtain at least one single sound signal; Identifying the at least one single sound signal and determining a frequency spectrum feature of the single interfering sound; The interfering sound feature is determined based on the frequency spectrum feature of the single interfering sound.

7. The system according to claim 5, characterized in that The third determination module is further configured to: determining whether the sound data satisfies the user preference data; and In response to no, determining the adjustment parameter based on the interfering sound feature, the frequency spectrum feature, and the user preference data.

8. The system according to claim 7, characterized in that The system further includes a fourth determining module, wherein the fourth determining module is configured to: Based on the user's historical adjustment data and its corresponding historical spectrum characteristics and historical interference sound characteristics, determining reference preference data corresponding to the user, and building a preference database; The user preference data is determined based on the frequency spectrum characteristics of the audio to be adjusted, the interfering sound characteristics of the sound data, and the preference database.

9. A responsive audio equalizer adjustment device, characterized in that: The apparatus comprises at least one processor and at least one memory; The at least one memory is used to store computer instructions; The at least one processor is configured to execute at least part of the computer instructions to implement the responsive audio equalizer adjustment method according to any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that: The storage medium stores computer instructions. When the computer reads the computer instructions in the storage medium, the computer executes the responsive audio equalizer adjustment method according to any one of claims 1 to 4.