An earphone personalization equalization method and system based on artificial intelligence

CN122602037APending Publication Date: 2026-08-18SHENZHEN LESHENG ACOUSTIC TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610875446.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0005]缺乏精准的自然语言个性化调音能力:现有EQ方案多基于固定曲线(如HarmanTarget)或简单的滑块调节,部分高端产品支持基于听力测试的个性化调音,但无法通过自然语言精准捕捉用户的主观听音偏好,难以满足用户对音色的精细化需求

Benefits of technology

通过局域网Web界面实现用户便捷交互,将自然语言听音偏好转化为精准的均衡参数,结合多维度声学分析生成适配参数,依托DSP模块实现实时滤波,同时支持本地离线存储调用,兼顾了操作便捷性、均衡精准度与使用场景的广泛性,为用户提供了高效、灵活的耳机个性化均衡体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122602037A_ABST
    Figure CN122602037A_ABST
Patent Text Reader

Abstract

The application provides an earphone personalized equalization adjustment method and system based on artificial intelligence, and belongs to the technical field of audio signal processing. The method receives an earphone model and a listening preference input by a user through a local area network Web interface, generates parameter equalizer adjustment parameters adaptive to the user's preference through multi-dimensional acoustic analysis, realizes real-time equalization in combination with adaptive DSP filtering and phase correction, and supports local offline storage calling. The application can effectively improve the equalization effect and audio fidelity, and takes into account the convenience of operation and the universality of use scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of audio signal processing technology, and in particular to a personalized equalization adjustment method and system for headphones based on artificial intelligence. Background Technology

[0002] With the development of digital audio technology, more and more audio devices are beginning to support the parametric equalizer (PEQ) function, which is used to adjust the gain of audio signals in different frequency bands to improve the sound performance of headphones or speakers.

[0003] However, existing technologies still have the following problems: The tuning threshold is relatively high: Traditional EQ adjustment requires users to have certain audio tuning experience. Ordinary users find it difficult to understand the role of parameters such as center frequency, Q value and gain, which makes it difficult to fully utilize the EQ function.

[0004] There are significant differences in headphones: different headphone models have significant differences in frequency response characteristics. Although software such as Sonarworks and EqualizerAPO have established a database of measured frequency responses of headphones, users still need to manually select the headphone model and make additional personalized adjustments, resulting in low overall adjustment efficiency.

[0005] Lack of precise natural language personalized tuning capabilities: Existing EQ solutions are mostly based on fixed curves (such as Harman Target) or simple slider adjustments. Some high-end products support personalized tuning based on hearing tests, but they cannot accurately capture users' subjective listening preferences through natural language, making it difficult to meet users' refined needs for timbre.

[0006] The offline availability of pure software EQ solutions is poor: the parameters of some pure software EQ solutions only take effect when the software is running, and cannot be used independently when the device is offline or the software is closed, which limits the breadth of application scenarios.

[0007] Therefore, this invention proposes a personalized equalization adjustment method and system for headphones based on artificial intelligence. Summary of the Invention

[0008] This invention provides a personalized equalization adjustment method and system for headphones based on artificial intelligence, in order to solve the aforementioned technical problems. This invention provides a personalized equalization adjustment method for headphones based on artificial intelligence, including: Step 1: The user terminal accesses the Web control interface through a local area network and establishes a two-way communication connection with the audio device; Step 2: Receive the model information of the currently used headphones, input or selected by the user, based on the Web control interface; Step 3: Receive the user's listening preference description information in natural language text format based on the Web control interface; Step 4: Send the model information and the listening preference description information to the artificial intelligence tuning module; Step 5: The AI ​​tuning module retrieves the measured frequency response data corresponding to the model information from the pre-stored headphone frequency response database, and performs multi-dimensional acoustic analysis by combining the pre-stored target acoustic reference curve, preset timbre model and the listening preference description information to generate parametric equalizer adjustment parameters containing at least one set of center frequency, gain value and quality factor Q value. Step 6: Convert the parameter equalizer adjustment parameters into DSP filter coefficients that can be recognized by the audio device's DSP processing module through the DSP parameter conversion module; Step 7: Load the DSP filter coefficients into the DSP processing module of the audio device, wherein the DSP processing module performs equalization filtering on the real-time input audio signal according to the DSP filter coefficients and outputs it to the headphones; Step 8: After receiving the save command sent by the user through the Web control interface, store the parametric equalizer adjustment parameters in the internal memory of the audio device, and generate a corresponding entry in the local EQ preset list of the audio device, so that the audio device can retrieve the parametric equalizer adjustment parameters from the internal memory and convert them into DSP filter coefficients to perform audio equalization processing when disconnected from the network connection.

[0009] Preferably, the process of performing multi-dimensional acoustic analysis based on the listening preference description information includes: The listening preference description information is input into the acoustic intent level decoupler, and decoupled into a discrete acoustic intent feature set according to the dimensions of low-frequency elasticity, mid-frequency density, high-frequency extension, sound field width, and harmonic saturation. High-dimensional semantic representations are output through BERT acoustic pre-trained models. ; Obtain the main preference concatenation path of the listening preference description information and the preference distribution of each main preference in the main preference concatenation path, and simulate the preference fluctuation diagram based on the main preference concatenation path and preference distribution, wherein the preference fluctuation diagram includes the preference strength of at least one newly added preference; Based on the preference strength of each new preference involved in the preference fluctuation graph, a high-dimensional semantic representation is generated. After correction, a preference-calibrated semantic representation is obtained. ; Collect user's historical listening frequency domain preference sequence, volume adaptation sequence, and track type time sequence to construct a cross-modal behavioral feature set. A contrastive learning framework is used to construct positive and negative sample pairs, which are then output as a noise reduction behavior representation via a distillation network. For new users without historical data, a pre-trained general behavioral feature template is loaded as the initial... ; Construct a Bayesian preference probability model, prior distribution Let be a uniform distribution Uniform(0.1, 2.0), and the likelihood function be... Given a diagonal Gaussian distribution, solve for the posterior probability. Output the expected dynamic preference weight coefficients: ,in, ; Will Substitute the basic frequency response compensation deviation The preference correction compensation gain is obtained. ,in, Acoustic intent - frequency mapping factor, To retrieve the pre-stored target acoustic reference curve; To retrieve measured frequency response data that matches the model information from the headphone frequency response database; Retrieve the standard gain curve corresponding to the preset tone model and adjust the gain to compensate for the preference. The fused gain is obtained by initially combining the fused gain with the center gain of the standard gain curve, and then embedded into the subsequent equalization parameter generation process.

[0010] Preferably, the simulation of preference fluctuation diagram based on the main preference concatenation path and preference distribution includes: Multi-dimensional semantic dependency parsing is performed on the listening preference description information, and preference intention, intensity modification, and scene constraint are mapped to standardized preference nodes, modification nodes and constraint nodes respectively, and a preference dependency directed acyclic graph containing node hierarchical relationship and association strength is constructed. The preference-dependent directed acyclic graph is parsed hierarchically to generate ordered main preference concatenation paths; For each valid main preference in the main preference concatenation path, intensity modification information is extracted from the text and associated with the adjustment records corresponding to the preference in the user's historical listening behavior data to construct the intensity distribution data of the corresponding valid main preference; Based on the correlation strength between nodes in the preference-dependent directed acyclic graph, a coupling correlation model between effective primary preferences is established to quantify the mutual influence between different effective primary preferences. The coupling correlation model is then adaptively corrected by combining user historical listening behavior data to eliminate the deviation between static text description and actual listening behavior. Using the main preference concatenation path as the transfer logic and the coupling correlation model as the constraint, the simulation interface is called to execute parallel simulation of the intensity distribution data of each effective main preference, generating a three-dimensional preference fluctuation model that includes the fluctuation range of preference intensity and the coupling effect of different preferences. When the preset simulation stop condition is reached, the simulation stop interface is called to terminate the simulation process. Based on the parallel simulation data, relevant initial boundary data is extracted and a preference fluctuation map is generated.

[0011] Preferably, relevant initial boundary data are extracted based on parallel simulation data, and a preference fluctuation diagram is generated, including: Based on the parallel simulation data, determine the first and last moments of the independent parallel simulation for each effective primary preference, construct a set of simulation segments for each effective primary preference, and calculate the mean and variance of preference intensity for each simulation segment in each simulation segment set, as well as the rate of change of intensity between different simulation segments. Based on the mean, variance, and rate of change of preference intensity of the same set of simulation segments, an intensity feature sequence is constructed. Based on the difference feature vectors between the intensity feature sequence and the intensity feature sequences of each other main preference, a preference fluctuation map is constructed.

[0012] Preferably, a preference fluctuation graph is constructed, including: For each effective principal preference intensity feature sequence within the same simulation segment set, convert it into a standardized feature vector representation, and then perform a difference operation with the feature vector of each other effective principal preference intensity feature sequence to generate the difference feature vector of the corresponding effective principal preference. The elements in the difference feature vector are arranged in ascending order of numerical value. Taking the first smallest difference feature element after sorting as the benchmark, the numerical similarity matching of the remaining difference feature elements is performed in turn. Feature elements whose numerical difference with the smallest difference feature element is within a preset threshold are selected to form a finite feature subset. Based on the element combination form of the finite feature subset, the feature variables corresponding to the same simulation segment set are determined. The simulation log data generated by parallel simulation is divided into blocks and indexed according to the simulation execution dimension. Index entries corresponding to different feature variable types are constructed, and the corresponding index entries are matched based on the identification information of the feature variable. The initial boundary data that matches the feature variable in the simulation log is retrieved. The initial boundary data is subjected to probability distribution fitting processing, and the confidence interval boundary value of the corresponding effective main preference is determined based on the fitted distribution shape, thereby generating the confidence interval of the effective main preference; Traverse all interval segments within the confidence interval, identify interval segments that exceed the preset value range as extreme fluctuation intervals, and adjust the extreme fluctuation intervals by linear interpolation replacement; for extreme fluctuation intervals located at both ends of the confidence interval, replace them with the mean of the nearest neighbor normal interval to generate corrected interval segments within the user's preferred acceptable range. The corrected interval segments are spliced ​​and integrated with the non-extreme interval segments within the confidence interval to form the effective intervals of the corresponding effective principal preferences. The effective intervals of all effective principal preferences are then structured and mapped according to the simulation segment set dimension to construct a complete preference fluctuation diagram. Based on the effective interval of the preference fluctuation graph, new preferences that users have not explicitly stated are identified through implicit semantic analysis, and the strength of the new preferences is determined according to the mean of the effective interval.

[0013] Preferably, the parameter equalizer adjustment parameters include at least one set of center frequency, gain value, and quality factor Q value, including: Based on the target equalization gain curve obtained from multi-dimensional acoustic analysis, the auditory perception characteristics of different frequencies are adapted and corrected by combining the human ear equal loudness curve. At the same time, the gain value that exceeds the user's preference acceptance range is limited according to the initial boundary data of the preference fluctuation diagram, so as to obtain a perception equalization compensation curve that is adapted to human ear perception and fits the user's preference range. The perceived equalization compensation curve is decomposed into a multi-scale frequency domain, and the peak points and valley points under different decomposition scales are extracted. For each extracted extreme point, the equalization gain change gradient and curvature characteristics of the frequency band are calculated. Extreme points whose change gradient and curvature characteristics meet the preset conditions are selected as the candidate center frequency set. For each center frequency in the candidate center frequency set, the corresponding gain value on the sensing equalization compensation curve is extracted as the initial gain. At the same time, based on the phase frequency response data at the center frequency, the group delay change caused by the filter under different quality factors is analyzed. According to the acceptable range of group delay change, the quality factor corresponding to the center frequency is adjusted. A simulation model of parametric equalizer filter superposition is constructed. All candidate center frequencies, initial gains and adjusted quality factors are input into the superposition simulation model to generate the composite frequency response curve of multiple filters superimposed. The synthesized frequency response curve and the perceived equalization compensation curve are compared point by point. If the deviation exceeds the preset range, the gain value or quality factor corresponding to the center frequency with the excessive deviation is adjusted until the deviation between the synthesized frequency response curve and the perceived equalization compensation curve is within the preset range. The adjusted center frequency, gain value, and quality factor Q value are grouped and packaged according to frequency band to generate multiple sets of equalizer adjustment parameters, with each set of parameters corresponding to equalization compensation for one frequency band.

[0014] Preferably, the real-time input audio signal is subjected to equalization filtering processing based on the DSP filter coefficients, including: The coefficients of the DSP filter are analyzed into a multi-order IIR filter bank that corresponds one-to-one with the equalizer adjustment parameters of each frequency band. The transfer function of each filter is defined by the corresponding center frequency, gain value and quality factor Q value. Based on the measured phase response data of the headphone's sound unit, a composite filtering unit is constructed by pre-configuring phase compensation coefficients for each filter. The real-time input audio signal is processed into frames to generate a sequence of consecutive short audio frames of fixed duration. Transient feature detection is performed on each audio frame to identify transient peak signal segments and steady-state signal segments. For transient peak signal segments, dynamically reduce the quality factor Q value of the corresponding frequency band filter to broaden the filtering bandwidth; For the steady-state signal segment, recover the preset Q value of the filter; The audio frame is input into the composite filtering unit for frequency band equalization filtering. At the same time, based on the group delay parameters of each filter, the filtered signals of different frequency bands are timestamped to compensate for the time domain offset caused by the difference in filter bandwidth. The time-aligned signals of each frequency band are synthesized to generate an equalized time-domain audio signal. Real-time intermodulation distortion detection is performed on the equalized time-domain audio signal. If the detected distortion exceeds the preset threshold, the gain value of the corresponding frequency band filter is finely adjusted based on the frequency distribution of the distortion component until the distortion falls back to the preset range. Finally, the equalized audio signal is output to the headphone driver unit.

[0015] This invention provides an artificial intelligence-based personalized equalization adjustment system for headphones, comprising: a communication establishment module: a user terminal accesses a Web control interface via a local area network and establishes a two-way communication connection with the audio device; The first interface receiving module is used to receive the model information of the currently used headphones input or selected by the user based on the Web control interface; The second interface receiving module is used to receive listening preference description information input by the user in the form of natural language text based on the Web control interface; The sending module is used to send the model information and the listening preference description information to the artificial intelligence tuning module; The acoustic analysis module is used by the artificial intelligence tuning module to retrieve the measured frequency response data corresponding to the model information from the pre-stored headphone frequency response database, and perform multi-dimensional acoustic analysis by combining the pre-stored target acoustic reference curve, preset timbre model and the listening preference description information to generate parametric equalizer adjustment parameters including at least one set of center frequency, gain value and quality factor Q value. The parameter conversion module is used to convert the parameter equalizer adjustment parameters into DSP filter coefficients that can be recognized by the audio device's DSP processing module through the DSP parameter conversion module. A coefficient loading module is used to load the DSP filter coefficients into the DSP processing module of the audio device, wherein the DSP processing module performs equalization filtering on the real-time input audio signal according to the DSP filter coefficients and outputs it to the headphones; The continuing execution module is used to receive a save command sent by the user through the Web control interface, store the parametric equalizer adjustment parameters in the internal memory of the audio device, and generate a corresponding entry in the local EQ preset list of the audio device. This allows the audio device to retrieve the parametric equalizer adjustment parameters from the internal memory and convert them into DSP filter coefficients to perform audio equalization processing when disconnected from the network.

[0016] Compared with the prior art, the beneficial effects of this application are as follows: The system enables convenient user interaction through a local area network web interface, transforms natural language listening preferences into precise equalization parameters, generates adaptation parameters by combining multi-dimensional acoustic analysis, achieves real-time filtering based on the DSP module, and supports local offline storage retrieval. It balances ease of operation, equalization accuracy, and wide applicability, providing users with an efficient and flexible personalized equalization experience for headphones.

[0017] Other features and advantages of the invention will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description and the accompanying drawings.

[0018] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart of a personalized equalization adjustment method for headphones based on artificial intelligence, as described in an embodiment of the present invention. Figure 2 This is a structural diagram of an artificial intelligence-based personalized equalization adjustment system for headphones, as described in an embodiment of the present invention. Detailed Implementation

[0020] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0021] This invention provides a personalized equalization adjustment method for headphones based on artificial intelligence, such as... Figure 1 As shown, it includes: Step 1: The user terminal accesses the Web control interface through the local area network and establishes a two-way communication connection with the audio device; A local area network (LAN) refers to a local network coverage environment built on the Wireless LAN protocol. It uses general-purpose wireless LAN communication modules to construct the network environment, such as wireless LANs in home or office settings. After the audio device is powered on, a 6-digit random token is generated and displayed on the device screen. The token is valid for 10 minutes and automatically refreshes after the timeout. After the user enters the token in the web interface, the server verifies it and establishes a persistent connection. Heartbeat packets are sent at 30-second intervals. If there is no response after the timeout, the connection is automatically disconnected. After disconnection, it automatically attempts to reconnect 3 times, with each reconnection 5 seconds apart.

[0022] Web control interface refers to a visual interactive interface built on web development technology, that is, a web interactive page written using conventional front-end development framework, such as the headphone adjustment web interface opened in a mobile phone or computer browser. A two-way communication connection refers to a communication link between a user terminal and an audio device that can simultaneously send and receive data. It uses standard network communication protocols to establish a data transmission channel, such as a connection channel between a terminal and an audio device for real-time transmission of control commands and status data.

[0023] Step 2: Receive the model information of the currently used headphones, input or selected by the user, based on the Web control interface; Headphone model information refers to the character or coded data used to identify the headphone product model. It is set in the web control interface with input boxes and selection lists, such as the identification information of specific headphone models like Sony WH-1000XM5 and Apple AirPods Pro.

[0024] The headphone frequency response database retrieval logic is as follows: it prioritizes exact matching of the model string, and if the match fails, it uses fuzzy matching with an edit distance ≤ 1. The fuzzy matching results are displayed in a maximum of 3, and the user is required to manually confirm the correct model to avoid model confusion.

[0025] Step 3: Receive the user's listening preference description information in natural language text form based on the Web control interface; Listening preference description information refers to the text of the user's auditory needs expressed in everyday language. The text input area is set in the Web control interface, such as natural language descriptions like "ample low-frequency elasticity, clear mid-frequency vocals, and smooth high-frequency extension".

[0026] For contradictory user input, such as "heavy low frequencies but not muddy", the system will automatically identify the conflict and pop up a prompt, requiring the user to further clarify the preference priority; for ambiguous input, such as "sounds better", the system will provide 3 common timbre styles for the user to choose from.

[0027] Step 4: Send the model information and the listening preference description information to the artificial intelligence tuning module; An AI-powered audio tuning module is a processing module that integrates semantic parsing, acoustic analysis, and parameter generation functions. It uses an embedded processor to build the data processing core, such as the processing chip with audio analysis capabilities inside an audio device.

[0028] Step 5: The AI ​​tuning module retrieves the measured frequency response data corresponding to the model information from the pre-stored headphone frequency response database, and performs multi-dimensional acoustic analysis by combining the pre-stored target acoustic reference curve, preset timbre model and the listening preference description information to generate parametric equalizer adjustment parameters containing at least one set of center frequency, gain value and quality factor Q value. In this embodiment, the headphone frequency response database refers to a database that stores measured acoustic data across the entire frequency band of multiple headphone models. It pre-collects and stores acoustic data from different headphones, using the IEC60268-7 standard. The testing environment is an anechoic chamber (background noise ≤15dB(A)), and the testing equipment includes a B&K4195 microphone, a B&K2734 power amplifier, and an APx500 audio analyzer. The sampling frequency range is 20Hz-20kHz, with a sampling accuracy of 1Hz. Three samples are collected for each headphone model, and the average value is used as the final measured frequency response data. The database is stored in JSON format, and each data entry includes fields such as headphone brand, model, frequency response array, phase response array, and release date.

[0029] The measured frequency response data refers to the actual sound response data of the headphones collected by professional acoustic testing equipment. The data is collected using an acoustic testing kit, such as the test data of a certain headphone with a sound amplitude of -2dB at a frequency of 1kHz. The target acoustic reference curve refers to the industry-standard headphone frequency response reference curve, which uses a recognized acoustic standard curve as a reference, such as the Harman target frequency response curve commonly used in the headphone acoustic field. Preset timbre models refer to frequency response gain models with different timbre styles that are set in advance. They have multiple preset gain parameter templates for fixed timbres, such as standard, transparent, and heavy preset timbre models. Multidimensional acoustic analysis refers to a comprehensive analysis process that integrates measured headphone data, target curves, timbre models, and user preferences. It uses algorithms to fuse and calculate multiple types of data, such as combining the difference between the measured frequency response and the target curve with user preferences for correction calculations. Parametric equalizer adjustment parameters refer to the core parameters used to control the operation of the parametric equalizer. These parameters generate a combination of parameters adapted to audio filtering. The center frequency refers to the center frequency point of each filter band of the parametric equalizer, such as 200Hz, 1kHz, 8kHz, etc.; the gain value refers to the signal amplification or attenuation value at the corresponding center frequency, such as +3dB, -2dB, etc.; and the quality factor Q value is a parameter that characterizes the width of the filter band of the parametric equalizer, such as Q values ​​of 1.4, 2.0, etc., which characterize the filter bandwidth.

[0030] Step 6: Convert the parameter equalizer adjustment parameters into DSP filter coefficients that can be recognized by the audio device's DSP processing module through the DSP parameter conversion module; In this embodiment, the DSP parameter conversion module refers to the functional module that realizes parameter format conversion. It uses digital signal conversion algorithm to realize parameter mapping, such as converting center frequency, gain, and Q value into filter operation coefficients. The DSP parameter conversion module uses a bilinear transform method to convert the PEQ parameters (center frequency) into PEQ parameters. The gain (G) and quality factor (Q) are converted into 4th-order IIR filter coefficients to adapt to the ARM Cortex-M7 DSP architecture. The conversion formula is as follows: Predistortion frequency: ,in, The sampling rate (default 48kHz); Intermediate variables: ; Filter coefficients: , , , , , ; Normalization: Divide all coefficients by This yields the final filter coefficients that the DSP can recognize; For example, for PEQ parameters with a center frequency of 20Hz, a gain of +4dB, and a Q value of 2, and a sampling rate of 48kHz, the calculated coefficients of a 4th-order IIR filter are: b=[1.0003,-1.9990,0.9987], a=[1.0000,-1.9990,0.9990].

[0031] DSP filter coefficients refer to the filtering operation parameters that the DSP processing module can directly call. They generate values ​​that conform to the DSP core operation rules, such as the filter operation coefficients adapted to ARM architecture DSPs.

[0032] Step 7: Load the DSP filter coefficients into the DSP processing module of the audio device, wherein the DSP processing module performs equalization filtering on the real-time input audio signal according to the DSP filter coefficients and outputs it to the headphones; The DSP processing module refers to the digital signal processing core inside an audio device. It uses a dedicated digital signal processing chip, such as the DSP chip in the audio device that is responsible for audio signal processing. Equalization filtering refers to the process of adjusting the gain of an audio signal in different frequency bands according to parameters. It uses filters to adjust the frequency band signal, such as boosting low-frequency signals and attenuating high-frequency signals.

[0033] Step 8: After receiving the save command sent by the user through the Web control interface, store the parametric equalizer adjustment parameters in the internal memory of the audio device, and generate a corresponding entry in the local EQ preset list of the audio device, so that the audio device can retrieve the parametric equalizer adjustment parameters from the internal memory and convert them into DSP filter coefficients to perform audio equalization processing when disconnected from the network connection.

[0034] Internal memory refers to the data storage device built into the audio device, which uses non-volatile memory chips, such as the flash memory module built into the audio device. The local EQ preset list refers to an indexed list that stores personalized equalization parameters, such as a parameter list containing entries like "User-defined 1, User-defined 2". When a preset is called, the system will automatically verify whether the currently connected headphone model matches the preset headphone model. If they do not match, the system will prompt the user to confirm whether to continue using the preset.

[0035] In this embodiment, the frequency response compensation deviation comparison is shown in Table 1: Table 1 Comparison of Frequency Response Compensation Deviations

[0036] Three headphones—Sony WH-1000XM5, Apple AirPods Pro, and Sennheiser HD660S—were used. The test frequency range was 20Hz-20kHz, with a total of 2000 test frequency points. Each solution was tested 10 times, and the average value was taken. Subjective listening tests showed that the user satisfaction rate of this invention's solution was 92%, significantly higher than other comparative solutions.

[0037] The beneficial effects of the above technical solution are: it enables convenient user interaction through a local area network web interface, transforms natural language listening preferences into accurate equalization parameters, generates adaptation parameters by combining multi-dimensional acoustic analysis, achieves real-time filtering by relying on the DSP module, and supports local offline storage retrieval. It takes into account the ease of operation, equalization accuracy, and wide applicability, providing users with an efficient and flexible personalized equalization experience for headphones.

[0038] This invention provides a personalized equalization adjustment method for headphones based on artificial intelligence. The method includes a multi-dimensional acoustic analysis process based on the listening preference description information, comprising: The listening preference description information is input into the acoustic intent level decoupler, and decoupled into a discrete acoustic intent feature set according to the dimensions of low-frequency elasticity, mid-frequency density, high-frequency extension, sound field width, and harmonic saturation. High-dimensional semantic representations are output through BERT acoustic pre-trained models. , where n is the total number of acoustic intent dimensions; The Acoustic Intent-Level Decoupling Unit is a parsing module that decomposes natural language preference text into standardized acoustic dimensional features. It builds its parsing logic based on rule matching and a predefined acoustic dimension dictionary. First, it identifies acoustically relevant keywords in the text and then maps them to preset acoustic dimensions. Specifically, the decoupling unit decouples the text according to five dimensions: low-frequency elasticity, mid-frequency density, high-frequency extension, soundstage width, and harmonic saturation. Low-frequency elasticity represents the user's preference for the rebound texture and power of sound in the 20Hz-500Hz low-frequency range; mid-frequency density represents the user's preference for the clarity and fullness of vocal and instrument details in the 500Hz-6kHz mid-frequency range; and high-frequency extension represents the user's preference for the 6kHz-200Hz mid-frequency range. The system considers user preferences for sound extension and detail in the kHz high-frequency range, soundstage width (representing the user's preference for sound space and openness), and harmonic saturation (representing the user's preference for sound overtone richness and warmth). Each dimension is assigned a quantization level of 0-10, with higher values ​​indicating stronger user preferences. For example, a user inputting "bounced low frequencies" corresponds to a feature value of 8 for the low-frequency bounce dimension, "clear vocals" to a feature value of 9 for the mid-frequency density dimension, "non-harsh high frequencies" to a feature value of 6 for the high-frequency extension dimension, "open soundstage" to a feature value of 7 for the soundstage width dimension, and "warm sound" to a feature value of 7 for the harmonic saturation dimension. These quantized feature values ​​form a discrete acoustic intent feature set. The feature values ​​of each dimension are arranged into a set in a preset order, such as the set... Subsequently, this feature set is input into the BERT acoustic pre-trained model, which outputs a high-dimensional semantic representation. .

[0039] We adopted the open-source BERT base model and fine-tuned it on a dataset containing a large amount of user audio preference description text. This enabled the model to better capture audio-related semantic information. We fine-tuned the Chinese BERT model on 100,000 headphone listening preference text data to obtain a model adapted for audio preference parsing. The batch size was 32, the learning rate was 2e-5, and the training epochs were 10. The high-dimensional semantic representation is a fixed-dimensional vector output by the model that can represent the overall semantic information of the listening preference text. The model outputs a 768-dimensional vector through the encoding layer. This vector contains the comprehensive semantic information of all acoustic dimension preferences in the text, such as encoding the user's input preference text into a 768-dimensional numerical vector.

[0040] Obtain the main preference concatenation path of the listening preference description information and the preference distribution of each main preference in the main preference concatenation path, and simulate the preference fluctuation diagram based on the main preference concatenation path and preference distribution, wherein the preference fluctuation diagram includes the preference strength of at least one newly added preference; The primary preference concatenation path is an ordered sequence formed by arranging the user's core listening preferences according to priority, extracted from the listening preference description information. By parsing the modifiers, positions and relationships of preferences in the text, the priority of the core preferences is determined and sorted. For example, if the user inputs "low frequency elasticity, clear vocals, wide soundstage", the parsed primary preference concatenation path is [low frequency elasticity, mid frequency density, soundstage width].

[0041] Preference distribution refers to the distribution of the intensity values ​​of a single primary preference. It can be combined with intensity modifiers in the text and the user's historical listening data to statistically determine the distribution range of the intensity values ​​of that primary preference. For example, the intensity distribution of the primary preference for low-frequency elasticity is 6-9, indicating that the user's preference intensity for low-frequency elasticity is mostly between 6 and 9. A preference fluctuation graph is a dynamic data graph representing the changes in the intensity of a user's listening preference as a function of different scenarios and music genres. Based on the primary preference's concatenation path and preference distribution, a fluctuation curve containing preference intensity and scenario dimensions is generated through simulation. For example, a curve with music genre as the horizontal axis and preference intensity as the vertical axis shows the intensity changes of the user's low-frequency elasticity preference. This fluctuation graph includes the strength of at least one newly added preference. The strength of a newly added preference is the intensity level of an additional preference implicit in the user's listening preference description information that is not explicitly listed in the primary preference. Newly added preferences are identified through implicit semantic analysis and assigned an intensity value of 0-10. For example, if a user inputs "low-frequency elasticity when listening to rock," the parsed result is "mid-frequency clarity suitable for rock tracks," with a preference strength of 7.

[0042] Based on the preference strength of each new preference involved in the preference fluctuation graph, a high-dimensional semantic representation is generated. After correction, a preference-calibrated semantic representation is obtained. ; Preference calibration semantic representation is a high-dimensional semantic representation vector that has been corrected by incorporating the strength information of newly added preferences. Based on the strength of the newly added preference, the weights of the corresponding dimensions in the original high-dimensional semantic representation vector are adjusted accordingly. For example, if the user's newly added preference "mid-frequency clarity suits rock music" has a preference strength of 7, the weights of the low-frequency elasticity correlation dimension in the original high-dimensional semantic representation are adjusted accordingly, resulting in a corrected 768-dimensional vector. This makes the vectors more closely match the user's actual preference scenarios.

[0043] Collect user's historical listening frequency domain preference sequence, volume adaptation sequence, and track type time sequence to construct a cross-modal behavioral feature set. A contrastive learning framework is used to construct positive and negative sample pairs, which are then output as a noise reduction behavior representation via a distillation network. ; A frequency domain preference sequence is a sequence of data on the user's adjustments to the gain of different frequency bands during their historical listening process. It records the user's adjustments to the gain of each frequency band of the headphone equalizer at different times and arranges them in chronological order to form a sequence. For example, the user's adjustment records for low-frequency gain in the past month: [+3dB,+2dB,+3dB,+2.5dB], which form a low-frequency domain preference sequence.

[0044] A volume adaptive sequence is a sequence of volume adjustment data from a user's historical listening process. It records the volume setting value of the user each time they listen to music and arranges them in chronological order to form a sequence. For example, a user's listening volume records for the past month: [60%, 65%, 70%, 65%], which form a volume adaptive sequence.

[0045] The track type time sequence is a sequence of track types that a user listens to over time. It records the track types that the user listens to each time and arranges them in chronological order to form a sequence. For example, the track types that a user has listened to in the past month: [rock, pop, classical, rock], which constitute the track type time sequence.

[0046] Cross-modal behavior feature set It is composed of the following three sequences: Frequency domain preference sequence: The user's gain adjustment records for 5 core frequency bands (20-200Hz, 200-1kHz, 1-4kHz, 4-8kHz, 8-20kHz) in the past 30 days. The average value of each frequency band is taken, with a dimension of 5. Volume adaptive sequence: The user's listening volume records over the past 30 days, averaged and normalized to [0,1], with a dimension of 1; Track type time sequence: Distribution of music genres listened to by users in the past 30 days (rock, pop, classical, jazz, electronic), percentage of each genre, 5 dimensions, total 11 dimensions, outputting a 256-dimensional noise reduction behavior representation after comparative learning distillation network. .

[0047] For new users with no historical data, a pre-trained general behavioral feature template is loaded as the initial... This template is generated based on the average listening behavior data of 10,000 ordinary users.

[0048] Contrastive learning is a machine learning framework for learning effective feature representations. It constructs pairs of similar and dissimilar samples, enabling the model to learn features that distinguish samples. Cross-modal behavioral features of the same user are used as positive sample pairs, and cross-modal behavioral features of different users are used as negative sample pairs to construct training samples. Positive and negative sample pairs are the combinations of samples used for training in contrastive learning; positive sample pairs are samples that are semantically or behaviorally similar, and negative sample pairs are samples that are semantically or behaviorally dissimilar. For example, the listening behavior features of user A in two separate instances form a positive sample pair, and the listening behavior features of users A and B form a negative sample pair. Distillation networks are network structures used to transfer knowledge from complex models to lightweight models. Here, they are used to extract denoised behavioral features from the contrastive learning model. A lightweight, fully connected network is built, and the output of the contrastive learning model is used as input to distill a more concise feature vector. Denoising behavioral representations are user behavior feature vectors after removing abnormal data and interference information from the user's historical behavior. Distillation networks remove noise data from cross-modal behavioral features to obtain feature vectors that better reflect the user's true listening behavior. For example, inputting a cross-modal behavioral feature set into a distillation network outputs a 256-dimensional denoised behavioral representation. This vector removes the impact of abnormal data caused by user errors.

[0049] Construct a Bayesian preference probability model, prior distribution For a uniform distribution Uniform(0.1, 2.0), the likelihood function is... Given a diagonal Gaussian distribution, solve for the posterior probability. Output the expected dynamic preference weight coefficients: ,in, ; The Bayesian preference probability model is a model that calculates the probability distribution of user preference weights based on Bayes' theorem. It constructs a probability model with user preference calibration semantic representation and noise reduction behavioral representation as inputs and preference weights as outputs. The prior distribution is an assumption about the probability distribution of user preference weights before observing user feature data; a uniform distribution (Uniform(0.1, 2.0)) is used, representing that there are no additional assumptions about user preference weights within a reasonable range initially. The likelihood function is the probability distribution of observed user preference calibration semantic representation and noise reduction behavioral representation given user preference weights; it assumes that user feature data follows a diagonal Gaussian distribution with a mean of [missing value]. The variance is 0.1. The expected value of the posterior probability is calculated using the Laplace approximation method. The posterior probability is the updated probability distribution of user preference weights obtained by combining user feature data, and is calculated using Bayes' theorem combined with the prior distribution and the likelihood function. The expected dynamic preference weight coefficient is the final value of the user preference weight obtained by expecting the posterior probability distribution. The mathematical expectation of the posterior probability distribution is calculated to obtain a value between 0.1 and 2.0. For example, the expected value of the posterior probability distribution obtained by the Laplace approximation is 1.2, which is the dynamic preference weight coefficient. =1.2, this coefficient is used for subsequent weighted adjustment of frequency response compensation deviation.

[0050] Will Substitute the basic frequency response compensation deviation The preference correction compensation gain is obtained. ,in, Acoustic intent - frequency mapping factor, To retrieve the pre-stored target acoustic reference curve; To retrieve measured frequency response data that matches the model information from the headphone frequency response database; The fundamental frequency response compensation deviation is the amplitude difference between the target acoustic reference curve and the measured frequency response data of the headphone at each frequency point. It is calculated by comparing the amplitude difference between the target curve and the measured curve at each frequency point in the 20Hz-20kHz range. For example, at 1kHz, the amplitude of the target acoustic reference curve is 0dB, while the amplitude of the measured frequency response data of the headphone is -3dB. Therefore, the fundamental frequency response compensation deviation... The target acoustic reference curve is the industry-recognized ideal headphone frequency response curve, pre-stored standard acoustic curves such as the Harman 2023 headphone target frequency response curve; the measured frequency response data is the actual frequency response data of a specific headphone model collected through professional acoustic testing equipment. Frequency response data of different headphone models are pre-collected and stored in a database, such as the amplitude data of the Sony WH-1000XM5 headphone in the 20Hz-20kHz frequency band stored in the database; the preference correction compensation gain is the final compensation gain curve for headphone equalization obtained by combining the user's dynamic preference weighting coefficient, the basic frequency response compensation deviation, and the acoustic intent-frequency mapping factor, according to the formula... Calculate the compensation gain for each frequency point. For example, at 1kHz, the dynamic preference weighting coefficient is 1.2, the fundamental frequency response compensation deviation is 3dB, and the acoustic intent-frequency mapping factor is 1.0. Therefore, the preference correction compensation gain is calculated. =4.6dB, which means that the gain at this frequency point needs to be increased by 4.6dB.

[0051] The acoustic intent-frequency mapping factor is a coefficient that maps the user's acoustic intent features to the corresponding frequency bands. It is implemented using a piecewise constant function. The values ​​in Table 2 are taken for the corresponding frequency bands of each acoustic dimension, and 1.0 is taken for the other frequency bands, as shown in Table 2: Table 2 Acoustic Intent - Frequency Mapping Factor Table

[0052] Retrieve the standard gain curve corresponding to the preset tone model and adjust the gain to compensate for the preference. The fused gain is obtained by initially combining the fused gain with the center gain of the standard gain curve, and then embedded into the subsequent equalization parameter generation process.

[0053] The preset tone models are frequency response gain templates with different pre-defined tone styles. Multiple preset tone models are available, including Standard, Transparent, Heavy, and Realistic. Each model corresponds to a fixed frequency response gain curve; for example, the "Heavy" tone model corresponds to a curve with low-frequency gain boost and high-frequency gain attenuation. The standard gain curve is the frequency response gain curve corresponding to the preset tone model. Gain data for the 20Hz-20kHz frequency band is pre-stored for each preset tone model; for example, the gain curve for the "Standard" tone model is a flat curve with 0dB gain at each frequency point. The center gain is the overall average gain value of the standard gain curve, calculated as the average gain of the standard gain curve in the 20Hz-20kHz frequency band; for example, the center gain of the standard gain curve for the "Heavy" tone model is +1dB. Weighted blending adds the center gain of the preference correction compensation gain curve and the standard gain curve proportionally. The default blending weight is set to 0.7 (preference correction compensation gain): 0.3 (standard gain curve center gain), and users can adjust this weight in the web interface. For example, at a frequency of 1kHz, the preference correction compensation gain is 3.8dB, the center gain of the standard gain curve is 0dB, and the blending gain = 0.7×3.8 + 0.3×0 = 2.66dB. This curve will serve as the basis for adjusting the parameters of the subsequent parameter equalizer.

[0054] The beneficial effects of the above technical solution are as follows: by decoupling the acoustic intent hierarchy, simulating preference fluctuations, fusing cross-modal behavioral features and Bayesian probabilistic modeling, the user's fuzzy natural language listening preferences are transformed into precise quantitative compensation gains and fused with the preset timbre model, which greatly improves the matching degree between user preferences and headphone equalization parameters, effectively eliminates the deviation between text description and actual auditory needs, and provides reliable basic data support for the subsequent generation of equalization parameters adapted to the user's listening habits.

[0055] This invention provides an artificial intelligence-based personalized equalization adjustment method for headphones, which simulates preference fluctuation diagrams based on the main preference concatenation path and preference distribution, including: The listening preference description information is subjected to multi-dimensional semantic dependency parsing, and the preference intention, intensity modification, and scene constraint are mapped to standardized preference nodes, modification nodes and constraint nodes, respectively, to construct a preference dependency directed acyclic graph containing node hierarchical relationships and association strength.

[0056] In this embodiment, multi-dimensional semantic dependency parsing is a method of hierarchically parsing user listening preference descriptions from the perspective of semantic association, identifying different types of semantic elements and their interrelationships. Semantic analysis tools are used to analyze the grammatical dependencies of each word in the text, identifying the modifying relationship between low frequencies and fullness / powerfulness, and the scene constraint relationship between noisy environments and a clear sound field. Preference intent refers to the user's core listening needs, such as the user's description of full, powerful low frequencies and prominent vocals. Keyword matching is used to extract core words related to the acoustic dimension from the text to determine the user's preference intent.

[0057] Intensity modifiers are descriptive words used by users to describe the strength of their preference intentions, such as "full and powerful" or "not too harsh or clear." A dictionary of intensity modifiers is constructed, and modifiers in the text are matched to determine the intensity level of the preference; for example, "full and powerful" corresponds to "high intensity." Contextual constraints are content in the user's description that limits the listening environment, such as listening to rock music in a noisy environment. Contextual keywords are used to identify contextual information in the text and determine the applicable scenarios for the preference.

[0058] A preference node is a standardized data unit that represents a user's core listening preference intent. A node is created for each identified preference intent, storing the corresponding acoustic dimension information, such as a low-frequency preference node.

[0059] Modifier nodes are standardized data units that represent the intensity of a user's preference intent. A node is created for each intensity modifier to store the corresponding intensity level. For example, a high intensity modifier node corresponds to an intensity level of 8.

[0060] A constraint node is a standardized data unit that represents the constraints of a user's listening scenario. A node is created for each scenario constraint to store the corresponding scenario information, such as a rock scene constraint node.

[0061] In this embodiment, the preference depends on the construction rules of the directed acyclic graph: Node creation: Map user text preferences to preference nodes (P), intensity modifiers to modifier nodes (M), and scenario constraints to constraint nodes (C). Edge creation: Modifier nodes point to the corresponding preference nodes, and constraint nodes point to all preference nodes that are constrained by them; Edge weight calculation: Based on the confidence of semantic dependency, the value range is [0,1]. For example, the edge weight for greatly modifying low frequencies is 0.9, and the edge weight for constraining low frequencies when listening to rock music is 0.8.

[0062] The preference-dependent directed acyclic graph is parsed hierarchically to generate ordered main preference concatenation paths.

[0063] The primary preference concatenation path is generated using a topological sorting algorithm, sorting nodes by their in-degree from smallest to largest. If the in-degrees are the same, they are sorted by the sum of their edge weights from largest to smallest. For example, if a user inputs "When listening to rock music, the low frequencies are very full and the vocals are clear," the constructed preference-dependent directed acyclic graph contains: preference nodes: P1 (low frequencies), P2 (vocals); modifier nodes: M1 (very), M2 (clear); constraint node: C1 (when listening to rock music); edges: M1→P1 (weight 0.9), M2→P2 (weight 0.8), C1→P1 (weight 0.8), C1→P2 (weight 0.5). After topological sorting, the primary preference concatenation path is obtained as: [P1 (low frequencies), P2 (vocals)].

[0064] In this embodiment, hierarchical parsing is a processing method that extracts the ordered sequence of user preferences according to the hierarchy and association strength of nodes in the preference-dependent directed acyclic graph, from core to auxiliary and from primary to secondary. The topological sorting algorithm is used to traverse the directed acyclic graph and the traversal order is determined according to the in-degree and association strength of the nodes.

[0065] An ordered primary preference concatenation path is generated through hierarchical parsing. The primary preference concatenation path is an ordered sequence formed by arranging the user's core listening preferences from high to low priority. The path is constructed by the node order obtained through hierarchical parsing. For example, if the user describes that the low frequencies are full and powerful with the highest priority, followed by prominent vocals, then non-harsh high frequencies, and finally clear sound field in noisy environments, then the primary preference concatenation path is [low frequency preference, vocal preference, high frequency preference, sound field preference], which reflects the priority order of the user's preferences.

[0066] For each valid main preference in the main preference concatenation path, intensity modification information is extracted from the text and associated with the adjustment records corresponding to the preference in the user's historical listening behavior data to construct the intensity distribution data of the corresponding valid main preference.

[0067] Effective main preferences are core preferences that are semantically unconflicting, scenario-compatible, and logically reasonable in the main preference chain path. In implementation, it is necessary to remove preferences that are contradictory or do not conform to scenario constraints in the path. For example, remove contradictory parts in the user's description of low frequency as weak and low frequency as strong, and retain effective preferences that conform to the scenario.

[0068] Intensity modifiers are the specific content in user text that describes the strength of effective primary preferences. Modifiers corresponding to effective primary preferences are extracted from the text, such as "full and powerful" for low-frequency preferences.

[0069] User listening history data records a user's past actions related to listening preferences while using headphones. This includes data on equalization and volume adjustments made at different times, for different tracks, and in different scenarios. For example, a user might have repeatedly increased the low-frequency gain while listening to rock music over the past month. Adjustment records show specific adjustments made to the headphone's equalization parameters within this historical listening history data. For instance, a user might have adjusted the low-frequency gain to +3dB while listening to rock music and the mid-frequency gain to +2dB while listening to vocal tracks.

[0070] Intensity distribution data represents the distribution of intensity values ​​for a single valid primary preference. It can be combined with intensity modification information in the text and user historical adjustment records to statistically analyze the intensity range and distribution of the primary preference. For example, the intensity distribution of a user's low-frequency preference is 6-9 (corresponding to a gain range of +2dB to +4dB). This data can be used to determine the distribution range by statistically analyzing the mean and variance of the user's historical adjustment data.

[0071] Based on the correlation strength between nodes in the preference-dependent directed acyclic graph, a coupling correlation model is established between effective primary preferences. The mutual influence relationship between different effective primary preferences is quantified, and the coupling correlation model is adaptively corrected by combining the user's historical listening behavior data to eliminate the deviation between static text description and actual listening behavior.

[0072] The correlation strength between nodes is a quantitative value of the degree of close association between different nodes in a preference-dependent directed acyclic graph. It is calculated based on the semantic dependency relationship between nodes and user historical data. For example, the correlation strength between low-frequency preference and rock scene constraint nodes is 0.9, and the correlation strength between low-frequency preference and human voice preference is 0.3.

[0073] The coupled correlation model is a model that characterizes the mutual influence relationship between different effective principal preferences, and the formula is: ,in, Let i be the coefficient of influence of preference i on preference j; Let be the association strength between nodes i and j; , Let i be the average intensity of preference i and j.

[0074] Adaptive correction involves comparing the deviation between the preference association predicted by the model and the actual adjustment data of the user, and adjusting the association strength coefficient. For example, the association strength between low frequency and human voice in the model is 0.3, but when the user adjusts the low frequency, the human voice will also be adjusted accordingly. Therefore, the association strength is corrected to 0.5 to make the model more in line with the user's true preferences.

[0075] Using the main preference concatenation path as the transfer logic and the coupling correlation model as the constraint, the simulation interface is called to execute parallel simulation of the intensity distribution data of each effective main preference, generating a three-dimensional preference fluctuation model that includes the fluctuation range of preference intensity and the coupling effect of different preferences. When the preset simulation stop condition is reached, the simulation stop interface is called to terminate the simulation process. Based on the parallel simulation data, relevant initial boundary data is extracted and a preference fluctuation map is generated.

[0076] The transition logic is a rule that processes each primary preference sequentially according to the chain path of the primary preference during the simulation, and uses the priority order of the primary preferences as the progression order of the simulation.

[0077] The constraints are based on the mutual influence limits between different primary preferences set by the coupled correlation model. The influence range of the intensity change of one primary preference on other primary preferences is set according to the correlation strength. For example, when the intensity change of low frequency preference is ±1, the intensity change of human voice preference is ±0.3.

[0078] The simulation interface is a program interface used to start and control the simulation of preference data. It builds a multi-process simulation call interface, receives parameters such as the intensity distribution data of the main preference and the coupling correlation model to start the simulation.

[0079] Parallel simulation is the process of simultaneously simulating the intensity distribution data of multiple valid primary preferences. It employs multi-threading or multi-processing, launching multiple simulation tasks concurrently to process the intensity distribution data of different primary preferences. The three coordinate axes of the three-dimensional preference fluctuation model are defined as follows: X-axis: Main preference type (low frequency, mid frequency, high frequency, sound field, harmonics); Y-axis: Preference intensity (0-10); Z-axis: Preference coupling influence coefficient (-0.5-0.5). The simulation was implemented using the Python SimPy library. The number of parallel simulation processes equals the number of primary preferences. The simulation stopped when the preference intensity change was less than 0.01 for 10 consecutive iterations, or when the total number of iterations reached 100.

[0080] Parallel simulation data comprises all data generated during the simulation process, including intensity fluctuation data for each primary preference and data on the coupling effects of different preferences, which can be stored as simulation logs. Initial boundary data consists of boundary values ​​such as extreme values ​​and mean values ​​representing the range of user preference intensity extracted from the parallel simulation data; that is, the maximum, minimum, mean, and standard deviation are extracted from the simulation data as initial boundary data.

[0081] The beneficial effects of the above technical solution are as follows: by constructing a directed acyclic graph of preference dependency through multi-dimensional semantic dependency parsing, extracting the main preference concatenation path and constructing the preference intensity distribution by combining it with user historical data, establishing a coupled correlation model and adaptively correcting it, and finally generating a three-dimensional preference fluctuation model and preference fluctuation graph through parallel simulation, the structured, correlated and dynamic modeling of user listening preferences is realized, which provides a reliable constraint basis for the subsequent generation of headphone equalization parameters that accurately adapt to user preferences, and effectively improves the matching degree between user preferences and equalization effect.

[0082] This invention provides an artificial intelligence-based method for personalized equalization adjustment of headphones, which extracts relevant initial boundary data based on parallel simulation data and generates a preference fluctuation diagram, including: Based on the parallel simulation data, determine the first and last moments of the independent parallel simulation for each effective master preference, construct a set of simulation segments for each effective master preference, and calculate the mean and variance of the preference intensity for each simulation segment in each simulation segment set, as well as the rate of change of intensity between different simulation segments.

[0083] Independent parallel simulation is a simulation process that is started separately for a single effective primary preference. Simulation processes for different primary preferences run independently of each other. For example, the simulation processes for low-frequency preference and human voice preference do not interfere with each other.

[0084] The first time point is the initial time point when the independent parallel simulation of a single effective master preference begins execution, which is recorded by the timestamp of the simulation log. For example, the first time point of the low-frequency preference simulation is time t01. The last time point is the termination time point when the independent parallel simulation of a single effective master preference completes execution. For example, the last time point of the low-frequency preference simulation is time t02.

[0085] A simulation segment set is a collection of all segments after the independent parallel simulation data of a single effective primary preference is divided into multiple segments according to the time dimension. In implementation, the simulation data can be segmented according to a preset time length (such as 5 minutes). For example, the simulation data of low-frequency preference can be divided into segments of 5 minutes each to obtain multiple simulation segments. These simulation segments together constitute the simulation segment set of low-frequency preference.

[0086] A simulation segment is a single time-segment of data divided from the simulation data of a single valid primary preference. For example, the 5-minute data from 10:00:00 to 10:05:00 in a low-frequency preference simulation is a simulation segment.

[0087] The mean preference intensity is the average of all preference intensity values ​​within a single simulation segment, the variance is the dispersion of preference intensity values, and the rate of change of intensity is the rate of change of the mean preference intensity between adjacent simulation segments.

[0088] An intensity feature sequence is constructed based on the mean, variance, and rate of change of preference intensity for the same set of simulation segments. The intensity feature sequence is a one-dimensional numerical sequence formed by arranging the mean, variance, and rate of change of preference intensity for all simulation segments in the set according to the chronological order of the simulations. For example, if the mean intensity of the low-frequency preference simulation segments is 7.2, 7.5, and 7.8, the corresponding variances are 0.5, 0.4, and 0.6, and the rate of change of intensity is 0.06 and 0.06, then the constructed intensity feature sequence is [7.2, 0.5, 0.06, 7.5, 0.4, 0.06, 7.8, 0.6, 0.06]. This sequence fully reflects the dynamic change characteristics of the preference intensity of a single effective primary preference.

[0089] Based on the difference feature vectors between the intensity feature sequence and the intensity feature sequences of each other main preference, a preference fluctuation map is constructed.

[0090] The difference feature vector refers to the feature vector calculated by numerically differentiating the intensity feature sequence of the current effective main preference from the intensity feature sequence of each other effective main preference. It is generated by element-wise difference. For example, the intensity feature sequence of low frequency preference is [7.2, 0.5, 0.06], and the intensity feature sequence of human voice preference is [6.8, 0.4, 0.05]. The difference [0.4, 0.1, 0.01] obtained by subtracting the corresponding elements is the difference feature vector between low frequency preference and human voice preference. This vector quantifies the dynamic difference in the preference intensity of the two main preferences.

[0091] The preference fluctuation chart is a visual data graph that reflects the fluctuation of users' listening preferences by combining the difference feature vectors of all effective main preferences. The effective main preferences are used as the horizontal axis and the values ​​of the difference feature vectors are used as the vertical axis to draw a line graph. For example, the horizontal axis represents the main preferences such as low frequency, human voice, and high frequency, and the vertical axis represents the difference feature values ​​between each preference. The fluctuation shape of the curve intuitively presents the difference in preference intensity between different main preferences, thereby reflecting the dynamic change characteristics of user preferences. This graph can directly provide a constraint basis for the subsequent generation of equilibrium parameters.

[0092] The beneficial effects of the above technical solution are as follows: by extracting time data from parallel simulation, constructing the duration sequence of a single effective main preference, and calculating the difference feature vector between different main preferences, a visual data graph reflecting the fluctuation of user preferences is finally generated. This realizes the time dimension quantitative analysis of the dynamic characteristics of user listening preferences, provides precise constraint support for the subsequent generation of headphone equalization parameters that are more in line with the actual listening habits of users, and effectively improves the matching degree between equalization effect and user preferences.

[0093] This invention provides an artificial intelligence-based method for personalized equalization adjustment of headphones, which constructs a preference fluctuation graph, including: For each effective principal preference intensity feature sequence within the same simulation segment set, convert it into a standardized feature vector representation, and then perform a difference operation with the feature vector of each other effective principal preference intensity feature sequence to generate the difference feature vector of the corresponding effective principal preference.

[0094] Standardized feature vectors convert intensity feature sequences into standardized vectors of fixed dimensions. In practice, sequence values ​​can be normalized to the [0,1] interval, such as mapping [7.2,0.5,0.06] to a three-dimensional vector of [0.8,0.1,0.02] to eliminate the influence of dimensions.

[0095] The difference operation is to subtract the corresponding elements of the feature vector of the current effective principal preference from the feature vectors of each other effective principal preference. For example, subtracting the feature vector of low frequency preference from the feature vector of human voice prominence preference [0.75, 0.15, 0.01] will result in [0.05, -0.05, 0.01].

[0096] The elements in the difference feature vector are arranged in ascending order of numerical value. Taking the first smallest difference feature element after sorting as the benchmark, the remaining difference feature elements are matched numerically to filter out the feature elements whose numerical difference with the smallest difference feature element is within a preset threshold (such as 0.02), forming a finite feature subset. Based on the element combination form of the finite feature subset, the feature variables corresponding to the same simulation segment set are determined.

[0097] Ascending sort rearranges the elements of the differential feature vector in ascending order, such as sorting [0.05, -0.05, 0.01] into [-0.05, 0.01, 0.05]. The smallest differential feature element is the first element after sorting, which is -0.05.

[0098] Numerical similarity matching determines whether the numerical difference between other elements and the benchmark element is within a preset threshold, thereby filtering out elements with similar values. For example, 0.01 differs from the benchmark by 0.06, and is therefore discarded if it exceeds the threshold. Similarly, 0.05 differs from the benchmark by 0.10, which also exceeds the threshold. The finite feature subset is the set of continuous elements selected, such as {-0.05}.

[0099] Feature variables are unique identifiers determined by the combination of elements in a subset. Integer codes are assigned to different subsets. For example, the feature variable corresponding to a subset is identified as 101, which is used for subsequent matching of simulation log data.

[0100] The simulation log data generated by parallel simulation is divided into blocks and indexed according to the simulation execution dimension. Index entries corresponding to different feature variable types are constructed. The key of the index entry is the primary preference ID + simulation segment ID + feature variable ID, and the value is the storage address of the corresponding initial boundary data. Based on the identification information of the feature variable, the corresponding index entry is matched, and the initial boundary data matching the feature variable in the simulation log is retrieved.

[0101] Simulation log data is structured data that records the parallel simulation process. Each log entry contains information such as the main preference, simulation segment, timestamp, and preference strength.

[0102] Simulation execution dimension is a classification dimension used to divide log data, such as by main preference, simulation segment, or feature variable type.

[0103] A block index divides log data into multiple data blocks according to the dimensions mentioned above, and creates an index for each block, storing the index in key-value pair format.

[0104] An index entry is an index item corresponding to each data block, which contains the feature variable identifier and the data storage path.

[0105] The initial boundary data is subjected to probability distribution fitting processing. Based on the fitted distribution shape, the boundary value of the corresponding effective main preference confidence interval is determined, and the confidence interval of the effective main preference is generated.

[0106] Probability distribution fitting involves using a mathematical model to fit the distribution pattern of the initial boundary data. A Gaussian distribution (normal distribution) is used to fit the data, and the mean and variance of the distribution are calculated through maximum likelihood estimation. For example, after fitting the initial boundary data with low frequency preference, a normal distribution with a mean of 7.5 and a variance of 0.5 is obtained.

[0107] The distribution shape is the fitted distribution curve, such as a bell-shaped normal distribution curve. The confidence interval boundary values ​​are the distribution quantiles at a preset confidence level (e.g., 95%). Calculate the 2.5% and 97.5% quantiles of the normal distribution, such as 6.5 and 8.5 quantiles respectively.

[0108] A confidence interval is the range between the upper and lower boundary values. For example, the confidence interval for low-frequency preference intensity is [6.5, 8.5], which means that there is a 95% probability that the user's low-frequency preference intensity is within this interval.

[0109] Traverse all interval segments within the confidence interval, identify interval segments that exceed the preset value range as extreme fluctuation intervals, and adjust the extreme fluctuation intervals by linear interpolation replacement; for extreme fluctuation intervals located at both ends of the confidence interval, replace them with the mean of the nearest neighbor normal interval to generate corrected interval segments within the user's preferred acceptable range.

[0110] Interval segments are data segments into which a confidence interval is divided according to a preset step size (e.g., 0.5). For example, [6.5, 8.5] is divided into four segments: [6.5, 7.0], [7.0, 7.5], [7.5, 8.0], and [8.0, 8.5]. The preset numerical range is the acceptable range of preference intensity determined based on the user's historical listening behavior data, such as [6, 9] set based on the user's past adjustment records.

[0111] Extreme fluctuation ranges are the portions of a confidence interval that exceed the preset range. For example, if a segment of [8.6, 8.8] exists in the confidence interval, exceeding the upper limit acceptable to the user, it will be identified as an extreme fluctuation range.

[0112] Linear interpolation replacement uses the values ​​of adjacent normal interval segments to perform linear interpolation to replace the values ​​of extreme intervals. That is, it takes the effective interval endpoints before and after the extreme interval. For example, linear interpolation is performed using [8.0, 8.5] and the upper limit of the user's acceptable range to generate a corrected segment [8.2, 8.5], which is within the user's acceptable range.

[0113] For extreme fluctuation ranges at the ends of the confidence interval, such as [5.5, 6.0], the mean of the nearest normal range [6.0, 6.5] (6.25) is used for replacement. The corrected interval segment is the interval data that conforms to the user's preferred range after replacement.

[0114] The corrected interval segments are spliced ​​and integrated with the non-extreme interval segments within the confidence interval to form the effective intervals of the corresponding effective principal preferences. The effective intervals of all effective principal preferences are then structured and mapped according to the simulation segment set dimension to construct a complete preference fluctuation diagram.

[0115] The concatenation and integration process involves concatenating the corrected interval segments with the segments within the original confidence interval that do not exceed the range, in chronological order. For example, concatenating [6.5,7.0], [7.0,7.5], [7.5,8.0], and the corrected [8.2,8.5] yields the effective interval [6.5,8.5] for low-frequency preferences. Structured mapping organizes the effective intervals of each effective primary preference according to the time dimension of the simulation segment set, such as arranging the effective intervals of each segment in chronological order.

[0116] Based on the effective interval of the preference fluctuation graph, new preferences that users have not explicitly stated are identified through implicit semantic analysis, and the strength of the new preferences is determined according to the mean of the effective interval.

[0117] Implicit semantic analysis infers users' implicit listening needs by analyzing the relationships and intensity distribution among their explicitly stated preferences. For example, if a user explicitly states that low frequencies are elastic, and the preference fluctuation plot shows that the effective interval of mid-frequency density is 7.5, which is higher than the average level of ordinary users, then a new preference for mid-frequency clarity can be identified, with a preference strength of 7.5.

[0118] The beneficial effects of the above technical solution are as follows: by converting the time sequence into a feature vector, filtering differential features, retrieving the simulation log index, fitting the probability distribution, and correcting extreme intervals, the simulation data of user listening preferences is refined, noise and outliers are eliminated, and an effective range that fits the actual user preferences is obtained. The constructed preference fluctuation diagram can accurately present the dynamic change characteristics of user preferences, providing reliable data support for the subsequent generation of headphone equalization parameters that are adapted to the user's listening habits, and effectively improving the matching degree between the equalization effect and user preferences.

[0119] This invention provides an artificial intelligence-based method for personalized equalization adjustment of headphones, generating parametric equalizer adjustment parameters including at least one set of center frequency, gain value, and quality factor Q value, including: Based on the target equalization gain curve obtained from multi-dimensional acoustic analysis, the auditory perception characteristics of different frequencies are adapted and corrected by combining the human ear equal loudness curve. At the same time, the gain value that exceeds the user's preference acceptance range is limited according to the initial boundary data of the preference fluctuation diagram, so as to obtain a perception equalization compensation curve that is adapted to human ear perception and fits the user's preference range. The perceived equalization compensation curve is decomposed into a multi-scale frequency domain, and the peak points and valley points under different decomposition scales are extracted. For each extracted extreme point, the equalization gain change gradient and curvature characteristics of the frequency band are calculated. Extreme points whose change gradient and curvature characteristics meet the preset conditions are selected as the candidate center frequency set. For each center frequency in the candidate center frequency set, the corresponding gain value on the sensing equalization compensation curve is extracted as the initial gain. At the same time, based on the phase frequency response data at the center frequency, the group delay change caused by the filter under different quality factors is analyzed. According to the acceptable range of group delay change, the quality factor corresponding to the center frequency is adjusted. A simulation model of parametric equalizer filter superposition is constructed. All candidate center frequencies, initial gains and adjusted quality factors are input into the superposition simulation model to generate the composite frequency response curve of multiple filters superimposed. The synthesized frequency response curve and the perceived equalization compensation curve are compared point by point. If the deviation exceeds the preset range, the gain value or quality factor corresponding to the center frequency with the excessive deviation is adjusted until the deviation between the synthesized frequency response curve and the perceived equalization compensation curve is within the preset range. The adjusted center frequency, gain value, and quality factor Q value are grouped and packaged according to frequency band to generate multiple sets of equalizer adjustment parameters, with each set of parameters corresponding to equalization compensation for one frequency band.

[0120] The target equalization gain curve is the curve data output by this multi-dimensional analysis that represents the target gain value at different frequencies. For example, in the 20Hz to 20kHz frequency band, the target gain is +3dB at 20Hz, 0dB at 1kHz, and -1dB at 10kHz. The equal loudness curves of the human ear are standard curves characterizing the auditory sensitivity of the human ear to sounds of different frequencies. ,in, For target gain; Reference sound pressure level (default 80dB); The sound pressure level at frequency f is the equal loudness curve of 80 dB.

[0121] The equal-loudness curves of the human ear are standard curves characterizing the auditory sensitivity of the human ear to sounds of different frequencies. Auditory perception characteristics refer to the subjective differences in how the human ear perceives sounds of different frequencies. Adaptation correction adjusts the target equalization gain curve based on these characteristics. For example, the human ear is not sensitive to the 20Hz low frequency; to achieve the same loudness as 80dB at 1kHz, a sound pressure level of 100dB is needed at 20Hz. Therefore, the target gain for the 20Hz low frequency is increased by another 20dB to compensate for the insufficient perception of the human ear. The preference fluctuation diagram is a data graph generated in the previous steps that characterizes the fluctuation of the user's listening preference intensity. Its initial boundary data is the extreme range of user preference intensity extracted from parallel simulation data. For example, the effective range for the user's low-frequency preference intensity is 6-9, corresponding to a gain value range of +2dB to +4dB. The user's preferred acceptable range is the gain value range corresponding to this effective interval. The limiting process involves adjusting the gain values ​​in the target equalization gain curve that exceed this range to the boundary value. For example, if the gain at 20Hz in the target curve is +5dB, which exceeds the upper limit of +4dB, then it is limited to +4dB. After the above adaptation correction and limiting process, the resulting perceptual equalization compensation curve is the final target curve that conforms to the characteristics of human auditory perception and does not exceed the user's preferred acceptable range.

[0122] Multi-scale frequency domain decomposition is a method of splitting the perceptual equalization compensation curve according to different frequency scales. Using wavelet decomposition algorithms, the curve can be decomposed into three scales, corresponding to the low-frequency band of 20-500Hz, the mid-frequency band of 500-6kHz, and the high-frequency band of 6-20kHz. The decomposition scale represents the different frequency levels after splitting. The peak point is the point on the perceptual equalization compensation curve with the largest gain value within a certain frequency band, such as +4dB at 20Hz, while the valley point is the point with the smallest gain value, such as -1dB at 10kHz. Both are collectively referred to as extreme points. The equalization gain gradient refers to the rate at which the gain value changes with frequency, calculated by dividing the gain difference between adjacent frequency points by the frequency difference. Within the 20Hz to 1kHz frequency band, the gain drops from +4dB to 0dB, with a gradient of approximately -0.004dB / Hz. Curvature characteristics refer to the degree of curvature of the perceived equalization compensation curve at extreme points. Calculated using the second derivative, a larger curvature indicates a sharper curve, while a smaller curvature indicates a smoother curve. Preset conditions are the criteria for selecting effective extreme points, setting a gradient ≥ 0.001dB / Hz and a curvature ≥ 0.0001 to ensure that the selected points are frequency bands with significant gain changes requiring focused equalization. The candidate center frequency set is the set of frequency values ​​of all extreme points that meet the conditions, such as [20Hz, 1kHz, 10kHz]. The relevant equalization parameters are shown in Table 3. Table 3 Equilibrium Parameter Table

[0123] Initial gain refers to the target gain value corresponding to the center frequency on the perceptual equalization compensation curve, such as +4dB corresponding to a center frequency of 20Hz; phase response data refers to the measured phase response data of the headphones at different frequencies, retrieved from the headphone frequency response database, for example, the phase shift of a certain headphone at 20Hz is -10 degrees; quality factor (Q value) is a parameter characterizing the bandwidth of the parametric equalizer filter. The larger the Q value, the narrower the filter bandwidth, and the more concentrated the adjustment of the center frequency, but it will lead to an increase in group delay; group delay variation refers to the time delay generated when the audio signal passes through the filter. The acceptable range is the upper limit of the delay that will not cause muddiness or inaccurate positioning, which can be set to ≤3ms; adjusting the quality factor is to optimize the Q value according to the group delay variation. For example, if the phase shift is large at 20Hz, setting the Q value to 10 will result in a group delay of 5ms, which exceeds the acceptable range. In this case, the Q value is adjusted to 2 to reduce the group delay to 1ms; if the phase shift is small at 1kHz, the Q value can be set to 4, and the group delay is 2ms, which meets the requirements. Finally, the adjusted Q value set is obtained, for example [2,4,3].

[0124] The parametric equalizer filter superposition simulation model is a simulation model that simulates the simultaneous operation of multiple parametric equalizer filters. It is built based on IIR filters, and each filter is defined by a center frequency, gain, and Q value. After inputting all candidate center frequencies, initial gains, and adjusted quality factors into the model, the model will calculate the total frequency response after the multiple filters are superimposed and generate a composite frequency response curve. This curve is the audio gain response under the combined action of all filters. For example, after superimposing three filters, the curves show a gain of +4dB at 20Hz, 0dB at 1kHz, and -1dB at 10kHz.

[0125] Point-by-point deviation comparison refers to calculating the difference in gain value between the synthesized frequency response curve and the perceived equalization compensation curve at each frequency point (e.g., one sampling point per 10Hz). The preset range is the acceptable deviation threshold, set at ±0.5dB. Deviation exceeding the limit means that the difference at a certain frequency point exceeds the preset range. For example, if the gain of the synthesized curve at 20Hz is +4.8dB, the deviation from the target curve's +4dB is 0.8dB, exceeding the ±0.5dB range. Adaptation adjustment involves fine-tuning the filter parameters for those with excessive deviations, such as adjusting the gain of the 20Hz filter from +4dB to +4.2dB, or reducing the Q value to widen the filter bandwidth and reduce the peak gain, until the deviation at that frequency point is reduced to within ±0.5dB. The comparison and adjustment process is repeated until the deviations at all frequency points are within the preset range.

[0126] Grouping and encapsulation refers to classifying parameters by frequency band, specifically low-frequency (20-500Hz), mid-frequency (500-6kHz), and high-frequency (6-20kHz). Each frequency band's parameters are encapsulated into a set of equalizer adjustment parameters. For example, the low-frequency group parameters are a center frequency of 20Hz, a gain of +4.2dB, and a Q value of 2; the mid-frequency group parameters are a center frequency of 1kHz, a gain of 0dB, and a Q value of 4; and the high-frequency group parameters are a center frequency of 10kHz, a gain of -1dB, and a Q value of 3. Each set of parameters corresponds to equalization compensation for one frequency band and can be directly loaded into the DSP processing module for use.

[0127] The beneficial effects of the above technical solution are as follows: by combining the characteristics of human hearing perception and the user's preference range, key equalization points are screened through multi-scale frequency domain decomposition, the filter group delay is optimized based on phase frequency response, and then through simulation fitting and deviation correction, the generated parametric equalizer adjustment parameters not only fit the user's listening preferences but also conform to the laws of human hearing perception, effectively improving the accuracy of headphone equalization effect and listening experience.

[0128] This invention provides a personalized equalization adjustment method for headphones based on artificial intelligence, which performs equalization filtering processing on the real-time input audio signal according to the DSP filter coefficients, including: The coefficients of the DSP filter are analyzed into a multi-order IIR filter bank that corresponds one-to-one with the equalizer adjustment parameters of each frequency band. The transfer function of each filter is defined by the corresponding center frequency, gain value and quality factor Q value. Based on the measured phase response data of the headphone's sound unit, a composite filtering unit is constructed by pre-configuring phase compensation coefficients for each filter. The real-time input audio signal is processed into frames to generate a sequence of consecutive short audio frames of fixed duration. Transient feature detection is performed on each audio frame to identify transient peak signal segments and steady-state signal segments. For transient peak signal segments, dynamically reduce the quality factor Q value of the corresponding frequency band filter to broaden the filtering bandwidth; For the steady-state signal segment, recover the preset Q value of the filter; The audio frame is input into the composite filtering unit for frequency band equalization filtering. At the same time, based on the group delay parameters of each filter, the filtered signals of different frequency bands are timestamped to compensate for the time domain offset caused by the difference in filter bandwidth. The time-aligned signals of each frequency band are synthesized to generate an equalized time-domain audio signal. Real-time intermodulation distortion detection is performed on the equalized time-domain audio signal. If the detected distortion exceeds the preset threshold, the gain value of the corresponding frequency band filter is finely adjusted based on the frequency distribution of the distortion component until the distortion falls back to the preset range. Finally, the equalized audio signal is output to the headphone driver unit.

[0129] In this embodiment, the DSP filter coefficients are computational parameters converted from the parameter equalizer adjustment parameters in the preceding steps and can be directly called by the DSP processing module of the audio device. Fourth-order IIR filter coefficients adapted to the ARM architecture are used, such as filter coefficients for the 20Hz low-frequency band. The analytical processing involves splitting these general coefficients by frequency band and mapping them to the corresponding low-frequency, mid-frequency, and high-frequency parameter equalizer adjustment parameters, forming a multi-order IIR filter bank that corresponds one-to-one with each frequency band. This filter bank consists of multiple high-order infinite impulse response (IIR) filters, such as a fourth-order IIR filter corresponding to each frequency band. The transfer function is a mathematical expression describing the relationship between the filter's input and output signals. It is generated using the standard bilinear transform method based on the center frequency (e.g., 20Hz), gain value (e.g., +4dB), and quality factor Q value (e.g., 2), ensuring that the transfer function of each filter accurately matches the frequency band parameters.

[0130] The headphone driver unit is the device that drives sound in the headphone, such as a dynamic driver unit; the measured phase response data is the phase shift data of the headphone at different frequencies collected by professional acoustic testing equipment, such as a phase shift of -10 degrees at 20Hz for a certain headphone; the phase compensation coefficient is the correction value used to compensate for this phase shift, such as configuring a +10 degree compensation coefficient for a 20Hz filter; the composite filter unit is a functional module that binds the IIR filter of each frequency band with the corresponding phase compensation coefficient. The compensation coefficient is configured through hardware registers, so that each composite unit has both equalization filtering and phase correction functions. For example, the low-frequency composite unit includes a 20Hz IIR filter and a +10 degree phase compensation coefficient.

[0131] Real-time input audio signals are the audio streams input during headphone playback, such as the audio signals of songs or videos; frame segmentation processes cut continuous audio signals into segments of fixed duration, typically 20ms per frame. For example, at a 48kHz sampling rate, each frame contains 960 sampling points; short-time audio frame sequences are ordered sets of these continuous segments, such as audio sequences generating 50 frames per second; transient feature detection identifies energy abrupt changes in the audio, achieved by calculating the peak factor (the ratio of peak value to root mean square value) for each frame, such as... ≥6 indicates a transient signal. <6 indicates a steady-state signal; the transient peak signal segment consists of segments with sudden energy changes, such as drum beats and percussion, while the steady-state signal segment consists of segments with stable energy, such as vocals and strings. The peak factor is: ,in, For audio frame sampling points; The value is the root mean square value. The dynamic Q-value adjustment rules are as follows: transient peak signal segment: Q-value is reduced by 50%; steady-state signal segment: the preset Q-value is restored.

[0132] The preset Q value is the initial quality factor obtained from the group delay optimization in the previous step, such as Q=2 in the low-frequency band; the recovery operation is to adjust the Q value of the filter back to the preset value when the signal returns to the steady state. For example, in the human voice signal segment, the Q value in the low-frequency band is restored to 2 to ensure the accuracy of the equalization filter and make the frequency response of the steady-state signal meet the requirements of the perceptual equalization compensation curve.

[0133] Frequency band equalization filtering divides audio frames into frequency bands and processes them through corresponding composite filtering units. For example, low-frequency frames are processed by a low-frequency composite unit, and mid-frequency frames are processed by an mid-frequency unit. The group delay parameter is the time delay generated when each filter processes the signal. For example, the group delay of the low-frequency filter is 1ms, and that of the mid-frequency filter is 0.5ms. Timestamp correction adds timestamps to the filtered signals of different frequency bands and adjusts the timestamps according to the group delay parameter. For example, the timestamp of the low-frequency signal is advanced by 0.5ms to compensate for the delay difference with the mid-frequency signal. Time domain offset is the time misalignment of signals of different frequency bands. After correction, the signals of each frequency band can be synchronously superimposed to avoid muddy sound or inaccurate positioning.

[0134] Signal synthesis is the process of superimposing time-aligned signals from different frequency bands into a complete audio signal. Real-time intermodulation distortion detection employs FFT analysis, performing a 1024-point FFT on the equalized audio signal to calculate the ratio of the power of the intermodulation distortion component (e.g., 1kHz + 2kHz = 3kHz) to the fundamental power. The preset threshold is set to ≤1%. If the threshold is exceeded, the main frequency band of the distortion component needs to be identified, such as the low frequency band. Fine-tuning the gain value is to slightly adjust the gain of the corresponding frequency band filter. For example, the low frequency gain is reduced from +4dB to +3.5dB to reduce intermodulation distortion. After the distortion drops back to within the threshold, the final equalized signal can be output to the headphone driver. If the distortion still exceeds the threshold after 5 consecutive fine-tunings, the system will automatically reduce the overall gain by 3dB and prompt the user on the web interface that the current gain is too high and has been automatically reduced to avoid distortion.

[0135] The beneficial effects of the above technical solution are as follows: by analyzing the DSP filter coefficients to construct a composite filter unit, and combining the dynamic Q value adjustment, timestamp correction and real-time distortion detection of transient / steady-state signals, accurate equalization filtering and phase and distortion correction of audio signals are achieved, effectively reducing transient distortion, phase shift and intermodulation distortion, and improving the fidelity and listening experience of headphone playback audio.

[0136] This invention provides an artificial intelligence-based personalized equalization adjustment system for headphones, such as... Figure 2 As shown, it includes: a communication establishment module: the user terminal accesses the Web control interface through the local area network and establishes a two-way communication connection with the audio device; The first interface receiving module is used to receive the model information of the currently used headphones input or selected by the user based on the Web control interface; The second interface receiving module is used to receive listening preference description information input by the user in the form of natural language text based on the Web control interface; The sending module is used to send the model information and the listening preference description information to the artificial intelligence tuning module; The acoustic analysis module is used by the artificial intelligence tuning module to retrieve the measured frequency response data corresponding to the model information from the pre-stored headphone frequency response database, and perform multi-dimensional acoustic analysis by combining the pre-stored target acoustic reference curve, preset timbre model and the listening preference description information to generate parametric equalizer adjustment parameters including at least one set of center frequency, gain value and quality factor Q value. The parameter conversion module is used to convert the parameter equalizer adjustment parameters into DSP filter coefficients that can be recognized by the audio device's DSP processing module through the DSP parameter conversion module. A coefficient loading module is used to load the DSP filter coefficients into the DSP processing module of the audio device, wherein the DSP processing module performs equalization filtering on the real-time input audio signal according to the DSP filter coefficients and outputs it to the headphones; The continuing execution module is used to receive a save command sent by the user through the Web control interface, store the parametric equalizer adjustment parameters in the internal memory of the audio device, and generate a corresponding entry in the local EQ preset list of the audio device. This allows the audio device to retrieve the parametric equalizer adjustment parameters from the internal memory and convert them into DSP filter coefficients to perform audio equalization processing when disconnected from the network.

[0137] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A personalized equalization adjustment method for headphones based on artificial intelligence, characterized in that, This includes: Step 1: The user terminal accesses the Web control interface via the local area network and establishes a two-way communication connection with the audio device; Step 2: Receive the model information of the currently used headphones, input or selected by the user, based on the Web control interface; Step 3: Receive the user's listening preference description information in natural language text form based on the Web control interface; Step 4: Send the model information and the listening preference description information to the artificial intelligence tuning module; Step 5: The AI ​​tuning module retrieves the measured frequency response data corresponding to the model information from the pre-stored headphone frequency response database, and performs multi-dimensional acoustic analysis by combining the pre-stored target acoustic reference curve, preset timbre model and the listening preference description information to generate parametric equalizer adjustment parameters containing at least one set of center frequency, gain value and quality factor Q value. Step 6: Convert the parameter equalizer adjustment parameters into DSP filter coefficients that can be recognized by the audio device's DSP processing module through the DSP parameter conversion module; Step 7: Load the DSP filter coefficients into the DSP processing module of the audio device, wherein the DSP processing module performs equalization filtering on the real-time input audio signal according to the DSP filter coefficients and outputs it to the headphones; Step 8: After receiving the save command sent by the user through the Web control interface, store the parametric equalizer adjustment parameters in the internal memory of the audio device, and generate a corresponding entry in the local EQ preset list of the audio device, so that the audio device can retrieve the parametric equalizer adjustment parameters from the internal memory and convert them into DSP filter coefficients to perform audio equalization processing when disconnected from the network connection.

2. The personalized equalization adjustment method for headphones based on artificial intelligence according to claim 1, characterized in that, The process of performing multi-dimensional acoustic analysis based on the aforementioned listening preference description information includes: The listening preference description information is input into the acoustic intent level decoupler, and decoupled into a discrete acoustic intent feature set according to the dimensions of low-frequency elasticity, mid-frequency density, high-frequency extension, sound field width, and harmonic saturation. High-dimensional semantic representations are output through BERT acoustic pre-trained models. ; Obtain the main preference concatenation path of the listening preference description information and the preference distribution of each main preference in the main preference concatenation path, and simulate the preference fluctuation diagram based on the main preference concatenation path and preference distribution, wherein the preference fluctuation diagram includes the preference strength of at least one newly added preference; Based on the preference strength of each new preference involved in the preference fluctuation graph, a high-dimensional semantic representation is generated. After correction, a preference-calibrated semantic representation is obtained. ; Collect user's historical listening frequency domain preference sequence, volume adaptation sequence, and track type time sequence to construct a cross-modal behavioral feature set. A contrastive learning framework is used to construct positive and negative sample pairs, which are then output as a noise reduction behavior representation via a distillation network. For new users without historical data, a pre-trained general behavioral feature template is loaded as the initial... ; Construct a Bayesian preference probability model, prior distribution Let be a uniform distribution Uniform(0.1, 2.0), and the likelihood function be... Given a diagonal Gaussian distribution, solve for the posterior probability. Output the expected dynamic preference weight coefficients: ,in, ; Will Substitute the basic frequency response compensation deviation The preference correction compensation gain is obtained. ,in, Acoustic intent - frequency mapping factor, To retrieve the pre-stored target acoustic reference curve; To retrieve measured frequency response data that matches the model information from the headphone frequency response database; Retrieve the standard gain curve corresponding to the preset tone model and adjust the gain to compensate for the preference. The fused gain is obtained by initially combining the fused gain with the center gain of the standard gain curve, and then embedded into the subsequent equalization parameter generation process.

3. The personalized equalization adjustment method for headphones based on artificial intelligence according to claim 2, characterized in that, Based on the simulated preference fluctuation diagram using the main preference concatenation path and preference distribution, the following is included: Multi-dimensional semantic dependency parsing is performed on the listening preference description information, and preference intention, intensity modification, and scene constraint are mapped to standardized preference nodes, modification nodes and constraint nodes respectively, and a preference dependency directed acyclic graph containing node hierarchical relationship and association strength is constructed. The preference-dependent directed acyclic graph is parsed hierarchically to generate ordered main preference concatenation paths; For each valid main preference in the main preference concatenation path, intensity modification information is extracted from the text and associated with the adjustment records corresponding to the preference in the user's historical listening behavior data to construct the intensity distribution data of the corresponding valid main preference; Based on the correlation strength between nodes in the preference-dependent directed acyclic graph, a coupling correlation model between effective primary preferences is established to quantify the mutual influence between different effective primary preferences. The coupling correlation model is then adaptively corrected by combining user historical listening behavior data to eliminate the deviation between static text description and actual listening behavior. Using the main preference concatenation path as the transfer logic and the coupling correlation model as the constraint, the simulation interface is called to execute parallel simulation of the intensity distribution data of each effective main preference, generating a three-dimensional preference fluctuation model that includes the fluctuation range of preference intensity and the coupling effect of different preferences. When the preset simulation stop condition is reached, the simulation stop interface is called to terminate the simulation process. Based on the parallel simulation data, relevant initial boundary data is extracted and a preference fluctuation map is generated.

4. The personalized equalization adjustment method for headphones based on artificial intelligence according to claim 3, characterized in that, Relevant initial boundary data are extracted based on parallel simulation data, and preference fluctuation plots are generated, including: Based on the parallel simulation data, determine the first and last moments of the independent parallel simulation for each effective primary preference, construct a set of simulation segments for each effective primary preference, and calculate the mean and variance of preference intensity for each simulation segment in each simulation segment set, as well as the rate of change of intensity between different simulation segments. Based on the mean, variance, and rate of change of preference intensity of the same set of simulation segments, an intensity feature sequence is constructed. Based on the difference feature vectors between the intensity feature sequence and the intensity feature sequences of each other main preference, a preference fluctuation map is constructed.

5. The personalized equalization adjustment method for headphones based on artificial intelligence according to claim 4, characterized in that, The resulting preference fluctuation graph includes: For each effective principal preference intensity feature sequence within the same simulation segment set, convert it into a standardized feature vector representation, and then perform a difference operation with the feature vector of each other effective principal preference intensity feature sequence to generate the difference feature vector of the corresponding effective principal preference. The elements in the difference feature vector are arranged in ascending order of numerical value. Taking the first smallest difference feature element after sorting as the benchmark, the numerical similarity matching of the remaining difference feature elements is performed in turn. Feature elements whose numerical difference with the smallest difference feature element is within a preset threshold are selected to form a finite feature subset. Based on the element combination form of the finite feature subset, the feature variables corresponding to the same simulation segment set are determined. The simulation log data generated by parallel simulation is divided into blocks and indexed according to the simulation execution dimension. Index entries corresponding to different feature variable types are constructed, and the corresponding index entries are matched based on the identification information of the feature variable. The initial boundary data that matches the feature variable in the simulation log is retrieved. The initial boundary data is subjected to probability distribution fitting processing, and the confidence interval boundary value of the corresponding effective main preference is determined based on the fitted distribution shape, thereby generating the confidence interval of the effective main preference; Traverse all interval segments within the confidence interval, identify interval segments that exceed the preset value range as extreme fluctuation intervals, and adjust the extreme fluctuation intervals by linear interpolation replacement; for extreme fluctuation intervals located at both ends of the confidence interval, replace them with the mean of the nearest neighbor normal interval to generate corrected interval segments within the user's preferred acceptable range. The corrected interval segments are spliced ​​and integrated with the non-extreme interval segments within the confidence interval to form the effective intervals of the corresponding effective principal preferences. The effective intervals of all effective principal preferences are then structured and mapped according to the simulation segment set dimension to construct a complete preference fluctuation diagram. Based on the effective interval of the preference fluctuation graph, new preferences that users have not explicitly stated are identified through implicit semantic analysis, and the strength of the new preferences is determined according to the mean of the effective interval.

6. The personalized equalization adjustment method for headphones based on artificial intelligence according to claim 1, characterized in that, Generate parametric equalizer adjustment parameters that include at least one set of center frequency, gain value, and quality factor Q value, including: Based on the target equalization gain curve obtained from multi-dimensional acoustic analysis, the auditory perception characteristics of different frequencies are adapted and corrected by combining the human ear equal loudness curve. At the same time, the gain value that exceeds the user's preference acceptance range is limited according to the initial boundary data of the preference fluctuation diagram, so as to obtain a perception equalization compensation curve that is adapted to human ear perception and fits the user's preference range. The perceived equalization compensation curve is decomposed into a multi-scale frequency domain, and the peak points and valley points under different decomposition scales are extracted. For each extracted extreme point, the equalization gain change gradient and curvature characteristics of the frequency band are calculated. Extreme points whose change gradient and curvature characteristics meet the preset conditions are selected as the candidate center frequency set. For each center frequency in the candidate center frequency set, the corresponding gain value on the sensing equalization compensation curve is extracted as the initial gain. At the same time, based on the phase frequency response data at the center frequency, the group delay change caused by the filter under different quality factors is analyzed. According to the acceptable range of group delay change, the quality factor corresponding to the center frequency is adjusted. A simulation model of parametric equalizer filter superposition is constructed. All candidate center frequencies, initial gains and adjusted quality factors are input into the superposition simulation model to generate the composite frequency response curve of multiple filters superimposed. The synthesized frequency response curve and the perceived equalization compensation curve are compared point by point. If the deviation exceeds the preset range, the gain value or quality factor corresponding to the center frequency with the excessive deviation is adjusted until the deviation between the synthesized frequency response curve and the perceived equalization compensation curve is within the preset range. The adjusted center frequency, gain value, and quality factor Q value are grouped and packaged according to frequency band to generate multiple sets of equalizer adjustment parameters, with each set of parameters corresponding to equalization compensation for one frequency band.

7. The personalized equalization adjustment method for headphones based on artificial intelligence according to claim 1, characterized in that, The real-time input audio signal is subjected to equalization filtering based on the DSP filter coefficients, including: The coefficients of the DSP filter are analyzed into a multi-order IIR filter bank that corresponds one-to-one with the equalizer adjustment parameters of each frequency band. The transfer function of each filter is defined by the corresponding center frequency, gain value and quality factor Q value. Based on the measured phase response data of the headphone's sound unit, a composite filtering unit is constructed by pre-configuring phase compensation coefficients for each filter. The real-time input audio signal is processed into frames to generate a sequence of consecutive short audio frames of fixed duration. Transient feature detection is performed on each audio frame to identify transient peak signal segments and steady-state signal segments. For transient peak signal segments, dynamically reduce the quality factor Q value of the corresponding frequency band filter to broaden the filtering bandwidth; For the steady-state signal segment, recover the preset Q value of the filter; The audio frame is input into the composite filtering unit for frequency band equalization filtering. At the same time, based on the group delay parameters of each filter, the filtered signals of different frequency bands are timestamped to compensate for the time domain offset caused by the difference in filter bandwidth. The time-aligned signals of each frequency band are synthesized to generate an equalized time-domain audio signal. Real-time intermodulation distortion detection is performed on the equalized time-domain audio signal. If the detected distortion exceeds the preset threshold, the gain value of the corresponding frequency band filter is finely adjusted based on the frequency distribution of the distortion component until the distortion falls back to the preset range. Finally, the equalized audio signal is output to the headphone driver unit.

8. A personalized equalization adjustment system for headphones based on artificial intelligence, characterized in that, include: Communication establishment module: The user terminal accesses the Web control interface through the local area network and establishes a two-way communication connection with the audio device; The first interface receiving module is used to receive the model information of the currently used headphones input or selected by the user based on the Web control interface; The second interface receiving module is used to receive listening preference description information input by the user in the form of natural language text based on the Web control interface; The sending module is used to send the model information and the listening preference description information to the artificial intelligence tuning module; The acoustic analysis module is used by the artificial intelligence tuning module to retrieve the measured frequency response data corresponding to the model information from the pre-stored headphone frequency response database, and perform multi-dimensional acoustic analysis by combining the pre-stored target acoustic reference curve, preset timbre model and the listening preference description information to generate parametric equalizer adjustment parameters including at least one set of center frequency, gain value and quality factor Q value. The parameter conversion module is used to convert the parameter equalizer adjustment parameters into DSP filter coefficients that can be recognized by the audio device's DSP processing module through the DSP parameter conversion module. A coefficient loading module is used to load the DSP filter coefficients into the DSP processing module of the audio device, wherein the DSP processing module performs equalization filtering on the real-time input audio signal according to the DSP filter coefficients and outputs it to the headphones; The continuing execution module is used to receive a save command sent by the user through the Web control interface, store the parametric equalizer adjustment parameters in the internal memory of the audio device, and generate a corresponding entry in the local EQ preset list of the audio device. This allows the audio device to retrieve the parametric equalizer adjustment parameters from the internal memory and convert them into DSP filter coefficients to perform audio equalization processing when disconnected from the network.