Audio equalization adjustment method and device, and storage medium

By acquiring and processing physiological motion and environmental audio signals, a target equalization parameter set is generated to intelligently adjust the audio of near-eye display devices, solving the problem of complex acoustic changes in motion scenarios and improving the stability and comfort of the auditory experience and audio output.

CN121985252APending Publication Date: 2026-05-05ZHUHAI MOJIE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHUHAI MOJIE TECH CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing near-eye display devices cannot effectively cope with complex acoustic changes in motion scenarios, resulting in a poor auditory experience and an inability to integrate user motion status and environmental sound field information for multi-dimensional and intelligent dynamic audio equalization adjustment.

Method used

By acquiring the user's physiological motion signals and environmental audio signals, processing them to generate physiological motion features and environmental noise features, making fusion decisions based on these features, generating a target equalization parameter set, and adjusting the audio output signal of the near-eye display device, including intelligent adjustment of frequency response gain, channel balance, and overall volume gain.

Benefits of technology

It achieves dual-dimensional perception and decision-making based on user physiological movement and environmental noise, improves the auditory experience of near-eye display devices in motion scenarios, and provides multi-dimensional intelligent dynamic audio equalization adjustment to ensure the stability and comfort of audio output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121985252A_ABST
    Figure CN121985252A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio signal processing, in particular to an audio equalization adjustment method and device and a storage medium. The audio equalization adjustment method comprises the following steps: acquiring a physiological movement signal and an environment audio signal of a user; respectively processing the physiological motion signal and the environment audio signal to obtain physiological motion features and environment noise features; performing fusion decision based on the physiological motion features and the environmental noise features to generate a target equilibrium parameter set; and based on the target equalization parameter set, performing equalization adjustment on the audio output signal of the near-to-eye display device. According to the method, the auditory experience of the near-to-eye display equipment in the motion scene can be intelligently optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio signal processing technology, and in particular to an audio equalization adjustment method, device and storage medium. Background Technology

[0002] With the increasing popularity of near-eye display devices such as Augmented Reality (AR) and Virtual Reality (VR), their applications in dynamic scenarios such as commuting, sports, and outdoor exploration are becoming more and more widespread. To balance wearing comfort and environmental awareness safety, these devices often use open or semi-open speakers for audio output.

[0003] In related technologies, most near-eye display devices use fixed equalizers (EQ), and only perform fixed wind noise suppression and simple automatic gain control based on the intensity of ambient noise, which cannot cope with the complex acoustic changes in motion scenarios.

[0004] Therefore, there is an urgent need to provide a solution that can integrate user motion status and ambient sound field information to perform multi-dimensional, intelligent, and dynamic audio equalization adjustment, so as to comprehensively improve the auditory experience of near-eye display devices in motion scenarios. Summary of the Invention

[0005] The embodiments of this application aim to provide an audio equalization adjustment method, device, and storage medium that can intelligently optimize the auditory experience of near-eye display devices in motion scenarios.

[0006] To address the aforementioned technical problems, this application provides the following technical solutions: In a first aspect, embodiments of this application provide an audio equalization adjustment method applied to a near-eye display device, the method comprising: Acquire the user's physiological motion signals and environmental audio signals; The physiological motion signal and the environmental audio signal are processed separately to obtain physiological motion characteristics and environmental noise characteristics; Based on the physiological motion characteristics and the environmental noise characteristics, a fusion decision is made to generate a target equilibrium parameter set; Based on the target equalization parameter set, the audio output signal of the near-eye display device is equalized and adjusted.

[0007] Optionally, the step of generating a target equilibrium parameter set by fusing the physiological motion characteristics and the environmental noise characteristics includes: When the physiological motion feature or the environmental noise feature switches from a normal state to a higher-order state, and remains in the higher-order state for at least a preset first switching duration, a first target equalization parameter set corresponding to the higher-order state is generated. When the physiological motion feature or the environmental noise feature switches back from the higher-order state to the normal state, and remains in the normal state for at least a preset second switching duration, a second target equalization parameter set corresponding to the higher-order state is generated, wherein the second switching duration is longer than the first switching duration.

[0008] Optionally, the near-eye display device integrates a dual-microphone array, and the environmental noise characteristics also include overall environmental noise energy and wind noise direction. The environmental audio signal is processed to obtain environmental noise characteristics including: The ambient audio signals collected by the dual microphone array on the left and right sides of the near-eye display device are processed using a noise power estimation algorithm to obtain the ambient noise energy of the left and right channels and the ambient noise energy of the right channel. The average value of the ambient noise energy in the left channel and the ambient noise energy in the right channel is calculated to obtain the overall ambient noise energy. The difference between the ambient noise energy of the left channel and the ambient noise energy of the right channel is calculated, and the difference is compared with a preset noise difference threshold to obtain the wind noise direction.

[0009] Optionally, the step of generating a target equilibrium parameter set by fusing the physiological motion characteristics and the environmental noise characteristics includes: The physiological motion characteristics and the environmental noise characteristics are used as input characteristics and matched with preset frequency response gain rule sets, channel balance rule sets and overall volume gain rule sets respectively to determine the matched frequency response gain rule, channel balance rule and overall volume gain rule. Based on the matched frequency response gain rules, channel balance rules, and overall volume gain rules, corresponding frequency response gain parameters, channel balance parameters, and overall volume gain parameters are generated to form a target equalization parameter set.

[0010] Optionally, the frequency response gain rule set includes a first frequency response gain rule and a second frequency response gain rule. The first frequency response gain rule is: when the duration of the heart rate exceeding a preset high heart rate threshold is greater than a first preset duration, a first frequency response gain parameter is generated to enhance the gain of a preset mid-to-high frequency band. The second frequency response gain rule is: when the duration of the step frequency exceeding a preset high step frequency threshold is greater than a second preset duration, a second frequency response gain parameter is generated to enhance the dynamic range of a preset low frequency band.

[0011] Optionally, the overall volume gain rule set includes the following overall volume gain rule: when the duration of the overall ambient noise energy exceeding the preset high noise threshold is greater than the third preset duration, an overall volume gain parameter for increasing the overall output volume is generated. The channel balance rule set includes the following channel balance rules: when the absolute value of the difference between the ambient noise energy of the left channel and the ambient noise energy of the right channel is greater than the duration of a preset noise difference threshold greater than a fourth preset duration, and the user's head posture is a non-turned head posture, channel balance parameters are generated to improve the gain of the speaker on the opposite side of the wind noise direction.

[0012] Optionally, the step of generating a target equilibrium parameter set by fusing the physiological motion characteristics and the environmental noise characteristics includes: Based on the aforementioned physiological movement characteristics, the user's current movement state is determined through a preset movement scene recognition model; Based on the current motion state and the environmental noise characteristics, a fusion decision is made to generate a target equalization parameter set.

[0013] Optionally, the step of equalizing and adjusting the audio output signal of the near-eye display device based on the target equalization parameter set includes: A parameter smoothing algorithm is applied to at least some parameters in the target equalization parameter set to obtain a parameter set sequence in which the at least some parameters gradually change within a fifth preset time period. Based on the parameter set sequence, the audio output signal of the near-eye display device is equalized and adjusted within the fifth preset time period.

[0014] Secondly, embodiments of this application provide a near-eye display device, including at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to perform any of the methods described above.

[0015] Thirdly, embodiments of this application provide a computer storage medium storing instructions or programs that, when executed by at least one processor, cause the at least one processor to perform any of the methods described above.

[0016] The beneficial effects of this application's embodiments are as follows: This application provides an audio equalization adjustment method. First, it acquires the user's physiological motion signal and the environmental audio signal. Then, it processes the physiological motion signal and the environmental audio signal separately to obtain physiological motion characteristics and environmental noise characteristics. Next, it performs a fusion decision based on the physiological motion characteristics and environmental noise characteristics to generate a target equalization parameter set. Finally, it equalizes the audio output signal of the near-eye display device based on the target equalization parameter set. This method upgrades the traditional adjustment mode, which only focuses on the external acoustic environment, to an intelligent adjustment system that simultaneously considers the user's internal physiological motion state and the external acoustic environment. It constructs a complete "user-environment" dual-dimensional perception and decision-making system, realizing multi-dimensional, intelligent, dynamic audio equalization adjustment based on motion scenarios, thus improving the auditory experience of near-eye display devices in motion scenarios. Attached Figure Description

[0017] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0018] Figure 1 An exemplary flowchart of an audio equalization adjustment method is shown; Figure 2 An example is shown in the logical architecture diagram of dynamic equilibrium decision-making and parameter generation; Figure 3 An exemplary schematic diagram illustrates a process for generating a target equalization parameter set based on motion state and environmental noise characteristics; Figure 4 An exemplary schematic diagram of the hardware structure of a near-eye display device is shown. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0020] Furthermore, the technical features involved in the various embodiments of this application described below can be combined with each other as long as they do not conflict with each other.

[0021] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0022] Please refer to Figure 1 , Figure 1 An audio equalization adjustment method is illustrated, which is applied to near-eye display devices such as AR glasses, VR glasses, AR helmets, and VR helmets. This method is implemented by the processor built into the near-eye display device executing program instructions stored in the storage medium. Figure 1 As shown, the method includes the following steps: Step S101: Acquire the user's physiological motion signals and environmental audio signals.

[0023] In one embodiment, the near-eye display device integrates a photoelectric heart rate sensor, an inertial measurement unit, and a dual-microphone array, among other multimodal sensors, to acquire the user's physiological motion signals and environmental audio signals. For example, the photoelectric heart rate sensor is located on the inside of the temples or helmet and in contact with the skin to continuously collect heart rate signals; the inertial measurement unit is located on the front frame of the near-eye display device to continuously collect inertial signals such as three-axis acceleration data and three-axis angular velocity data; and the dual-microphone array is located on the left and right temples or the left and right sides of the helmet to continuously collect environmental audio signals from the left and right sides of the near-eye display device.

[0024] Step S102: Process the physiological motion signal and the environmental audio signal respectively to obtain physiological motion characteristics and environmental noise characteristics.

[0025] In one embodiment, physiological motion signals are processed to obtain physiological motion characteristics such as heart rate, cadence, and head posture.

[0026] The method for obtaining the heart rate is as follows: Kalman filtering is applied to the heart rate signal collected by the photoelectric heart rate sensor. By utilizing its optimal estimation characteristics, motion artifacts caused by head shaking are effectively predicted and filtered out, and a stable and reliable heart rate is output.

[0027] The method for obtaining gait frequency is as follows: The triaxial acceleration data collected by the inertial measurement unit is processed using a gait frequency detection algorithm to obtain the gait frequency. For example, the vertical component or vector amplitude is selected from the triaxial acceleration data. Then, a Butterworth bandpass filter with a passband of 0.7Hz–2.5Hz (corresponding to 42–150 steps / minute) is used to filter out high-frequency noise above 2.5Hz (corresponding to sensor electronic noise or other non-gait rapid jitter) and low-frequency noise below 0.7Hz (corresponding to slow body tilting or posture adjustment). After filtering, a filtered signal covering the typical human gait frequency range from walking to running is obtained. Further, peak detection is performed on the filtered signal, and effective step points are selected by setting amplitude thresholds and minimum time interval thresholds. The time difference between adjacent effective step points is then calculated to obtain the gait cycle. Finally, the gait frequency is calculated based on the gait cycle.

[0028] The head attitude is obtained by processing the angular velocity and acceleration data collected by the inertial measurement unit to obtain the head attitude. For example, the angular velocity data is first integrated over time to obtain an attitude change estimate including accumulated error; then the acceleration data is processed to extract the reference attitude representing the direction of gravity under static conditions; then a complementary filtering algorithm is used to fuse the high-frequency part of the attitude change estimate with the low-frequency part of the reference attitude to output accurate head attitude data.

[0029] In one embodiment, the ambient audio signal is processed to obtain ambient noise characteristics such as the ambient noise energy of the left channel, the ambient noise energy of the right channel, the overall ambient noise energy, and the direction of wind noise. For example, the ambient audio signals collected by the dual-microphone array on the left and right sides of the near-eye display device are first framed and windowed, and then converted into frequency domain signals. Based on these frequency domain signals, the estimated values ​​of the ambient noise energy of the left channel and the right channel are estimated using a noise estimation method based on minimum statistics or a noise power estimation method based on spectral flatness measurement, respectively. Then, the estimated values ​​of the ambient noise energy of the left channel and the right channel are smoothed and filtered within a preset time window (e.g., 50 milliseconds to 200 milliseconds) to obtain the smoothed ambient noise energy of the left channel. and right channel ambient noise energy .

[0030] Furthermore, the ambient noise energy of the left channel is calculated. and right channel ambient noise energy The average value is used to obtain the overall environmental noise energy. Calculate the ambient noise energy in the left channel. and right channel ambient noise energy The difference .like If so, the wind noise is determined to be coming from the left; if If so, the wind noise is determined to be coming from the right; if If so, it can be determined that the wind noise direction has not shifted. This is a preset noise difference threshold, which is greater than zero, for example, 6dB.

[0031] Step S103: Based on physiological motion characteristics and environmental noise characteristics, a fusion decision is made to generate a target equilibrium parameter set.

[0032] In one embodiment, the target equalization parameter set includes at least one of frequency response gain parameters, channel balance parameters, and overall volume gain parameters. The frequency response gain parameters include at least one of a first frequency response gain parameter for boosting or reducing the volume of a certain frequency band and a second frequency response gain parameter for enhancing or weakening the dynamic range of a certain frequency band. The channel balance parameters are used to control the volume of the left and right channels and include parameters with two dimensions: left-side gain and right-side gain. The overall volume gain parameters are used to control the volume of the overall output audio and include parameters with one dimension: volume gain.

[0033] In one embodiment, the first frequency response gain parameter includes parameters in three dimensions: center frequency, frequency response gain, bandwidth, or Q value. The Q value is a parameter used to control the "width" or "narrowness" of the frequency range affected by the audio equalizer; it, along with the center frequency, determines the target frequency band for frequency response gain adjustment. The second frequency response gain parameter includes parameters in seven dimensions: center frequency, bandwidth or Q value, level threshold, compression ratio, attack time, release time, and overall gain compensation. The level threshold is a preset level limit value in the dynamic range compressor, used to determine when the dynamic range compressor is triggered. The compression ratio refers to the ratio between the change in input signal level after exceeding the level threshold and the corresponding change in output signal level in the dynamic range compressor, used to determine the degree to which the dynamic range is reduced. The attack time refers to the time required for the dynamic range compressor's gain reduction to reach its full target value calculated based on the compression ratio after the input signal level exceeds the level threshold, used to control the dynamic range compressor's response speed to transient components of the signal. Release time refers to the time required for the gain reduction of a dynamic range compressor to return to zero after the input signal level falls below the threshold level from a state exceeding the threshold. Overall gain compensation refers to the fixed gain boost applied to the output signal after the dynamic range compressor has processed the input signal. This boost is used to compensate for the overall level drop caused by dynamic compression processing and to ensure that the processed signal reaches the target loudness.

[0034] This step can be implemented using either a rule-based expert system for generating target equilibrium parameter sets, or a trained lightweight machine learning model for generating target equilibrium parameter sets. When using a rule-based expert system for generating target equilibrium parameter sets, as shown below... Figure 2 As shown, the expert system can include a three-layer architecture: a feature input layer, a fusion decision layer, and a parameter output layer. The feature input layer acquires various input features, including physiological motion features such as heart rate, cadence, and head posture, and environmental noise features such as left channel environmental noise energy, right channel environmental noise energy, overall environmental noise energy, and wind noise direction. The fusion decision layer matches the input features with preset target equalization rule sets of various types to determine a subset of matching target equalization rules. The parameter output layer generates corresponding target equalization parameters based on each target equalization rule in the subset of matching target equalization rules, and combines these target equalization parameters into a target equalization parameter set.

[0035] In one embodiment, the target equalization rules include frequency response gain rules, channel balance rules, and overall volume gain rules. The target equalization rule set includes a frequency response gain rule set consisting of at least one frequency response gain rule, a channel balance rule set consisting of at least one channel balance rule, and an overall volume gain rule set consisting of at least one overall volume gain rule. The matched subset of target equalization rules includes a frequency response gain rule matched from the frequency response gain rule set, a channel balance rule matched from the channel balance rule set, and an overall volume gain rule matched from the overall volume gain rule set.

[0036] In one embodiment, the frequency response gain rule set may include the following first frequency response gain rule: when the duration of the heart rate exceeding a preset high heart rate threshold is greater than a first preset duration, a first frequency response gain parameter is generated to enhance the gain of a preset mid-to-high frequency band. For example, audio is divided into seven frequency bands according to frequency range: sub-low frequency (20Hz-60Hz), low frequency (60Hz-250Hz), mid-low frequency (250Hz-500Hz), mid frequency (500Hz-2kHz), mid-high frequency (2kHz-8kHz), high frequency (8kHz-16kHz), and ultra-high frequency (16kHz-20kHz+). The sub-low frequency is the frequency band containing the lowest bass note of a pipe organ, seismic sounds, thunder, and ultra-low bass in electronic music. The low frequency is the frequency band containing the impact of a bass drum, the fundamental tone of a bass, and the fullness of a cello and bass brass, representing fundamental tones and rhythms. The low-mid frequency range encompasses the chest resonance of the human voice, the mid-low range of pianos and guitars, and the fullness and muddiness of saxophones. The mid-frequency range contains the fundamental tone and core harmonics of the human voice and most instruments; for example, telephone sounds (300Hz-3kHz) fall within this range. This frequency range is crucial for clarity and intelligibility. The mid-high frequency range reflects the brightness, clarity, and impact of sound, and is also the core frequency range for the highlights of human voices and instruments. The high frequency range reflects the airiness and spatiality of sound, such as high-frequency harmonics, breath sounds, the hissing sound of cymbals, and spatial reflections from the recording environment. The ultra-high frequency range represents the limits of human hearing, including the overtones of some instruments and the high frequencies of electronic synthesizers.

[0037] For example, when a heart rate exceeding 120 beats per minute lasts for more than 2 seconds, it's assumed the user is engaging in moderate to high-intensity exercise. This typically involves rapid breathing, body tremors, and the presence of footsteps and airflow in the environment. These noises can mask sound details in the mid-to-high frequency range (2kHz-8kHz) of the audio (such as voice navigation commands and musical melodies). Therefore, it's necessary to increase the volume of this mid-to-high frequency range by a certain decibel (e.g., 3dB) to counteract the masking effect of motion noise. This allows users to clearly hear the key details of navigation voices, coach instructions, or music without manually increasing the volume, preventing repeated operations due to poor hearing and allowing them to focus on their exercise.

[0038] In one embodiment, the frequency response gain rule set may further include a second frequency response gain rule as follows: when the duration of the step frequency exceeding a preset high step frequency threshold is greater than a second preset duration, a second frequency response gain parameter is generated to enhance the dynamic range of the preset low frequency band. For example, when the duration of the step frequency exceeding 180 steps / minute is greater than 3 seconds, it is inferred that the user is engaged in fast-paced running or rope skipping, and the user needs a strong sense of rhythm to match the movement frequency. At this time, the low frequencies of ordinary audio will seem weak and unable to drive the movement. 60Hz-200Hz is the core of low frequencies in music (such as drum beats and bass). Enhancing the dynamic range of this frequency band can make the "contrast between strong and weak low frequencies more obvious"—drum beats are more impactful, bass is more elastic, which perfectly matches the rhythm of fast step frequency. For example, enhancing the dynamic range of this frequency band can be achieved by setting a dynamic range adjustment parameter of low level threshold + low compression ratio + fast attack time + slow release time for this frequency band. The specific implementation process can be referred to in related technologies, and will not be elaborated here.

[0039] In one embodiment, the overall volume gain rule set may include the following overall volume gain rule: when the duration of the overall ambient noise energy exceeding a preset high noise threshold is greater than a third preset duration, an overall volume gain parameter for increasing the overall output volume is generated. For example, when the overall ambient noise energy continuously exceeds the high noise threshold, a corresponding overall volume gain parameter is generated based on the noise level corresponding to the overall ambient noise energy. For instance, when the duration of the detected overall ambient noise energy exceeding 65dB is greater than 2 seconds, if the noise is at a moderate level (e.g., 65~75dB), the overall volume gain parameter is set to 3dB; if the noise is high (e.g., >75dB), the overall volume gain parameter is set to 5dB.

[0040] In one embodiment, the channel balance rule set may include the following channel balance rules: when the absolute value of the difference between the ambient noise energy of the left channel and the ambient noise energy of the right channel is greater than a preset noise difference threshold for a duration greater than a fourth preset duration, and the user's head posture is not a head-turning posture, channel balance parameters for increasing the gain of the speaker on the opposite side in the wind noise direction are generated. For example, when the difference between the ambient noise energy of the left channel and the ambient noise energy of the right channel is greater than 6dB for a duration greater than 2 seconds and the user's head posture is not a head-turning posture, the gain of the right speaker is increased by 1~3dB; when the difference between the ambient noise energy of the left channel and the ambient noise energy of the right channel is less than -6dB for a duration greater than 2 seconds and the user's head posture is not a head-turning posture, the gain of the left speaker is increased by 1~3dB; if the difference between the ambient noise energy of the left channel and the ambient noise energy of the right channel is within ±6dB, the left and right balance is maintained to avoid excessive offset. This strategy can maintain a centered overall listening experience and natural voice positioning even when there is a strong crosswind, a passing vehicle, or the user's head is turned to the side, resulting in a significant increase in noise on one side.

[0041] Step S104: Based on the target equalization parameter set, the audio output signal of the near-eye display device is equalized and adjusted.

[0042] To avoid frequent equalization adjustments to the audio output signal of the near-eye display device due to brief fluctuations in parameters such as heart rate, cadence, and wind noise, a time threshold hysteresis and parameter smoothing mechanism is introduced in one embodiment. Specifically, time threshold hysteresis refers to the following steps during step S103: when the physiological motion feature or environmental noise feature switches from a normal state to a higher-order state and remains in that higher-order state for at least a preset first switching duration, a first target equalization parameter set corresponding to that higher-order state is generated; when the physiological motion feature or environmental noise feature switches back from the higher-order state to the normal state and remains in that normal state for at least a preset second switching duration, a second target equalization parameter set is generated to remove the higher-order state. Here, a normal state refers to a state with low motion intensity or low wind noise, a higher-order state refers to a state with high motion intensity or high wind noise, and the second switching duration is longer than the first switching duration. For example, mid-to-high frequency enhancement is triggered only when the heart rate exceeds 120 beats / minute (high-intensity exercise) and lasts for ≥2 seconds, and the mid-to-high frequency enhancement mode is deactivated only when the heart rate drops to <110 beats / minute (low-intensity exercise) and remains there for ≥3 seconds. This mechanism effectively avoids repeated audio equalization adjustments caused by fluctuations at the edge of the exercise state. It is understandable that different higher-order states correspond to different first and second switching durations.

[0043] The parameter smoothing mechanism refers to applying a parameter smoothing algorithm to at least a portion of the parameters in the target equalization parameter set to obtain a sequence of parameters that gradually change over a fifth preset time period. Based on this parameter sequence, the audio output signal of the near-eye display device is then equalized within the fifth preset time period. For example, first-order low-pass or exponential smoothing (e.g., a time constant of 100–300ms) is applied to frequency response gain parameters (such as frequency response gain, center frequency, bandwidth, or Q value) to ensure a gradual change during audio equalization adjustment based on the frequency response gain parameter, guaranteeing a continuous and natural listening experience. This strategy improves the rhythmic stability of audio equalization adjustment, avoiding repeated audio equalization adjustments caused by short-term noise, motion jitter, or instantaneous heart rate fluctuations, making dynamic audio equalization smoother, more reliable, and more consistent with human hearing characteristics.

[0044] Furthermore, when boosting the overall output volume based on the overall volume gain parameter, an automatic gain control strategy with fast attack and slow release is employed to ensure continuity and abruptness. For example, if the duration of detected overall ambient noise energy exceeding the high noise threshold is longer than 2 seconds, the overall output volume is boosted within 200ms; if the duration of detected overall ambient noise energy not exceeding the high noise threshold is longer than 3 seconds, the output volume returns to normal within a longer period of 500–800ms.

[0045] To better suit motion scenarios, this application also provides a method for audio equalization adjustment based on motion state and environmental noise characteristics. Please refer to... Figure 3 , Figure 3 A method for generating a target equalization parameter set based on motion state and environmental noise characteristics is shown. One embodiment of step S103 of the method includes the following steps: Step S1031: Based on physiological movement characteristics, determine the user's current movement state through a preset movement scene recognition model.

[0046] For example, physiological motion features such as heart rate, cadence, and head posture, obtained after processing physiological motion signals, are input into a preset motion scene recognition model to determine the user's current motion state. In one embodiment, the motion scene recognition model can be a rule-based motion scene recognition expert system or a trained lightweight machine learning model for recognizing motion scenes. For example, a rule-based motion scene recognition classifier can be constructed as follows: Input parameters: Heart rate (HR), cadence (CAD), head posture (vertical acceleration features, head pitch features). Output: Motion state (stationary, walking, running, cycling, climbing stairs, descending stairs) Core recognition rules (priority from high to low): 1) Rule 0: State Preservation and Hysteresis First, when switching between any activity states, the new activity state must be maintained for at least a preset first switching duration (e.g., 2.0 seconds). Second, when switching back from a higher-order activity state to a lower-order activity state, the lower-order activity state must be maintained for at least a preset second switching duration (e.g., 3.0 seconds), and the second switching duration must be longer than the first switching duration. For example, stationary and walking can be set as lower-order activity states, while running, cycling, climbing stairs, and descending stairs can be set as higher-order activity states.

[0047] 2) Rule 1: Identify "Stillness" If the cadence is less than the minimum cadence threshold and the heart rate is less than the minimum heart rate threshold (e.g., CAD < 70 and HR < 100), then the current exercise state is determined to be a stationary state.

[0048] 3) Rule 2: Recognize "going up and down stairs" If the minimum step frequency threshold is less than or equal to the step frequency threshold (e.g., 70 ≤ CAD ≤ 90) and the vertical acceleration feature is the stair type and the head pitch feature is oscillation, then the current motion state is determined to be going up or down stairs.

[0049] If further determination of the direction of going up or down stairs is needed, a barometer can be used. If the air pressure continues to drop slowly, it indicates going up stairs; if it continues to rise, it indicates going down stairs.

[0050] 4) Rule 3: Recognize "cycling" If the cadence is less than the minimum cadence threshold and the heart rate is greater than or equal to the running heart rate threshold (e.g., CAD < 70 and HR ≥ 130), then the current exercise state is determined to be cycling.

[0051] 5) Rule 4: Recognize "Running" If the cadence is greater than or equal to the running cadence threshold and the heart rate is greater than or equal to the running heart rate threshold (e.g., CAD ≥ 140 and HR ≥ 130), then the current exercise state is determined to be running.

[0052] 6) Rule 5: Recognize “walking” If the minimum cadence threshold is less than or equal to the cadence threshold and the minimum heart rate threshold is less than or equal to the heart rate threshold and the running heart rate threshold (e.g., 70 <= CAD <= 130 and 100 <= HR < 130), then the current exercise state is determined to be walking.

[0053] In addition, the current motion state is set to walking by default. If the conditions of all the more specific states mentioned above are not met, the current motion state can be classified as walking.

[0054] It is understood that, in addition to the aforementioned states of exercise, other states of exercise may also be included, such as rope skipping, aerobics, etc. Those skilled in the art can set corresponding rules according to their needs.

[0055] Step S1032: Based on the current motion state and environmental noise characteristics, a fusion decision is made to generate a target equilibrium parameter set.

[0056] In one embodiment, using and Figure 2 A similar rule-based expert system for generating target equilibrium parameter sets is used to make fusion decisions based on the current motion state and environmental noise characteristics, and to generate target equilibrium parameter sets. Figure 2Unlike other embodiments, in this embodiment, each target equalization rule in the target equalization rule set is set based on the current motion state and environmental noise characteristics. For example, the target equalization rule set in this embodiment may include the following target equalization rule: if the current motion state is running and the overall environmental noise energy continuously exceeds a preset high noise threshold, a first frequency response gain parameter for increasing the preset mid-to-high frequency band gain, a second frequency response gain parameter for enhancing the preset low-frequency band dynamic range, and an overall volume gain parameter for increasing the overall output volume are generated.

[0057] In summary, the audio equalization adjustment method provided in this application has the following beneficial effects: 1) Achieved a leap from "environmental perception" to "user state perception": By processing physiological motion signals, multi-dimensional physiological motion characteristics such as heart rate, cadence, and head posture are obtained, enabling the system to accurately understand user behavior (such as high-intensity running and leisure cycling), thereby upgrading the audio adjustment strategy from passively responding to the environment to actively adapting to the user's state, achieving true personalized audio optimization. 2) A sophisticated multi-dimensional environmental noise countermeasure system has been constructed: it not only intelligently increases the overall volume to counteract general noise, but also innovatively uses a dual-microphone array to analyze the spatial distribution of noise and accurately compensates the masked channels through dynamic channel balancing technology, thereby maintaining stable sound field positioning and natural listening balance in noisy and asymmetrical acoustic environments. 3) Ensuring the accuracy and reliability of regulation through rule fusion decision-making: By matching complex multimodal features with a pre-set set of rules that have been validated acoustically and physiologically, exemplary audio processing parameters are generated. This method is transparent, controllable, and computationally efficient, ensuring the accuracy of regulation decisions and the real-time nature of system response. 4) Ensures an extremely smooth and seamless transition experience: By introducing a parameter smoothing algorithm, the audio parameters can change gradually with a smooth curve when switching states, completely eliminating the abruptness of the sound caused by parameter jumps, and ensuring the continuity and comfort of the user experience.

[0058] According to an embodiment of this application, a near-eye display device is provided, such as... Figure 4 The diagram shown is a hardware structure schematic of a near-eye display device according to an embodiment of this application. The near-eye display device 100 includes a processor 10, a memory 20, and a communication interface 30. The processor 10, memory 20, and communication interface 30 are connected via lines. Figure 4 In the embodiment shown, the processor 10, memory 20, and communication interface 30 are connected to each other via a bus.

[0059] The memory 20 is used to store software programs, computer-executable program instructions, etc. The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the device 100, etc.

[0060] The memory 20 can be a read-only memory (ROM), or other types of static storage devices that can store static information and instructions, or random access memory (RAM), or other types of dynamic storage devices that can store information and instructions, or electrically erasable programmable read-only memory (EEPROM). The specific type is not limited here.

[0061] For example, the aforementioned memory 20 can be Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM). This memory 20 can exist independently but is connected to the processor 10. Optionally, the memory 20 can also be integrated with the processor 10, for example, integrated within one or more chips.

[0062] In some embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and this remote memory may be connected to the device 100 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0063] The processor 10 connects various parts of the entire device 100 using various interfaces and lines. By running or executing software programs stored in the memory 20 and calling data stored in the memory 20, it performs various functions of the device 100 and processes data, such as implementing the methods described in any embodiment of this application.

[0064] The processor 10 can be a field programmable gate array (FPGA), a digital signal processor (DSP), a central processing unit (CPU), or the like.

[0065] Processor 10 can be a single-core processor or a multi-core processor. For example, processor 10 can be composed of multiple FPGAs or multiple DSPs. Furthermore, processor 10 can refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). Processor 10 can be a standalone semiconductor chip or integrated with other circuits into a single semiconductor chip. For example, it can form a system-on-a-chip (SoC) with other circuits (such as encoding / decoding circuits, hardware acceleration circuits, or various bus and interface circuits), or it can be integrated as a built-in processor within an application-specific integrated circuit (ASIC). This ASIC with integrated processor can be packaged separately or together with other circuits.

[0066] The communication interface 30 can use a transceiver device, such as a transceiver, to enable communication between the device 100 and other devices or communication networks.

[0067] This application also provides a computer storage medium storing instructions or programs that are executed by one or more processors, for example... Figure 4 One of the processors 10 may enable the one or more processors to perform the methods in any of the above method embodiments.

[0068] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general-purpose hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.

[0069] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of variations or substitutions within the technical scope disclosed in this application. Therefore, any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. An audio equalization adjustment method, characterized in that, Applied to near-eye display devices, the method includes: Acquire the user's physiological motion signals and environmental audio signals; The physiological motion signal and the environmental audio signal are processed separately to obtain physiological motion characteristics and environmental noise characteristics; Based on the physiological motion characteristics and the environmental noise characteristics, a fusion decision is made to generate a target equilibrium parameter set; Based on the target equalization parameter set, the audio output signal of the near-eye display device is equalized and adjusted.

2. The method according to claim 1, characterized in that, The step of generating a target equilibrium parameter set by fusing the physiological motion characteristics and the environmental noise characteristics includes: When the physiological motion feature or the environmental noise feature switches from a normal state to a higher-order state, and remains in the higher-order state for at least a preset first switching duration, a first target equalization parameter set corresponding to the higher-order state is generated. When the physiological motion feature or the environmental noise feature switches back from the higher-order state to the normal state, and remains in the normal state for at least a preset second switching duration, a second target equalization parameter set corresponding to the higher-order state is generated, wherein the second switching duration is longer than the first switching duration.

3. The method according to claim 1, characterized in that, The near-eye display device integrates a dual-microphone array. The environmental noise characteristics also include overall environmental noise energy and wind noise direction. The environmental audio signal is processed to obtain the following environmental noise characteristics: The ambient audio signals collected by the dual-microphone array on the left and right sides of the near-eye display device are processed using a noise power estimation algorithm to obtain the ambient noise energy of the left channel and the ambient noise energy of the right channel. The average value of the ambient noise energy in the left channel and the ambient noise energy in the right channel is calculated to obtain the overall ambient noise energy. The difference between the ambient noise energy of the left channel and the ambient noise energy of the right channel is calculated, and the difference is compared with a preset noise difference threshold to obtain the wind noise direction.

4. The method according to claim 1, characterized in that, The step of generating a target equilibrium parameter set by fusing the physiological motion characteristics and the environmental noise characteristics includes: The physiological motion characteristics and the environmental noise characteristics are used as input characteristics and matched with preset frequency response gain rule sets, channel balance rule sets and overall volume gain rule sets respectively to determine the matched frequency response gain rule, channel balance rule and overall volume gain rule. Based on the matched frequency response gain rules, channel balance rules, and overall volume gain rules, corresponding frequency response gain parameters, channel balance parameters, and overall volume gain parameters are generated to form a target equalization parameter set.

5. The method according to claim 4, characterized in that, The frequency response gain rule set includes a first frequency response gain rule and a second frequency response gain rule. The first frequency response gain rule is: when the duration of the heart rate exceeding a preset high heart rate threshold is greater than a first preset duration, a first frequency response gain parameter is generated to enhance the gain of a preset mid-to-high frequency band. The second frequency response gain rule is: when the duration of the step frequency exceeding a preset high step frequency threshold is greater than a second preset duration, a second frequency response gain parameter is generated to enhance the dynamic range of a preset low frequency band.

6. The method according to claim 4, characterized in that, The overall volume gain rule set includes the following overall volume gain rule: when the duration of the overall ambient noise energy exceeding the preset high noise threshold is greater than the third preset duration, an overall volume gain parameter for improving the overall output volume is generated. The channel balance rule set includes the following channel balance rules: when the absolute value of the difference between the ambient noise energy of the left channel and the ambient noise energy of the right channel is greater than the duration of a preset noise difference threshold greater than a fourth preset duration, and the user's head posture is a non-turned head posture, channel balance parameters are generated to improve the gain of the speaker on the opposite side of the wind noise direction.

7. The method according to claim 1, characterized in that, The step of generating a target equilibrium parameter set by fusing the physiological motion characteristics and the environmental noise characteristics includes: Based on the aforementioned physiological movement characteristics, the user's current movement state is determined through a preset movement scene recognition model; Based on the current motion state and the environmental noise characteristics, a fusion decision is made to generate a target equalization parameter set.

8. The method according to any one of claims 1 to 7, characterized in that, The process of equalizing and adjusting the audio output signal of the near-eye display device based on the target equalization parameter set includes: A parameter smoothing algorithm is applied to at least some parameters in the target equalization parameter set to obtain a parameter set sequence in which the at least some parameters gradually change within a fifth preset time period. Based on the parameter set sequence, the audio output signal of the near-eye display device is equalized and adjusted within the fifth preset time period.

9. A near-eye display device, characterized in that, The method includes at least one processor and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor to enable the at least one processor to perform the method as described in any one of claims 1 to 8.

10. A computer storage medium, characterized in that, The computer storage medium stores instructions or programs that, when executed by at least one processor, cause the at least one processor to perform the method as described in any one of claims 1 to 8.