Volume adjustment method and device, head-mounted audio device, and storage medium

CN122554760APending Publication Date: 2026-08-11ZHUHAI MOJIE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-09
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本申请提供了一种音量调节方法及头戴音频设备,旨在解决相关设备在近场主动语音交互场景下音量调节的场景自适应性低下的技术问题

Benefits of technology

[0010] This application provides a volume adjustment method, apparatus, head-mounted audio device, and storage medium. The method acquires multimodal near-field perception data through a multimodal near-field perception component, fuses multimodal perception information, and achieves multi-dimensional perception of the near-field environment of the user, thereby improving the perception accuracy of near-field active voice interaction events. When the multimodal near-field perception data meets the triggering conditions for near-field active voice interaction events for the user, at least one audio collaborative adjustment operation, including media volume adjustment and dialogue voice adjustment, is executed to achieve dynamic balance adjustment of media volume and dialogue voice. This allows the volume adjustment to automatically adapt to interaction needs, thereby improving the scene adaptability of the head-mounted audio device's volume adjustment in near-field active voice interaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122554760A_ABST
    Figure CN122554760A_ABST
Patent Text Reader

Abstract

This application relates to the field of speech processing technology, and provides a volume adjustment method, device, head-mounted audio device, and storage medium. The method collects multimodal near-field perception data through a multimodal near-field perception component, fuses multimodal perception information, and achieves multi-dimensional perception of the near-field environment of the wearer, thereby improving the perception accuracy of near-field active voice interaction events. When the multimodal near-field perception data meets the triggering conditions for near-field active voice interaction events for the wearer, at least one audio collaborative adjustment operation, including media volume adjustment and dialogue voice adjustment, is executed to achieve dynamic balance adjustment of media volume and dialogue voice. This allows the volume adjustment to automatically adapt to the interaction needs, thereby improving the scene adaptability of the head-mounted audio device's volume adjustment in near-field active voice interaction scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a volume adjustment method, device, head-mounted audio device, and storage medium. Background Technology

[0002] With the widespread adoption of head-mounted audio devices such as smart glasses and smart headphones, users often play music, videos, or voice content while interacting with their surroundings during commutes, work, and daily life.

[0003] The audio processing solutions of related head-mounted audio devices mainly employ one or more of the following technologies: manual or semi-automatic volume control, adaptive adjustment of ambient noise, audio switching triggered by specific conditions, and voice enhancement in remote scenarios. Although the above technologies improve the audio experience to some extent, in real-life face-to-face interaction scenarios, media playback and voice enhancement in these technologies are usually separate modules. This results in users having to manually adjust the volume when someone suddenly speaks, which is cumbersome, interrupts the immersive experience, and even poses safety hazards in scenarios such as commuting.

[0004] Therefore, improving the scene adaptability of volume adjustment for head-mounted audio devices in near-field active voice interaction scenarios has become an urgent technical problem to be solved. Summary of the Invention

[0005] This application provides a volume adjustment method and a head-mounted audio device, aiming to solve the technical problem of low scene adaptability of volume adjustment in near-field active voice interaction scenarios.

[0006] Firstly, this application provides a volume adjustment method including: The multimodal near-field sensing component collects multimodal near-field sensing data of the current environment of the user wearing the head-mounted display device. When the multimodal near-field perception data satisfies the triggering conditions for a near-field active voice interaction event for the wearer, the audio co-adjustment operation of the head-mounted audio device is executed in response to the near-field active voice interaction event, wherein the audio co-adjustment operation includes at least one of media volume adjustment operation and conversational voice adjustment operation.

[0007] Secondly, the volume adjustment device provided in this application includes: The near-field sensing module is used to collect multimodal near-field sensing data of the current environment of the user wearing the head-mounted display device through the multimodal near-field sensing component; An audio adjustment module is configured to perform audio collaborative adjustment operations of the head-mounted audio device in response to the near-field active voice interaction event when the multimodal near-field perception data meets the triggering conditions for the near-field active voice interaction event for the wearer. The audio collaborative adjustment operations include at least one of media volume adjustment operations and dialogue voice adjustment operations.

[0008] Thirdly, the head-mounted audio device provided in this application includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the volume adjustment method as described above.

[0009] Fourthly, this application also provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the volume adjustment method described above.

[0010] This application provides a volume adjustment method, apparatus, head-mounted audio device, and storage medium. The method acquires multimodal near-field perception data through a multimodal near-field perception component, fuses multimodal perception information, and achieves multi-dimensional perception of the near-field environment of the user, thereby improving the perception accuracy of near-field active voice interaction events. When the multimodal near-field perception data meets the triggering conditions for near-field active voice interaction events for the user, at least one audio collaborative adjustment operation, including media volume adjustment and dialogue voice adjustment, is executed to achieve dynamic balance adjustment of media volume and dialogue voice. This allows the volume adjustment to automatically adapt to interaction needs, thereby improving the scene adaptability of the head-mounted audio device's volume adjustment in near-field active voice interaction scenarios. Attached Figure Description

[0011] Figure 1 This is a flowchart illustrating a first embodiment of a volume adjustment method provided in this application. Figure 2 This is a flowchart illustrating a second embodiment of a volume adjustment method provided in this application. Figure 3 This is a schematic diagram of the structure of a first embodiment of a volume adjustment device provided in this application; Figure 4 This is a schematic block diagram of the structure of a head-mounted audio device provided in an embodiment of this application. Detailed Implementation

[0012] The following detailed description of some embodiments of this application is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.

[0013] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a first embodiment of a volume adjustment method provided in this application.

[0014] like Figure 1 As shown, the volume adjustment method includes steps S101 to S102.

[0015] S101. Collect multimodal near-field sensing data of the current environment of the user through the multimodal near-field sensing component.

[0016] The volume adjustment method provided in this application is mainly used in head-mounted audio devices, such as smart glasses and smart helmets. The head-mounted audio device is equipped with a multimodal near-field sensing component to collect and analyze multimodal near-field sensing data of the user's environment.

[0017] Multimodal refers to the simultaneous use of multiple different types of perception channels or sensors (such as audio, vision, and motion) to acquire environmental data. By fusing data from multiple different modalities (such as audio data, visual data, and motion data), a more comprehensive and accurate understanding of the scene can be achieved.

[0018] Near field is a technical term in acoustics, referring to the region where the sound source is relatively close to the microphone array, as opposed to the far field. The boundary between the near and far fields is determined by the critical distance, which is calculated by multiplying the square of the array aperture by two and then dividing by the wavelength of the sound wave. When the sound source distance is less than the critical distance (e.g., 1 meter), the sound wave exhibits spherical wave characteristics, with both amplitude and phase depending on the distance. When the sound source distance is greater than or equal to the critical distance, the sound wave approximates a plane wave, with only the phase depending on the direction.

[0019] A multimodal near-field sensing component is a hardware module that integrates multiple sensors specifically designed to sense a user's near-field environment. Specifically, the multimodal near-field sensing component includes at least one of the following sensing devices: a microphone array, an inertial measurement unit (IMU), and a camera.

[0020] Specifically, the microphone array includes at least two microphone units, which are spatially distributed on the head-mounted audio device to collect audio data from the near-field environment of the user. Taking smart glasses as an example, microphone units can be arranged at the front, middle, and rear of the temples / ear hooks to form a multi-point layout for forward pickup, lateral pickup, and rear reference. Alternatively, microphone units can be symmetrically arranged on the left and right sides of the front frame of the device to form a forward stereo pickup pair. Microphone units can also be arranged at the center of the nose pad or front frame to form a near-field voice reference point.

[0021] An inertial measurement unit (IMU) can be positioned at the center of gravity of the main frame of the head-mounted audio device or near the center of the user's head. For example, in smart glasses, an IMU module can be installed within the central crossbeam of the frame / headband. Generally, the IMU includes a three-axis accelerometer and a three-axis gyroscope, which are aligned with the coordinate system of the head-mounted audio device to accurately collect attitude data and motion status of the device or the user's head.

[0022] The camera system may include at least one main camera and one or more optional auxiliary cameras. The main camera can be positioned on the front frame of the head-mounted audio device (e.g., at the center of the front frame), with its optical axis substantially parallel to the user's line of sight, enabling the main camera to capture the user's image. Additionally, auxiliary cameras such as infrared cameras and eye-tracking cameras can be selected according to actual needs. Infrared cameras can be positioned on either side of the main camera for image data acquisition in low-light conditions; eye-tracking cameras can be positioned inside the front frame of the device, facing the user's eyes, for eye tracking and identifying the user's line of sight.

[0023] Understandably, each component in a multimodal near-field sensing assembly has known physical positional relationships and orientation parameters. For example, the three-dimensional coordinates of each microphone unit in the microphone array relative to the device coordinate system of the head-mounted audio device; the orientation angle of the camera's optical axis relative to the device coordinate system; the transformation relationship between the inertial measurement unit's coordinate system and the device coordinate system; and the time synchronization mechanism of all components to ensure the time alignment of the multimodal near-field sensing data.

[0024] Multimodal near-field sensing data refers to a time-synchronized, spatially aligned, multi-dimensional data set collected by multimodal near-field sensing components and preprocessed. This multi-dimensional data set includes at least multi-channel audio data collected by a microphone array, device attitude data collected by an inertial measurement unit, and visual image data collected by a camera.

[0025] Furthermore, the microphone array collects multi-channel audio data of the user's current environment; the inertial measurement unit collects device attitude data of the head-mounted audio device; and the camera collects visual image data of the user's current environment.

[0026] In one embodiment, based on the actual application requirements, multiple microphone units (e.g., 2, 3, 4, etc.) are selected to form a microphone array. These microphone units are spatially distributed in an asymmetrical or symmetrical manner on the head-mounted audio device, forming a multi-dimensional spatial acquisition capability. The microphone array can acquire multi-channel audio data at a first sampling frequency. This first sampling frequency can be dynamically selected according to the audio processing requirements of the application scenario and the power consumption limitations of the head-mounted audio device. For example, the value range of the first sampling frequency can be set to 16kHz, 24kHz, 32kHz, 48kHz, etc.

[0027] Multi-channel audio data acquired by the microphone array is used for sound source localization and speech activity detection. Before sound source localization and speech activity detection, the acquired multi-channel audio data can be preprocessed, including one or more methods such as bandpass filtering, automatic gain control, and frame segmentation and windowing.

[0028] For example, bandpass filtering uses digital filters to filter the original audio signal of multi-channel audio data in the frequency domain, retaining the frequency band range of 80Hz to 8kHz. This frequency band range covers the main energy distribution area of ​​human speech, while suppressing low-frequency wind noise and high-frequency electronic noise.

[0029] Automatic gain control can independently apply an adaptive gain adjustment algorithm to each audio channel based on the current ambient noise level and the energy distribution of each channel signal, so that while preserving the relative phase information, the signal amplitude of each audio channel is also within the preset target dynamic range.

[0030] Framing and windowing are used to divide the continuous audio signal after filtering and gain control into short frames of fixed duration. The duration of each short frame can be 10 to 30 milliseconds, and the inter-frame overlap rate can be 30% to 50%. Then, a Hamming window or Hanning window function is applied to each frame of audio data to reduce spectral leakage.

[0031] In one embodiment, the inertial measurement unit (IMU) includes at least one three-axis accelerometer and one three-axis gyroscope. The IMU can continuously acquire device attitude data of the head-mounted audio device at a second sampling frequency. The second sampling frequency can be set according to the actual application scenario requirements; for example, a sampling frequency within the range of 50Hz to 200Hz can be selected as the second sampling frequency, which can balance power consumption requirements while meeting real-time requirements.

[0032] Device attitude data includes three-axis acceleration data, three-axis angular velocity data, and device orientation data obtained by fusing acceleration and angular velocity data. The device orientation data is used to determine the real-time orientation of the user's head in three-dimensional space. Specifically, the three-axis accelerometer in the inertial measurement unit (IMU) measures the linear acceleration of the head-mounted audio device in three orthogonal directions (X, Y, and Z axes in the device coordinate system), including gravitational acceleration and motion acceleration components, thus obtaining three-axis acceleration data. The three-axis gyroscope in the IMU measures the rotational angular velocity of the head-mounted audio device around the X, Y, and Z coordinate axes, thus obtaining three-axis angular velocity data.

[0033] The equipment attitude data also includes equipment orientation data obtained by fusing acceleration data and angular velocity data.

[0034] Specifically, the short-term changes in the attitude of the head-mounted audio device are calculated by integrating angular velocity data. Specifically, a three-axis gyroscope measures the rotational angular velocity of the head-mounted audio device around each axis of the device's coordinate system, and the incremental change in attitude is obtained through integration. Assume the angular velocity data output by the three-axis gyroscope is... ,in, Represents the angular velocity data vector. Represents the angular velocity component about the X-axis. Represents the angular velocity component about the Y-axis. Represents the angular velocity component about the Z-axis. This is the transpose of the matrix.

[0035] During the sampling period Assuming a constant angular velocity, calculate the incremental changes of each attitude angle (pitch, roll, and yaw):

[0036] in, This represents the pitch angle increment (around the Y-axis). This indicates the roll angle increment (around the X-axis). This represents the yaw angle increment (around the Z-axis).

[0037] The attitude angle at the current moment is obtained by adding the increment to the attitude angle at the previous moment:

[0038] in, Represents attitude quaternions, To represent quaternion multiplication, Indicates the current moment. Indicates the sampling period. The attitude increment quaternion is obtained by integrating the angular velocity.

[0039] In static or low-speed motion states, the triaxial accelerometer primarily measures the components of gravitational acceleration along each axis of the device's coordinate system, which can be used to calculate the pitch and roll angles of the head-mounted audio device. The direction of gravity measured by the accelerometer is used as an absolute reference. When the head-mounted audio device moves relatively smoothly, the gyroscope drift error is calculated by comparing the gravity components in the current posture with the true direction of gravity, and the integral results of the angular velocity data are corrected.

[0040] Assuming the acceleration data output by the triaxial accelerometer is ,in, Represents an acceleration data vector. Represents the X-axis acceleration component. Represents the Y-axis acceleration component. Represents the Z-axis acceleration component. This is the transpose matrix. The normalized acceleration direction vector is:

[0041] The pitch and roll angles of a head-mounted audio device can be expressed as:

[0042]

[0043] in, The pitch angle represents the rotation angle of the head-mounted audio device around the Y-axis; The roll angle represents the angle of rotation of the head-mounted audio device around the X-axis.

[0044] To combine the long-term stability of acceleration data with the short-term dynamic accuracy of gyroscope data, a complementary filtering algorithm is used for data fusion. The complementary filtering algorithm uses a low-pass filter to process the acceleration data (suppressing high-frequency noise and motion interference) and a high-pass filter to process the gyroscope integral data (suppressing low-frequency drift). The two are then complementaryly superimposed to obtain the fused attitude.

[0045] For pitch and roll angles, the complementary filtering formula is:

[0046]

[0047] in, The pitch angle, This is the roll angle. This represents the pitch angle increment (around the Y-axis). This indicates the roll angle increment (around the X-axis). This represents the combined pitch angle (at the current moment). Indicates the roll angle after merging (at the current moment). This indicates the fused pitch angle at the previous moment. This indicates the fusion roll angle at the previous moment. These are complementary filter coefficients, representing the degree of trust in the gyroscope data. The closer the result is to 1, the more dependent the fusion result is on the dynamic response of the gyroscope; The closer it is to 0, the more it depends on the static reference of acceleration.

[0048] For yaw angle, since acceleration data cannot provide an absolute reference (gravity direction is aligned with the Z-axis and is not sensitive to rotation around the Z-axis), it mainly relies on gyroscope integration and can be corrected by combining magnetometer data (if available):

[0049] in, This represents the merged yaw angle (at the current moment). This indicates the fused yaw angle at the previous moment. This represents the yaw angle increment (around the Z-axis). The yaw angle measured by the magnetometer. This is the correction factor.

[0050] The final fused device orientation data can be represented in the following form (Eulerian angle representation):

[0051] in, This is a vector representation of the fused device orientation data. Indicates the combined pitch angle. Indicates the roll angle after fusion. This represents the yaw angle after fusion. Equipment orientation data can also be represented in other forms, such as quaternion representation or rotation matrix representation.

[0052] This embodiment uses a complementary filtering algorithm to fuse acceleration and angular velocity data, making comprehensive use of the long-term stable attitude reference provided by the three-axis accelerometer and the short-term dynamic rotation information provided by the three-axis gyroscope. This eliminates the drift error accumulated by the gyroscope integral over time and suppresses the interference of the accelerometer under dynamic motion, thereby obtaining stable, accurate and low-latency real-time head orientation data of the wearer. This can improve the environmental perception accuracy and interactive response accuracy of head-mounted audio devices in complex dynamic scenarios.

[0053] In one embodiment, a camera can be integrated into the front end of a head-mounted audio device to acquire visual image data of the area facing the user's head at a third sampling frequency, used for object detection within this area. The third sampling frequency can be dynamically adjusted based on actual application requirements, camera hardware parameters, and the power consumption limitations of the head-mounted audio device. For example, setting the third sampling frequency to 15fps means the camera acquires visual image data at a sampling frequency of 15 frames per second, or five image frames per second. If the third sampling frequency does not meet the actual application requirements, it can be appropriately reduced or increased.

[0054] Furthermore, since the sampling frequencies of the microphone array, inertial measurement unit, and camera are different, and each sensor has its own independent hardware clock, a unified time synchronization mechanism needs to be established to achieve the fusion analysis of multimodal data. This mechanism includes hardware synchronization, software synchronization, data buffering and alignment, and spatial alignment.

[0055] Specifically, hardware synchronization refers to sending a unified trigger signal or timestamp to multimodal proximity sensing components such as microphone arrays, inertial measurement units, and cameras through the main control chip of the head-mounted audio device, so that the multimodal proximity sensing data collected by each component has the same starting time reference.

[0056] Software synchronization refers to attaching a high-precision timestamp (e.g., millisecond or microsecond level) to each frame of data at the data acquisition layer, and aligning multimodal data based on the timestamp. Considering the differences in sampling frequencies of various components, nearest neighbor interpolation or linear interpolation methods can be used to unify data with different sampling frequencies onto the same time axis.

[0057] Data buffering and alignment temporarily store multimodal near-field perception data collected by each component by setting up a circular buffer. Using timestamps as indexes, multimodal near-field perception data frames within the same time moment or time window are extracted to form a synchronized multimodal near-field perception data group for subsequent near-field speech perception and interaction intent judgment.

[0058] Spatial alignment is based on the pre-calibrated installation position and orientation parameters of each component relative to the device coordinate system. It transforms the multimodal near-field sensing data collected by different components into a unified device coordinate system or world coordinate system, ensuring the spatial consistency of sound source direction, device orientation, and visual target position.

[0059] By collecting, preprocessing, synchronizing and aligning the multimodal near-field perception data, we can obtain comprehensive environmental perception information that includes audio features, posture features and visual features. This provides a complete data foundation for subsequent steps such as near-field effective sound source corresponding user detection, interaction intent judgment and adaptive volume adjustment.

[0060] After completing the aforementioned data processing of the multimodal near-field perception data, data analysis is performed on the multimodal near-field perception data to determine whether it meets the triggering conditions for a near-field active voice interaction event for the wearer. If the multimodal near-field perception data meets the triggering conditions for a near-field active voice interaction event for the wearer, it is determined that a near-field active voice interaction event for the wearer currently exists; if the multimodal near-field perception data does not meet the triggering conditions for a near-field active voice interaction event for the wearer, it is determined that a near-field active voice interaction event for the wearer currently does not exist.

[0061] Multimodal near-field perception data includes multi-channel audio data, device posture data, and visual image data.

[0062] After collecting and synchronizing multimodal near-field perception data, the process proceeds to determine near-field active voice interaction events. By analyzing multi-dimensional information such as multi-channel audio data, device posture data, and visual image data, the audio signals in the user's current environment are identified and distinguished, thereby accurately recognizing near-field active voice interaction events directed at the user.

[0063] Furthermore, such as Figure 2 As shown, the method further includes: Step S201: Based on the multi-channel audio data, perform sound source localization and voice activity detection to identify the sound source location information of effective sound sources in the current environment of the wearer. The sound source location information includes the sound source direction and the sound source distance.

[0064] In one embodiment, voice activity detection (VAD) is performed on the preprocessed multi-channel audio data to identify human voice signals within the multi-channel audio data, thereby identifying valid sound sources in the user's current environment. Simultaneously, direction of arrival (DoA) estimation is performed on the preprocessed multi-channel audio data to obtain the direction of valid sound sources in the environment. The preprocessing includes the aforementioned bandpass filtering, automatic gain control, and framing and windowing processes.

[0065] Speech activity detection aims to distinguish speech segments from non-speech segments (silence, noise, music, etc.) in a continuous audio stream, providing reliable speech feature data for sound source localization while reducing interference from non-speech sound sources (such as knocking sounds, wind sounds).

[0066] First, speech activity detection is performed on the multi-channel audio data. By analyzing the time-frequency characteristics of the multi-channel audio data (such as short-time energy, zero-crossing rate, spectral entropy, etc.), it is determined whether the current frame contains human voice signals, and a binary flag is output. And its confidence level.

[0067] Specifically, time-domain and frequency-domain features are extracted from each frame of audio data in the multi-channel audio data (e.g., frame length 20ms, frame shift 10ms).

[0068] Time-domain characteristics include short-time energy, which reflects signal strength, and short-time zero-crossing rate, which reflects signal frequency characteristics.

[0069] The formula for calculating short-time energy can be expressed as:

[0070] in, Indicates the first Frame audio data The amplitude of each sampling point This indicates the number of sampling points in each frame of audio data. Indicates the first Short-time energy of frame audio data.

[0071] The formula for calculating the short-time zero crossing rate can be expressed as:

[0072] in, Indicates the first Short-time zero-crossing rate of frame audio data Indicates the first Frame audio data The amplitude of each sampling point Indicates the first Frame audio data The amplitude of each sampling point This indicates the number of sampling points for each frame of audio data.

[0073] After extracting time-domain features, a Fast Fourier Transform (FFT) is performed on the preprocessed and framed multi-channel audio data to transform it from the time domain to the frequency domain. Frequency domain features such as the spectral centroid, spectral entropy, and Mel-frequency cepstral coefficients (MFCCs) are then extracted to more precisely characterize the spectral properties of the multi-channel audio data. Specifically, the spectral centroid reflects the location of concentrated spectral energy, with speech typically concentrated in the low-frequency band; the spectral entropy reflects the uniformity of the spectral distribution, with speech exhibiting lower spectral entropy; and the Mel-frequency cepstral coefficients (MFCCs) typically extract 12-13 dimensional MFCC features and their first and second-order differences to characterize the spectral envelope properties of the multi-channel audio data. It is understood that the frequency domain feature extraction methods are conventional techniques in this field, and will not be elaborated upon in this embodiment.

[0074] After extracting the temporal and frequency domain features of each frame of multi-channel audio data, an adaptive dual-threshold detection algorithm can be used to determine whether each frame of multi-channel audio data is speech or non-speech. The adaptive dual-threshold detection algorithm includes a first-level threshold (energy threshold) and a second-level threshold (spectral threshold).

[0075] Specifically, in the first-level threshold, the short-time energy of the current frame audio data is calculated and compared. Energy estimation for environmental noise (Updated via voice pauses). If ( If the energy threshold coefficient is 2-4, then the detection proceeds to the second-level threshold; if If so, the audio data of the current frame is determined to be a non-speech frame.

[0076] Among them, environmental noise estimation energy Update using the current frame energy only when a non-speech frame is detected. The initial value is obtained through speech gap detection. Then, when a non-speech frame is detected, the energy information of the non-speech frame is used to estimate the energy of the environmental noise using a first-order recursive average (exponential weighted moving average) algorithm. Update:

[0077] in, For the first Environmental noise energy estimation after frame audio data update For the first Noise energy estimation of frame audio data (the initial value can be set as the average energy of the previous few frames). For the first Measured short-time energy of frame audio data This is the smoothing coefficient.

[0078] After the current frame audio data passes the first-level threshold, it enters the second-level threshold detection, and the spectral entropy of the current frame audio data is calculated. and zero crossing rate .like and If the condition is met, the current frame audio data can be determined to be a speech frame; otherwise, the current frame audio data is determined to be a non-speech frame. , , This is a threshold preset based on statistical learning. Spectral entropy. This reflects the uniformity of the spectral energy distribution. Speech exhibits lower spectral entropy due to its harmonic structure, while noise exhibits higher spectral entropy due to its uniform energy distribution. Therefore, spectral entropy... It can be used to distinguish speech signals with harmonic structures from noise signals with uniform energy distribution. Spectral entropy. The Shannon entropy can be obtained by normalizing the power spectrum by probability distribution. The calculation process is a prior art in this application and will not be described in detail here.

[0079] Multi-channel audio data often contains weak consonant endings (such as aspirated sounds and nasal finals) at the end of words and sentences. Strictly adhering to energy thresholds for termination would truncate these endings, affecting speech integrity and subsequent recognition accuracy. Therefore, after speech is detected to have ended, the speech state is extended for K frames (K=20-40 frames, corresponding to 200-400ms) to avoid truncation of the final consonant. (Non-audio frame,) (Based on the VAD determination result), the state is not switched immediately. Instead, a timer is used to keep track until the duration from the detection of a non-speech frame to the preset K frames, and then the speech activity detection operation on the multi-channel audio data is terminated.

[0080] After performing speech activity detection on multi-channel audio data, a VAD flag is marked for each frame of audio data in the multi-channel audio data according to the detection result. That is, the speech activity detection judgment result (speech / non-speech) of each audio data is marked. In this way, the multi-channel audio data can be converted into a VAD flag sequence.

[0081] Median filtering or a state machine is used to smooth the VAD flag sequence in time, eliminating misclassifications caused by isolated noise points. For example, a sliding window median filter is applied to the VAD flag sequence to eliminate isolated noise points:

[0082] in, Indicates the first The VAD flag after median filtering of the frame audio data. Indicates the VAD flag sequence of the first The original VAD flag bit of the frame audio data, i.e., the first VAD determination result (speech / non-speech) of frame audio data. This represents the half-window size of the median filter; the total length of the median filter window is 2. +1. Median filtering can effectively remove burst noise that lasts for 1-2 frames and misclassify speech, while preserving true short speech segments (such as monosyllabic words). This represents the median filtering operation function.

[0083] After performing temporal smoothing on the VAD flag sequence to eliminate isolated noise points, the VAD flag bits of each frame of audio data in the multi-channel audio data are output. ∈{0,1}, and the start and end timestamps of the determined speech segment.

[0084] This embodiment detects speech activity by using time-domain and frequency-domain features to perform dual-threshold decision-making on multi-channel audio data, thereby detecting speech segments and effectively distinguishing speech from non-speech signals. This provides reliable speech activity information for subsequent sound source localization, near-field distance estimation, and interaction intent judgment, and can improve the ability of head-mounted audio devices to perceive user speech corresponding to effective near-field sound sources in complex acoustic environments.

[0085] Based on the audio frames where speech activity was detected, a microphone array signal processing algorithm is used to estimate the direction of the sound sources. Specifically, methods such as generalized cross-correlation phase transform or beamforming spatial spectrum estimation are used to calculate the azimuth information of one or more main sound sources in the environment relative to the microphone array for each audio frame where speech activity was detected, and the azimuth angles of each main sound source relative to the microphone array are output.

[0086] For example, coarse localization can be performed using the Time Difference of Arrival (TDOA) for microphone pairs in a microphone array. Calculate the generalized cross-correlation function:

[0087] in, Indicates microphone pair The generalized cross-correlation function, , They represent the first Fourier transform of the audio signal of the first channel and the... Fourier transform of the audio signals of each channel Indicates complex conjugation. For frequency variables, Represents the time delay variable. It is a negative exponential phase factor used to achieve time shifting in the frequency domain.

[0088] The peak position is extracted by obtaining the time-domain cross-correlation function through inverse Fourier transform. This is the arrival time difference between the two channels.

[0089] Geometric position based on microphone pair , and speed of sound (Approximately 343 m / s), establish the equation:

[0090] in, This is the azimuth angle (horizontal plane, 0° is directly in front). The pitch angle is (vertical plane, horizontal is 0°). For the speed of sound, For the first The audio signal of the first channel and the first The arrival time difference of the audio signals in each channel.

[0091] For multiple microphone pairs, an overdetermined system of equations can be constructed, and the optimal direction estimate can be solved using the least squares method or spherical interpolation method. .

[0092] Based on the coarse positioning of TDOA, precise positioning is achieved by using delay-sum beamforming or minimum variance distortionless response (MVDR) beamforming.

[0093] Taking delay-sum beamforming as an example, time delay compensation is applied to the signals of each channel of the microphone array. , making the direction from In-phase superposition of signals:

[0094]

[0095] in, Indicates the beamforming output signal. It is the azimuth angle. The pitch angle is M, and M represents the total number of microphone units in the microphone array. For the first The weighting factor for each microphone (usually 1 / M, where M represents the total number of microphone units in the microphone array). Indicates the first The geometric position of each microphone. It is a time variable, representing continuous time.

[0096] In the direction of A grid search is performed in the space (azimuth -90° to 90°, elevation -45° to 45°, step size 1°-5°) to calculate the beam output power in each direction:

[0097] in, Indicates the beamforming output signal. Indicates direction as The beam output power.

[0098] The peak position of the beam output power is the precise positioning result:

[0099] in, Indicates direction as beam output power, This indicates the result of locating the direction of the sound source.

[0100] This embodiment utilizes a sound source localization algorithm based on generalized cross-correlation phase transformation and delay-sum beamforming. By leveraging time difference of arrival (TDOA) coarse localization and beamforming, it achieves accurate localization of effective sound source directions in the environment. This effectively overcomes the ambiguity of single-microphone localization and environmental noise interference, providing reliable sound source location information for subsequent near-field distance calculation, interactive intent judgment, and directional speech enhancement. This improves the accuracy of spatial perception and near-field speech interaction recognition of head-mounted audio devices in complex acoustic scenarios.

[0101] After completing the speech activity detection and sound source direction estimation for a single frame, the detection results in the time series are analyzed for continuity and stability to eliminate instantaneous noise interference and sudden non-speech sound sources, ensuring that subsequent distance estimation and interaction intent judgment are based on reliable and stable targets.

[0102] Specifically, a sliding time window is used to perform cumulative analysis on the VAD results. Assume the window length of the sliding window is... =300ms~800ms (corresponding to 30 to 80 frames, frame length 10ms), sliding step =100ms.

[0103] Within the sliding window, count the number of audio frames. and total number of frames Calculate the percentage of speech:

[0104] in, Indicates the percentage of speech, This indicates the number of audio frames containing human voice. This indicates the total number of audio frames.

[0105] Stable human voice activity is considered detected when the following conditions are met: 1) Voice portion: > , The threshold representing the proportion of speech is usually taken as... =0.6~0.8 indicates that at least 60%~80% of the sliding window consists of audio frames; 2) There exists at least one continuous speech frame, the length of which is... > , Indicates the minimum acceptable length of consecutive speech frames, such as =200ms, to eliminate intermittent noise; 3) The start and end points of the effective speech segment were detected using the aforementioned dual-threshold detection algorithm, and the duration of the speech activity was recorded. .

[0106] In addition, for microphone arrays, the VAD results of each channel are required to have a high degree of consistency (consistency ratio > 0.9) to avoid misjudgment caused by single-channel anomalies.

[0107] After completing the continuity analysis of each speech frame in the sliding window, output the stable speech activity flag corresponding to each sliding window. , =0 indicates that the speech segments in the corresponding sliding window are not continuous and stable. =1 indicates that the corresponding audio segments in the sliding window are continuous and stable. Simultaneously, the start and end timestamps of each continuous audio segment are output. , ) and duration .

[0108] Within a sliding window confirming stable speech activity, the source direction estimation results for each speech frame are extracted to form a direction sequence:

[0109] in, This indicates the number of audio frames containing human voices within the sliding window. For the first The azimuth angle of each audio frame For the first The pitch angle of each audio frame.

[0110] The calculation of the degree of change in the direction of the sound source within the window mainly includes the range of azimuth angle change, the range of elevation angle change, and the variance of the direction vector:

[0111] in, Indicates the range of azimuth variation. This indicates the range of pitch angle variation. This represents the variance of the direction vector. This indicates the number of audio frames containing human voices within the sliding window. It is the first Azimuth angle of each audio frame and pitch angle Convert to a unit vector in three-dimensional space. Represents the transpose matrix. This indicates the average direction of the vector. Indicates the first The azimuth and elevation angles of each audio frame. Indicates the first The azimuth and elevation angles of each audio frame.

[0112] When the azimuth angle corresponding to consecutive speech segments within a certain sliding window changes Within the azimuth change threshold range (e.g., 0°~30°), pitch angle change Within the pitch angle variation threshold range (e.g., 0°~20°), and the vector variance When the variance is less than a preset variance threshold (e.g., 0.2), the direction of the sound source in the continuous speech segment can be determined to be stable.

[0113] If there are continuous speech segments within a continuous sliding window, and the direction of the sound source of the continuous speech segments is determined to be stable, then it can be determined that there is a potential speaking target in the direction of the sound source corresponding to the continuous speech segments.

[0114] This embodiment effectively eliminates transient noise interference and sudden non-speech sound sources through a continuity analysis and stability verification process based on a sliding time window, ensuring the authenticity and continuity of detected speech activities. Simultaneously, variance analysis of the direction sequence confirms the spatial stability of the sound source, providing a reliable potential speaking target for subsequent near-field distance estimation and interaction intent judgment. This improves the accuracy and robustness of head-mounted audio devices in perceiving realistic near-field speech interactions in complex dynamic environments.

[0115] For a confirmed potential speaking target, in order to distinguish between far-field background speech and near-field active communication speech, the distance between the potential speaking target and the user (the user wearing the head-mounted audio device) can be calculated by comprehensively considering information such as the amplitude difference of signals in each channel of the microphone array, the model of sound pressure level attenuation with distance, and the phase and amplitude nonlinear characteristics of the microphone array under near-field conditions.

[0116] In one embodiment, due to differences in spatial position, the signal amplitude received by each microphone unit in the microphone array from the same sound source varies. This difference is closely related to the distance from the sound source, and is particularly significant under near-field conditions.

[0117] Assume the microphone array contains The microphone unit, the first The microphone unit positions are as follows: The location of the sound source is ,in, To estimate the distance, The direction of the sound source has been estimated. The distance from the sound source to the... The distance between the microphone units is:

[0118] in, Indicates the distance from the sound source to the first The distance between microphone units, No. The location of each microphone unit Indicates the location of the sound source. To estimate the distance, The direction of the sound source has been estimated.

[0119] Under the assumption of free-field spherical wave propagation, the first... The sound pressure level received by each microphone unit is inversely proportional to the distance:

[0120] in, Indicates the first The sound pressure level received by each microphone unit For the amplitude of the sound source, For the first The directional gain of each microphone unit (for an omnidirectional microphone, ≈1). Indicates the distance from the sound source to the first The distance between each microphone unit.

[0121] Calculate the relative amplitude ratio using a reference microphone (such as the microphone at the geometric center of the array or the microphone with the strongest signal) as a reference. Let the reference microphone be... Distance to the sound source is Then the first The relative amplitude ratio of each microphone unit to the reference microphone is:

[0122] in, Indicates the first The sound pressure level received by each microphone unit Indicates the distance from the sound source to the first The distance between microphone units, Reference microphone Received sound pressure level, Reference microphone Distance to the sound source.

[0123] right Each microphone unit can be used to construct Several independent magnitude ratio equations. The above equations are rearranged to represent the distance to be estimated. Nonlinear equations:

[0124] in, Indicates the first The nonlinear equations corresponding to each microphone Indicates the distance to be estimated. Indicates the first The sound pressure level received by each microphone unit Reference microphone The received sound pressure level. Indicates the first The distance from each microphone to the sound source, This indicates the distance from the reference microphone to the sound source.

[0125] The least squares method is used for optimization, and the cost function is defined as follows:

[0126] in, Let the cost function be the magnitude ratio method. This indicates the total number of microphones. Indicates the distance to be estimated. Indicates the first The sound pressure level received by each microphone unit Reference microphone The received sound pressure level. Indicates the first The distance from each microphone to the sound source, This indicates the distance from the reference microphone to the sound source.

[0127] Solve for the optimal distance estimate using numerical optimization methods (such as Newton's iteration method or grid search):

[0128] in, This indicates the optimal distance estimated by the magnitude ratio method. This is the cost function for the magnitude ratio method. This indicates the distance to the sound source to be estimated. Indicates the distance to the search range.

[0129] In one embodiment, the distance estimation method based on the sound pressure level attenuation model with distance utilizes the physical law of sound pressure level attenuation with distance and combines it with the energy characteristics of the audio signal received by the microphone array to establish a distance estimation model.

[0130] Specifically, for the propagation of a point sound source in a free field, the sound pressure level changes with distance according to the following:

[0131] in, Distance to be estimated Sound pressure level (in dB) at the location. For reference distance (e.g.) =1m), It is the air absorption coefficient (related to frequency, temperature, and humidity, approximately 0.001~0.01dB / m in the speech frequency range).

[0132] Let the sound pressure level of the signal received by the microphone array be... It is necessary to establish the relationship with the sound pressure level of the sound source. Considering the positional differences of each channel in the microphone array, the equivalent received sound pressure level of the array is calculated:

[0133] in, For the first The relative sound pressure levels of each channel, among which Indicates the first The sound pressure level received by each microphone unit Reference microphone The received sound pressure level. This represents the equivalent received sound pressure level of the microphone array. This indicates the total number of microphones.

[0134] Assuming the sound source is at the reference distance The typical sound pressure level at that location is (Normal speech at 1 meter is approximately 60-70 dBSPL), then:

[0135] in, This is a correction term for environmental reflection and scattering. This indicates the sound pressure level of the sound source at a reference distance. This indicates the diffusion and attenuation of spherical waves. This represents the air absorption coefficient. The distance to be estimated. This represents the equivalent received sound pressure level of the microphone array.

[0136] Ignore air absorption (near field) ), simplified to:

[0137] in, This indicates the distance to the sound source estimated by the sound pressure level method. This represents the equivalent received sound pressure level of the microphone array. Indicates the sound source at the reference distance The sound pressure level at that location. Indicates the reference distance.

[0138] In one embodiment, a distance estimation method based on near-field phase and amplitude nonlinear characteristics can effectively distinguish near-field sound sources and estimate their distances by analyzing the nonlinear characteristics of phase and amplitude, based on the different sound wave reception characteristics exhibited by the microphone array in the near and far fields.

[0139] For the boundary between the near and far fields: critical distance .in, For the wavelength of sound waves (voice segment) =1kHz, ≈0.343m); The array aperture (maximum microphone spacing) refers to the maximum physical distance between any two microphone units in a microphone array. When the sound wave front exhibits spherical wave characteristics, a near-field model is required.

[0140] For the plane wave assumption (far field), the first The phase difference between each microphone and the reference microphone is:

[0141] in, This represents the phase difference under the far-field plane wave assumption. Represents frequency variables. Indicates the speed of sound. This represents the position difference vector between the microphone pairs. Indicates the direction of the sound source.

[0142] For spherical waves (near field), the phase difference is distance-dependent:

[0143] in, This represents the phase difference under near-field spherical waves. Represents frequency variables. It indicates the speed of sound. Indicates the first The distance from each microphone to the sound source, Indicates the first The distance from each microphone to the sound source, This indicates the difference in the distance sound waves travel.

[0144] Calculate the deviation between the measured phase difference and the theoretical phase difference of the plane wave:

[0145] in, This represents the weighted squared error of the phase fit. This represents the measured phase difference. This represents the variance of the phase measurement. This represents the phase difference under near-field spherical waves.

[0146] By minimizing Find the optimal distance:

[0147] in, This represents the weighted squared error of the phase fit. This indicates the sound source distance estimated by the phase method.

[0148] By combining the estimation results of the above three methods (based on the amplitude difference of signals in each channel of the microphone array, the model of sound pressure level attenuation with distance, and the phase and amplitude nonlinear characteristics of the microphone array under near-field conditions), an adaptive weighted fusion strategy is adopted to obtain the final sound source distance estimate.

[0149] Weights are dynamically assigned based on the confidence level of each method:

[0150] in, Indicates the first Fusion weights for various sound source distance estimation methods Indicates the first The confidence level of several sound source distance estimation methods. For the summation index, Indicates the first Confidence of various sound source distance estimation methods. This represents a set of methods for estimating the distance to a sound source.

[0151] Overall distance estimation:

[0152] in, This indicates the distance between the sound sources after fusion. Indicates the first The distance is estimated by this method. Indicates the first Fusion weights for various sound source distance estimation methods.

[0153] This embodiment uses three sound source distance estimation methods—multi-channel signal amplitude difference, sound pressure level attenuation model with distance, and near-field phase and amplitude nonlinearity—to calculate the distance between the potential speaking target and the user wearing the head-mounted audio device. This can effectively distinguish between far-field background speech and near-field active communication speech, overcoming the limitations of a single method in complex acoustic environments. It provides a reliable quantitative basis for subsequent near-field determination and interaction intent judgment, thereby improving the accuracy of the head-mounted audio device in perceiving the spatial location of the user corresponding to the effective sound source in real interactive scenarios.

[0154] The calculated sound source distance Compare with a preset near-field range threshold (e.g., approximately 1 meter). When the sound source distance... If the sound source is less than the preset near-field range threshold, it is determined to be a valid sound source, that is, a valid near-field sound source corresponds to the sound source emitted by the user. At this time, the interaction intent judgment process can be triggered.

[0155] Step S202: When the distance to the sound source is less than a preset near-field range threshold, calculate the interaction intent parameter corresponding to the effective sound source based on the sound source direction, the device posture data, and the visual image data.

[0156] When the calculated sound source distance When the sound source is below a preset near-field range threshold, it is identified as a valid sound source, meaning it is a sound source emitted by the user corresponding to a valid near-field sound source. However, the presence of near-field speech does not equate to the user of the valid sound source having an active interactive intention towards the wearer—the user of the valid sound source may be facing away from the wearer, walking alongside the wearer, or simply conversing with others within the near-field range. Therefore, it is necessary to further integrate multimodal information such as sound source direction, device posture, and visual images to quantitatively assess the intensity of the interaction intention of the user of the valid sound source, generate interaction intention parameters, and provide a basis for subsequent decisions on whether to trigger volume adjustment.

[0157] Further, based on the sound source direction and the device posture data, a horizontal angle evaluation component is calculated between the sound source direction and the user's head orientation; based on the visual image data, the facial orientation angle of the user corresponding to the effective sound source is detected, and an orientation consistency evaluation component is calculated between the facial orientation angle and the user's facial orientation angle; temporal continuity analysis is performed on the multi-channel audio data to normalize the speech duration in the multi-channel audio data into a speech continuity evaluation component; the horizontal angle evaluation component, the orientation consistency evaluation component, and the speech continuity evaluation component are weighted and summed to obtain the interaction intent parameter corresponding to the effective sound source.

[0158] First, for the device attitude data acquired by the inertial measurement unit (IMU), the orientation of the user's head in three-dimensional space can be obtained through an attitude calculation algorithm. Specifically, assuming that in the device coordinate system, the orientation of the user's head is determined by the yaw angle... (In the horizontal plane, 0° is directly in front, and rightward deviation is positive; range []) 180°, 180°) and pitch angle (In the vertical plane, horizontal is 0°, and upward is positive) Description.

[0159] Head pose is represented using quaternions or rotation matrices. Let the pose quaternion corresponding to the device pose data output by the IMU be... Then, the unit vector of the user's face orientation angle is:

[0160] in, This represents the device's attitude data. This represents a unit vector indicating the angle of the user's face. This represents the transpose of the matrix.

[0161] Yaw and pitch angles are calculated as follows:

[0162] in, This indicates the yaw angle in the angle of the user's face. This indicates the pitch angle in the angle in which the user's face is facing. This represents the projection of the user's face orientation angle onto the X, Y, and Z axes of the world coordinate system.

[0163] The direction of the sound source obtained from the aforementioned calculation is: The projection azimuth angle of the sound source direction onto the horizontal plane can be obtained as follows: .

[0164] Calculate the angle between the direction of the sound source and the direction of the user's head in the horizontal plane. :

[0165] in, This indicates the angle between the direction of the sound source and the direction the user's head is facing in the horizontal plane. This indicates the angle of the user's face when wearing the device. Indicates the direction of the sound source.

[0166] The smaller the horizontal angle, the closer the effective sound source is to the user directly in front of the wearer, and the higher the likelihood of interaction. For example, a Gaussian or piecewise linear function can be used to adjust the angle. Mapped to horizontal angle evaluation components :

[0167] in, Indicates the horizontal angle evaluation component. To control the decay rate, a range of 30° to 45° is typically used. This indicates the angle between the direction of the sound source and the orientation of the user's head in the horizontal plane. When the angle is greater than 90°, the user corresponding to the effective sound source is considered to be located behind and to the side of the user, which basically eliminates the possibility of active interaction, and the score is set to 0.

[0168] Secondly, visual image data provides direct observation of the user's face orientation angle corresponding to the effective sound source, which is a key modality for verifying interaction intent. By detecting the image region corresponding to the sound source direction, the user's face corresponding to the effective sound source is identified and its orientation is estimated, thereby calculating the consistency with the head orientation of the user wearing the device.

[0169] Specifically, a deep learning face detection algorithm is used to detect the user's face corresponding to the valid sound source direction in image frames of visual image data, and outputs the face bounding box and confidence score. Within the detected face region, facial key points are extracted, including the centers of the eyes, the tip of the nose, and the corners of the mouth. Based on the detected facial key points corresponding to the valid sound source, the facial orientation of the user corresponding to the valid sound source is calculated using methods such as geometric analysis or deep learning regression. For example, the deep learning regression method uses a pre-trained orientation estimation network, directly taking the face image as input, and regresses three orientation angles (face pitch angle). Yaw angle and roll angle ).

[0170] The effective sound source is represented by yaw angle and pitch angle, corresponding to the angle of the user's face orientation. This is used to reflect the orientation of the user's face relative to the camera, corresponding to the effective sound source. Since the camera and the user's head are essentially aligned (optical axes parallel), the angle of the user's face orientation corresponding to the effective sound source can be directly compared to the user's head orientation angle. The angle between the user's face yaw angle and the user's head yaw angle is calculated as follows:

[0171] in, This indicates the yaw angle in the angle of the user's face orientation corresponding to the effective sound source. This indicates the yaw angle in the angle in which the user's face is facing. This indicates the angle between the yaw angle of the user's face corresponding to the effective sound source and the yaw angle of the user's head.

[0172] Similarly, a threshold function can be defined to map the included angle. For the consistency assessment components:

[0173] in, To move towards a consistent evaluation component, To ensure the effective sound source corresponds to the angle between the user's face yaw angle and the user's head yaw angle, This is the maximum tolerance deviation angle (e.g., 45°).

[0174] Furthermore, the temporal continuity of speech reflects the continuity and stability of user communication corresponding to an effective sound source, and is an important indicator for distinguishing active conversation from brief environmental speech (such as coughing or short shouts).

[0175] Set an observation time window (e.g., 1.5 seconds). Within this window, the percentage of frames that VAD identifies as having "voice activity" out of the total number of frames is recorded as the voice activity ratio. Compare voice activity to After smoothing and normalization, we get [0 Speech continuity evaluation component [1] .

[0176] Finally, by combining the above-mentioned horizontal angle evaluation components, orientation consistency evaluation components, and speech continuity evaluation components, the final interaction intent parameter is obtained through weighted summation.

[0177] Based on the reliability and scenario adaptability of the scores for each dimension, dynamic or static weights are assigned to the scores for the three dimensions mentioned above: the weight corresponding to the horizontal angle evaluation component. The weights corresponding to the consistency evaluation components And the weights corresponding to the speech continuity evaluation components. The weights satisfy the normalization condition: 1.

[0178] The weights can be dynamically adjusted based on the current environmental conditions. For example, when visual information is lacking (dark light, occlusion), the weight of the orientation consistency evaluation component can be reduced. Increase the weight of the horizontal angle evaluation component. Weights corresponding to speech continuity evaluation components When the head rotates rapidly, the uncertainty in orientation estimation increases, so the weight of the horizontal angle evaluation component can be reduced. Increase the weight of the component corresponding to the consistency assessment. Weights corresponding to speech continuity evaluation components When the speech signal-to-noise ratio is low, the weight of the speech continuity evaluation component can be reduced. Increase the weight of the horizontal angle evaluation component. Weights corresponding to the orientation consistency assessment components .

[0179] Based on the weights of the horizontal angle evaluation component, the orientation consistency evaluation component, and the speech continuity evaluation component, the scores are weighted and calculated to determine the interaction intent parameter.

[0180] in, The parameter represents the interaction intent. Indicates the horizontal angle evaluation component. Indicates the direction towards consistency assessment component, This represents the speech continuity evaluation component. This indicates the weight corresponding to the horizontal angle evaluation component. This indicates the weight corresponding to the consistency assessment component. This represents the weight corresponding to the speech continuity evaluation component.

[0181] The interaction intent parameter can be represented by a score; alternatively, it can be represented by a segmented approach, rating the weighted results of the horizontal angle evaluation component, orientation consistency evaluation component, and speech continuity evaluation component, and using the rating result to represent the interaction intent parameter. For example, if the weighted result is represented by a score, with 90 points or above being the first level, 75 to 90 points being the second level, 60 to 75 points being the third level, and below 60 points being the fourth level, then the interaction intent parameter is represented by the rating result corresponding to the segment in which the weighted result of the horizontal angle evaluation component, orientation consistency evaluation component, and speech continuity evaluation component falls. Similarly, other methods can also be used to represent the interaction intent parameter, such as probability distributions or confidence intervals. This application does not impose any restrictions on these methods; it is sufficient to implement the discrimination of the interaction intent parameter.

[0182] This embodiment calculates the final interaction intent parameter by weighting and summing the horizontal angle evaluation component, orientation consistency evaluation component, and speech continuity evaluation component through multimodal fusion weighting. This enables a quantitative judgment on whether there is an active interaction intent between the wearer and the user corresponding to the effective sound source. It can effectively eliminate interference scenarios such as background sound sources, occasional speech, and users corresponding to the effective sound source from the side / back, and reduce the false trigger rate of audio adjustment operations.

[0183] Step S203: When the interaction intent parameter meets the preset conditions, determine that the multimodal near-field perception data meets the triggering conditions for near-field active voice interaction events for the wearer.

[0184] In one embodiment, one or more intent scoring thresholds can be pre-set to construct preset conditions for distinguishing the degree of interaction intent between the user corresponding to the detected valid sound source and the user wearing the head-mounted audio device. For example, two intent scoring thresholds can be set: a high intent threshold and a low intent threshold. and low intent threshold At this point, the following preset conditions can be constructed: the interaction intent parameter is greater than or equal to the high intent threshold, the interaction intent parameter is less than the high intent threshold but greater than or equal to the low intent threshold, and the interaction intent parameter is less than the low intent threshold. Thus, based on the pre-constructed preset conditions, the interaction intent parameter can be divided into interaction state, no interaction intent, and fuzzy interaction state.

[0185] For example, the interaction intent parameter is evaluated based on a preset intent scoring threshold:

[0186] in, Indicates the interaction intent parameter, Indicates a high intent threshold. This indicates a low intent threshold.

[0187] In the above formula, if the interaction intent parameter Greater than or equal to the high intent threshold If the multimodal near-field perception data meets the triggering conditions for a near-field active voice interaction event for the user, then the near-field active voice interaction event for the user is triggered, which means that the user corresponding to the effective sound source and the user wearing the head-mounted audio device are currently in an interactive state.

[0188] If the interaction intent parameter Less than the high intent threshold And greater than or equal to the low intent threshold If the interaction is fuzzy, it means that the user corresponding to the valid sound source and the user wearing the head-mounted audio device are currently in a state of ambiguous interaction. That is, there may be voice interaction, but it cannot be accurately determined. In this case, the interaction status of the two can be continuously monitored and the interaction intentions of the two can be continuously evaluated. Then, a judgment can be made based on the subsequent evaluation of the interaction intentions.

[0189] If the interaction intent parameter Less than the low intent threshold If the result is negative, it means that the user corresponding to the valid sound source and the user wearing the head-mounted audio device currently have no intention to interact. It is determined that the multimodal near-field perception data does not meet the triggering conditions for near-field active voice interaction events for the user wearing the device, that is, there are currently no near-field active voice interaction events for the user wearing the device.

[0190] This embodiment achieves hierarchical judgment of the interaction state between the wearer and the user corresponding to the effective sound source by setting a multi-threshold interaction intent evaluation mechanism. This can avoid misjudgment caused by fluctuations in a single score and effectively filter out non-interactive scenarios. It not only improves the reliability and real-time performance of interaction judgment, but also provides a flexible and accurate decision-making basis for subsequent differentiated audio processing for different interaction states, thereby optimizing the wearer's user experience.

[0191] S102. When the multimodal near-field perception data meets the triggering conditions for a near-field active voice interaction event for the wearer, the audio coordination adjustment operation of the head-mounted audio device is executed to respond to the near-field active voice interaction event, wherein the audio coordination adjustment operation includes at least one of a media volume adjustment operation and a dialogue voice adjustment operation.

[0192] When a near-field active voice interaction event is triggered for the user, it indicates that the user corresponding to the valid sound source and the user wearing the head-mounted audio device are currently in an interactive state. At this time, the audio collaborative adjustment mechanism is immediately activated. This audio collaborative adjustment mechanism does not simply execute two independent operations of reducing media volume and enhancing conversational speech in sequence. Instead, it performs smooth suppression of media volume and signal enhancement of conversational speech in parallel within a coordinated time window. This achieves a dynamic balance between media content and near-field speech, ensuring that the user can clearly perceive the conversational content without experiencing discomfort due to sudden changes in media volume.

[0193] Generally, in related technologies, head-mounted audio devices typically employ one or more of the following techniques: manual or semi-automatic volume control, adaptive adjustment to ambient noise, audio switching triggered by specific conditions, and voice enhancement for distant scenes. For example, most smart glasses or headphones only allow users to manually adjust the media volume via physical buttons, touch controls, or voice commands, or perform simple automatic gain control (AGC) based on the overall ambient noise level to achieve a fixed or semi-automatic volume control mechanism. However, when someone suddenly speaks, it is often necessary to manually pause or lower the volume, which is cumbersome, interrupts the immersive experience, and poses safety hazards in scenarios such as commuting. Some products detect ambient sound pressure levels through microphones, increasing the media volume when ambient noise increases and decreasing the volume when noise decreases to ensure the audibility of the media content. However, such products usually only focus on the overall ambient noise intensity and cannot distinguish whether there is a nearby speaker (about one meter away) actively communicating with the user. Even if a voice is detected, it is impossible to determine whether the speaker is facing the user or has an interactive intention, easily misinterpreting background pedestrians' voices or distant sounds as the person being communicated with. In addition, some products use conditional triggering mechanisms such as calls or voice assistants to trigger and operate volume adjustment. For example, when a phone call or wake word is detected, the current media playback will be paused or suppressed, and the system will switch to call or voice assistant mode. Other related products use basic voice enhancement and noise reduction technologies. These products typically integrate modules such as noise suppression, echo cancellation, and beamforming, but media playback and voice enhancement usually exist as independent modules, primarily serving remote calls or voice input scenarios rather than local face-to-face communication. They lack the ability to coordinate media volume suppression and near-field voice enhancement within the same interactive scenario.

[0194] To address the aforementioned audio adjustment issues in related head-mounted audio devices, this application further proposes the following: within a preset first time period, based on a first smoothing control strategy, the volume of the media currently played by the head-mounted audio device is adjusted from a first volume level to a second volume level; within the first time period, signal enhancement processing is performed on the dialogue speech in the near-field active voice interaction event to enhance the voice signal of the user corresponding to the effective sound source from the ambient sound.

[0195] The first smooth control strategy refers to a control method that, when it is determined that there is a near-field active voice interaction event facing the user, smoothly reduces the volume of the media currently being played by the head-mounted audio device from a first volume level to a second volume level within a first time period.

[0196] Triggered by near-field active voice interaction event Establish the first time period with the time origin as the origin. A unified timeline (e.g., 150ms to 300ms). The first time segment... Divided into Equal interval control cycles ( ), in each control cycle The media volume gain parameters and conversational speech enhancement parameters are updated synchronously to ensure precise alignment in the time domain.

[0197] Define the normalized time variable:

[0198] in, This refers to the trigger moment of a near-field active voice interaction event. ∈[0,1]. Based on this normalized time variable, a media suppression curve is constructed. and speech enhancement curve The two satisfy the energy complementarity constraint:

[0199] in, For media signal power, For voice signal power, A constant power reference value is used to maintain a stable overall perceived loudness.

[0200] First volume level The original loudness level of the current media playback, expressed in digital gain or decibels. Second volume level. The loudness level after target suppression is dynamically calculated based on the preset target speech and media signal-to-noise ratio.

[0201] Preset target signal-to-noise ratio For example, 6dB to 12dB, based on the currently estimated effective sound source corresponding to the loudness of the user's speech. Calculate the target media loudness:

[0202] in, Indicates the target signal-to-noise ratio. This indicates the user speech estimation corresponding to the valid sound source. This indicates the target media loudness (i.e., the second volume level).

[0203] Convert to linear gain to obtain media volume suppression factor :

[0204] in, This indicates the media volume suppression factor. This indicates the target media loudness (i.e., the second volume level). This indicates the original media loudness (i.e., the first volume level).

[0205] In practical applications, media volume suppression factor It can adaptively adjust based on constraints such as the clarity of the user's speech corresponding to the effective sound source, the level of ambient noise, and the personalized preferences of the wearer. For example, if the speech clarity is low (e.g., high background noise), the media volume suppression factor can be appropriately reduced. (That is, suppressing more media volume) to improve dialogue intelligibility. If the ambient noise is high, the target signal-to-noise ratio can be appropriately increased. The value of this reduces the media volume suppression factor. In addition, users can preset preferences (such as "mild suppression" or "deep suppression"), or adjust the media volume suppression factor by learning the user's historical adjustment behavior. The baseline value.

[0206] In one embodiment, to avoid discomfort caused by sudden volume changes, media volume adjustment can be achieved using a smoothing function. Specifically, a smoothing function such as a linear ramp, exponential decay, or S-curve (such as the sigmoid function) is used to adjust the media volume within a preset first time period. Within 150-300 milliseconds, the media volume will be increased from the first volume level. Gradually and continuously adjust to the second volume level. This achieves a natural transition of media volume from the first volume level to the second volume level. During the adjustment process, the media volume suppression factor is updated every frame or every few milliseconds. To ensure a smooth and natural transition.

[0207] In addition, in the first time period Inside, the error between the actual output loudness and the target loudness can be monitored in real time, and the gain can be finely adjusted through a proportional-integral-derivative controller to ensure that the actual loudness closely tracks the target curve.

[0208] This embodiment achieves a natural transition of media volume from the first level to the second level by normalizing time variables and using smoothing functions (such as S-curves), improving the smoothness and comfort of the user experience. By constructing an audio collaborative adjustment mechanism, it enhances conversational speech while suppressing media volume, maintaining the stability of the overall perceived loudness and avoiding the illusion of fluctuating volume during interaction. By updating media volume and conversational speech in parallel on a unified timeline, it provides robust and intelligent audio support for near-field active voice interaction.

[0209] While smoothly reducing the media volume, the voice signal of the user corresponding to the effective sound source in the near field is enhanced to amplify the voice signal of the user corresponding to the effective sound source from the ambient sound, so as to ensure that the voice conversation is clearly identifiable in the media background.

[0210] Specifically, based on the calculated direction of the sound source The beamforming weight vector of the microphone array is dynamically calculated. For the [missing information - likely a specific microphone element] in the array... One microphone is used to calculate the time delay relative to a reference point:

[0211] in, For the first The spatial coordinates of each microphone The speed of sound. Indicates the first The latency of each microphone.

[0212] The beamout is a weighted sum of the signals from each microphone after time delay compensation:

[0213] in, Indicates the beam output signal. Based on sound source distance The Near-field amplitude compensation corresponding to the distance from the microphone to the sound source. Indicates the distance to the sound source. This indicates the total number of microphones. Indicates the first The weighting coefficients for each microphone. Indicates the first The latency of each microphone. Indicates the number of delay compensations. road signal. It is a time variable.

[0214] The beamwidth of the calculated output beam can be adaptively adjusted according to the distance to the sound source. In the near field, a wide beam can be used to avoid speech loss due to slight head movements, while in the far field, a narrow beam is used to improve directional selectivity. This ensures that the direction of the maximum gain of the main lobe of the formed pickup beam is precisely aligned with the azimuth angle of the user corresponding to the effective near-field sound source, thereby achieving spatially selective sound pickup. Simultaneously, the null or sidelobe suppression characteristics of the beamforming automatically suppress environmental noise and residual media sound from other directions, improving the signal-to-noise ratio of the target speech. Furthermore, in complex noisy environments, a minimum variance distortionless response beamformer can be used to achieve adaptive noise suppression through continuous updates of the noise covariance matrix.

[0215] Traditional beamforming algorithms are typically based on the far-field assumption, which assumes that the sound source is far enough away from the microphone array that the sound waves are treated as plane waves when they reach the array elements. This means that the time difference (delay) of the sound waves arriving at different elements depends only on the direction of the sound source and is independent of the distance. However, under near-field conditions, since the wavefront of a near-field sound source (e.g., <1 meter) is a spherical wave (spherical wave characteristics), the amplitude of a spherical wave attenuates with increasing distance during propagation. Furthermore, the phase and amplitude changes of sound waves of different frequencies along the propagation path are more complex, resulting in differences in the amplitude and phase relationships of different frequency components.

[0216] To address the spherical wave characteristics of sound wave propagation under near-field conditions, a near-field gain compensation model is introduced to perform amplitude and phase compensation on the speech signal after beamforming, making the reconstructed speech signal spectrum more natural and avoiding "hollowness" or distortion.

[0217] Specifically, the amplitude compensation gain is proportional to the distance:

[0218] in, Indicates amplitude compensation gain. Indicates the distance to the sound source. For reference distance, The decay exponent, This is the reference gain.

[0219] Phase compensation is achieved through a near-field focusing filter to compensate for the phase difference between spherical waves and plane waves.

[0220] in, This indicates a near-field phase compensation filter. Indicates the distance to the sound source. It represents the equivalent distance of a plane wave. It indicates the speed of sound. Represents a frequency variable. It is the imaginary unit.

[0221] The speech enhancement curve and media suppression curve are designed to complement each other to ensure rapid speech prominence.

[0222] in, Indicates speech enhancement gain. This indicates the elapsed time, that is, the time variable from the start of the speech. To rapidly increase the time constant, ensure that speech is clearly identifiable in the early stages of media suppression. Indicates the target gain. This represents the initial gain.

[0223] Building upon beamforming and near-field gain compensation, the enhanced speech signal can be further post-processed to improve speech clarity and naturalness. Specifically, spectral subtraction is used to reduce noise in the beamout.

[0224] in, This represents the amplitude spectrum of the enhanced signal. This represents the amplitude spectrum of the beam output signal. This is the estimated noise power spectrum. This is the over-subtraction factor, used to control the noise suppression intensity. This is the lower limit of the spectrum, used to prevent excessive noise reduction from causing noise.

[0225] Spectral subtraction estimates the power spectrum of ambient noise and subtracts noise components from the spectrum of noisy speech, while setting a lower spectral limit to prevent the generation of musical noise. The over-subtraction factor and the lower spectral limit are dynamically adjusted according to the current signal-to-noise ratio (SNR). At high SNR, a milder noise reduction strategy is used to preserve speech details, while at low SNR, the noise reduction is enhanced to suppress background noise.

[0226] Through the above post-processing, the voice signal corresponding to the effective sound source in the environment can be effectively enhanced, making the dialogue voice clear and prominent while the media volume is suppressed, ensuring that the wearer can have a natural and smooth face-to-face communication.

[0227] Understandably, both media volume smoothing and interactive voice signal enhancement are completed within the first time period, but these two operations are independent. The two operations can be executed simultaneously or in stages; that is, the media volume smoothing operation can be performed first, followed by the interactive voice signal enhancement operation, or vice versa.

[0228] This embodiment achieves spatially selective sound pickup, distance adaptive gain compensation, and spectral reduction and noise reduction post-processing of user speech corresponding to effective near-field sound sources by combining adaptive beamforming with near-field spherical wave compensation and dynamic gain control. This effectively suppresses environmental noise and residual media sound, ensuring that the dialogue speech stands out quickly and clearly in complex acoustic environments, thus improving the user experience of near-field voice interaction.

[0229] During the continuous maintenance of audio coordination, the operating system of the head-mounted audio device also needs to monitor the termination conditions of near-field active voice interaction events in real time to determine whether the near-field active voice interaction events have ended. When the termination of the near-field active voice interaction event is detected, the media volume needs to be promptly restored to its initial state to avoid prolonged suppression that could affect the user's media experience.

[0230] In one embodiment, upon detection of the end of the near-field active voice interaction event, the media volume is smoothly adjusted so that the media volume returns to its initial state before the audio collaborative adjustment operation.

[0231] The conditions for detecting the end of the near-field active voice interaction event include at least one of the following: no dialogue voice directed at the wearer is detected within a preset time window; the deviation angle between the direction of the effective sound source in the near-field active voice interaction event and the direction of the wearer's head is greater than a preset angle threshold, and the duration of the deviation is greater than or equal to a preset first duration threshold; no facial image of the user corresponding to the effective sound source is detected in the visual image data, and the duration is greater than or equal to a preset second duration threshold; and the duration of the continuous deviation of the wearer's head direction is greater than a preset third duration threshold, based on the device posture data.

[0232] Specifically, no conversational voice directed at the wearer was detected within the preset time window. Specifically, voice activity detection can be continuously performed on multi-channel audio data; if within a continuous time window... If the VAD flag remains at 0 for 1 to 2 seconds (e.g., no valid voice activity is detected), it is determined that the user corresponding to the valid sound source has stopped speaking, and the near-field active voice interaction event ends. This condition ensures that short pauses (such as breathing between sentences) are not misinterpreted as the end of the near-field active voice interaction event, while triggering a callback in a timely manner during prolonged silence.

[0233] In near-field active voice interaction events, the deviation angle between the direction of the effective sound source and the orientation of the user's head is greater than a preset angle threshold, and the duration of the deviation is greater than or equal to a preset first duration threshold. Specifically, the horizontal angle between the direction of the sound source and the orientation of the user's head is calculated in real time. ,like If the angle is greater than a preset threshold (e.g., 60°) and the duration of the deviation is greater than or equal to a preset first duration threshold (e.g., 3 seconds), it is determined that the user corresponding to the valid sound source has left the wearer's field of vision, and the near-field active voice interaction event ends. This near-field active voice interaction event condition is used to capture scenarios where the user corresponding to the valid sound source turns away or the wearer's head continuously moves away from the user corresponding to the valid sound source.

[0234] If no facial image of the user corresponding to the valid sound source is detected in the visual image data, and the duration is greater than or equal to a preset second duration threshold, then the near-field active voice interaction event ends. Specifically, the image area corresponding to the direction of the sound source can be continuously detected by the camera. If no facial image of the user corresponding to the valid sound source is detected within the second duration threshold (e.g., 2 seconds), or if the angle of the face of the user corresponding to the valid sound source is away from the user, it is determined that the user corresponding to the valid sound source is no longer in the interaction position or has lost the intention to interact, and the near-field active voice interaction event ends. This near-field active voice interaction event end condition can serve as a visual verification for audio determination, improving the reliability of near-field active voice interaction event end detection.

[0235] Based on device posture data, the duration of a continuous head-turning event detected by the device exceeds a preset third duration threshold. Specifically, if the user's head-turning event continues for more than the third duration threshold (e.g., 3 seconds), it is determined that the user of the head-mounted audio device has actively turned away, and the near-field active voice interaction event ends. This near-field active voice interaction event termination condition is used to capture the user's intention to actively end the interaction.

[0236] The termination conditions for the aforementioned near-field active voice interaction events can be one or more of the conditions combined as the final condition for determining the termination of the near-field active voice interaction event. Alternatively, other conditions can be used, or other conditions combined with one or more of the aforementioned termination conditions can be used as the termination conditions for the near-field active voice interaction event. Furthermore, the termination conditions for each near-field active voice interaction event can be combined using "OR" logic, meaning that the termination determination of the near-field active voice interaction event is triggered when any condition is met.

[0237] This embodiment uses a multimodal fusion detection mechanism (including voice activity detection, sound source direction and head orientation deviation analysis, visual facial recognition, and device posture monitoring) to accurately determine the termination conditions of near-field active voice interaction events. This ensures that a smooth media volume callback operation is triggered in a timely manner when the near-field active voice interaction event ends, avoiding the impact of suppressing the media volume for a long time on the user's media experience. It can effectively prevent misjudgments caused by short pauses or instantaneous deviations, thereby ensuring natural and smooth interaction while achieving intelligent and seamless switching of audio collaborative adjustment states.

[0238] In one embodiment, upon determining that the near-field active voice interaction event has ended, within a preset second time period, the media volume is adjusted back from the second volume level to the first volume level based on a second smoothing control strategy.

[0239] The second smooth control strategy refers to a control method that, upon detecting the end of a near-field active voice interaction event, smoothly restores the currently suppressed media volume of the head-mounted audio device from the second volume level to the first volume level within a second time period.

[0240] Second time period The settings need to take into account both the required callback speed and listening comfort. If the media volume callback is too fast, the sudden increase in volume may cause discomfort to the user; if the media volume callback is too slow, it will affect the continuous experience of media content.

[0241] Understandably, the media volume callback trigger requires the continuous confirmation of the near-field active voice interaction event termination condition. When any near-field active voice interaction event termination condition is met for the first time, it enters the "pre-release" state and starts a confirmation timer. If the near-field active voice interaction event termination condition is met continuously within the preset duration, the media volume callback process is formally triggered; if the near-field active voice interaction event termination condition disappears within the preset duration, it returns to the near-field active voice interaction maintenance state to avoid frequent volume fluctuations caused by misjudgment.

[0242] The second smoothing control strategy can employ a gain change curve that is symmetrical to but inversely related to the first smoothing control strategy, ensuring a natural and smooth callback process. For example, defining a normalized callback time variable... ,in, Indicates the current moment. This is the callback trigger time for media volume. This represents the second time period (total callback duration). The media volume callback curve is represented as follows:

[0243] in, This represents the media volume callback gain function. This represents the normalized callback time variable. This is the second volume level currently being suppressed. The first volume level after the callback. This is a callback transition function.

[0244] The callback transition function uses an S-curve symmetrical to the suppression phase, but adjusts the midpoint to slow down the initial callback speed.

[0245] in, This is a callback transition function. This represents the normalized callback time variable. The steepness of the pullback curve (compared to the suppression phase) Slightly smaller, making the pullback smoother); This is the midpoint offset (shifted forward from 0.5 in the suppression phase to make the initial pullback slower and avoid a sudden increase).

[0246] In one embodiment, during the callback process, the second time period can be... It is divided into multiple sub-stages, using a tiered callback mechanism. For example, the second time period... The process is divided into three phases: an initial gradual increase phase, a stable transition phase, and a final completion phase. In the initial gradual increase phase, the media volume can be slowly reduced, allowing the user to gradually perceive the volume reduction. In the stable transition phase, linear or smooth acceleration can be used to reduce the media volume at a faster rate, stabilizing the reduction and accelerating the process to avoid excessively long reduction times. In the final completion phase, the remaining reduction margin is quickly utilized, bringing the media volume back to its initial state, i.e., the first volume level, ensuring a complete media experience is restored within the second time period.

[0247] During the callback process, environmental changes are monitored in real time. If new potential voice activity is detected (VAD flag changes from 0 to 1), the callback is paused, the current volume level is maintained, and the interaction state is reassessed. If the ambient noise increases significantly, the callback speed is appropriately slowed down to avoid the media volume being too abrupt against a noisy background.

[0248] Furthermore, based on historical interaction patterns, the urgency of wearers to revisit media experiences can be predicted. If historical data shows that wearers typically focus on media content immediately after an interaction, a faster replay speed is used; if wearers tend to continue with environmental awareness after an interaction, a slower replay speed is used.

[0249] This embodiment achieves a smooth, intelligent, and personalized recovery process of media volume from a suppressed state to an initial state by using a pre-release confirmation mechanism and a multi-stage step-by-step callback strategy, combined with a symmetrical S-curve transition function and dynamic parameter adjustment. This effectively avoids the listening discomfort caused by sudden volume increases and the fluctuations caused by frequent misjudgments. Furthermore, the adaptive adjustment ensures the best balance between callback efficiency and continuous media content experience, thereby improving the naturalness of head-mounted audio devices in interactive scene switching and the user experience.

[0250] This embodiment provides a volume adjustment method. This method collects environmental data through a multimodal near-field perception component, achieving accurate perception and judgment of near-field active voice interaction events. This enables the head-mounted audio device to adaptively perceive the occurrence of near-field active voice interaction events, improving the adaptability of scene perception. By judging near-field active voice interaction events directed at the user, the method achieves accurate understanding of the interaction intent, allowing the head-mounted audio device to adaptively trigger an audio collaborative adjustment mechanism. By executing collaborative adjustment operations of media volume and dialogue voice, simultaneously completing media volume suppression and dialogue voice enhancement, dynamic balance adjustment of media volume and dialogue voice is achieved, enabling volume adjustment to automatically adapt to interaction needs. By detecting the end of a near-field active voice interaction event and smoothly restoring the media volume, a natural transition in volume recovery during scene switching is achieved, allowing the head-mounted audio device to adaptively restore its initial state. This enables intelligent and continuous adaptive volume adjustment, improving the scene adaptability of volume adjustment in near-field active voice interaction scenarios.

[0251] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of a first embodiment of a volume adjustment device provided in this application, which is used to perform the aforementioned volume adjustment method.

[0252] like Figure 3 As shown, the volume adjustment device 300 includes: a near-field sensing module 301 and an audio callback module 302.

[0253] The near-field sensing module 301 is used to collect multimodal near-field sensing data of the current environment of the user wearing the head-mounted display device through the multimodal near-field sensing component; The audio adjustment module 302 is used to execute the audio collaborative adjustment operation of the head-mounted audio device in response to the near-field active voice interaction event when the multimodal near-field perception data meets the triggering conditions of the near-field active voice interaction event for the wearer. The audio collaborative adjustment operation includes at least one of the media volume adjustment operation and the dialogue voice adjustment operation.

[0254] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the device and each module described above can be referred to the corresponding processes in the aforementioned volume adjustment method embodiments, and will not be repeated here.

[0255] The apparatus provided in the above embodiments can be implemented as a computer program, which can be used in, for example... Figure 4 It runs on the head-mounted audio device shown.

[0256] Please see Figure 4 , Figure 4 This is a schematic block diagram of a head-mounted audio device provided in an embodiment of this application. The head-mounted audio device may be a server.

[0257] See Figure 4 The head-mounted audio device includes a multimodal near-field sensing component, as well as a processor, memory, and network interface connected via a system bus, wherein the memory may include non-volatile storage media and internal memory.

[0258] The multimodal near-field sensing component is communicatively connected to the processor. The multimodal near-field sensing component includes at least one sensing device such as a microphone array, an inertial measurement unit (IMU), and a camera. The microphone array is used to collect audio data of the near-field environment of the user wearing the device; the IMU is used to collect posture data and motion state of the head-mounted audio device or the user's head; and the camera is used to collect visual image data. Audio data, posture data, motion state, and visual image data together constitute multimodal near-field sensing data. The multimodal near-field sensing data is transmitted to the processor via a data channel between the multimodal near-field sensing component and the processor, so that the processor can execute any volume adjustment method based on the multimodal near-field sensing data.

[0259] Non-volatile storage media can store operating systems and computer programs. These computer programs include program instructions that, when executed, cause the processor to perform any volume adjustment method.

[0260] The processor provides computing and control capabilities to support the operation of the entire head-mounted audio device.

[0261] Internal memory provides an environment for the execution of computer programs stored in non-volatile storage media. When these computer programs are executed by the processor, the processor can perform any volume adjustment method.

[0262] This network interface is used for network communication, such as sending assigned tasks. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the head-mounted audio device to which the present application is applied. A specific head-mounted audio device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0263] It should be understood that the processor can be a Central Processing Unit (CPU), but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among these, a general-purpose processor can be a microprocessor or any conventional processor.

[0264] In one embodiment, the processor is configured to run a computer program stored in memory to perform the following steps: The multimodal near-field sensing component collects multimodal near-field sensing data of the current environment of the user wearing the head-mounted display device. When the multimodal near-field perception data satisfies the triggering conditions for a near-field active voice interaction event for the wearer, the audio co-adjustment operation of the head-mounted audio device is executed in response to the near-field active voice interaction event, wherein the audio co-adjustment operation includes at least one of media volume adjustment operation and conversational voice adjustment operation.

[0265] The embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions, and the processor executing the program instructions to implement any of the volume adjustment methods provided in the embodiments of this application.

[0266] The computer-readable storage medium can be an internal storage unit of the head-mounted audio device described in the foregoing embodiments, such as the hard drive or memory of the head-mounted audio device. Alternatively, the computer-readable storage medium can be an external storage device of the head-mounted audio device, such as a plug-in hard drive, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard equipped on the head-mounted audio device.

[0267] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A volume adjustment method, characterized in that, For a head-mounted audio device, the head-mounted audio device being provided with a multimodal near-field sensing component, the method includes: The multimodal near-field sensing component collects multimodal near-field sensing data of the current environment of the user wearing the head-mounted display device. When the multimodal near-field perception data satisfies the triggering conditions for a near-field active voice interaction event for the wearer, the audio co-adjustment operation of the head-mounted audio device is executed in response to the near-field active voice interaction event, wherein the audio co-adjustment operation includes at least one of media volume adjustment operation and conversational voice adjustment operation.

2. The volume adjustment method according to claim 1, characterized in that, The multimodal near-field sensing component includes at least a microphone array, an inertial measurement unit, and a camera; The step of collecting multimodal near-field perception data of the current environment of the user wearing the head-mounted display device through the multimodal near-field perception component includes: The microphone array is used to collect multi-channel audio data of the user's current environment. The device attitude data of the head-mounted audio device is acquired through the inertial measurement unit; The camera captures visual image data of the user's current environment.

3. The volume adjustment method according to claim 2, characterized in that, Before executing the audio coordination adjustment operation of the head-mounted audio device when the multimodal near-field perception data meets the triggering conditions for a near-field active voice interaction event for the wearer, the method further includes: Based on the multi-channel audio data, sound source localization and voice activity detection are performed to identify the sound source location information of effective sound sources in the current environment of the wearer. The sound source location information includes the sound source direction and the sound source distance. When the distance to the sound source is less than a preset near-field range threshold, the interaction intent parameter of the user corresponding to the effective sound source is calculated based on the sound source direction, the device posture data, and the visual image data. When the interaction intent parameter meets the preset conditions, it is determined that the multimodal near-field perception data meets the triggering conditions for near-field active voice interaction events for the wearer.

4. The volume adjustment method according to claim 3, characterized in that, The step of calculating the user's interaction intent parameters corresponding to the effective sound source based on the sound source direction, the device posture data, and the visual image data includes: Based on the sound source direction and the device posture data, calculate the evaluation component of the horizontal angle between the sound source direction and the direction of the user's head; Based on the visual image data, the facial orientation angle of the user corresponding to the effective sound source is detected, and the orientation consistency evaluation component between the facial orientation angle of the user corresponding to the effective sound source and the facial orientation angle of the wearer is calculated. Temporal continuity analysis is performed on the multi-channel audio data to normalize the speech duration in the multi-channel audio data into a speech continuity evaluation component. The horizontal angle evaluation component, the orientation consistency evaluation component, and the speech continuity evaluation component are weighted and summed to obtain the interaction intent parameter corresponding to the effective sound source.

5. The volume adjustment method according to claim 1, characterized in that, Performing audio coordination adjustment operations on the head-mounted audio device includes: Within a preset first time period, based on a first smoothing control strategy, the volume of the media currently being played by the head-mounted audio device is adjusted from a first volume level to a second volume level; and, During the first time period, signal enhancement processing is performed on the dialogue speech in the near-field active voice interaction event to enhance the speech signal of the user corresponding to the effective sound source from the ambient sound.

6. The volume adjustment method according to claim 5, characterized in that, The step of smoothly adjusting the media volume so that the media volume returns to its initial state before the audio collaborative adjustment operation includes: During a preset second time period, based on a second smoothing control strategy, the media volume is adjusted back from the second volume level to the first volume level.

7. The volume adjustment method according to claim 2, characterized in that, The conditions for detecting the end of the near-field active voice interaction event include at least one of the following: No conversational voice messages directed at the wearer were detected within the preset time window; In the near-field active voice interaction event, the direction of the effective sound source deviates from the orientation of the user's head by an angle greater than a preset angle threshold, and the duration of the deviation is greater than or equal to a preset first duration threshold. The visual image data does not detect a valid sound source corresponding to the user's facial image, and the duration is greater than or equal to a preset second duration threshold. Based on the device posture data, it was detected that the duration of continuous head tilt of the user was greater than a preset third duration threshold.

8. A volume control device, characterized in that, The volume adjustment device includes: The near-field sensing module is used to collect multimodal near-field sensing data of the current environment of the user wearing the head-mounted display device through the multimodal near-field sensing component; An audio adjustment module is configured to perform audio collaborative adjustment operations of the head-mounted audio device in response to the near-field active voice interaction event when the multimodal near-field perception data meets the triggering conditions for the near-field active voice interaction event for the wearer. The audio collaborative adjustment operations include at least one of media volume adjustment operations and dialogue voice adjustment operations.

9. A head-mounted audio device, characterized in that, The head-mounted audio device includes a multimodal near-field sensing component, a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, implements the steps of the volume adjustment method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the volume adjustment method as described in any one of claims 1 to 7.